Jeremy Nwachukwu // Field notes

Gemini 3.6 Flash Review: Not Frontier, But a Step in the Right Direction

103 views July 22, 2026

I wrote about the end of Google recently, focusing on their loss of talent and shifting company culture. I said Google has to impress everyone with Gemini 4 so they don't fall behind and die in this AI race, and I stand by everything I said. But while everyone was waiting for Google's next fabled frontier model, Google quietly dropped a few new releases: Gemini 3.6 Flash, 3.5 Flash Lite, and Gemini 3.5 Cyber. This post is a quick first-impression write-up and a defense of Google, not one of my in-depth model reviews like the one I wrote for Muse Spark 1.1. But just to touch on it: Gemini 3.5 Flash Lite is surprisingly capable for its price range. I wouldn't use it to write an entire app, but for building single features? Absolutely. I don't have an opinion on the Cyber model yet because I haven't had access to it. Also, Google hasn't given up on Gemini 3.5 Pro—they stated it's currently in early testing with trusted testers. (And Google, why am I not a trusted tester? I can test, I'm unbiased, and I write solid model reviews!) Anyway, let's dive into why people are wrong about this release. Google Gemini 3.6 Flash has roughly the same intelligence score as Gemini 3.5 Flash, beating out GPT 5.6 Luna, Grok 4.5, and Muse Spark 1.1. But that’s not what Google used this release for. This was not a step up in raw intelligence—it was a step up in efficiency. Let's compare this model to the 3.5 Flash model that came before it. This chart from DeepSWE shows what Google is targeting. The x-axis represents cost, and the y-axis represents the DeepSWE score. Here you can see it scores higher and is cheaper than any other Google model. It’s not the absolute best model overall, but it's a solid step forward. This release wasn't meant to be the absolute best or a frontier-defining model; it was meant to be an efficiency step, and it gives me hope for the future.

Cost

When it comes to pure cost, it’s not the absolute most efficient option out there. There are models with better cost-per-token metrics. To name a few:

  • Muse Spark 1.1
  • Grok 4.5
  • GPT 5.6 Luna Are these models more cost-efficient than Gemini 3.6 Flash? Yes. But on the flip side, there's also Sonnet 5, which might be one of the most token-inefficient models ever. Let's look at the Artificial Analysis cost benchmark:

This measures cost per task. As you can see, Sonnet costs significantly more, and Mistral Medium 3.5 is even more expensive despite being a worse model. To be fair, if I had to be in a tough spot in the AI race, I'd rather be Google than Mistral. That's just my opinion—even though Mistral hasn't dropped anything relevant lately. Honestly, they should probably stick to making DeepSeek distills and leave frontier modeling to the big boys.

This chart shows us that 3.6 Flash is more efficient than several competing models, even if some models in its price bracket edge it out. This isn't a life-changing update, but it's a step in the right direction.

Conclusion

The model benchmarks at or slightly above the same level as its previous version because it’s fundamentally the same base weight with more Reinforcement Learning (RL) applied. We need to stop "benchmaxxing" and actually focus on real-world usability. I give this model a 5/10. It’s not mind-blowing, but if this is what holds us over until the next major Gemini release, I won't complain. Gemini 3.5 Pro is in early testing, and the Gemini 4 pre-training run has officially started. Remember to follow me on social media:

  • X:
  • GitHub:
  • Bluesky: You can read more of my writing here:

And as always... Stay coding.

© 2026 Ifeanyichukwu Jeremy Nwachukwu // Tactical Terminal