Bye Bye Google: How Brain Drain is Killing Gemini AI
Recently, I wrote a blog post about what I think the future of AI holds. In that post, I said the Gemini 3.5 series should be the last Google model that cannot do tool calls. Since we have seen the future of agents, Google models are in the worst position because they struggle to call tools properly. I think I need to go more in-depth about what I think about Google right now.
Bye Bye Google
Recently, Noam Shazeer of the Google DeepMind team left. Why is this important? Because he's one of the people who helped work on the Transformer paper the paper that made modern AI possible. It's pretty bad that he left for OpenAI. Then there is John Jumper this guy worked on AlphaFold and is a Nobel Prize winner. That guy also recently left Google for Anthropic. And then there is the guy who made the Google Workspace CLI, Justin Poehnelt. He worked on an interesting program that could help agents access your Google Workspace for the agentic future that everyone is preaching about. But what happened? He got fired. This is a company that is firing people for making good projects, but we think this company is going to innovate? I think innovation happens when a problem meets an interesting thinker. Most products we use today are just people brain dumps , and if your company rewards good brain dumps , there will be more innovation there than if it doesn't. Google's culture right now is just not rewarding brain dumps .
Left Behind
Google's latest model, Gemini 3.5 Flash, is a good chatter it's really smart and fast, but it is not good for the agentic era. Gemini overthinks. The best place to see this phenomenon is on the DeepSWE benchmark, where it takes 200k tokens the most on that benchmark for half the score of Claude Opus 4.8 and GPT 5.5. The Gemini Flash model is comparable to Claude Sonnet 4.6, which is an old-ass model, and that's what the model is being compared to on DeepSWE, and it's still not winning that benchmark. Even when they stop overthinking, they can't call tools consistently. While 3.5 Flash is supposed to be better at calling tools than on other benchmarks, I have seen Gemini fail at calling tools completely. Gemini models are serious work to get working in a proper harness. I use T3 Chat as a chat app, and there is a search feature where you can specify if you want to search once or twice. With other models, when you pick "once," they search once. Google models are the only ones that try to search twice. Google models are too eager to use tools, but when they actually have the tools, they fail at calling them.
Can Google Still Build Good Dev Tools?
During their Google I/O event, Google released Antigravity 2.0, which has a CLI, SDK, app, and IDE. First, they messed up their launch of the app; when people updated, they found themselves using the app when they wanted to use the IDE. They fixed that, but then they fucked up by killing the Gemini CLI. What I thought was that they would merge the Gemini CLI, Antigravity, and maybe the Jules team to make the best AI coding system. But they did not. Instead, what they did is release a bad model that can barely call tools half the time and overthinks the other half, paired with a crappy CLI that sucks. Instead of trying to fix it, they would be better off asking Gemini to port the Gemini CLI to Go and change the name, or just leave Gemini as-is and change the name to the Antigravity CLI. And the SDK is good when someone needs it, I guess, but no, they did not do that. They made a worse product and called it a better product. And it's not even open source! Why close source it? I was not going to contribute, but I would have loved to fork it and build things that work with my workflow. But no, they close-sourced it.
Pay More, Get Less
When Google released the 2.0 series, Flash was positioned as balanced intelligent enough for the price. But now, Gemini 3.5 Flash is 3x the price of 3.0 Flash, while 3.0 Flash was 1.5x the price of 2.5 Flash, which was 2x the price of the input and 6x the price of the output compared to 2.0 Flash. Then, if you want to see the price jump from 3.5 Flash to 2.0 Flash, its input is 15x and output is 22.5x the price of 2.0 Flash. And the reason for this pricing is because they said the model is good for agentic stuff, but as we can see, it is not. It's crazy how Cursor, which is not an AI lab, could make a model orchestration better than Gemini 3.5 Flash. It's funny that GLM 5.2 is better than every other model that has been released to date for agentic stuff, and you might say it's new and you don't see it beating Opus 4.8 or GPT 5.5 on benchmarks that matter, but it's out there.
Stop Benchmaxxing, Start Engineering
- Stop benchmaxxing: Make a model that is actually usable not necessarily cheap, but worth its price.
- Fix the overthinking: Streamline the reasoning tracks so it doesn't waste tokens.
- Train with modern techniques: Leverage reinforcement learning effectively to make the model better.
If you want to learn about reinforcement learning, I wrote a blog post about AI, so click here.
- Fix the culture: Change the culture to make people enjoy testing and creating better things.
- Give early access: Give me and other trusted people early access to the models. To give credit where it's due, I think they have started doing some of these things. The only thing they haven't done yet is give early access, but I'm patiently waiting for my invite. So Google, you know my email!
Conclusion
Gemini 4 is the last dance. Not because Google doesn't have the money, talent, or resources, but just because if that drop is bad, they will end up like Mistral in the sense that when they drop something, no one will care. Once agents become fully commercialized, Google needs a killer model to win back the public.
You can read more of my blog posts on my website, or just click here. If you want to reach out, I will link my X and other platforms below:
- X: @JemoLife0213
- GitHub: Jemo69
- BlueSky: @jemolife