The AI Firehose Podcast
← All episodes
Episode 1 · May 17, 2026 · 10:47

NEW Google Gemini DESTROYS GPT-5.5?

Google Gemini 3.2 Flash Leaked: The GPT-5.5 Killer?


A massive new leak reveals Google's Gemini 3.2 Flash, a model that reportedly matches GPT-5.5 performance at a fraction of the cost. Discover how the new liquid glass design and upcoming AI agents could revolutionize your business automation workflows. We break down the benchmarks, technical specs, and what to expect from Google IO.


00:00 - Intro

00:31 - Gemini 3.2 Flash Leak Spotted

01:10 - Benchmarks: Beating Pro Models

02:04 - Price Performance vs GPT 5.5

04:15 - The Power of Distillation

05:41 - AI Workflows for Small Business

07:41 - Leaked: AI Agents Beta Feature

09:01 - Preparing for Google I/O

Full transcript

New Google Gemini destroys GPT-55. A new Google AI model called Gemini 3.2 Flash has been spotted hiding inside the iOS Gemini app and from what early testers have seen it looks like it's going to wreck GPT-55. Note before we go deeper, everything I'm about to share is based on leaked screenshots, leaked metadata and early benchmark sightings. Google has not officially confirmed any of it yet.

Treat this as a rumor breakdown, not a press release. Google I.O. is just days away. We'll know for sure on May 19th.

The leaks are big enough to walk through right now. Here's what happened. A new Google AI model called Gemini 3.2 Flash has been spotted hiding inside the iOS Gemini app and from what early testers have seen it looks like it's going to wreck GPT-55. Let me walk you through it.

On May 5th, a Reddit user opened the Gemini app on their iPhone. The model picker started cycling on its own. Gemini 3 Flash landed on something nobody had ever seen. Gemini 3.2 Flash.

A model Google hasn't even announced yet. The same user noticed a brand new look in the app. A pill-shaped prompt box. Glowing gradient background.

A model picker moved to the top left corner. Google's reportedly calling this design liquid glass. It looks gorgeous. The design isn't the real story.

The real story is what this model can apparently do. A tracker on X named Waguri Kawaruko grabbed a screenshot from iOS build 1.2026.1775. He's known for spotting Google's internal builds before they go public. Within hours, another tracker, AI Battle, caught a mystery model running silent tests on LM Arena.

That's the blind benchmark site where Google likes to stress test new models before launch. The mystery model crushed it. It built a full-screen ASCII animation of a city on a hill in under two minutes. Gemini 3.1 Pro, Google's current flagship, took five and a half minutes and produced broken code that didn't even run.

These leaks hold up. A small, cheap, fast model is now beating Google's own flagship at creative coding. I'll say that one more time, slower. Tiny Flash model appears to be beating their own giant pro model.

That changes everything. And here's why it matters for anyone running a business with AI. Bindu Reddy, she's the CEO of Abacus AI, confirmed the rumors this week. According to her sources, Gemini 3.2 Flash hits 92% of GPT 5.5's performance on coding and reasoning tasks while running at roughly 1 15th to 1 20th the cost of GPT 5.5 millisecond response times on most queries.

That's faster than you can blink. Three things appear to line up at once in this leak. Speed, quality, and cheap to run. That hasn't happened in AI before.

You used to pick two of those, now you might get all three. In one model from one company, let me show you what that actually means in real life. Imagine you want to send 200 personalized first touch messages to local businesses who could use AI automation. With a top tier model that runs slow and costs a lot, the tiny cheap model, the messages come out generic and bad.

With a leak like Gemini 3.2 Flash, the messages would be smart, personalized, written in under three minutes total at a tiny fraction of the cost. That's the unlock. A small business can now do enterprise level outbound every single day. Quick pause here.

Look, Google is shipping new models like software updates now. Gemini 3 Flash in late 2025. Gemini 3.1 Flash in March. And now this leaked Gemini 3.2 Flash showing up this week.

That's three big flash steps in seven months. The pace is only going up. Keeping up with all of it on your own is basically impossible. Inside the AI profit boardroom, the second Gemini 3.2 Flash goes live.

We're walking you through exactly how to plug it into your business step by step. We've already built a 30-day Gemini automation roadmap. How to use it for lead gen, cold outreach, content, customer support, and follow-ups. Four live coaching calls every single week where we go deep on Gemini workflows.

Daily new tutorials. A prompt library full of Gemini tested prompts for getting more customers. Since owners inside, a lot of them already running Gemini for client work right now. And a member map so you can DM and meet up with other Gemini users near you.

Link in the description or go to AIprofitboardroom.com. Now back to the leak. Let me explain how Google might have pulled this off. Because the technical story matters.

Reports say Google is using something called distillation plus sparsity techniques. Big words. Here's what they actually mean in plain English. Imagine you have a brilliant teacher with 30 years of experience.

That teacher knows everything in their field. They're expensive to hire. Kind of slow because they carry so much in their head. Now imagine you make a smaller faster version of that teacher.

Same answers. Same instincts. Just strip down to the stuff that actually matters for the questions you usually ask. That's distillation.

Google reportedly takes their biggest most expensive model. Strip it down. What comes out is a model that gives you near pro quality at flash speed at flash cost. Big model trained the small model.

And the small model now does most of the work just as well. This is why OpenAI may have a real problem on its hands right now. GPT 5.5 came out April 23rd. OpenAI charges premium prices for it.

Roughly double what GPT 5.4 cost. And Greg Brockman, OpenAI's president, called it a step toward agentic computing on the launch core. Big language. Big price tag.

Then Gemini 3.2 flash apparently shows up two weeks later. A fraction of the cost with 92% of the quality if the leaks are right. For most business use cases that could be a game over moment. Think about it for a second.

If a model gives you 92% of GPT 5.5 at a tiny fraction of the cost, why would you keep paying full price for GPT 5.5? For most things you wouldn't. Real talk. Most people using AI for their business are doing stuff like this.

Writing sales emails. Researching leads. Summarizing customer calls. Building landing pages.

Generating social posts. Answering inbound questions. Drafting proposals. Running follow-up sequences.

None of that needs the absolute smartest model on earth. It needs a model that's smart enough, fast enough, and cheap enough to run all day every day without burning a hole in your budget. Gemini 3.2 flash, if these leaks hold, is that model. Example.

You could build a workflow where every new lead that fills out a form on your site gets researched, profiled, and sent a custom first email all in under 60 seconds. No human in the loop. Running on a leaked style price like this, it would cost almost nothing per lead. That's the kind of automation that used to take a team of three people to run.

Or another one. You could run a content engine that pulls trending topics in your space, writes five LinkedIn posts a day in your voice, and queues them for review every morning. Running on a small, fast, cheap model. It just works.

The knowledge cutoff rumor matters too. Leaks suggest Gemini 3.2 flash is updated through January 2026. Most older AI models are stuck two years back. They give you outdated info.

They don't know about new tools. They don't know what's trending now. Gemini 3.2 flash, if the rumor holds, knows what's happening in your industry this quarter. For example, if you ask it about the latest AI tools to plug into your customer onboarding, it would actually name the new ones from 2026, not the dead tools from 2023.

Google is also reportedly pushing hard on grounding. That means the model checks search results before answering. Less hallucination. Fewer made-up facts.

More real-world reliability. For a business owner, that's the difference between an AI you can trust with client work, and an AI that embarrasses you with confident lies. Okay, so when does this drop? Google IOO is May 19th and 20th.

That's just days away from when I'm filming this. Most analysts think Gemini 3.2 flash gets its official announcement at the keynote on the 19th. Some think Google might announce it one or two days early to build hype. Either way, if the leaks are real, this is happening this week.

And there's one more thing in the leak that almost nobody is talking about. The same Reddit user who found Gemini 3.2 flash also spotted a brand new sidebar tab in the iOS app. It says agents beta. It's greyed out right now, waiting.

So Google might not just be shipping a new model. They could be shipping AI agents inside Gemini. Agents that take real actions. Browse the web for you, book stuff, fill out forms, manage your inbox, the whole thing.

This is the race open AI was hinting at when Greg Brockman called GPT 5.5 a step toward a new way of getting work done on a computer race. Google might be pulling ahead. And here's why Gemini 3.2 flash is probably the engine that runs those agents. Because agents are token monsters.

A single agent might make 50 model calls to finish one task. If each call burns through compute, you can't run it long. If each call costs basically nothing, you can run an army of agents. This leak isn't a small thing.

Google may have just handed every business owner a cheap engine to run AI agents on at scale and then teased that the agents are coming in the same week. If Gemini 3.2 flash launches the way the leaks suggest, it's going to change how normal businesses use AI. The reason is simple. It would be the first time smart, fast and cheap have all lined up in one model.

When that happens, AI stops being a nice to have tool. Becomes the default way work gets done. Two days from now, Google is probably going to announce a model that hits 92% of GPT 5.5 quality at a tiny fraction of the cost. Runs in under 200 milliseconds.

Powers an entire new agent system that takes actions for you. Business owners who plug this into their workflows in the first 30 days are going to get a massive head start. The ones who wait six months are going to be playing catch up for the rest of the year. Inside the AI profit boardroom, the moment Gemini 3.2 flash goes live, confirmed or rebranded, we're running live coaching calls walking you through exactly how to use it to land more customers, fill your pipeline, automate your follow-ups and free up your time.

Not just any coaching calls, Gemini specific ones. Four every single week. Daily new tutorials built around Gemini workflows for getting more leads and customers. 30-day Gemini implementation roadmap so you know exactly what to do in week 1, week 2, week 3, week 4.

A prompt library full of Gemini tested prompts for real business use cases like cold email, sales calls, content, lead research and customer support. Since owners in there right now, a lot of them already running Gemini for client work, automation builds and lead gen. And a member map so you can DM and meet up with other Gemini users in your city or your industry. There's always someone online 24x7 to help you when you get stuck.

Link in the description or go to aiprofitboardroom.com. And if you want the full process, all the SOPs plus 100 plus AI use cases like this one, totally free, join the AI Success Lab. It's our free AI community. You'll get all the video notes from this episode in there, plus access to 67,000 members already crushing it with AI.

Link in the comments and description. That's the leak. That's what it could mean for your business. And that's why the next 48 hours are going to be the biggest AI release window of 2026.

I'll see you in the next one.

More episodes

Browse all episodes →