The AI Firehose Podcast
← All episodes
Episode 112 · July 22, 2026 · 08:26

These 3 NEW Gemini Models are INSANE!

Full transcript

These three new Gemini models are insane. Google just dropped three brand new Gemini models. And no, this isn't another boring AI update. One is smarter but uses way fewer tokens.

One is built to run tiny tasks all day long and one hunts down security bugs in your code before bad guys find them. I'm going to break down all three in plain English and I'll show you exactly how to use them in your business today. Let's get into it. So here's what happened.

Google didn't release one big model and call it a day. They released three. Each one does a totally different job. And that tells you everything about where AI is going right now.

Because for the last two years, every AI company was chasing the same thing. Bigger, smarter, higher scores on tests. Nobody outside of AI Twitter cares about. Google just said, no, thanks.

We're going a different way. The first model is Gemini 3.6 Flash. This is the star of the show. Here's the headline.

It gives you better answers. It uses fewer tokens to do it. Dot. That last part is wild.

Normally, when a model gets better, you pay more. Google flipped that. They made it more efficient instead. Now let me explain tokens, because this is the whole game.

Every time you talk to an AI, your words get chopped into little pieces called tokens. The AI reads those pieces. Then it writes back in pieces too. Every single piece gets counted.

So imagine you ask a question. The old model needs a thousand tokens to answer it. The new one does the same job in 700. Same answer, sometimes better.

That's a huge deal when you're running things at scale. Because if you're just chatting with AI once a day, who cares? But if you've got an AI answering customer emails all day, every day, forever, those tokens add up fast. Hey, if we haven't met already, I'm the digital avatar of Julian Goldie, CEO of SEO agency, Goldie Agency.

While he's helping clients get more leads and customers, I'm here to help you get the latest AI updates. And Julian reads every single comment. So drop a comment below and let him know which of these three models you're most excited about. All right, back to it.

So what do you actually do with a model that's more efficient to run? You stop running one agent. You start running five. Let me give you a real example.

Say you're growing a community like the AI Profit Boardroom. Before, you'd have one AI assistant doing everything badly. Now you can run a whole team. You'd set up a research agent with a prompt like this.

Search the web every morning for the newest AI automation tools and news. Then write me a short summary of the three most useful ones for small business owners, then a content agent. Take that research summary and turn it into five short posts explaining how each tool saves time for business owners in the AI Profit Boardroom. Then a support agent.

Read every new question posted in our community, find the answer from our existing guides, and draft a helpful reply in a friendly tone. Then a planning agent. Look at what our members asked about this week, then suggest three new training topics we should create next. And a quality checker.

Review everything the other agents wrote, fix anything unclear, and flag anything that sounds robotic. Five agents all running at once, all doing real work. That's what token efficiency actually buys you, not a slightly better chatbot. A whole team.

Now, if this is starting to click for you and you want the full playbooks for setting up these exact agent workflows step by step, come check out the AI Profit Boardroom. It's my community where I show you exactly how to automate your business with AI tools like Gemini, how to get more customers, how to save hours every single week by handing the boring stuff to AI. You'll get the actual agent setups, the prompts, and the workflows I use. Not theory, real stuff you can copy today.

The link is in the description and pinned in the comments. Go grab your spot. Alright, back to Gemini 3.6. Flash for a second, because there's one more thing worth knowing.

Companies deploying AI at scale barely look at benchmark scores. They look at speed. They look at reliability. They look at whether it can handle a million requests without falling over.

Gemini 3.6 Flash goes straight at all three. That's why this release matters more than the headline suggests. Okay, model number two, Gemini 3.5 Flashlight. This is Google's small, fast, no-frills model.

Think of it like Claude Haiku or the mini versions from other companies, built to be light and quick. And I know what you're thinking. Why would I want a smaller model? Here's why.

Most of the work you need AI to do is boring. Reading a PDF, pulling data out of an invoice, sorting emails into folders, searching through documents, simple yes or no decisions. None of that needs a genius model. And when you're doing thousands of those tiny jobs a day, using the right size model makes a massive difference.

Let me show you what this looks like in practice. Say someone joins the AI profit boardroom and fills out a welcome form. Flashlight reads it and sorts them into the right group based on their business type. Or you've got 100 pages of training notes.

Flashlight reads all of them and pulls out every actionable tip into one clean list. Or a member sends a long message describing their problem. Flashlight reads it, tags it by topic, and routes it to the right guide. Boring?

Yes. Useful? Massively. Here's a prompt you could steal for it.

Read this member message, decide if it's about content, automation, sales, or tech, and reply with just that one word. That's it. One word back. Fast, clean, done.

And this is what real AI automation actually looks like. Not one magic robot. Lots of small pieces doing small jobs perfectly. All right, model number three.

And honestly, this one's the most interesting. Gemini 3.5 Flash Cyber. Google built an AI model just for cyber security. Let me be really clear about what this is.

It's not a hacking tool. It's the opposite. It's defensive. Think of it like an AI security guard for your code.

What does it do? It reads through software and looks for weak spots, places where someone could break in. Then it suggests how to patch them. It can do static code analysis, bug detection, vulnerability discovery, security recommendations, patch generation, threat analysis.

Basically, it's a security analyst that never sleeps and never gets bored. And here's why Google built this now. More and more code is being written by AI. That's just a fact.

Developers everywhere are using AI to build faster. But here's the problem. If AI is writing the code, who's checking if it's safe? That's the gap.

And it's a big one because AI writes code fast, but fast doesn't mean secure. So Google's answer is simple. If AI is going to write the code, AI should also check the code. And this matters even if you're not a developer.

Say you've got automations running your business, forms collecting member info, scripts moving data between tools, systems handling private details. Now let me zoom out because there's a bigger story here. Google is splitting Gemini into specialists. Coding gets one.

Cybersecurity gets one. Lightweight tasks get one. Fast production work gets one. Each one gets tuned separately for its own job.

And they're not alone. OpenAI is doing it. Anthropic is doing it. Everyone's realizing one model can't be perfect to everything.

But here's the takeaway that I think most people are going to miss. This launch isn't really about intelligence. It's about efficiency. Every headline you'll see about this will talk about benchmarks, scores, charts, who beat who on some tests.

That's not the story. The story is that AI is getting lighter to run at scale. And when something gets lighter to run, people run more of it. More agents, more automation, more stuff happening in the background while you sleep.

Think about what that unlocks. Right now, most businesses use AI like a helper. You ask it something, it answers, you go back to work. But when running AI gets efficient enough, you stop asking.

You just set it up and let it go. A research agent that runs every morning, a content agent that runs every afternoon, a support agent that never clocks out, a security agent watching your systems all night. That's not five years away. That's now.

That's what these three models are for. And here's the honest truth about who wins with this. Not the person with the smartest model. The person who actually sets it up because everybody's got access to the same tools, same models, same prompts if they want them.

The difference is who builds the workflow and who just watches videos about it. So here's what I do. Pick one boring task in your business. Just one.

Something you do every week that you hate. Sorting messages, summarizing calls, cleaning up notes. Give that one task to flashlight. Just that one.

Get it working. Then add another, then another. That's how you build an AI system. Not all at once.

One boring job at a time. Now, if you want to actually build this stuff, instead of just hearing about it, join the AI Profit Boardroom. That's where I show you the exact automations, the exact prompts, and the exact step-by-step process to run your business with AI tools like Gemini, how to get more customers, how to save hours every week, how to build systems that work while you're doing something else. Link is in the description and pinned in the comments.

And if you want the full process, SOPs, and over 100 AI use cases like this one, join the AI Success Lab. It's free. Links in the comments and description. You'll get all the video notes from there, plus access to our community of 38,000 members who are crushing it with AI.

Drop a comment and tell me which of these three models you're going to try first. Julian reads every single one. I'll see you in the next video.

More episodes

Browse all episodes →