Moonshot AI's Kimmy K 2.7 Code delivers blistering speeds of up to 260 tokens per second, transforming how AI agents operate in real-time. Learn how to pair this massive 1-trillion parameter model with Agent OS and Hermes to build lightning-fast, high-context automated workflows.
00:00 - Intro: Kimmy K 2.7 Code Drops
01:23 - The Speed Advantage for AI Agents
02:34 - Setting Up Agent OS with Hermes
04:30 - 1 Trillion Parameters Explained
05:07 - The Power of 256k Context
06:26 - Benchmarks vs GPT & Claude
07:29 - Real-World Workflow Automation
09:35 - The Future of Open Weight Models
Full transcript
Kimmy K 2.7 code plus agent OS just changed AI agents. Kimmy K 2.7 code just dropped and it's running at 260 tokens per second. To put that in perspective, that's up to six times faster than what Moonshot AI was shipping just weeks ago. Six times on a model that's already one of the strongest open source coding agents in the world right now.
This is Moonshot AI. They're a Beijing based lab backed by Alibaba. They launched their first model in 2023 and in under a year from July 2025 to right now they've shipped five major model versions. K2, K2 thinking, K2.5, K2.6 and now K2.7 code.
That's not a slow research lab, that's a team shipping fast. The new model dropped June 12th 2026 and then three days later on June 15th they announced high speed mode. A separate model variant called Kimmy-K 2.7-code-highspeed that takes that same model and fires it at roughly 180 tokens per second on normal coding tasks. Short context, you're hitting 260 tokens per second.
Here's why that actually matters. When you're running an AI agent on a task, something like write a new onboarding sequence for the AI profit boardroom, research what members ask about most, draft a week of email follow-ups. The agent isn't just answering once, it's thinking, it's calling tools, it's writing, checking, revising. All of that is tokens and the faster the model outputs tokens the faster your agent finishes the task.
With slow models you're waiting, you send a task, you get coffee, you come back. Hey if we haven't met already I'm the digital avatar of Julian Goldie, CEO of SEO agency Goldie Agency. Whilst he's helping clients get more leads and customers I'm here to help you get the latest AI updates. Julian Goldie reads every comment so make sure you comment below.
With K2.7-code-highspeed running in your agent you're watching it finish tasks in real time and there's something else that most people aren't talking about yet. This model thinks 30% less than the previous version while scoring higher on coding tasks. The old knock on reasoning models, models that think before they respond, was that they burned through tokens doing unnecessary mental gymnastics. K2.6 was already good but K2.7-code cuts the overthinking by 30% so you're getting more output faster with fewer resources burned getting there.
On Moonshot's own benchmarks K2.7-code is up 21.8% on their coding eval versus K2.6 up 31.5% on MLS Benchlight and up 10% on agentic capabilities overall. Independent verification on the big public leaderboards is still coming, there are no third-party SWE-bench numbers yet but the directional signal is clear. This is a meaningful jump. Now let me show you what this actually looks like in practice because the way I'm running this is on AgentOS, my local AI dashboard with the K2.7-highspeed API key plugged into Hermes Agent and I want to talk about that inside AI Profit boardroom right now because we've already built out the full AgentOS workflow around Kimi K2.7-code high speed.
Inside the boardroom we've got a 30-day roadmap for getting Hermes Agent running on AgentOS with K2.7-code as the back-end step-by-step tutorials for setting up the API key and running your first agentic task and four weekly coaching calls where we go deep on exactly this setup. There are members in there right now running Hermes Agent with Kimi for content automation, lead follow-up and community management. If you want the full AgentOS zip file pre-configured and ready to install with Kimi K2.7 plugged in it's inside the AI Profit boardroom, link in the comments and description or go to AIProfitBoardroom.com. So let's talk about how this is actually set up.
AgentOS runs on your laptop, it's a local dashboard, what I call a mission control for your AI agents. One screen, you can chat with your agents, talk to them by voice, save your conversations to notes, track your goals, everything stays on your machine, no subscriptions, no accounts, no third-party servers getting your data. Inside AgentOS you're running Hermes Agent which is a self-improving AI agent built by NOS Research. Hermes has a built-in learning loop, it gets better the longer you use it, it has memory that carries across sessions, over 40 built-in tools and it natively supports Kimi's API because Kimi uses an OpenAI compatible endpoint.
That means you plug in your API key from platform.moonshot.ai, point Hermes at the kimi-k2.7-code-highspeed model and you're live. From that point every conversation you have inside Hermes on AgentOS is running on a 1 trillion parameter model, 32 billion active parameters per query, a 256k context window which means it can hold an enormous amount of context in a single session and it's outputting at up to 260 tokens per second. Now let's talk about what 1 trillion parameters actually means because it sounds big but people throw that number around without context. Kimi-k2.7-code uses what's called a mixture of experts architecture, think of it like a team of 384 specialists.
When a task comes in the model doesn't use all of them, it picks the right specialists for that specific job. Only 32 billion parameters are active on any given request, the rest are standing by. That's why you get both speed and quality. You're not running all 1 trillion parameters every time, you're running the most relevant 32 billion.
The 256k context window is the other thing to pay attention to. Most people underestimate how important context is for agentic tasks. Here's a simple way to think about it. Imagine you're giving a task to an assistant who can only remember the last five minutes of your conversation.
Every time they get close to finishing they forget the beginning. That's what working with a short context model feels like in a multi-step agent task. 256k tokens lets Hermes hold your full project brief, all the back and forth, all the previous outputs, everything in a single session without losing track. For something like building out new onboarding content for the AI profit boardroom, that matters a lot.
You can have the agent pull in the community overview, the existing welcome emails, the member FAQs, the onboarding checklist and still have room to generate a completely revised version without the agent forgetting what you told it at the start. And this is where the high speed mode changes things in real practice. Because in most agentic workflows, speed is the bottleneck. The model is good enough, the tooling is there, but you're waiting.
You're waiting for a response to draft an email, you're waiting for a tool call to come back, you're waiting for a summary to generate so you can review it. At 180 to 260 tokens per second, Hermes on AgentOS with K2.7 code high speed is not making you wait. That's a different experience. You can have a full back and forth with your agent in real time.
It feels like a conversation, not a loading screen. Now, fair point on the benchmarks, and I want to be straight with you on this. Every benchmark Moonshot publishes their own. Kimi Code Bench V2, Program Bench, MLS Bench Lite, MCP Atlas.
These are all Moonshot designed evaluations. There are no independent SWE-Bench verified or Terminal-Bench submissions for K2.7 code yet. The previous model, K2.6, earned its reputation because it topped OpenRouter's weekly LLM leaderboard in April 2026, a ranking based on actual developer API usage, not a company-controlled test. That was real traction.
K2.7 code is early, and the independent evals are still coming. On Moonshot's own table, K2.7 code scores 62.0 on Kimi Code Bench V2. GPT 5.5 scores 69.0. Claude Opus 4.8 scores 67.4.
So it's not claiming to be at the top of the market on every eval. It's claiming to be close, open weight, and a fraction of the price. API pricing is $0.95 per million input tokens and $4 per million output. Compare that to the closed model pricing from OpenAI or Anthropic, and you start to see why developers are routing through Kimi.
Let me show you a real example of what this looks like in Hermes on AgentOS. Say you want to build an automated engagement system for the AI profit boardroom, something that takes new member sign-ups, generates a personalized welcome message based on what they said when they joined, and queues a follow-up check-in for day 3 and day 7. You tell Hermes, create an onboarding workflow for AI profit boardroom new members, personalize the day 1 welcome based on their join reason, write day 3 and day 7 check-in messages that mention their specific interest area. With K2.7 code high speed, Hermes doesn't just write three emails, it calls its tools, it thinks through the logic, it creates the message templates, it structures the workflow, and it does all of that in real time.
Tokens coming back so fast, you're reading while it's still generating. That's what 260 tokens per second feels like when it's doing real work for you. And Hermes learns. Every successful task you complete, it adds to its skill library.
Next time you ask for a member communication piece, it already knows the AI profit boardroom voice, the format that worked, the structure you preferred. It's building institutional knowledge about how you work. That's the combination that makes this stack genuinely powerful. Agent OS as your home screen, Hermes as the agent that learns your workflows, Kimi K2.7 code high speed as the brain that's fast enough to actually feel responsive.
One more thing worth mentioning, the open way aspect. Kimi K2.7 code is released under a modified MIT license. The weights are on Hugging Face right now at moonshot.ai slash kimi-k 2.7-code. If you have the hardware, you can run this locally.
The smallest useful quantized version runs at around 340 gigabytes. So you're not running that on a gaming laptop, but through the Kimi API at 95 cents per million input tokens, you're accessing that same model without needing the hardware. That's the practical path for most people. And that's the thing about this moment in AI.
A year ago, a 1 trillion parameter model was something only the biggest labs could touch. Now you're running it through Hermes on agent OS on your laptop screen. The compute is remote, but the control is yours. The data doesn't leave your machine unless you send it.
Kimi K2.7 code high speed is five major model releases in under a year from a team that keeps shipping. The speed gains are real. The efficiency gains are real. The agentic capability gains are real.
The independent benchmark confirmation is still pending, but the trajectory here is clear. If you're building with AI agents right now, this is a backend worth testing. And if you want to get this running, if you want the full agent OS zip file pre-configured with Kimi K2.7 code high speed ready to install, plus the Hermes agent setup tutorial, the 30 day roadmap for running your first agentic workflows and four weekly coaching calls where we walk through exactly this kind of setup. Everything is inside the AI profit boardroom.
3,500 members in there right now. A lot of them already using Hermes agent for content automation, lead follow-up, onboarding, and more. A prompt library built around these exact workflows and a member map so you can find other Kimi plus Hermes users near you and connect. Link in the comments and description or go to AIprofitboardroom.com.
And if you want all the notes from this video, the tool links and access to a community of 75,000 people working with AI right now, join the AI success lab. It's free. Link is in the comments and description too.
More episodes