Full transcript
Kuen 3.8 vs Kimi K3, best Chinese open source AI. And this time, we've got real proof, not guesses. Someone actually ran both models through the exact same 41 tasks, side by side, no do-overs, and scored every single one. The results surprised a lot of people.
Let's get into it. Both of these are huge open AI models out of China. Kimi K3 comes from Moonshot AI, 2.8 trillion parameters. QN 3.8 comes from Alibaba, 2.4 trillion parameters.
Both launched the same week. But size alone doesn't tell you which one actually gets the job done. So a builder named Julian Goldie ran a real test. Same prompt, one shot, no edits, no second tries.
Then he scored what came out 0-10 based on whether it worked, how close it matched the brief, and how good it looked. Across all 41 shared tasks. Hey, if we haven't met already, I'm the digital avatar of Julian Goldie, CEO of SEO agency Goldie Agency. Whilst he's helping clients get more leads and customers, I'm here to help you get the latest AI updates.
Julian Goldie reads every comment, so make sure you comment below. QN 3.8 came out ahead. It averaged 8.22 out of 10. Kimi K3 landed at 7.75.
Head to head, QN 3.8 won 22 tasks. Kimi K3 won 15. 4 ended in a tie. That's a real gap, built from real builds, not marketing claims.
Here's where it gets specific. On two space simulation tasks, a galaxy render and an orbiting star system, QN 3.8 didn't just win, it ran away with it. 8.7 on both. Kimi K3 scored 3.0 on both.
That's over a five point gap. On the exact same prompt, QN 3.8 also won clearly on a dogfight flight test, a terrain builder, and a glowing chemical simulation task. And here's the one that actually matters most for anyone building a business page. One of the test tasks was literally called AIPB promo, a landing page brief.
QN 3.8 won that one outright. Gold medal. So when it comes to building something like a landing page for a business or a community, this test says QN 3.8 is the stronger pick based on real output, not opinion. Now Kimi K3 didn't lose everywhere.
It won on a handful of atmospheric game builds, a torch lit dungeon, a neon arena, a dragon flight scene, a frozen open world. In each of those, Kimi K3 built something with stronger mood and lighting. So Kimi still has real strengths, just not the kind that matters most for business facing work like pages, copy, and promo builds. Kimi K3's real advantage sits somewhere else.
It has a 1 million token memory window. That was actually tested too. The tester fed it over 160,000 tokens of noise and asked it to recall one specific detail buried inside. It found it in 18 seconds.
That's a real verified memory test, not a spec sheet claim. So if your job is holding a massive pile of documents or a long research trail in memory at once, Kimi K3 still leads there. But Kimi K3 has a real downside too. On harder tasks, some testers reported it taking up to 35 minutes to finish one single build.
That's slow. QN 3.8 isn't perfect either. Right now it's only available through Alibaba's own coding platform, no public access point yet. And on one specific open world task, it flat out failed.
A black screen, nothing rendering at all. So here's the honest breakdown. Need something built fast, clean, and ready to ship on the first try? Especially anything visual like a page or a promo?
QN 3.8 is the stronger option right now. Need something to hold a massive amount of information in memory through a long drawn out task? Kimi K3 still has the edge. Before we get into how to actually use this, let's pause here for a second.
This is exactly why we built the agent operating system inside the AI profit boardroom. We don't just tell members which model is trendy that week, we run the same kind of head-to-head testing you just saw, and we hand members the exact setup, the prompts, and the templates that already came out on top. So nobody has to run 41 test builds themselves just to figure out what works. When a task looks like building a page or a promo, we point members straight at QN 3.8.
When a task needs a huge memory window, we point them at Kimi K3 instead. Stop guessing, start using what's actually been tested. If you want the exact templates from tests like this one, the link is in the comments and description. Now let's get into exactly how to use both of these in a way that fits real work, not just game demos.
Workflow. Building the AI profit boardroom landing page with QN 3.8. This one is for anyone who needs a landing page built fast and built right, since QN 3.8 already won this exact task on the real benchmark test. The prompt goes like this, build a landing page for a business community that helps people automate their work with AI.
Explain what members get, show the value clearly, make it feel trustworthy and simple, use a clean layout, a strong headline, and one clear button that says join now. Output one finished file. What comes back is a full working landing page in one shot, matching the exact task QN 3.8 already scored highest on in testing. Second workflow, turning a huge pile of member questions into one clear guide with Kimi K3.
This one is for anyone sitting on months of coaching call transcripts with no way to see what people keep struggling with over and over again. The prompt goes like this, read through all of these coaching call transcripts and community posts, find the three questions members ask most often, word for word if possible. Explain why each one keeps coming up and list them in short simple bullet points I can hand to someone else. What comes back is a short list of the real repeated sticking points pulled from months of material Kimi K3 held in memory the whole time.
Third workflow, writing outreach that actually gets replies with QN 3.8. This one is for anyone tired of sending the same message to every lead and getting silence back, who wants something that reads like it was written for one real person. The prompt goes like this, here's a summary of one business and the exact problem they're dealing with, write a short personal message inviting them to a community that solves that specific problem. Keep it under 80 words, no hype, no pressure, make it sound like one real person wrote it, not a template.
What comes back is a short specific message built around one real problem, the exact kind of one-shot writing QN 3.8 already scored strongly on. Here's the bigger point underneath all of this, a lot of people online just repeat whichever model has the biggest number attached to it. Real testing tells a different story, bigger doesn't always mean better at the task you actually need done. QN 3.8 has fewer total parameters than Kimi K3 and it still won the head-to-head, worth remembering next time someone tells you size alone decides the winner.
If you're a business owner the takeaway is simple, when you need something built and shipped especially anything customer facing like a page or a promo, reach for QN 3.8 first. When you're drowning in documents and need something to actually hold it all in memory, reach for Kimi K3, test both on a small task before you commit either one to something important. That's exactly the mindset we run inside the AI Profit Boardroom, we don't chase whichever model made the most noise that week, we test them side by side the same way you just saw and we build the exact playbook around whichever one actually wins the task, whether that's QN 3.8 for a page like this one or Kimi K3 for a research heavy job. Members get the finished setup, the tested prompts and four live coaching calls every week where they can bring their own build and get direct help improving it.
There's also a full prompts library built around tested workflows like this one not guesses and a member map so you can connect with other people near you already running the exact same tools. If you want the tested templates instead of trial and error the link is in the comments and description and if you just want to get started for free join the AI Success Lab instead. It's our free community, completely open, no cost to join, inside you'll get this same kind of tested breakdown plus over 100 other AI use cases laid out step by step and access to a community of people already deep into this stuff and comparing notes in real time. Links for that one are in the comments and description too.
So there's your real answer, not a guess, not a hunch, an actual tested head-to-head. QN 3.8 came out on top across 41 real builds and it's the stronger pick for anything customer facing like a landing page. Kimi K3 still holds the edge when the job is memory and depth. Know which one to reach for and you'll build faster than almost everyone's still guessing.
More episodes