← All posts
Website / Vibe Code··6 min read

Kimi K3 Just Launched at $3/$15 per Million Tokens: Here Is What We Are Testing This Week

Kimi K3 launched this week with a 1 million token context window and pricing cheap enough that Moonshot AI had to pause new signups. We ran it through a full client website build to see what actually changed.

Kimi K3 Just Launched at $3/$15 per Million Tokens: Here Is What We Are Testing This Week
Answer

Kimi K3 is Moonshot AI's new flagship model: 1M token context, $3 per million input tokens and $15 per million output tokens, open weights due by July 27, 2026. We are testing it now for fast, multi-page client website builds, but holding production bets until the weights and benchmarks are independently verified.

Kimi K3 launched on July 16, 2026, and by July 20 Moonshot AI had already paused new subscriptions because demand outran capacity. That is not a story about a chatbot getting popular. It is a story about the cost of building software dropping fast enough that the pricing page itself became the news.

We run a 40-person real estate group and we build AI systems for operators like us. Every time a new model lands with numbers this aggressive, we run it against real client work before we write a word about it. This week we ran Kimi K3 through a full website build, from a blank brief to a live multi-page site and a reskinned dashboard app, and we are writing down exactly what we found.

What Kimi K3 actually is

Kimi K3 is Moonshot AI's new flagship model, built on a hybrid linear attention architecture the company calls Kimi Delta Attention, paired with what it calls Attention Residuals. According to Moonshot's own API documentation, the model carries roughly 2.8 trillion parameters, a 1,048,576 token context window, and native support for text, image, and video input.

The pricing is the part that matters for operators. Per Moonshot's official docs, Kimi K3 costs $3 per million input tokens and $15 per million output tokens, with a discounted $0.30 per million rate on cache hits. There is no separate tier for the long context. You get the full 1M token window at the same flat rate. Full model weights are due for public release by July 27, 2026, so as of today this is API-only.

Moonshot claims the model performs competitively with Anthropic's Fable 5 and beats Opus 4.8 and GPT 5.6 Sol on its benchmark suite. Those are the vendor's own numbers. We treat every first-party benchmark claim as a marketing document until an independent group replicates it. Our advice below assumes you do the same.

Why the context window changes what "vibe coding" can do

A 1 million token context window is not a spec sheet number, it is a workflow change. You can hand a coding model an entire brand system, a full page-by-page copy deck, and a client's asset library in a single prompt and it still has room left to reason about the layout. That means fewer round trips, fewer half-finished pages, and fewer places where the model forgets what section three looked like by the time it builds section seven.

Combined with pricing this low, the economics of a full custom website build shift. A multi-page site that used to require a template and a week of manual assembly can now be generated, styled, and iterated on inside a single working session. That is the shift we tested this week.

The build we ran

Here is the actual playbook, step by step, the way we ran it on a real client brief. Step two used Figma's AI image editing to keep the reference art on-brand before any code got generated.

StepWhat we didTool
1. ReferenceSourced a visual reference that matched the client's brand directionManual search
2. Style passEdited the reference image with a text prompt to match brand color and textureFigma AI image editing
3. Layout generationGenerated the first landing page from a UI reference prompt describing fonts, spacing, and section orderKimi K3
4. Media swapReplaced placeholder video and images with the client's real assets, switched to dark mode textKimi K3
5. InteractionAdded a mouse hover reveal effect to the hero sectionKimi K3
6. Page expansionGenerated services, pricing, and download pages using the same design language, no new briefKimi K3
7. App variantRebuilt the same content structure as a dashboard app, restyled from a second design referenceKimi K3
8. Motion polishAnimated a static hero image into a short looping video for the backgroundKling AI

Step 8 used Kling AI's image-to-video generation, which prices a five-second 1080p clip at roughly $0.70 through its API. That is cheap enough to generate three or four variations and pick the one that actually fits the section, instead of settling for the first pass.

The step that surprised us most was step 7. Once the landing page existed with real content in it, restyling the entire thing into a different product, a dashboard instead of a marketing site, took one prompt: keep the structure and the copy, change only the visual language to match a second reference. No rebuild. No re-briefing. That is the part of this workflow that actually threatens the old template-and-theme model of building software fast.

The other piece: skills, not just prompts

The same week Kimi K3 launched, we also watched interest spike in Claude's Agent Skills, packaged folders of instructions an agent loads dynamically for a specific task. Anthropic shipped Skills across Claude.ai, Claude Code, and the API in October 2025, and the pattern is showing up in every serious agent stack now: instead of one long system prompt trying to cover every case, you build small, swappable skill packages for each job. A skill that rewrites replies for clarity, a skill that formats client reports, a skill that reviews a contract clause. It is the same instinct as our build playbook above: reusable, composable pieces instead of one-off prompts.

What this means if you run a 10 to 50 person company

Three things, in order of how fast they matter.

  • Website and internal tool costs are compressing further. A custom multi-page site plus a reskinned internal app is now a same-day build, not a two-week engagement. If your agency is still quoting template timelines, ask why.
  • Do not bet production on unverified claims. Kimi K3's weights are not public until July 27, 2026, and its benchmark wins are self-reported. Test it through the API this week. Hold off on self-hosting or routing production traffic to it until independent evals and the open weights are both out.
  • Ownership still matters more than speed. Whatever model you use to build a client site or internal tool, the client should walk away owning the code and the files, not renting access to a black box. That is a baseline we hold on every build we ship, and it is worth writing into any agency contract you sign this quarter.

What we are doing about it this quarter

We are folding this workflow into how we scope new builds. When a client comes through our AI Concierge assessment, part of that engagement now includes testing two or three current models against their actual brief, not just defaulting to whichever one we used last time. The assessment costs €999 and that fee is credited to the build if they move forward, so the model testing is not an extra line item, it is part of what they are already paying for.

On the delivery side, this is exactly the kind of workflow we run inside automation builds and through our broader product line. The pattern holds across every build we ship: a first working system live in days to weeks, not quarters, running on infrastructure the client actually owns. You can see how that plays out in practice in our case studies, and we will keep writing up what we find as new models land on our blog.

The headline is not that a Chinese lab shipped a big model. Labs ship big models every month now. The headline is that the price of turning a client brief into a live, multi-page, on-brand website or internal app just dropped again, and the agencies still quoting template timelines are going to feel that shift before their clients tell them about it.

// Next move

See where AI pays you back first.

A free 10-minute assessment. Your top AI quick win plus the hours and money it returns. No cost, no pitch.