A new AI model seems to launch every few weeks, each with a benchmark chart saying it beats everything else. Kimi K3 is harder to ignore than most. Moonshot AI shipped a 2.8-trillion-parameter model in July, put the weights online 11 days later, and priced it at $15 per million output tokens.
I wanted a simpler answer: is Kimi K3 worth trying, and what's the easiest way to start?
So I went through Moonshot's tech blog, the model card on GitHub and Hugging Face, the API docs, and the license text, then checked which platforms host K3 today.
Our team at Lindy also ran two earlier Kimi models through evals this year while picking a production model, and that experience shaped how I read the numbers.
The fine print mattered as much as the headline. "Open" at this size means a cluster of 64-plus accelerators, the license has two revenue triggers for big companies, and the model gets unstable if your tools drop its earlier thinking.
Here's what changed from K2.6, how the benchmarks hold up, what the license allows, and five ways to start using K3 today.
Kimi K3 is Moonshot AI's flagship large language model, a 2.8-trillion-parameter mixture-of-experts model with native vision and a 1-million-token context window. Moonshot built it for long-horizon coding, knowledge work, and reasoning.
It's the model behind Kimi's chat app, its desktop app, and its coding agent, and Moonshot calls it the world's first open 3T-class model. Anyone can download the weights.
Moonshot unveiled K3 on July 16, 2026, and the full weights landed on Hugging Face on July 27.
Specs come from Moonshot's model card, API pricing page, and K3 quickstart.
A quick note on units, because the two figures get mixed up a lot. The 104B is how many parameters fire for each token, while "16 of 896" is how many experts get picked. They describe the same sparsity from two angles.
Kimi K2.6 was already a big model, and K3 is roughly three times bigger on almost every line of the spec sheet. Here's how the two compare on the figures Moonshot publishes:
Sources: the Kimi K2.6 model card, Moonshot's model list, and the pricing page.
On Humanity's Last Exam, which both model cards report, K2.6 scored 34.7 (full set, no tools) and K3 scores 43.5, a nearly nine-point climb on a test where no model in Moonshot's table breaks 55 without tools.
If you were on K2.5, Moonshot retired it on August 31, 2026. K2.6 is still on the API at roughly a third of K3's per-token price, which keeps it useful for lighter jobs.
Three changes do most of the work, and each one is easier to follow than its name suggests.
Moonshot says these changes, plus better training and data recipes, give K3 about 2.5 times the overall scaling efficiency of Kimi K2. It also trained K3 with quantization baked in from the fine-tuning stage onward (MXFP4 weights with MXFP8 activations) for broad hardware compatibility.
If the model is the brain, the loop around it (memory, tools, and the logic that decides what happens next) is what gets work finished, which is the whole idea behind AI agent architecture.
Moonshot published a long benchmark table comparing K3 at max effort against Claude Fable 5, GPT-5.6 Sol, Claude Opus 4.8, GPT-5.5, and GLM-5.2. Here's a slice covering reasoning, coding, agentic work, and computer use:
All scores are Moonshot's reported numbers from the Kimi K3 model card, with Fable 5 listed as "max, w/ fallback."
K3 leads on the long-running tests (SWE-Marathon and BrowseComp), sits right in the pack on Terminal-Bench and computer use, and trails both rivals on DeepSWE, plus Fable 5 on the hardest reasoning test and on GDPval.
Moonshot's own launch post admits K3's overall performance still trails Claude Fable 5 and GPT-5.6 Sol, even while it beats the other models it tested.
Artificial Analysis runs its own evaluations, and its read is a little more mixed. K3 at max effort scores 44 on its Intelligence Index, far above the median of 18 for comparable models.
Where it loses points is cost. K3 is somewhat verbose (160M tokens on the index against a 140M median), but the bigger driver is price: $3.00 and $15.00 per million input and output tokens, against medians of $0.30 and $1.15.
That's why Artificial Analysis calls it "particularly expensive" next to open-weight models of similar size, at about $2.00 per index task.
Moonshot's table is upfront about its setup, as long as you read the footnotes before the bold numbers:
Who serves the model can move the score too. When we swapped models at Lindy this year, Kimi K2.5 did well in offline evals and then flopped on real usage, and in a later round of testing the same nominal model scored differently depending on who served it.
DeepSeek V4 Flash came out on top in that later round. K3 is available from Moonshot plus at least five other platforms, so that lesson carries straight over. Test on the host you'll run in production.
Moonshot launched K3 in its chat app, desktop app, coding agent, and API on day one, so the right door depends on what you want to do with it.
The quickest way in is the Kimi chat app. Download or update it on iOS, Android, or HarmonyOS, or open kimi.ai in a browser, and you can start chatting with K3 right away.
The app also has Slides, Docs, Sheets, and Deep Research features, so it's a fair test if you're shopping for chat apps beyond ChatGPT and want to see how K3 handles research and documents.
Kimi Work is Moonshot's desktop app for longer knowledge work, and you'll need version 3.1.0 or later on Windows or an Apple silicon Mac to use K3.
K3 arrived there with two new features. Widgets lets you build interactive components inside a chat, and Dashboard pins the widgets you care about into one persistent view for a project or goal.
Kimi Code is Moonshot's terminal coding agent, and it's the harness Moonshot recommends for K3. Here's the setup from the Kimi Code docs:
Start a fresh session when you switch to K3, because Moonshot warns that moving an ongoing session from another model can make the output unstable. If you're testing K3 side by side with other AI coding agents, give each one its own clean session.
For developers, the model name is kimi-k3 on the Kimi API Platform. The API works with the OpenAI SDK, and K3 unlocks once you top up at least $1.
from openai import OpenAI
client = OpenAI(api_key="YOUR_MOONSHOT_API_KEY", base_url="https://api.moonshot.ai/v1")
resp = client.chat.completions.create(
model="kimi-k3",
reasoning_effort="high",
messages=[{"role": "user", "content": "Summarize this README in five bullets."}],
)
print(resp.choices[0].message.content)
A few details will save you a confusing afternoon:
Moonshot also has setup guides for running K3 inside Codex and OpenCode, if you'd like to keep your current coding tool.
If you want to stay on a platform you already pay for, K3 is live on several of them:
You'll find each listing on OpenRouter, NVIDIA, Together AI, and Fireworks. Amazon Bedrock added K3 on September 18, 2026, according to AWS's Bedrock model card.
Limits and premium routes vary by host, and so can quality (see the provider lesson above), so run the same test on each before you commit.
{{templates}}
Kimi K3 is open-weight, released under a custom Kimi K3 License. The license lets you download, use, modify, fine-tune, and sell the model much like an MIT license, with two conditions aimed at very large players.
For most companies, that means you can build on K3 freely. If you run a large AI platform, read the license with your lawyer (it's about a page long).
You can, as long as you have a lot of hardware. Moonshot recommends deploying K3 on supernode configurations with 64 or more accelerators, and its model card recommends vLLM, SGLang, and TokenSpeed as inference engines.
NVIDIA's model card lists Blackwell as the supported hardware. So "open" here means your company's GPU cluster can run it, and your gaming PC should sit this one out.
On Moonshot's own API, K3 costs $0.30 per million cached input tokens, $3.00 per million uncached input tokens, and $15.00 per million output tokens. Writing a prompt into the cache costs $3.00 per million tokens for a 5-minute cache or $6.00 for a 1-hour cache, per the pricing page.
Caching does a lot of the heavy lifting. According to Moonshot, its API hits the cache more than 90% of the time on coding workloads, which can pull your effective input price toward the $0.30 rate.
Third-party hosts can charge differently, and Kimi's app plans are priced separately from the API, so price out your real workload on the host you plan to use before you budget a big project. For the app plans and a host-by-host breakdown, see our Kimi K3 pricing guide.
Moonshot is pretty upfront about where K3 struggles. Here's what its tech blog and API docs flag:
Kimi K3 is the biggest open-weight model released so far, and on long agentic runs its scores sit right next to the best closed models. On the hardest reasoning tests and on cost per task, it still has some catching up to do.
The fastest way to know is to test it on your own work. Pick one real task you do every week, run it through Kimi Code or the API at high effort, and compare the result, the token count, and the bill with your current model.
If K3 holds up on your work (and on the host you'd use in production), you've found a strong model with weights you can keep. If it wanders off, you've learned that for the price of a coffee.
{{cta}}
Kimi K3 is good for long-horizon coding, agentic research, and knowledge work over large documents. Moonshot built it for multi-hour engineering sessions, big codebases, terminal tools, and visual tasks like frontend work, and its 1M-token context fits long reports and repositories.
Kimi K3 costs $0.30 per million cached input tokens, $3.00 per million uncached input tokens, and $15.00 per million output tokens on Moonshot's Kimi API.
Together AI and Fireworks currently match those rates on their standard serverless endpoints, though Fireworks charges more for its Fast and US-only routes. Kimi's app plans are billed separately, and our Kimi K3 pricing guide lists them.
Yes, Kimi K3's full weights are public. Moonshot published them on Hugging Face on July 27, 2026, under the Kimi K3 License, which allows commercial use with extra conditions for very large platforms.
Yes, Moonshot AI is a Chinese AI company headquartered in Beijing. Yang Zhilin, Zhou Xinyu, and Wu Yuxin founded it in March 2023, and its backers include Alibaba and Tencent.
Yes, Kimi K3 is on Amazon Bedrock as of September 18, 2026, under the model ID moonshotai.kimi-k3, according to AWS's Bedrock documentation.
