1. Home
  2. Blog
  3. AI Tools

Kimi K3: What's New, Benchmarks, and 5 Ways to Use It

Marvin Aziz
Marvin Aziz
Growth Engineer
Marvin is a Growth Engineer at Lindy focused on AI agents, automation, and product-led growth.
Marvin Aziz
Written by
Marvin Aziz
Flo Crivello
Flo Crivello
Founder and CEO of Lindy
Flo Crivello is the founder and CEO of Lindy. Before that, he founded Teamflow and was a product manager at Uber. He writes about technology, startups, and the future of work on his blog.
Flo Crivello
Reviewed by
Flo Crivello
Last Updated:
September 30, 2026
Expert Verified

A new AI model seems to launch every few weeks, each with a benchmark chart saying it beats everything else. Kimi K3 is harder to ignore than most. Moonshot AI shipped a 2.8-trillion-parameter model in July, put the weights online 11 days later, and priced it at $15 per million output tokens.

I wanted a simpler answer: is Kimi K3 worth trying, and what's the easiest way to start?

So I went through Moonshot's tech blog, the model card on GitHub and Hugging Face, the API docs, and the license text, then checked which platforms host K3 today.

Our team at Lindy also ran two earlier Kimi models through evals this year while picking a production model, and that experience shaped how I read the numbers.

The fine print mattered as much as the headline. "Open" at this size means a cluster of 64-plus accelerators, the license has two revenue triggers for big companies, and the model gets unstable if your tools drop its earlier thinking.

Here's what changed from K2.6, how the benchmarks hold up, what the license allows, and five ways to start using K3 today.

What is Kimi K3?

Kimi K3 is Moonshot AI's flagship large language model, a 2.8-trillion-parameter mixture-of-experts model with native vision and a 1-million-token context window. Moonshot built it for long-horizon coding, knowledge work, and reasoning.

It's the model behind Kimi's chat app, its desktop app, and its coding agent, and Moonshot calls it the world's first open 3T-class model. Anyone can download the weights.

Moonshot unveiled K3 on July 16, 2026, and the full weights landed on Hugging Face on July 27.

Kimi K3 at a glance

📋 Spec 🔍 Kimi K3
Maker Moonshot AI (Beijing)
Released July 16, 2026
Open weights July 27, 2026, on Hugging Face
Total parameters 2.8 trillion
Active parameters 104 billion per token
Experts 16 of 896 picked per token
Context window 1,048,576 tokens (1M)
Input Text and images (video via the Kimi API)
Output Text
License Kimi K3 License
API model name kimi-k3
API price per 1M tokens $0.30 cached input, $3.00 input, $15.00 output
Thinking Always on, with low, high, or max effort

Specs come from Moonshot's model card, API pricing page, and K3 quickstart.

A quick note on units, because the two figures get mixed up a lot. The 104B is how many parameters fire for each token, while "16 of 896" is how many experts get picked. They describe the same sparsity from two angles.

What's new in Kimi K3 compared to Kimi K2.6

Kimi K2.6 was already a big model, and K3 is roughly three times bigger on almost every line of the spec sheet. Here's how the two compare on the figures Moonshot publishes:

📏 Spec Kimi K2.6 Kimi K3
Total parameters 1T 2.8T
Active parameters 32B 104B
Context window 256K tokens 1M tokens
Thinking Thinking and non-thinking modes Always thinking (low, high, max)
License Modified MIT Kimi K3 License
API input / output per 1M $0.95 / $4.00 $3.00 / $15.00

Sources: the Kimi K2.6 model card, Moonshot's model list, and the pricing page.

On Humanity's Last Exam, which both model cards report, K2.6 scored 34.7 (full set, no tools) and K3 scores 43.5, a nearly nine-point climb on a test where no model in Moonshot's table breaks 55 without tools.

If you were on K2.5, Moonshot retired it on August 31, 2026. K2.6 is still on the API at roughly a third of K3's per-token price, which keeps it useful for lighter jobs.

The architecture in plain words

Three changes do most of the work, and each one is easier to follow than its name suggests.

  • Kimi Delta Attention (KDA): a hybrid linear attention mechanism that keeps long inputs affordable. Picture a reader who keeps a running summary as they go, so page 900 doesn't mean re-reading pages 1 through 899.
  • Attention Residuals (AttnRes): Moonshot says it "selectively retrieves representations across depth." In plain terms, deeper layers can reach back and grab the useful work of earlier layers when they need it.
  • Stable LatentMoE: the model holds 896 specialist sub-networks, called experts, and picks 16 of them for each token. That's how a 2.8T model gets by with 104B active parameters per token.

Moonshot says these changes, plus better training and data recipes, give K3 about 2.5 times the overall scaling efficiency of Kimi K2. It also trained K3 with quantization baked in from the fine-tuning stage onward (MXFP4 weights with MXFP8 activations) for broad hardware compatibility.

If the model is the brain, the loop around it (memory, tools, and the logic that decides what happens next) is what gets work finished, which is the whole idea behind AI agent architecture.

Kimi K3 benchmarks: what the scores say and what they leave out

Moonshot published a long benchmark table comparing K3 at max effort against Claude Fable 5, GPT-5.6 Sol, Claude Opus 4.8, GPT-5.5, and GLM-5.2. Here's a slice covering reasoning, coding, agentic work, and computer use:

🧪 Benchmark Kimi K3 Claude Fable 5 GPT-5.6 Sol
GPQA Diamond 93.5 92.6 94.1
HLE-Full (no tools) 43.5 53.3 44.5
DeepSWE 67.5 70.0 73.0
Terminal-Bench 2.1 88.3 88.0 88.8
SWE-Marathon 42.0 35.0 39.0
BrowseComp 91.2 88.0 90.4
GDPval-AA v2 (Elo) 1686 1747 1736
OSWorld-Verified 84.8 85.0 83.0

All scores are Moonshot's reported numbers from the Kimi K3 model card, with Fable 5 listed as "max, w/ fallback."

K3 leads on the long-running tests (SWE-Marathon and BrowseComp), sits right in the pack on Terminal-Bench and computer use, and trails both rivals on DeepSWE, plus Fable 5 on the hardest reasoning test and on GDPval.

Moonshot's own launch post admits K3's overall performance still trails Claude Fable 5 and GPT-5.6 Sol, even while it beats the other models it tested.

What independent testing says

Artificial Analysis runs its own evaluations, and its read is a little more mixed. K3 at max effort scores 44 on its Intelligence Index, far above the median of 18 for comparable models.

Where it loses points is cost. K3 is somewhat verbose (160M tokens on the index against a 140M median), but the bigger driver is price: $3.00 and $15.00 per million input and output tokens, against medians of $0.30 and $1.15.

That's why Artificial Analysis calls it "particularly expensive" next to open-weight models of similar size, at about $2.00 per index task.

How to read vendor benchmark claims

Moonshot's table is upfront about its setup, as long as you read the footnotes before the bold numbers:

  • Harnesses differ: K3 ran in Moonshot's own Kimi Code harness on several coding tests, while rivals ran in Claude Code, Codex, or Terminus 2. The harness shapes the score, so two cells in the same row can reflect different setups.
  • K3 runs at max effort: Moonshot reports K3 at its highest reasoning setting. On low effort, the Artificial Analysis figures mirrored on OpenRouter show K3 at 84.2% on GPQA Diamond, down from 93.5% at max.
  • Fallbacks count: Moonshot notes Claude Fable 5 hit fallbacks on 35% of SWE-Marathon tasks in its runs, which may have dragged Fable's score down.
  • Hardware and dates vary: some runs used H20 GPUs where the official setup calls for H100s, and several scores are "as of" a specific July date.
  • Some tests are homegrown: Kimi Code Bench 2.0 and PerceptionBench are Moonshot's own in-house benchmarks.

Who serves the model can move the score too. When we swapped models at Lindy this year, Kimi K2.5 did well in offline evals and then flopped on real usage, and in a later round of testing the same nominal model scored differently depending on who served it.

DeepSeek V4 Flash came out on top in that later round. K3 is available from Moonshot plus at least five other platforms, so that lesson carries straight over. Test on the host you'll run in production.

How to use Kimi K3: 5 ways to get started

Moonshot launched K3 in its chat app, desktop app, coding agent, and API on day one, so the right door depends on what you want to do with it.

1. The Kimi app and website

The quickest way in is the Kimi chat app. Download or update it on iOS, Android, or HarmonyOS, or open kimi.ai in a browser, and you can start chatting with K3 right away.

The app also has Slides, Docs, Sheets, and Deep Research features, so it's a fair test if you're shopping for chat apps beyond ChatGPT and want to see how K3 handles research and documents.

2. Kimi Work on your desktop

Kimi Work is Moonshot's desktop app for longer knowledge work, and you'll need version 3.1.0 or later on Windows or an Apple silicon Mac to use K3.

K3 arrived there with two new features. Widgets lets you build interactive components inside a chat, and Dashboard pins the widgets you care about into one persistent view for a project or goal.

3. Kimi Code in your terminal

Kimi Code is Moonshot's terminal coding agent, and it's the harness Moonshot recommends for K3. Here's the setup from the Kimi Code docs:

  1. Install it: on macOS or Linux, run curl -fsSL https://code.kimi.com/kimi-code/install.sh | bash. You can also use npm install -g @moonshot-ai/kimi-code with Node.js 22.19.0 or later.
  2. Start it: move into your project folder and run kimi.
  3. Log in: type /login and follow the device-code flow.
  4. Pick K3: type /model and select Kimi K3.

Start a fresh session when you switch to K3, because Moonshot warns that moving an ongoing session from another model can make the output unstable. If you're testing K3 side by side with other AI coding agents, give each one its own clean session.

4. The Kimi API

For developers, the model name is kimi-k3 on the Kimi API Platform. The API works with the OpenAI SDK, and K3 unlocks once you top up at least $1.

from openai import OpenAI

‍

client = OpenAI(api_key="YOUR_MOONSHOT_API_KEY", base_url="https://api.moonshot.ai/v1")

‍

resp = client.chat.completions.create(

    model="kimi-k3",

    reasoning_effort="high",

    messages=[{"role": "user", "content": "Summarize this README in five bullets."}],

)

print(resp.choices[0].message.content)

A few details will save you a confusing afternoon:

  • Reasoning effort: set reasoning_effort to low, high, or max. Max is the default, and thinking can't be turned off.
  • Pass the whole message back: in multi-turn chats and tool calls, return the complete assistant message, including reasoning_content and tool_calls.
  • Sampling is locked: temperature, top_p, and a few other settings are fixed, so leave them out of your requests.
  • Images need uploading: vision input doesn't accept public image URLs, so send base64 or upload the file first.

Moonshot also has setup guides for running K3 inside Codex and OpenCode, if you'd like to keep your current coding tool.

5. Third-party hosts

If you want to stay on a platform you already pay for, K3 is live on several of them:

🏢 Host 🔑 Model ID
OpenRouter moonshotai/kimi-k3
NVIDIA NIM moonshotai/kimi-k3
Together AI moonshotai/Kimi-K3
Fireworks AI accounts/fireworks/models/kimi-k3
Amazon Bedrock moonshotai.kimi-k3

You'll find each listing on OpenRouter, NVIDIA, Together AI, and Fireworks. Amazon Bedrock added K3 on September 18, 2026, according to AWS's Bedrock model card.

Limits and premium routes vary by host, and so can quality (see the provider lesson above), so run the same test on each before you commit.

‍

{{templates}}

‍

Is Kimi K3 open source? What the license allows

Kimi K3 is open-weight, released under a custom Kimi K3 License. The license lets you download, use, modify, fine-tune, and sell the model much like an MIT license, with two conditions aimed at very large players.

  • Big model-as-a-service businesses need a deal: if you or your affiliates run a model-as-a-service business (giving third parties API-style access to the model) and your combined revenue tops $20 million in any 12 consecutive months, you need a separate agreement with Moonshot first.
  • Huge products have to show the name: if a commercial product built on K3 has more than 100 million monthly active users or more than $20 million in monthly revenue, it must display "Kimi K3" prominently in its interface.
  • Two exemptions: neither rule applies to internal use (nothing exposed to third parties) or to access through Moonshot's own products or certified inference partners.

For most companies, that means you can build on K3 freely. If you run a large AI platform, read the license with your lawyer (it's about a page long).

Can you self-host Kimi K3?

You can, as long as you have a lot of hardware. Moonshot recommends deploying K3 on supernode configurations with 64 or more accelerators, and its model card recommends vLLM, SGLang, and TokenSpeed as inference engines.

NVIDIA's model card lists Blackwell as the supported hardware. So "open" here means your company's GPU cluster can run it, and your gaming PC should sit this one out.

How much does Kimi K3 cost?

On Moonshot's own API, K3 costs $0.30 per million cached input tokens, $3.00 per million uncached input tokens, and $15.00 per million output tokens. Writing a prompt into the cache costs $3.00 per million tokens for a 5-minute cache or $6.00 for a 1-hour cache, per the pricing page.

Caching does a lot of the heavy lifting. According to Moonshot, its API hits the cache more than 90% of the time on coding workloads, which can pull your effective input price toward the $0.30 rate.

Third-party hosts can charge differently, and Kimi's app plans are priced separately from the API, so price out your real workload on the host you plan to use before you budget a big project. For the app plans and a host-by-host breakdown, see our Kimi K3 pricing guide.

Kimi K3 limitations worth knowing

Moonshot is pretty upfront about where K3 struggles. Here's what its tech blog and API docs flag:

  • It's sensitive to thinking history: if a harness drops earlier reasoning, or you switch to K3 mid-session, output quality "may become highly unstable."
  • It can get too proactive: K3 was trained on long, hard tasks, so with fuzzy instructions it may make unexpected decisions on your behalf. Moonshot suggests spelling out boundaries in the system prompt or an AGENTS.md file.
  • The experience still trails the top closed models: by Moonshot's own account, K3 shows a noticeable difference in user experience compared with Claude Fable 5 and GPT-5.6 Sol.
  • It burns tokens: max effort is the default, and Artificial Analysis measured it as wordier than the median model.
  • Some API features are still settling: the built-in web search is being updated and isn't recommended for production for now, and image input needs an upload (see the API tips above).
  • Self-hosting takes a cluster: see the 64-accelerator recommendation above.

Who should use Kimi K3 (and who should skip it)

Kimi K3 is a good fit if you:

  • Run long coding sessions: multi-hour agentic work across big repositories is where K3's scores look strongest.
  • Work with huge inputs: the 1M-token context handles long reports, contracts, and codebases in one go.
  • Want weights you can own: you can start on the API and move to your own infrastructure or a fine-tuned version later (Fireworks already offers LoRA fine-tuning).
  • Are shopping for a coding model outside Anthropic: K3 belongs on your shortlist of Claude alternatives for development work.

Skip Kimi K3 if you:

  • Need the top reasoning scores today: Fable 5 leads K3 by about 10 points on Humanity's Last Exam in Moonshot's own table.
  • Run simple, high-volume tasks: K2.6 costs roughly a third as much per token on the API, and lighter jobs rarely need a 2.8T model.
  • Want an assistant that handles your workday: a model is one piece of that, which is why the agentic tools people use every day wrap models with memory, integrations, and approvals.

How to find out if Kimi K3 fits your work

Kimi K3 is the biggest open-weight model released so far, and on long agentic runs its scores sit right next to the best closed models. On the hardest reasoning tests and on cost per task, it still has some catching up to do.

The fastest way to know is to test it on your own work. Pick one real task you do every week, run it through Kimi Code or the API at high effort, and compare the result, the token count, and the bill with your current model.

If K3 holds up on your work (and on the host you'd use in production), you've found a strong model with weights you can keep. If it wanders off, you've learned that for the price of a coffee.

‍

{{cta}}

‍

Frequently asked questions

What is Kimi K3 good for?

Kimi K3 is good for long-horizon coding, agentic research, and knowledge work over large documents. Moonshot built it for multi-hour engineering sessions, big codebases, terminal tools, and visual tasks like frontend work, and its 1M-token context fits long reports and repositories.

How much does Kimi K3 cost?

Kimi K3 costs $0.30 per million cached input tokens, $3.00 per million uncached input tokens, and $15.00 per million output tokens on Moonshot's Kimi API.

Together AI and Fireworks currently match those rates on their standard serverless endpoints, though Fireworks charges more for its Fast and US-only routes. Kimi's app plans are billed separately, and our Kimi K3 pricing guide lists them.

Has Kimi K3 been released as open weights?

Yes, Kimi K3's full weights are public. Moonshot published them on Hugging Face on July 27, 2026, under the Kimi K3 License, which allows commercial use with extra conditions for very large platforms.

Is Moonshot AI a Chinese company?

Yes, Moonshot AI is a Chinese AI company headquartered in Beijing. Yang Zhilin, Zhou Xinyu, and Wu Yuxin founded it in March 2023, and its backers include Alibaba and Tencent.

Is Kimi K3 available on Amazon Bedrock?

Yes, Kimi K3 is on Amazon Bedrock as of September 18, 2026, under the model ID moonshotai.kimi-k3, according to AWS's Bedrock documentation.

Save 2 Hours Every Day
Lindy is your ultimate AI assistant that manages inbox, meetings, and follow-ups—so you stay ahead of the chaos.
Try Lindy for Free
About the editorial team
Marvin Aziz
Marvin Aziz
Growth Engineer

Marvin is a Growth Engineer at Lindy focused on AI agents, automation, and product-led growth.

Flo Crivello
Flo Crivello
Founder and CEO of Lindy

Flo Crivello is the founder and CEO of Lindy. Before that, he founded Teamflow and was a product manager at Uber. He writes about technology, startups, and the future of work on his blog.

Ready when you are.

Free to try. In your Slack in two minutes.

Try for free
7-day free trial • Cancel anytime