⌂ AiTimeline

Technology · AI Models · Timeline

Kimi K3 History and Timeline: How Moonshot AI Went From an Unknown Startup to a 2.8-Trillion-Parameter Open Frontier

📅 Updated 18 July 2026🧠 Moonshot AI · Hugging Face · Independent evals📊 Confirmed specs vs community reports separated

Late on a Thursday in July 2026, a developer in Berlin refreshed a Hugging Face page she had been watching all week. A machine-learning engineer in Bengaluru opened a terminal. A research lead in San Francisco pulled up a benchmark leaderboard and blinked at the ranking. The cause of all three reactions was the same: Moonshot AI, a Beijing startup that barely existed three years earlier, had just released Kimi K3—a 2.8-trillion-parameter model it called the largest open-weight system ever built, and one that independent blind tests placed shoulder to shoulder with the strongest closed models from OpenAI and Anthropic. Only a few years ago, a Chinese lab shipping a frontier-class model that anyone could download would have sounded implausible. In the summer of 2026, it was the biggest story in AI. This is the timeline of how Kimi got here—and an honest reading of what the model can and cannot do.

🚀 Now live: Moonshot AI announced Kimi K3 on 16 July 2026, days ahead of the World Artificial Intelligence Conference in Shanghai. The company has said full open weights will land on Hugging Face around 27 July 2026 under a Modified MIT license. This timeline is updated as official specs, independent benchmarks and the weight release are confirmed.
🧭How this page handles facts: Specifications and release details attributed to Moonshot AI are treated as confirmed. Benchmark rankings from independent evaluators such as LMArena are labelled as third-party results. Anything circulating without official confirmation is marked as a community report. This is an explainer and historical timeline, not investment advice or a formal benchmark study.
At a GlanceIn One MinuteQuick AnswersWhy It MattersKimi EvolutionFull TimelineThe Tech ExplainedComparisonsPros & LimitsFAQ
Kimi K3 At a Glance
DeveloperMoonshot AI (Beijing)
Announced16 July 2026
Model familyKimi (Chat, K1.5, K2, K3)
ArchitectureMixture-of-Experts + KDA
Parameters2.8T total, 16 of 896 experts active
Context window~1 million tokens
LicenseModified MIT (open-weight)
VariantsK3 Max, K3 Swarm Max
API pricing$3 in / $15 out per M tokens
Primary usesCoding, agents, long-context, vision

⚡ In One Minute

Kimi K3 is the flagship large language model from Moonshot AI, announced on 16 July 2026. It is a Mixture-of-Experts model with 2.8 trillion total parameters—of which only a small fraction activate for any given token—paired with a roughly one-million-token context window, native vision, and an always-on reasoning mode Moonshot calls thinking mode. Its headline claim is scale with openness: Moonshot says it is the largest model ever released with downloadable weights, under a permissive Modified MIT license.

In independent blind tests reported at launch, developers favoured K3 for front-end coding over several leading US systems, and it ranked at or near the top on general text quality. It ships in two flavours—K3 Max for chat and single-agent work, and K3 Swarm Max for large multi-agent workloads—with an OpenAI-compatible API. Benchmarks aside, its real significance is competitive: a serious, open, frontier-class alternative to closed models.

Quick Answers

Kimi K3 in Plain Terms

What is it?
Kimi K3 is Moonshot AI’s flagship open-weight large language model: a 2.8-trillion-parameter Mixture-of-Experts system with a one-million-token context window, native vision and built-in reasoning, aimed at coding, agents and long-document work.
Who built it?
Moonshot AI, a Beijing startup founded in 2023 by Yang Zhilin and two Tsinghua University classmates, and backed by investors including Alibaba. Kimi is the name of both its chatbot and its model family.
When was it released?
Moonshot announced Kimi K3 on 16 July 2026, just before the World Artificial Intelligence Conference in Shanghai. The company has said full open weights will follow on Hugging Face around 27 July 2026 under a Modified MIT license.
Why does it matter?
It is the largest open-weight model yet, and independent tests place it near the top of the field. That means organisations can run frontier-level AI on their own infrastructure, intensifying competition between open and closed models.
How is it different?
It combines frontier scale with open weights, a very long context, native vision and new attention techniques (Kimi Delta Attention) that aim to make long-context inference more efficient than standard Transformers.
How do I use it?
You can use Kimi K3 through Moonshot’s OpenAI-compatible API, or, once weights are public, download and self-host it on suitable GPU infrastructure. It ships as K3 Max for chat and K3 Swarm Max for multi-agent workloads.
Key Takeaways

What to Remember

Why Kimi K3 Matters

Different audiences care for different reasons.

To grasp why a single model release dominated a week of AI headlines, it helps to see who was paying attention and why. Developers cared because a model that tops coding leaderboards and speaks the OpenAI API dialect is something they can adopt in an afternoon, not a quarter. Enterprises cared because open weights mean data can stay inside their own walls—no sending sensitive prompts to a third-party endpoint, and no vendor able to deprecate the model out from under them.

Researchers cared because a 2.8-trillion-parameter model with published weights and novel attention mechanisms is a rare gift: a frontier-scale system they can actually inspect, fine-tune and study rather than probe through a locked API. Open-source communities cared because each capable open release resets expectations for what should be freely available, pressuring the whole industry. And investors cared because Kimi K3 is evidence that the economics and geography of frontier AI are shifting—that leadership is contestable, and that a well-funded lab outside Silicon Valley can ship at the frontier.

None of this means Kimi K3 is the best model at everything. It means the model matters as a marker: proof that open, frontier-class AI is no longer a contradiction in terms.

💻 Developer Insight — K3 vs Earlier Kimi Generations

Each Kimi generation solved a different bottleneck. Kimi Chat (2023) was about raw context—reading enormous documents. K1.5 (2025) added structured reasoning. K2 (2025) went big and agentic with a trillion-parameter Mixture-of-Experts built for tool use and coding. K2.5 (early 2026) bolted on native vision. K3 fuses all of it—scale, reasoning, agents, vision and a million-token window—into one system, then adds new attention machinery to keep that long context affordable. If you used K2 for coding agents, K3 is a drop-in step up rather than a new paradigm to relearn.

The Evolution of Kimi Models

From a long-context chatbot to an open frontier model, in five steps.

GenerationWhenProblem it solvedSignature capability
Kimi ChatOct 2023Reading very long inputs200K-character context, later 2M
Kimi K1.5Jan 2025Weak step-by-step reasoningReasoning to rival OpenAI o1 (Moonshot claim)
Kimi K2Jul 2025Scale and tool use1T-parameter MoE, open weights, agentic coding
Kimi K2.5Jan 2026No image or video understandingNative vision via the MoonViT encoder
Kimi K3Jul 2026Unifying scale, reasoning, agents, vision2.8T MoE, ~1M context, KDA, thinking mode

The Full Kimi Timeline

Newest first. Tags mark confirmed facts, milestones and community or third-party reports.

Jul 2026

Kimi K3 Is Announced

Confirmed16 July 2026 · Moonshot AI

What happened: Moonshot AI unveiled Kimi K3, a 2.8-trillion-parameter Mixture-of-Experts model with a roughly one-million-token context window, native vision and an always-on reasoning mode. The company positioned it as the largest open-weight model ever released, timed just before the Shanghai AI conference.

Technical breakthrough: K3 introduces Kimi Delta Attention, a hybrid linear-attention mechanism, and Attention Residuals, described by Moonshot as a drop-in improvement on standard residual connections that scales more reliably.

Industry impact: A downloadable frontier-scale model reframed the open-versus-closed debate overnight and put direct pressure on both US labs and other Chinese competitors.

Timeline takeaway: K3 is less a single breakthrough than the convergence of everything Kimi learned across three years.
2.8T parameters~1M contextOpen-weight
Jul 2026

Benchmarks, Weights and Reception

Third-party / communityMid-July 2026 · Independent evals

What happened: In independent blind tests reported at launch, developers preferred Kimi K3 for front-end coding over several leading US models, and it ranked at or near the top on general text quality. Moonshot said full weights would reach Hugging Face around 27 July under a Modified MIT license.

Developer impact: The API launched priced at about $3 per million input tokens and $15 per million output, with OpenAI-SDK compatibility that let teams swap it in with minimal code changes.

Current relevance: Independent rankings can shift week to week, so treat leaderboard positions as a snapshot, not a verdict.

Timeline takeaway: Winning a blind coding test is a real signal, but it measures preference on specific tasks, not universal superiority.
Weights ~27 JulOpenAI-compatible APILeaderboard snapshot
Jan 2026

Kimi K2.5 Adds Native Vision

ConfirmedJanuary 2026 · Multimodal upgrade

What happened: Moonshot released Kimi K2.5, a multimodal upgrade to K2 that added native vision through a roughly 400-million-parameter encoder called MoonViT, letting the model process both images and video.

Technical breakthrough: Video understanding enabled genuinely agentic behaviour—reproducing a website user journey from a screen recording alone, a task that stumps text-only models.

Developer impact: Teams building visual agents and document-understanding tools gained an open option that could see, not just read.

Timeline takeaway: K2.5 turned Kimi from a text specialist into a multimodal system—the groundwork for K3’s native vision.
MoonViT encoderImage + video
Jul 2025

Kimi K2 Goes Big and Open

ConfirmedJuly 2025 · Open-weight flagship

What happened: Moonshot released the weights for Kimi K2, a one-trillion-parameter Mixture-of-Experts model with about 32 billion parameters active per token, trained on 15.5 trillion tokens and published under a Modified MIT license.

Technical breakthrough: K2 was engineered for agentic work—tool calling, multi-step coding and autonomous task execution—and briefly became the top-ranked open model on community hubs.

Industry impact: It established Moonshot’s open-weight strategy and proved a Chinese lab could ship a trillion-parameter model that developers worldwide actually wanted to use.

Timeline takeaway: K2 set the template—big, open, agentic—that K3 would scale up a year later.
1T params, 32B activeModified MITAgentic coding
Jan 2025

Kimi K1.5 Learns to Reason

Confirmed20 January 2025 · Reasoning model

What happened: Moonshot released Kimi K1.5, a reasoning-focused model the company said matched OpenAI’s o1 in mathematics, coding and multimodal reasoning.

Technical breakthrough: K1.5 leaned into reinforcement learning and long chains of thought—the shift from models that answer instantly to models that deliberate before responding.

Developer impact: It signalled that Moonshot could compete on reasoning quality, not just context length, narrowing the gap with the leading US labs.

Timeline takeaway: K1.5 is where Kimi stopped being only a long-context tool and started thinking.
Reasoningo1-level claim
2024

Scaling, Funding and a 2-Million-Character Window

Milestone2024 · Growth and infrastructure

What happened: In early 2024 Moonshot said Kimi could handle up to two million Chinese characters in a single prompt, a tenfold jump. The startup also attracted major investment, including backing led by Alibaba, that pushed its valuation into the billions.

Industry impact: Capital and compute let Moonshot scale training infrastructure and user numbers as Kimi became one of China’s most-used consumer AI apps.

Developer impact: The extreme context window found a real audience among users feeding in whole books, codebases and legal files.

Timeline takeaway: 2024 gave Moonshot the money and scale to graduate from a clever startup to a serious lab.
2M charactersAlibaba-led funding
Oct 2023

The Kimi Chatbot Launches

ConfirmedOctober 2023 · First product

What happened: Moonshot launched Kimi, a chatbot that could process up to 200,000 Chinese characters in a single conversation—far beyond what most rivals handled at the time.

Technical breakthrough: Long-context processing became Kimi’s identity. While others optimised for chat, Moonshot bet that the ability to read enormous inputs would be the differentiator.

Industry impact: Kimi quickly gained a following among students, researchers and professionals who needed to work through long documents.

Timeline takeaway: Kimi’s founding obsession—long context—still defines the family three years later.
200K-char contextLong-context bet
Mar 2023

Moonshot AI Is Founded

MilestoneMarch 2023 · Origins

What happened: Yang Zhilin and two Tsinghua University classmates founded Moonshot AI. The name nods to Pink Floyd’s The Dark Side of the Moon, released fifty years earlier—a moonshot bet on building frontier AI from Beijing.

Industry impact: Moonshot launched into a crowded field of Chinese AI startups, but its focus on long context and, later, open weights set it apart.

Current relevance: The founding team’s research pedigree helps explain how a young lab produced original architecture work rather than just fine-tuning others’ models.

Timeline takeaway: A three-year path from founding to the largest open-weight model is extraordinarily fast by any measure.
Founded 2023Yang Zhilin
💡 Did You Know? Open-weight models let organisations download and run the AI on their own servers, instead of relying entirely on a hosted API. That means sensitive data never has to leave your infrastructure, you are not exposed to a vendor changing or retiring the model, and you can fine-tune it for your own needs—the core reason enterprises and researchers watch open releases so closely.

The Technology, in Plain English

The concepts behind Kimi K3, without the jargon.

Mixture of Experts (MoE)

A dense model runs every parameter for every word. A Mixture-of-Experts model instead splits its network into many specialised sub-networks—experts—and a router picks a handful for each token. Kimi K3 has 2.8 trillion parameters spread across 896 experts, but only 16 fire for any given token. You get the knowledge of a huge model at the running cost of a much smaller one. That is how a 2.8-trillion-parameter system can be practical to serve at all.

Long-context processing and Kimi Delta Attention

A model’s context window is how much it can hold in mind at once. K3’s roughly one million tokens is enough for entire codebases or stacks of documents. The catch is that standard attention gets expensive as context grows. Moonshot’s answer is Kimi Delta Attention, a hybrid linear-attention design intended to keep very long context affordable rather than letting cost explode with length.

Native vision, reasoning and agents

Native vision means the model understands images and video directly, not through a bolted-on add-on. Reasoning is the ability to work through a problem step by step before answering—K3’s always-on thinking mode. Agentic AI is when a model plans, uses tools and takes multi-step actions toward a goal; K3 Swarm Max is built to run many such agents in parallel.

Training, inference, memory and open weights

Training is the expensive one-time process of teaching the model from trillions of tokens. Inference is running the finished model to answer a prompt—the cost you pay per use. Memory, in practice, is a mix of the context window and external systems that feed relevant information back in. Open weights means Moonshot publishes the trained parameters so anyone can download them—which brings us to cost and deployment.

Inference cost and deployment

Open weights are free to download, but not free to run. A 2.8-trillion-parameter model demands serious GPU memory and engineering to self-host, even with MoE keeping active compute modest and MXFP4 quantization shrinking the footprint. For most teams, the practical choices are Moonshot’s API for convenience, or self-hosting when data control, customisation or scale justify the infrastructure.

🔬 Research Insight — Why Long-Context Models Matter

A bigger context window is not just a bigger inbox. When a model can hold an entire codebase, a full contract set or a year of research notes in mind at once, it can reason across that material rather than through a narrow keyhole. That reduces the need for elaborate retrieval pipelines, cuts the errors that come from chopping documents into fragments, and unlocks tasks—whole-repository refactoring, long-horizon agents—that shorter models simply cannot attempt. The hard part has always been doing it affordably, which is exactly what Kimi Delta Attention targets.

Comparisons: The Kimi Family

How K3 stacks up against earlier generations.

DimensionKimi K1.5Kimi K2Kimi K3
ReleasedJan 2025Jul 2025Jul 2026
ArchitectureReasoning model1T MoE (32B active)2.8T MoE (16 of 896 experts)
Context window~128K tokens~128K tokens~1M tokens
ReasoningCore focusStrongAlways-on thinking mode
CodingCompetitiveStrong, agenticTop-tier in blind tests
VisionLimitedAdded in K2.5Native
Agent supportBasicAgentic tool useMulti-agent (Swarm Max)
OpennessPartly openOpen weightsOpen weights (Modified MIT)

Comparisons: K3 vs the Frontier

Only publicly documented traits are compared. Closed-model internals are not disclosed by their makers.

ModelMakerWeightsContext (public)Notes
Kimi K3Moonshot AIOpen (Modified MIT)~1M tokens2.8T MoE, native vision, agents
GPT-5.6OpenAIClosedLarge (undisclosed internals)Frontier proprietary system
Claude (Opus 4.8 / Fable 5)AnthropicClosedLong contextStrong reasoning and coding
GeminiGoogle DeepMindClosedVery long contextDeeply multimodal
DeepSeek V4DeepSeekOpen weightsLong contextEfficient open MoE
QwenAlibabaOpen weightsLong contextBroad open model family
LlamaMeta AIOpen weightsLong contextWidely adopted open baseline

Two things stand out in that table. First, on the axis that Moonshot most wants to win—openness at frontier scale—K3’s main rivals are not the closed US giants but other open labs like DeepSeek, Qwen and Llama, and K3 is the largest of them. Second, direct capability comparisons with closed models are genuinely hard: OpenAI, Anthropic and Google do not publish parameter counts or architectures, so any head-to-head rests on benchmarks and blind tests rather than spec sheets. That is why independent evaluations matter—and why they should be read as evidence, not gospel.

What the Benchmarks Actually Measure

Reading leaderboards like a researcher, not a fan.

Coding benchmarks such as blind front-end tests ask developers which model’s output they prefer for a real task. That captures something genuine—usable code, sensible structure—but it rewards style and first-attempt polish, and preferences vary by language and framework. A win on front-end UI does not automatically transfer to systems programming or debugging a legacy codebase.

Reasoning benchmarks test math, logic and multi-step problem solving. High scores show a model can follow a chain of thought, but many public benchmarks are near saturation and vulnerable to contamination—test questions leaking into training data—so small gaps at the top are often noise. Agent benchmarks measure whether a model can plan, call tools and complete multi-step tasks; they are the most realistic and the most fragile, because a single wrong tool call can sink an otherwise strong run.

The honest reading: Kimi K3’s results are strong and independently reported, which is meaningful. But a leaderboard rank is a measurement of specific tasks on a specific day, not a permanent statement that one model is better than another for your work.

AI Timeline Takeaway

From Followers to Frontier

Pros and Limitations

An honest ledger, based on current evidence.

Strengths

  • Largest open-weight model to date, with a permissive Modified MIT license.
  • Roughly one-million-token context for whole-codebase and long-document work.
  • Strong, independently reported coding and front-end results.
  • Native vision plus always-on reasoning in a single system.
  • OpenAI-compatible API and self-hosting both available.
  • Multi-agent support via K3 Swarm Max.

Limitations

  • Self-hosting 2.8T parameters demands heavy GPU infrastructure.
  • Inference cost and latency at long context can be significant.
  • Closed models may still lead on specific reasoning or safety tasks.
  • Benchmark rankings are snapshots and can shift quickly.
  • Full weights arrive after the announcement, so early claims precede scrutiny.
  • Real-world reliability for autonomous agents remains unproven at scale.

By the Numbers

Key Kimi K3 figures. Third-party results are marked.

MetricFigure
Total parameters2.8 trillion (MoE)
Active experts per token16 of 896
Context window~1,048,576 tokens
LicenseModified MIT (open-weight)
Announced16 July 2026
Weights release (stated)Around 27 July 2026
API input price~$3 per million tokens
API output price~$15 per million tokens
Cache-hit input price~$0.30 per million tokens
Default max output tokens131,072 (configurable)
Predecessor (K2)1T MoE, 32B active, Jul 2025
Coding blind tests (third-party)Preferred over leading US models
🔮 Future Watch: Based on Moonshot’s trajectory rather than promises, watch for deeper multimodal capability, more reliable agent workflows, efficiency gains that lower the cost of serving huge contexts, and clearer enterprise deployment tooling. None of this is guaranteed, and roadmaps change—treat future capability as direction of travel, not fact.

Key Entities and Terms

The people, organisations and concepts behind Kimi K3.

Model

Kimi K3

Moonshot AI’s 2.8-trillion-parameter open-weight Mixture-of-Experts model with a one-million-token context window and native vision.

Company

Moonshot AI

The Beijing startup, founded in 2023 and backed by investors including Alibaba, that builds the Kimi chatbot and model family.

Founder

Yang Zhilin

Co-founder and CEO of Moonshot AI, a Tsinghua University researcher who set the lab’s long-context and open-weight direction.

Product

Kimi

Both Moonshot’s consumer chatbot and the name of its model line, known first for extreme long-context reading.

Concept

Mixture of Experts

An architecture that activates only a few specialised sub-networks per token, giving large capacity at lower running cost.

Technique

Kimi Delta Attention

Moonshot’s hybrid linear-attention method designed to keep very long context affordable to process.

Concept

Open-weight AI

Models whose trained parameters are published for download, enabling self-hosting, fine-tuning and independent study.

Concept

AI Agents

Systems that plan, use tools and take multi-step actions toward a goal; K3 Swarm Max targets many agents in parallel.

Concept

Long Context

The amount of text a model can hold at once; K3’s ~1M tokens spans whole codebases and document sets.

Concept

Reasoning Model

A model that works through problems step by step before answering, rather than replying instantly.

Rival

DeepSeek and Qwen

Other Chinese labs (DeepSeek; Alibaba’s Qwen) whose open models are Kimi K3’s closest open-weight competition.

Rivals

OpenAI, Anthropic, Google DeepMind, Meta AI

The US labs behind GPT, Claude, Gemini and Llama—K3’s closed and open benchmarks of comparison.

Confirmed Facts vs Community Reports

Confirmed by Moonshot AI: the 2.8-trillion-parameter MoE design, ~1M-token context, native vision, thinking mode, Kimi Delta Attention and Attention Residuals, the K3 Max and Swarm Max variants, and the Modified MIT open-weight license.

Third-party / independent: blind-test coding preference and general leaderboard rankings come from evaluators such as LMArena, not Moonshot, and can change over time.

Community reports: the exact 27 July weights date and some deployment details circulated via researchers and press ahead of the full public release, so treat precise figures as provisional until the weights and technical report are out.

Lesser-Known Facts

Explore More AI Timelines

Related reading from AiTimeline.

Frequently Asked Questions

Thirty clear answers on Kimi K3 and Moonshot AI.

What is Kimi K3?
Kimi K3 is Moonshot AI’s flagship large language model, announced on 16 July 2026. It is an open-weight Mixture-of-Experts model with 2.8 trillion total parameters, a roughly one-million-token context window, native vision and an always-on reasoning mode, designed for coding, agents and long-context work.
Who created Kimi K3?
Kimi K3 was built by Moonshot AI, a Beijing-based startup founded in 2023 by Yang Zhilin and two Tsinghua University classmates. The company is backed by investors including Alibaba and develops both the Kimi chatbot and the Kimi model family.
When was Kimi K3 released?
Moonshot AI announced Kimi K3 on 16 July 2026, shortly before the World Artificial Intelligence Conference in Shanghai. The company has stated that full open weights will be published on Hugging Face around 27 July 2026 under a Modified MIT license.
Is Kimi K3 open source?
Kimi K3 is open-weight rather than fully open-source. Moonshot releases the trained model weights under a Modified MIT license, so anyone can download, run and fine-tune the model, but the full training data and pipeline are not released. This is the same approach Moonshot used for Kimi K2.
What is open-weight AI?
Open-weight AI means the developer publishes the trained parameters of a model for download, so organisations can run it on their own infrastructure, fine-tune it and study it. It differs from fully open-source, which would also release training data and code, and from closed models available only through a hosted API.
How many parameters does Kimi K3 have?
Kimi K3 has 2.8 trillion total parameters, which Moonshot says makes it the largest open-weight model released to date. Because it uses a Mixture-of-Experts design, only 16 of its 896 experts activate for any given token, so the active compute per token is far smaller than the total suggests.
What is the context window of Kimi K3?
Kimi K3 supports a context window of roughly one million tokens, about 1,048,576. That is large enough to hold entire codebases, long contract sets or stacks of research papers in a single prompt, which is central to its use for long-document analysis and whole-repository coding tasks.
What license is Kimi K3 released under?
Kimi K3 is released under a Modified MIT license, the same permissive open-weight license Moonshot used for the Kimi K2 family. It allows commercial use, self-hosting and fine-tuning, with some conditions, which is why enterprises can deploy it on their own infrastructure.
What is Mixture of Experts?
Mixture of Experts is a model architecture that divides the network into many specialised sub-networks, or experts, and uses a router to activate only a few per token. It gives a model very large capacity while keeping the compute cost per token low, which is how a 2.8-trillion-parameter model can be practical to run.
What is Kimi Delta Attention?
Kimi Delta Attention is a hybrid linear-attention mechanism developed by Moonshot AI for Kimi K3. It is designed to keep very long context affordable to process, addressing the way standard attention grows expensive as the context window gets larger. Moonshot also introduced a technique called Attention Residuals.
Does Kimi K3 support vision?
Yes. Kimi K3 has native vision, meaning it can understand images directly as part of the model rather than through a separate add-on. Moonshot introduced native vision in Kimi K2.5 in early 2026 through an encoder called MoonViT, and K3 builds that capability into the flagship model.
Does Kimi K3 support coding?
Coding is one of Kimi K3’s strongest areas. In independent blind tests reported at launch, developers preferred K3 for front-end coding over several leading US models. It is also built for agentic coding, meaning it can plan, call tools and work through multi-step programming tasks rather than just generating single snippets.
What is the thinking mode in Kimi K3?
Thinking mode is Kimi K3’s always-on reasoning capability. Instead of answering instantly, the model works through a problem step by step before responding, which tends to improve accuracy on math, logic and complex coding. It reflects the wider industry shift toward reasoning models that deliberate before they answer.
What are K3 Max and K3 Swarm Max?
They are the two variants Moonshot shipped at launch. K3 Max is tuned for chat and single-agent tasks, while K3 Swarm Max is built for large-scale parallel processing across multi-agent workloads, where many agents run at once. Teams pick the variant that matches whether they need one strong agent or many.
How much does the Kimi K3 API cost?
At launch, Moonshot priced the Kimi K3 API at roughly $3 per million input tokens and $15 per million output tokens, with cache-hit input around $0.30 per million. The default maximum output is about 131,072 tokens, configurable up toward the full context window. Prices can change, so check Moonshot’s official pricing.
How does Kimi K3 compare with GPT and Claude?
In independent blind tests reported at launch, Kimi K3 was competitive with and sometimes preferred over leading systems from OpenAI and Anthropic, especially for front-end coding. Direct spec comparison is limited because those models are closed, so any comparison rests on benchmarks rather than published parameters and should be read as evidence, not a final ranking.
How does Kimi K3 compare with DeepSeek and Qwen?
DeepSeek and Alibaba’s Qwen are Kimi K3’s closest open-weight rivals. All three publish downloadable weights, but K3 is the largest, at 2.8 trillion parameters. The open Chinese labs increasingly compete with each other as much as with US firms, and the best choice depends on your task, budget and hardware rather than size alone.
Can businesses deploy Kimi K3 locally?
Yes. Because Kimi K3 is open-weight, businesses can download and self-host it on their own infrastructure, keeping sensitive data in-house and customising the model. The trade-off is that running a 2.8-trillion-parameter model requires substantial GPU resources, so many teams start with the API and self-host only when data control or scale justifies it.
What hardware do you need to run Kimi K3?
Self-hosting Kimi K3 needs a multi-GPU server with large total memory, since even a Mixture-of-Experts model of this size must hold all its parameters in memory. MXFP4 quantization reduces the footprint, but this is data-centre-class hardware, not a laptop. Teams without that infrastructure typically use Moonshot’s hosted API instead.
What is Moonshot AI?
Moonshot AI is a Beijing-based artificial intelligence startup founded in 2023, known for the Kimi chatbot and the Kimi model family. Backed by investors including Alibaba, it focused early on long-context models and later on open-weight releases, becoming one of China’s most prominent frontier AI labs.
Who is Yang Zhilin?
Yang Zhilin is a co-founder and the CEO of Moonshot AI. A Tsinghua University researcher with a strong machine-learning background, he helped set the company’s direction toward long-context processing and open-weight models, and is one of the most watched figures in China’s AI industry.
How is Kimi K3 different from Kimi K2?
Kimi K2, from July 2025, was a one-trillion-parameter Mixture-of-Experts model with about 32 billion active parameters and a 128K context window. Kimi K3 scales this to 2.8 trillion parameters, expands context to around one million tokens, adds native vision and always-on reasoning, and introduces new attention techniques for efficiency.
What is an AI agent and does Kimi K3 support agents?
An AI agent is a system that plans, uses tools and takes multiple steps to complete a goal, rather than answering a single prompt. Kimi K3 is built for agentic work, and its K3 Swarm Max variant is specifically designed to run many agents in parallel across large multi-agent workloads.
What languages does Kimi K3 support?
Kimi K3 is a multilingual model with particular strength in Chinese and English, reflecting Moonshot’s origins and training data, and it also handles many other languages and, crucially for developers, a wide range of programming languages. As always, quality varies by language, so test it on the specific languages you need.
What are the limitations of Kimi K3?
Key limitations include the heavy GPU infrastructure needed to self-host 2.8 trillion parameters, meaningful inference cost and latency at very long context, and the fact that closed models may still lead on some reasoning or safety tasks. Benchmark rankings are also snapshots, and real-world agent reliability at scale remains to be proven.
Is Kimi K3 better than closed models?
On some tasks, independent tests suggest it is competitive with or preferred over leading closed models, especially for coding. But better is task-specific: closed systems may still lead elsewhere, and openness itself is a distinct advantage for data control and customisation. The honest answer is that it depends on what you are building and how you measure quality.
When will the Kimi K3 weights be available?
Moonshot has said full open weights for Kimi K3 will be published on Hugging Face around 27 July 2026, shortly after the 16 July announcement, under a Modified MIT license. The exact date has been reported ahead of release, so it is best confirmed against Moonshot’s official channels once the weights go live.
What is a reasoning model?
A reasoning model is one trained to work through a problem step by step, often generating an internal chain of thought, before giving a final answer. This tends to improve performance on math, logic and complex coding compared with models that respond instantly. Kimi K3’s always-on thinking mode is an example of this approach.
Why do open-weight models matter?
Open-weight models let organisations run AI on their own infrastructure instead of depending entirely on hosted APIs. That protects sensitive data, avoids lock-in to a single vendor, enables fine-tuning for specific needs, and lets researchers study the model directly. Each capable open release also pressures the wider industry to keep more capability freely available.
What makes Kimi K3 different from other AI models?
Kimi K3 combines frontier scale with openness: 2.8 trillion parameters released as downloadable weights, paired with a one-million-token context, native vision, always-on reasoning and new attention techniques for efficient long-context processing. Few models offer that mix, and none at this size are openly available, which is what set it apart at launch.

Why Kimi K3 Matters Beyond Benchmarks

It is tempting to reduce a model like Kimi K3 to a row on a leaderboard. Resist it. Benchmark scores are useful signals, but they are measured on fixed tasks under fixed conditions, and they say little about whether a model will hold up inside your codebase, your workflow or your regulatory environment. A model that wins a blind coding test can still frustrate a team that needs reliability over flair, and a model that ranks second can be the right choice because you can run it on your own servers.

Kimi K3’s deeper significance is what it represents. It widens the competition at the frontier of AI, pushing both open and closed labs to move faster. It advances the open-weight movement by proving that scale and openness can coexist, giving developers and enterprises more genuine deployment choices—API or self-hosted, closed or open, US or Chinese. And it marks a shift in who gets to define the frontier, with a three-year-old Beijing lab shipping the largest open model in the world.

What it does not do is settle anything. The AI landscape is moving too quickly for any single release to be the last word, and next month’s model—from Moonshot or anyone else—will rewrite parts of this page. That is precisely why a model’s real value shows up not in the week of its launch, but in the months afterward, in the hands of the people who actually build with it.

📚On sourcing: Confirmed specifications and release details are attributed to Moonshot AI. Benchmark and ranking claims come from independent evaluators such as LMArena and are labelled as third-party. Some figures, including the exact weight-release date, were reported ahead of the full public release and should be treated as provisional until Moonshot’s weights and technical report are out. This is an educational timeline, not investment advice. Last reviewed against current sources: 18 July 2026.