Available Through the Kimi API

Kimi K3

Moonshot AI's most capable flagship model is a 2.8-trillion-parameter, native-vision model with a 1-million-token context window, long-horizon coding strength, and always-on reasoning.

2.8T
Total parameters
1M
Context window
Native
Image and video understanding
Official Evaluations

Kimi K3 Benchmark Results

The official Kimi evaluation suite covers coding, terminal work, browsing, office automation, knowledge, reasoning, and visual tasks. The two official charts below preserve Kimi's published comparison context.

Swipe horizontally to inspect the full benchmark charts.

Coding Benchmarks
Official Kimi K3 coding benchmark comparison chart
Agentic and Vision Benchmarks
Official Kimi K3 general-agent and visual-agent benchmark comparison chart
Availability and Pricing

Official Kimi K3 Pricing

Kimi K3 uses flat pay-as-you-go token pricing across its 1M context window, with separate input rates for automatic cache hits and misses.

Cache-Hit Input
$0.30 / 1M tokens
Cache-Miss Input
$3.00 / 1M tokens
Output
$15.00 / 1M tokens

Official Kimi API pricing checked on July 17, 2026. Kimi reports a cache hit rate above 90% for coding workloads on its disaggregated Mooncake inference architecture.

Frontier Capabilities

Why Kimi K3 Matters for Hosted Agents

Kimi K3 combines long-horizon software engineering, production-oriented knowledge work, and native visual understanding in one model built for extended agent workflows.

Long-Horizon Coding

Kimi reports that K3 can sustain long engineering sessions, navigate massive repositories, coordinate terminal tools, optimize kernels, and build complete systems with limited human oversight.

Agentic Knowledge Work

K3 is designed to turn research, evidence, spreadsheets, documents, and web retrieval into interactive reports, visualizations, and production-ready outputs.

Native Visual Reasoning

Native image and video understanding lets K3 inspect screenshots and live visual feedback while iterating on frontend, game, CAD, research, and media workflows.

Official Case Studies

Coding and Long-Horizon Creation

Kimi's launch materials show K3 moving beyond short coding prompts into extended engineering, compiler, game, chip, and computational research projects.

Native Vision in the Loop

From Visual Feedback to a Playable 3D World

This official Kimi case shows K3 turning a concept into a procedural browser-based 3D exploration game with forests, villages, snowy mountains, dynamic weather, a rider, and a horse.

The local video is copied from Kimi's official K3 technical blog so the demonstration remains available without depending on a third-party embed.

Kernel Optimization

In identical 24-hour sandboxes, K3 optimized four GPU kernel tasks spanning AttnRes, KDA, and MLA, performing competitively with Kimi's strongest comparison model and ahead of several other frontier systems.

GPU Compiler Development

K3 built MiniTriton from scratch with a tile-level IR over MLIR, optimization passes, PTX code generation, and end-to-end nanoGPT training that tracked the reference loss curve.

Game Development and Digital Creation

K3 combined 3D reasoning, code, and visual feedback to create playable browser experiences, including a procedural open world, emulator, fighting game, FPS arena, and scientific visualization.

Chip Design

In a 48-hour autonomous run, K3 designed and verified a 4 mm² chip for a nano model, closing timing at 100 MHz and simulating more than 8,700 tokens per second of decode throughput.

Coding for Research

K3 reproduced I–Love–Q relations by reviewing more than 20 papers, evaluating over 300 equations of state, writing over 3,000 lines of Python, and producing an interactive research dashboard.

Knowledge Work

Research, Visualization, and Persistent Workflows

Kimi K3's knowledge-work cases combine long context, web and terminal retrieval, subagents, data analysis, presentation design, dashboards, and video editing.

Interactive Industry Research

K3 produced a drill-down research site spanning 42 years of the AI ASIC industry after 120+ self-improvement rounds, 2.8K+ web searches and fetches, 1.1K+ terminal pulls, and 11K+ pages.

Financial and Scientific Visualization

Official examples include a fusion-industry consulting report with timelines and Gantt charts plus a GWTC-5 analysis of 391 gravitational-wave events using more than 20 concurrent subagents.

Widgets and Dashboards

Kimi Work can turn K3 outputs into interactive widgets connected to local data or plugins, then organize those components into persistent dashboards around a project or topic.

Motion Design and Video Editing

K3 created an architecture explainer with animated diagrams and edited its own teaser from 56 source clips, handling selection, matched cuts, beat synchronization, audio, and revision.

Model Scale

A 2.8T Open Frontier Model

Kimi K3 is the first announced open model in the 3-trillion-parameter class. Moonshot says Kimi models held the open-model scale frontier for nine of the twelve months leading into the K3 launch, with full weights scheduled for release by July 27, 2026.

Official Kimi chart showing open-source frontier model scale over time
Architecture

Architecture Built for Trillion-Scale Efficiency

Kimi K3 combines Kimi Delta Attention, Attention Residuals, Gated MLA, and Stable LatentMoE to improve information flow across long sequences and deep networks.

Only 16 of 896 experts are activated for each token. Moonshot reports that the architecture, training methods, and data recipe deliver roughly 2.5× the overall scaling efficiency of Kimi K2. K3 also uses quantization-aware training with MXFP4 weights and MXFP8 activations.

Official Kimi K3 architecture diagram showing KDA, Attention Residuals, Gated MLA, and Stable LatentMoE
MyClaw Agent Fit

Where Kimi K3 Fits in MyClaw

K3 is most useful when a hosted agent can combine its long context and reasoning with persistent files, terminal tools, browsers, and scheduled execution.

Repository Engineering

Use K3 for large codebase inspection, terminal-driven implementation, tests, profiling, refactoring, and long-running repair loops that need a persistent workspace.

Research to Deliverable

Give K3 browser, document, spreadsheet, and presentation workflows so it can turn source gathering into a report, dashboard, decision memo, or implementation plan.

Visual Build Loops

Pair native vision with screenshots and local previews for frontend, game, design, CAD, chart, and media workflows where the model must inspect and refine its own output.

Cost-Aware Long Context

Keep stable prefixes for automatic cache reuse, route high-value long-context tasks to K3, and reserve simpler requests for lower-cost models in the MyClaw directory.

MyClaw Skills

Use Kimi K3 With Focused Agent Skills

Pair Kimi K3's long context, native vision, and coding strength with reusable MyClaw skills for repository work, research, SEO, and persistent agent operations.

Related Models

Compare Kimi K3 With Other Agent Models

Compare Kimi K3 with other frontier models for long-context coding, multimodal work, reasoning, price, and production agent routing.

Kimi K3 FAQs

Run Kimi K3 in a Hosted MyClaw Agent

Put Kimi K3's 1M context, native vision, long-horizon coding, and knowledge-work capabilities inside a managed agent workspace with files, tools, browsers, and persistent execution.

Run Kimi K3 in MyClaw