
Qwen 3.8 vs Kimi K3: Coding, Price & Agent Tests
By Emma Reed
MyClaw Editorial
MyClaw
Run Best-in-Class AI Agents Now
Run OpenClaw or Hermes Agent on managed MyClaw hosting, so your AI agent stays online, updated, and ready for real work.
Updated July 20, 2026. Both models are changing quickly.
AI Takeaway
- Which model is the safer choice today? Kimi K3. It has published API pricing, documented benchmarks, a confirmed 1M-token context window, and broader public access. Qwen3.8-Max-Preview is promising, but it is still changing.
- Which performed better in a direct coding comparison? In one matched 269-file repository test, Kimi scored 83/100 and Qwen scored 80/100. Kimi finished faster and used fewer tokens; Qwen produced cleaner system boundaries and stronger replay metadata. One test is useful evidence, not a universal verdict.
- Which is cheaper? Kimi charges $0.30 per 1M cache-hit input tokens, $3 per 1M cache-miss input tokens, and $15 per 1M output tokens. Qwen currently uses promotional Token Plan credits, so the prices are not directly comparable.
- Can either model be self-hosted now? Not in a practical, generally available form. Both open-weight releases are still pending, and trillion-parameter models require serious infrastructure.
- Bottom line: Pick Kimi for a better-documented production evaluation. Pick Qwen for controlled, low-cost preview testing, then test it again after the production release.
Qwen 3.8 vs Kimi K3 at a Glance
Released days apart, the models are not equally mature: Kimi can be evaluated today, while Qwen remains a changing preview.
| Category | Qwen3.8-Max-Preview | Kimi K3 |
|---|---|---|
| Current status | Preview | Public model and API |
| Total parameters | 2.4T, vendor-reported | 2.8T |
| Active parameters | Not disclosed | 16 of 896 experts per token |
| Context window | About 1M reported in integration metadata; not yet confirmed in a full model card | 1M tokens |
| Input | Text and visual input | Text, images, and video |
| Published benchmarks | No complete official benchmark table yet | Broad coding, agent, reasoning, and vision results |
| Pricing | Token Plan credits and temporary discounts | Standard pay-as-you-go API rates |
| Open weights | Promised, no release date | Full weights scheduled for July 27, 2026 |
| Best current use | Controlled experiments | Production-oriented evaluation |
The naming deserves attention: Qwen3.8-Max-Preview is not Qwen3-8B. Several automatically generated comparison pages currently confuse the two. For a closer look at Kimi's architecture, scores, and official prices, the Kimi K3 model guide keeps the details in one place.
Specs, Context, and Release Status
Qwen3.8-Max-Preview Is Still Moving
Alibaba's Token Plan documentation confirms the qwen3.8-max-preview model ID, reasoning, visual understanding, and text generation. It also warns that the endpoint will keep changing before being retired or replaced.
Alibaba reports 2.4 trillion total parameters and plans open weights, but active parameters, the final license, standard API pricing, and complete benchmarks are missing. A roughly 1M-token context window appears in integration metadata; treat it as provisional until the model card arrives.
Kimi K3 Has Fewer Blank Fields
Kimi K3 has 2.8 trillion total parameters and activates 16 of 896 experts for each token. The official Kimi K3 technical blog confirms native image and video understanding, a 1M-token context window, and always-on reasoning—useful for sessions combining code, documents, screenshots, and tools.
Full weights are scheduled for July 27, not available today. The final license and serving requirements remain unknown, and ordinary laptops are out.
Benchmark Claims vs Real-World Coding
Kimi Has More Scores, but the Harness Matters
Moonshot reports strong Terminal-Bench 2.1, FrontierSWE, Program Bench, BrowseComp, and visual reasoning results. But Kimi used Kimi Code on some tasks while comparison models used Claude Code, Codex, or other harnesses. Prompts, retries, context management, tools, and safety behavior can all move the score. Treat a narrow lead with different harnesses as a near-tie. The recent Kimi K3 vs GPT-5.6 Sol test also shows how token use can erase an apparent price advantage.
The Current Head-to-Head Test Was Close
A matched StackPerf experiment asked both models to inspect 269 frozen files, define an integration, cite evidence, and produce a migration plan. Kimi scored 83 after factual penalties; Qwen scored 80.
The three-point gap is less interesting than how they worked:
- Qwen drew cleaner boundaries between systems, used fewer tools, and recorded stronger replay metadata.
- Kimi completed the job faster, used fewer tokens, and handled revisions, regeneration, and lifecycle state more completely.
- Both reached broadly similar architectural conclusions and still made claims that required correction.
One run per model cannot settle every coding task, but it shows that Qwen's preview can sustain repository exploration and function tools.
Which Model Is Better for Coding and AI Agents?
Choose Qwen for Controlled Experiments
Qwen3.8-Max-Preview makes sense when low-cost access matters more than stability. Early results encourage repository analysis, function calling, architecture work, and multimodal input. The endpoint can change, pricing is promotional, and independent evidence is thin. Save every test's files, endpoint, date, and settings.
Choose Kimi for Long, Tool-Heavy Work
Kimi is easier to recommend for extended coding, browser research, visual build loops, and knowledge work. Its context, API prices, reasoning behavior, and benchmark setup are documented, and it showed better speed and token efficiency in the matched test. It is not automatically cheaper: long reasoning and review time can consume the headline discount.

For either model, the surrounding workflow matters as much as the first response. A useful coding agent workflow includes repository access, terminal execution, tests, logs, and a definition of success—not just a prompt asking for code.
Pricing and Access Are Not Apples to Apples
Kimi Uses Standard API Rates
Kimi charges $0.30 per million cache-hit input tokens, $3 per million cache-miss input tokens, and $15 per million output tokens. Caching matters when the same repository instructions and files appear across runs. A monthly Kimi Code plan is different from API billing and depends on allowances, reset rules, and throttling.
Qwen Uses Preview Credits
Qwen3.8-Max-Preview comes through Alibaba's Token Plan and supported coding tools. Temporary multipliers make it inexpensive to explore, but they are not a permanent price per million tokens. Standard production pricing is still missing.
Token Plan is also limited to interactive use in approved programming and agent tools. Alibaba prohibits automated scripts, application backends, and non-interactive batch processing. A scheduled production agent therefore needs an endpoint whose terms allow that workload.
The most useful calculation is:
Completed task cost = model spend + retries + elapsed time + human review + failure cleanup
A low token rate loses its appeal after three failed attempts. Count accepted results rather than generated tokens. The guide to choosing the best model for OpenClaw applies the same task-first approach to cloud and local options.
Turn Your Chosen Model Into a 24/7 Agent With MyClaw

Choosing a model does not provide persistent files, browser access, terminal tools, schedules, integrations, or recovery. MyClaw runs OpenClaw and Hermes Agent in managed, always-on workspaces with supported providers and bring-your-own-key connections.
Step 1: Launch an Agent, Not Another Chat Tab
Choose OpenClaw or Hermes Agent and start a private workspace. The runtime stays online with persistent storage, monitoring, updates, and backups, so a long job does not disappear when a laptop sleeps.
Step 2: Bring the Model You Chose
Connect a supported provider or add an API key. Select Kimi K3 or a compatible Qwen3.8 endpoint when available, with terms that match the task. Switch models at task boundaries: Moonshot warns that K3 can become unstable when moved into a session without its full thinking history.
Step 3: Give It a Real Job
Connect the repository, files, browser, or communication tools the job needs. Try a stubborn bug, a competitor report due Friday, or a recurring research brief. Set a time limit and define three checks that determine whether the result is usable. Compare completion quality, elapsed time, token use, and required corrections; keep the model that performs better on that job.
Qwen 3.8 or Kimi K3: What Should You Choose?
- Pick Kimi K3 for a production-oriented evaluation today, especially for long, visual, or tool-heavy work.
- Pick Qwen3.8-Max-Preview for inexpensive experimentation and architecture-focused coding tests.
- Wait if open weights, stable pricing, or reproducible third-party benchmarks are mandatory.
Whichever model wins, API access still needs an environment that can hold context and keep working. A managed OpenClaw workspace is more practical than self-hosting a trillion-parameter model and gives the model files, tools, memory, and schedules without tying the workflow to one provider.
Conclusion: Qwen 3.8 vs Kimi K3
Kimi K3 currently wins on decision confidence: it has clearer documentation, transparent API pricing, more benchmark evidence, and a confirmed context window. Qwen3.8-Max-Preview is the more uncertain bet, but its early repository performance and discounted access make it worth testing.
Qwen still needs a model card, standard pricing, and independent benchmarks. For now, the better model is the one that completes a real job with fewer retries, lower total cost, and less cleanup.
Skip the Setup, Run Best-in-Class AI Agents Now
Launch a managed OpenClaw or Hermes Agent workspace in minutes, with always-on hosting, updates, and support handled by MyClaw.