Signals

Kimi K3 Makes Frontier-Class Agent Intelligence Open-Weight

Published July 16, 2026Update Released Jul 16, 2026
Comic cover showing a chain labeled '1M TOKEN CONTEXT' breaking a moat wall, a toll bridge to 'AI GATEWAY' with '$3/M input, $15/M output', and a developer holding an 'OPEN WEIGHTS' lantern. Title 'Kimi K3' and verdicts as specified.

Moonshot released Kimi K3 as a 2.8T open-weight model for long-horizon coding, tool use, and knowledge work. Moonshot's own evaluations put it near Fable 5 and GPT-5.6 Sol on several agent benchmarks, while Vercel AI Gateway serves it for $3 per million input tokens and $15 per million output tokens.

Primary SourceView Full Update In Vercel ChangelogJul 16, 2026

What Changed

  • Kimi K3 is a 2.8T mixture-of-experts model with 104B active parameters, native text and image input, a 1M-token context window, and thinking enabled by default.
  • Moonshot released the complete model weights under the Kimi K3 License, giving teams a deployment and customization path that closed frontier models do not offer.
  • Moonshot reports Kimi K3 at 88.3 on Terminal-Bench 2.1, 77.8 on ProgramBench, and 94.5 on MCPMark Verified, close to or above Fable 5 and GPT-5.6 Sol on those specific evaluations.
  • Vercel AI Gateway lists Kimi K3 at $3 per million input tokens, $15 per million output tokens, and $0.30 per million cache-read tokens across Moonshot and several third-party providers.
  • That is 70% below Fable 5's $10 input and $50 output list price per million tokens. It does not prove a 70% lower cost per successful task, which depends on quality, token use, latency, and retries.
  • US-based providers on AI Gateway add optional Zero Data Retention, US inference routing, automatic failover, and a faster serving tier.
  • The benchmark table is Moonshot-reported and mixes harnesses across models. The weights are open, but the Kimi K3 License includes commercial conditions for very large products and model-as-a-service businesses.

Who Should Care

  • Agent builders comparing Fable 5, GPT-5.6 Sol, and other premium models for coding, tool use, or long-horizon work.
  • Teams whose model volume makes a $3 input and $15 output API route materially different from premium closed-model pricing.
  • Products that want an open-weight deployment option without giving up access to managed US inference, failover, and Zero Data Retention.
  • Builders already using OpenAI- or Anthropic-compatible clients who can run a bounded Kimi K3 evaluation without rebuilding their model layer.

Who Should Not Care

  • Small workloads where premium-model spend is negligible and another provider route would add more complexity than value.
  • Teams that require independently reproduced benchmarks, mature enterprise support, or a simple unrestricted open-source license before evaluation.
  • Workloads that cannot use Moonshot, third-party Kimi providers, or self-hosted infrastructure because of privacy, residency, procurement, or policy constraints.
  • Products without a real coding, multimodal, knowledge-work, or long-horizon agent workload to benchmark.

The Verdict

This is one of July's structural Signals, not another model-release headline. A model competing in the Fable 5 and GPT-5.6 Sol performance tier is now available as weights and through production APIs at $3 input and $15 output per million tokens. That shrinks both the capability gap and the closed-provider pricing moat. The benchmarks remain vendor-reported, so evaluate cost per successful task, hosting requirements, provider boundaries, and license fit before adopting it.

Take This To Your Agent

Copy a ready-to-paste investigation handoff that asks your coding agent to check this update against your product and recommend what to do next.

From Signal To Advantage

Make Your Stack's Acceleration Your Advantage

ShipFoundry monitors the tools your product depends on, checks important updates against your stack and codebase, and surfaces the opportunities worth acting on, without you living in changelogs or X.