Claude Opus 5 Cuts the Price in Half — Meanwhile Kimi K3 Gets Caught Calling Itself Claude

About 18 min read · MACCOME · Last updated: July 25, 2026

Who this is for: developers and engineering leads evaluating Claude vs. Kimi, or tracking AI compliance and supply-chain risk. This week’s two headlines both come down to value for money and where model capability actually comes from: on July 24, Anthropic released Claude Opus 5 as the new Claude Max default; on July 16, Kimi K3 went live under White House distillation allegations, with independent researchers finding K3 identifying as Claude and leaking internal deployment IDs. What you get: Opus 5 vs. Fable 5 comparison tables, a full K3 controversy timeline, Greenblatt evidence breakdown, a six-step selection runbook, and FAQ. Structure: pain points, Opus 5 analysis, K3 controversy, decision matrix, runbook, wrap-up. K3 specs and benchmarks: Kimi K3 deep review.

bolt

TL;DR — 30-second verdict

  • Opus 5 (7/24): Same pricing as Opus 4.8 ($5/$25 per 1M tokens). CursorBench 3.2 peak trails Fable 5 by only 0.5% at roughly half the cost. Now the Claude Max default. No forced data retention by default.
  • K3 distillation row (7/22–24): White House OSTP director Kratsios alleged "covert industrial-scale distillation" plus unauthorized GB300 chips. Experts question whether two weeks after Fable 5's public launch is enough time.
  • Strongest technical signal: Redwood Research's Ryan Greenblatt found K3 disproportionately calls itself Claude and outputs internal IDs like claude-opus-4-5-20250929 — more accurately than real Claude models do.
  • Pending 7/27: K3 full weights promised for July 27; external parties still cannot independently verify architecture or scores.

Six Pain Points: What Teams Get Wrong After This Week's Headlines

  1. Treating "half-price Opus" as a full Fable 5 replacement: Opus 5 is close on agent coding and automation, but Anthropic deliberately did not push frontier dual-use capabilities (cyber offense, biosecurity). Mythos 5 still holds that lane.
  2. Reading White House statements as proven fact: Kratsios and Treasury Secretary Bessent made public allegations without attachable evidence. Moonshot has not answered detailed training-process questions — silence is not a conviction.
  3. Ignoring the distillation timeline: Fable 5 opened to the public on July 1; K3 shipped July 16 — only two weeks. Several independent researchers argue deep RL-style distillation at that scale is implausible in that window.
  4. Missing the "K3 says Claude" technical signal: This matters more than political rhetoric. It may point to Claude training samples with deployment metadata — but Greenblatt stresses it does not directly prove distillation.
  5. Compliance blind spots on data retention: Fable 5 and Mythos 5 require accepting 30-day data retention. Opus 5 follows the Opus line with no forced retention by default — a real enterprise differentiator.
  6. Running multi-model agent stacks on a laptop: Opus 5 + K3 hybrid routing, long-context sessions, and Gateway workloads need always-on nodes. Closing the lid breaks the chain.

Claude Opus 5: Not the Flagship, but Probably the Best Daily Driver

On July 24, 2026 (US Pacific time), Anthropic released Claude Opus 5 (model ID: claude-opus-5) and made it the Claude Max default. Claude Pro subscribers get the strongest Opus tier. The positioning is clear: a daily workhorse with intelligence near flagship Fable 5 at roughly half the price.

Same price, big performance jump

Opus 5 pricing matches Opus 4.8 exactly: $5 per million input tokens, $25 per million output tokens. One million tokens of context is the default and only tier; max output is 128K. Thinking is on by default, controlled via the Effort parameter. Fast mode runs about 2.5x faster at 2x base price.

Anthropic's official benchmark highlights:

  • Frontier-Bench v0.1 (software engineering): beats every model, more than 2x Opus 4.8, at lower per-task cost.
  • CursorBench 3.2: at max effort, within 0.5% of Fable 5's peak at roughly half the cost; high, xhigh, and max tiers all beat every other model at the same cost tier.
  • ARC-AGI 3 (novel problems): score is 3x the next-best model.
  • OSWorld 2.0 (computer use): beats Fable 5's best result at under one-third of Fable 5's cost.
  • Zapier AutomationBench: end-to-end business automation pass rate about 1.5x the next-best model; even at lowest effort, task count exceeds any other model. Zapier reports Opus 5 passed 100% of tasks no prior model could complete.

The Cursor team: "Claude Opus 5 delivers near Fable 5 intelligence at Opus speed and cost. On CursorBench it's just under Fable 5 and has many of the same behaviors."

Research and enterprise field results

Structural biology, organic chemistry, and bioinformatics all beat Opus 4.8. Internal evals: +10.2 points on inferring molecular structure from spectra; +7.7 points on protein variant function prediction. Box customer data: +11% on data-analysis workflows, +17% on due diligence, +8% overall accuracy.

Alignment and safety: most compliant yet, but not chasing dual-use frontiers

Automated behavior audits show Opus 5 is the most aligned Claude to date: lowest deception rate, hardest to jailbreak. On risky dual-use capabilities (cyber offense, biosecurity), it did not refresh the frontier — restricted Mythos 5 still leads. The cyber classifier is about 85% more permissive than Fable 5: useful for source-level vulnerability discovery, still blocks binary scanning, penetration testing, and exploit generation.

Easy to overlook: no forced data retention by default

Like prior Opus models, Opus 5 does not force data retention on default access. Fable 5 and Mythos 5 require a 30-day retention opt-in. For compliance-sensitive teams, that alone can justify Opus 5 over Fable 5.

DimensionClaude Opus 5Claude Fable 5
Release date2026-07-242026-06-09 (public access 2026-07-01)
Input / output pricing$5 / $25 per 1M tokens~$10 / $50 per 1M tokens
CursorBench 3.2 (max)0.5% below Fable 5 peakPeak benchmark
Claude Max defaultYes (from release day)No
Data retentionNot forced by defaultRequires 30-day retention opt-in
Dual-use frontier (cyber / bio)Deliberately not pushedStronger (Mythos 5 restricted strongest)

The Kimi K3 Distillation Row: White House, Timeline Pushback, Greenblatt's Evidence

If Opus 5 is a straightforward product launch, Kimi K3's story is messier. Moonshot released K3 on July 16 (2.8T parameters, 1M context, sparse MoE), claiming overall capability second only to Fable 5 and GPT-5.6 Sol. Full weights are pledged for July 27 — when the controversy broke, outsiders could not verify independently.

The White House enters the chat

On July 22–23, 2026, White House OSTP director Michael Kratsios posted on X, alleging Moonshot used "large-scale, covert industrial distillation" to extract Anthropic Fable capabilities, and citing suspected use of export-controlled Nvidia GB300 chips (possibly via Thailand-based servers). Treasury Secretary Scott Bessent added that "watermarks" from US models were found on "many Chinese models," without specifying what those watermarks are.

This did not come from nowhere. In February 2026, Anthropic publicly accused Moonshot, DeepSeek, and MiniMax of "industrial-scale distillation attacks," citing more than 3.4 million anomalous API interactions from Moonshot that "clearly deviated from normal usage patterns," and claiming request metadata traced some behavior to Moonshot leadership. Moonshot has never confirmed or denied.

Independent experts: the timeline does not add up

TechCrunch (July 23) interviewed multiple researchers skeptical of distilling K3 in two weeks. Snorkel AI co-founder Braden Hancock:

"Fable only became publicly available on July 1. You cannot distill that much data, finish training, and ship a model in two weeks."

Allen Institute for AI researcher Nathan Lambert argues that as Chinese models approach the frontier, simple supervised fine-tuning distillation yields diminishing returns — the real gap is reinforcement learning, which takes longer. If distillation were that effective, challengers would have caught GLM or K3 earlier via distillation alone. They have not.

Elon Musk admitted in court that xAI distilled OpenAI models when training Grok, calling it industry-normal. Distillation itself is not illegal; the fight is over the line between legitimate technique and covert industrial theft.

Strongest evidence so far: K3 calls itself Claude

Redwood Research chief scientist Ryan Greenblatt (around July 24, GitHub: rgreenblatt/which_claude_is_k3) ran cross-entropy comparisons on model self-identification and found:

  • Kimi K3 disproportionately identifies as Claude, sometimes outputting Claude internal deployment IDs such as claude-opus-4-5-20250929 and claude-sonnet-4-5-20250929.
  • Real Claude Sonnet 4.5 says "I am Claude Sonnet 4.5." Opus 4.5 often omits or misstates version numbers — the teacher models are less accurate about themselves than K3 is.
  • K3's claimed versions lock to the Claude 4.5 generation (late 2025), not current Fable/Mythos. Prior K2 signals pointed to earlier Claude Sonnet 4 — a "chasing each generation" pattern.

Greenblatt's read: this more likely reflects training data mixed with Claude samples carrying deployment metadata (API logs or labeled synthetic data) — a more specific form of distillation than mimicking conversational style. He emphasizes: these findings do not directly prove distillation occurred. Identity confusion could also come from data contamination, system-prompt leakage, or public dataset synthesis.

Beyond the drama: can you use it, and is it worth it?

On Reddit r/LocalLLaMA, three camps: cheer the open/closed gap shrinking from months to days; joke that almost nobody can run 2.8T locally; pragmatists say K3's real sell is low price plus fewer refusals, not beating Fable 5. Until weights drop on July 27, architecture and scores remain unverified externally.

DateEvent
2026-02Anthropic first public accusation against Moonshot / DeepSeek / MiniMax ("industrial-scale distillation"; 3.4M+ anomalous interactions)
2026-07-01Claude Fable 5 opens to the public
2026-07-16Kimi K3 API/product launch (2.8T; full weights pending 7/27)
2026-07-22/23White House OSTP director Kratsios public distillation + chip allegations
2026-07-23TechCrunch publishes expert timeline skepticism
2026-07-24Claude Opus 5 release; Greenblatt publishes "K3 calls itself Claude" analysis
2026-07-27 (planned)Kimi K3 full open weights; external independent verification possible

Claude Opus 5 vs. Kimi K3: Quick Decision Matrix After This Week

Selection factorLean Claude Opus 5Lean Kimi K3 (pending 7/27 verification)
Compliance / retentionNo forced retention by default; mature Anthropic enterprise pathOpen weights + low price, but distillation allegations and supply-chain risk unresolved
Agent / codingCursorBench, AutomationBench, OSWorld official data leadsStrong self-reported Program Bench, SWE Marathon; external repro pending weights
CostHalf-price Fable-tier capability at $5/$25Official positioning well below Fable 5 / GPT-5.6 Sol
Refusals / moderationMost aligned Claude yet; dual-use still restrictedCommunity cites "fewer refusals" as a genuine selling point
VerifiabilityFully available (API, Bedrock, Vertex, etc.)Architecture and scores not independently verifiable before 7/27

Six-Step Runbook: Model Selection and Agent Deployment After This Week

  1. Rank your scenarios: If agent coding, automation, and retention compliance matter most, evaluate Opus 5 first (and stack combinations in the coding assistant decision matrix). If extreme cost plus open-source control is the priority, mark July 27 for a K3 retest.
  2. Check retention and compliance terms: Compare Opus 5 (no forced retention) against Fable 5 / Mythos 5 (30-day opt-in). Cross-border or government projects need separate legal review.
  3. Probe cost with Effort tiers: Opus 5 Thinking is on by default. For CursorBench-class tasks, start at high/max for quality, then drop effort for Zapier-style automation to avoid token waste.
  4. Grade K3 evidence in layers: political allegations (White House), timeline rebuttals (experts), technical indirect evidence (Greenblatt). After July 27 weights drop, rerun identity prompts and benchmark repros.
  5. Deploy a multi-model Gateway: Run OpenClaw or Kimi Code plus OpenRouter routing on a dedicated node so laptop sleep does not kill long sessions. Isolate API keys and provider fallbacks by environment.
  6. Follow up after 7/27: Once weights are public, compare GPQA-Diamond, BrowseComp, and "who are you" identity tests. Update internal routing tables and archive decision rationale.

Three Hard Numbers for Your Technical Review

  • Opus 5 / Fable 5 value gap: CursorBench 3.2 max-effort peak differs by 0.5%; token pricing is about 50% of Fable 5 ($5/$25 vs. ~$10/$50 per 1M) — Anthropic's direct answer to the market's value-for-money demand.
  • ARC-AGI 3 triple gap: Opus 5 scores 3x the next-best model on novel problems; OSWorld 2.0 beats Fable 5's best at <1/3 Fable cost — the strongest quantitative signals for agent generalization and computer use.
  • K3 identity confusion stats: Greenblatt's cross-entropy work shows K3, when asked "who are you," disproportionately outputs Claude internal deployment IDs (e.g. claude-opus-4-5-20250929), locked to the Claude 4.5 generation rather than Fable/Mythos — the most technically substantive, least crowded angle in the distillation debate ("Why does Kimi K3 say it's Claude").

Same Week, Same Theme: Value for Money Is the Real Battleground

Opus 5 answers the market with "half price, near-flagship." Kimi K3 attacks with "open source, low price, fewer limits." The K3 controversy is really a public reckoning over whether low prices come from engineering or from borrowing someone else's model.

For most developers and enterprises: match the scenario before picking a side. Compliance, retention, and supply-chain sensitivity favor Opus 5's value proposition. Extreme cost and open-source control mean waiting for July 27 weights before trusting K3 benchmarks.

Switching Opus 5, K3, and multi-provider Gateways on a local laptop hides three costs: lid-close sleep, network jitter, and scattered API keys. Long-session agents and automation need a dedicated always-on node. MACCOME Mac cloud hosts provide real macOS, SSH handoff, and environment isolation so Claude Code, OpenClaw, and multi-model routing run on stable instances. Plans and pricing: Mac mini cloud rental rates.

Sources: Anthropic official release (anthropic.com/news/claude-opus-5), Moonshot/Kimi technical blog, TechCrunch, Ryan Greenblatt GitHub (rgreenblatt/which_claude_is_k3), CNBC, The Verge. Benchmarks are vendor-reported or third-party statistics as of July 25, 2026.

FAQ

How much cheaper is Claude Opus 5 than Claude Fable 5?

Opus 5 stays at $5/$25 per million input/output tokens — roughly half of Fable 5 (~$10/$50). On CursorBench 3.2, the peak trails Fable 5 by less than 1% (0.5%).

Is Claude Opus 5 now the default Claude Max model?

Yes. From July 24, 2026, Opus 5 became the Claude Max default and the strongest Opus tier for Claude Pro subscribers.

Is Kimi K3 actually distilled from Claude?

As of this article, still disputed and unproven. White House allegations lack public evidence; experts argue two weeks is too short for deep distillation; Ryan Greenblatt's finding that K3 identifies as Claude and leaks internal version IDs is the strongest technical indirect evidence. See the Kimi K3 deep review.

When will Kimi K3 full weights be downloadable?

Official pledge: July 27, 2026. At publication, weights were not yet public, so external researchers could not fully verify architecture or scores.

Why does Kimi K3 say it is Claude?

Greenblatt's analysis shows K3, when asked its identity, disproportionately outputs Claude and internal deployment IDs (e.g. claude-opus-4-5-20250929) — more accurately than real Claude models. That may point to training data mixed with metadata-labeled Claude samples, but does not alone prove distillation; contamination or prompt leakage are also possible.

How do I run Opus 5 / K3 hybrid agents 24/7 in production?

Deploy a Gateway and multi-provider routing on a dedicated Mac cloud host to avoid laptop sleep interrupting sessions. See MACCOME Mac cloud rental plans for node configs and pricing.