Kimi K3・Claude Fable 5・GPT-5.6徹底比較:ベンチマークと価格で読むフロンティア三強
この6週間で、フロンティアモデルの「最強」が3回入れ替わりました。6月9日にClaude Fable 5、6月26日にGPT-5.6 Sol、そして7月16日にKimi K3。私事ですが、第38回まで3回にわたって書いた自社エージェントアプリValuScopeはモデルを固定して運用しており、この世代交代で「どのエージェントをどのモデルに載せ替えるか」という宿題を突きつけられました。本稿はその宿題を片付けるために、独立系評価サイトArtificial Analysisのデータを軸に三強を比較したノートです。
第10回でGPT-5.5とClaude Mythosを比較したとき、私は「当面は米国2強の綱引きが続く」と書きました。7ヶ月経って、この見立ては半分外れた。中国Moonshot AIのオープンウェイトモデルが、総合指標で2強の背中に手をかけるところまで来たからです。
本稿は2本立ての第1弾(性能編)です。第2弾(第40回)では、この3モデルを生んだ3社の創業背景と、創業者たちの研究系譜を掘ります。
バージョン注記: 数値は2026年7月23日時点。Artificial Analysis Intelligence Index v4.1、各社公式発表、報道ベースで確認しています。この分野は数週間で序列が入れ替わるため、最新値は文末の参考資料から必ず確認してください。
| テーマ | 学べること |
|---|
| 総合序列 | AA Intelligence Index上位:Fable 5(60)・GPT-5.6 Sol(59)・Kimi K3(57)の現在地 |
| ベンチ深掘り | GDPval-AA v2・AA-Briefcase・Frontend Code Arenaで勝者が三つに割れる構造 |
| 価格構造 | $3/$15 vs $10/$50 vs $5/$30 — 「知能単価」で見ると序列が逆転する |
| 実装 | エージェント用途別のモデルルーティング基準とコード例・落とし穴 |
まず全体の地図から。正直に言うと、6月の時点でこの地図をひと月半で描き直すことになるとは思っていませんでした。
Artificial Analysis Intelligence Index(v4.1)は、推論・数学・コーディング・実務タスクなど複数の評価セットを合成した独立系の総合指標です。2026年7月23日時点の上位は次の通りで、以下にClaude Opus 4.8、GPT-5.5(xhigh)、Claude Sonnet 5、GLM-5.2が続きます。

| モデル | 開発元 | リリース | Index(最高設定) | API価格($/M、入力/出力) |
|---|
| Claude Fable 5 | Anthropic | 2026/6/9 | 60 | $10 / $50 |
| GPT-5.6 Sol | OpenAI | 6/26プレビュー→7/9 GA | 59 | $5 / $30 |
| Kimi K3 | Moonshot AI | 7/16 | 57(189モデル中4位) | $3 / $15 |
3点差の中に、性格の違いがぎっしり詰まっています。順位だけ見れば「米2強+中国の追走」ですが、Kimi K3は重みの公開を前提としたオープンウェイトモデルで、この形式のモデルがフロンティア総合指標の4位まで来たのは初めてのことです。AIニュース番組AI QUESTで東大出身のAI研究者・今井翔太氏が「Kimi K3の開発者はAI研究で世界一」と紹介していましたが、これはMoonshot創業者・楊植麟(Yang Zhilin)氏の研究被引用実績を指しています——この人物の話は第2弾で詳しくやります。
ちなみに同じ週、GoogleはGemini 3.5 Proの投入を延期し、SNS上ではGemini 4の「匂わせ」投稿が話題になりました。番組内ではこれを「焦りの表れ」と読んでいて、私も同感です。本稿では、リリース済みの三強に対象を絞ります。
総合Indexだけでモデルを選ぶのは、偏差値だけで採用面接を通すようなものだと私は思っています。
私がCVCに入った当時、投資チームにはDD(デューデリジェンス)の品質基準が統一されておらず、案件ごとに「見る軸」がバラバラでした。評価軸を先に固定してテンプレート化し、ようやく案件同士が比較可能になった。モデル評価もまったく同じです。「何を測るベンチか」を先に固定しないと、各社が都合のいい数字だけを並べるリリースノートに流されます。ここでは軸を三つに固定します。①実務タスクの再現度(GDPval-AA v2)、②長時間のエージェント遂行(AA-Briefcase)、③フロントエンド実装(Frontend Code Arena)。

| ベンチマーク | 測るもの | 1位 | 2位 | 3位 |
|---|
| GDPval-AA v2 | 44職種・9業種の実務タスク | Fable 5 Max(1,815) | GPT-5.6 Sol Max(1,747.8) | Kimi K3(1,687) |
| AA-Briefcase | 長時間ナレッジワークの遂行 | Fable 5 Max(1,587) | Kimi K3(1,527) | GPT-5.6 Sol Max(1,495) |
| Frontend Code Arena | フロントエンド実装の対戦評価 | Kimi K3(1,679) | Fable 5(1,631) | GPT-5.6 Sol(1,618) |
この表で私がいちばん驚いたのは、AA-BriefcaseでK3がGPT-5.6 Solを上回ったことです。長時間エージェントはOpenAIが作り込んできた領域で、そこでオープンウェイトモデルが2位に入るとは予想していませんでした。Frontend Code ArenaのK3首位も、Tom's Hardwareが「Fable 5超え」と報じた通りです。ただし——ここは正直に書きますが——対戦型Arenaでの50点差は、私の体感では「明確に勝った」ではなく「並んだ」程度です。
個別ベンチも並べておきます。SWE-Bench ProはFable 5が80.3%でトップ。Anthropicが「本番のソフトウェア開発に最も近い」と位置づける指標で、タスクが長く複雑になるほど自社のOpus 4.8との差が開くと説明されています。Terminal-Bench 2.1はOpenAI公称でSolが88.8%、Sol Ultraが91.9%と、Claude Mythos 5とGPT-5.5の88.0%を上回りました。ベンチごとに勝者が違う——これが2026年7月の実態です。
スコアの次は中身と値札です。私はここが今回の世代交代でいちばん面白い部分だと思っています。

Kimi K3は総パラメータ2.8兆の史上最大オープンウェイトMoE(Mixture of Experts)です。896個のエキスパートから毎トークン16個だけをアクティブ化し、KDAとAttention Residualsで計算効率を稼ぐ。コンテキストは1Mトークン、ネイティブマルチモーダル対応。ただしアクティブパラメータ数は非公開で、「2.8Tがそのまま毎回動く」わけではない点は割り引いて読む必要があります。
Claude Fable 5は、Opusの上に新設された「Mythosクラス」の第1号です。研究機関などに限定公開されるClaude Mythos 5と同一の重みを共有し、安全分類器が高リスクと判定したクエリをOpus 4.8へ再ルーティングする(発動は平均でセッションの5%未満)。コンテキスト1M、出力は最大128kトークン。
GPT-5.6はSol/Terra/Lunaの3階層構成です。フラッグシップのSolに加え、半額のTerra、高速・最安のLunaを同時提供する「価格の階段」を最初から設計に組み込んでいます。
| 項目 | Kimi K3 | Claude Fable 5 | GPT-5.6 Sol / Terra / Luna |
|---|
| アーキテクチャ | 2.8T MoE(896エキスパート/16アクティブ)+KDA | Mythosクラス(Mythos 5と同一重み+安全分類器) | 非公開・3階層構成 |
| コンテキスト/出力 | 1M / 非公開 | 1M / 128k | 公表値未確認 |
| 入力価格($/M) | $3(キャッシュヒット$0.30) | $10 | $5 / $2.5 / $1 |
| 出力価格($/M) | $15 | $50 | $30 / $15 / $6 |
| 重み | オープンウェイト(段階公開) | クローズド | クローズド |
「知能単価」で割り算をすると景色が変わります。Index 1ptあたりで見ると、Fable 5の首位(60)はSol(59)の約2倍、K3(57)の約3.3倍の出力単価で買っている計算になる。──いや、正確に言えば、この割り算自体が乱暴です。1Mコンテキストを使い切る長時間エージェントなら、キャッシュヒット$0.30のK3と$10のFable 5で入力コストは30倍以上開く。逆に1タスクの成功率が数%違うだけで、リトライ込みの実効コストは逆転しうる。単価の議論は「どのベンチの勝者が自分のワークロードに近いか」とセットでしか意味を持ちません。

もう一つ、the-decoderが指摘するように、K3の$3/$15は「激安中国AI」時代の終わりも告げています。K2.6世代までの中国勢は1桁安い価格が武器でしたが、K3は品質で上位に並んだ分、価格も米国勢の下限に寄せてきた。安さだけを理由に中国オープンモデルを選ぶ時代は、この四半期で終わったと私は見ています。
結論を先に言うと、私は「1モデルに全部任せる」のをやめました。
ValuScopeのエージェント群での使い分け方針(2026年7月版)はこうです。第38回で書いた通り、エージェントごとに求める性質が違うので、割り当てもベンチの軸に合わせています。
| タスク | 割り当て | 根拠 |
|---|
| 長時間のDD調査エージェント | Claude Fable 5 | AA-Briefcase 1位・SWE-Bench Pro 80.3% |
| フロントエンド実装・プロトタイプ | Kimi K3 | Frontend Code Arena 1位(1,679) |
| 汎用ドラフト・要約・分類 | GPT-5.6 Sol / Terra | Index 59を約1/3の価格で・階層で微調整可 |
| 大量バッチ処理 | GPT-5.6 Luna / K3キャッシュ | $1/$6の最安階層と$0.30キャッシュヒット |
ルーティング自体は薄い関数で十分です。OpenRouter経由なら三強を同一インターフェースで叩けるので、最小実装はこれだけになります。
// model-router.ts: route tasks across the July 2026 frontier trio via OpenRouter
const MODELS = {
frontend: 'moonshotai/kimi-k3', // Frontend Code Arena #1 (1,679)
longAgent: 'anthropic/claude-fable-5', // AA-Briefcase #1 (1,587)
general: 'openai/gpt-5.6-sol', // Index 59 at ~1/3 of Fable 5 cost
bulk: 'openai/gpt-5.6-luna', // $1/$6 for high-volume drafts
} as const
type Task = keyof typeof MODELS
export async function complete(task: Task, prompt: string) {
const res = await fetch('https://openrouter.ai/api/v1/chat/completions', {
method: 'POST',
headers: {
Authorization: `Bearer ${process.env.OPENROUTER_API_KEY}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({
model: MODELS[task],
messages: [{ role: 'user', content: prompt }],
// cap output tokens to keep cost predictable ($15-$50 per 1M output)
max_tokens: 4096,
}),
})
if (!res.ok) throw new Error(`router: ${res.status} for ${MODELS[task]}`)
const json = await res.json()
return json.choices[0].message.content as string
}
動かす前に、ハマりどころを4つ挙げておきます。数えたら、どれも「ベンチの数字と本番の間」に落ちている穴でした。
- ベンチ最適化の疑い: TechTimesはGPT-5.6のレビューで「ベンチマーク問題」を指摘しています。上位3モデルは主要ベンチへの最適化が進んでおり、1〜3pt差は実務の体感差とほぼ相関しません。自社タスクでの小規模evalを省略しないこと。
- K3の重み公開は段階的: 「オープンウェイト」と報じられていますが、7月23日時点で完全な重み公開はAPI提供に遅行しています。セルフホスト前提の計画は、公開時期とライセンス条件を確認してからにすべきです。
- Fable 5の分類器再ルーティング: セッションの5%未満とはいえ、Opus 4.8への差し替えが自動で起こります。evalの再現性が要る用途では、レスポンスのモデルIDをログで必ず確認してください。私はこれに気づかず、評価スコアのブレを半日疑いました。
- データレジデンシー: この使い分けが向かないケースを1つ挙げるなら、機密データを扱う規制業種です。K3のAPIは中国系ホスティングが基本で、社内規程によっては選択肢から外れます。2.8T MoEのセルフホストは複数ノードのGPUクラスタが前提で、個人や小規模チームの現実解ではありません。
最後に一枚で締めます。
| 観点 | 現時点の勝者 |
|---|
| 総合Index | Claude Fable 5(60) |
| コストパフォーマンス | GPT-5.6 Sol(59を約1/3の価格で) |
| フロントエンド実装 | Kimi K3(Arena 1位) |
| 長時間エージェント | Fable 5、次点でKimi K3 |
書き終えた今も、この使い分けが3ヶ月後に有効かは正直自信がありません。それくらい入れ替わりが速い。ただ「評価軸を先に固定してから数字を見る」という手順だけは、序列が何度入れ替わっても使い回せます。
第2弾(第40回)では、この3モデルを生んだMoonshot AI・Anthropic・OpenAIの創業背景と創業者の研究系譜を掘ります。Transformer-XLの筆頭著者がなぜ1Mコンテキストにこだわるのか——モデルの設計思想は、創業者の研究史から読めるからです。
次号の記事案
- 案1:Kimi K3セルフホストの現実解を計算する|2.8T MoEの推論インフラ要件とクラウド比較 — 重み公開後のK3を自前で動かす場合のGPUクラスタ構成・月額コスト・スループットを、API利用との損益分岐まで計算する。
- 案2:エージェントCLIハーネス三番勝負|Claude Code・Codex CLI・Kimi CLIで同一Issueを消化させる — モデルではなくハーネス側の差を、同一リポジトリの同一Issueで実測する。ツール呼び出し回数・所要時間・成功率を比較。
- 案3:自社タスク用ミニevalの作り方|ベンチに騙されないための30問設計 — 公開ベンチの点差が体感と相関しない問題に対し、自社ワークロードから評価セットを30問切り出す手順を実装付きで示す。
本文の数値・事実関係は、読者が確認できる以下の一次情報・報道に基づいています(2026年7月23日時点)。
本記事は情報提供を目的としたものであり、特定のサービス、銘柄、契約条件の推奨や投資助言ではありません。ベンチマークスコアと価格は執筆時点の公表値であり、変更される可能性があります。調査、翻訳、校正の一部に生成AIを利用していますが、最終的な内容はZYL0が確認しています。詳細は免責事項をご確認ください。
Kimi K3 vs Claude Fable 5 vs GPT-5.6: Benchmarks, Pricing, and When to Use Which
In the past six weeks, the title of "strongest frontier model" changed hands three times: Claude Fable 5 on June 9, GPT-5.6 Sol on June 26, and Kimi K3 on July 16. This hit home for me personally — the agent fleet in ValuScope, the in-house app I documented through No. 38, runs on pinned models, and this generational shift handed me homework: which agent moves to which model? This post is my working note for that homework, built around data from the independent evaluation site Artificial Analysis.
When I compared GPT-5.5 and Claude Mythos in No. 10, I wrote that "the tug-of-war between the two US leaders will continue for now." Seven months later, that call is half wrong. A Chinese open-weight model from Moonshot AI has climbed to within arm's reach of the two leaders on the aggregate index.
This is part one of a two-part series (the performance piece). Part two (No. 40) digs into the three labs behind these models — their founding stories and the research lineages of their founders.
Version note: All figures are as of July 23, 2026, verified against Artificial Analysis Intelligence Index v4.1, official announcements, and press reports. Rankings in this field flip within weeks — always check the latest numbers via the references at the end.
| Topic | What you'll learn |
|---|
| Overall standings | The AA Intelligence Index top tier: Fable 5 (60), GPT-5.6 Sol (59), Kimi K3 (57) |
| Benchmark deep dive | How the winner splits three ways across GDPval-AA v2, AA-Briefcase, and Frontend Code Arena |
| Pricing structure | $3/$15 vs $10/$50 vs $5/$30 — the ranking inverts on "cost per intelligence point" |
| Implementation | A task-routing policy with code, and the pitfalls between benchmark and production |
Start with the map. Honestly, back in June I did not expect to be redrawing it within a month and a half.
The Artificial Analysis Intelligence Index (v4.1) is an independent composite of multiple evaluation sets spanning reasoning, math, coding, and real-world tasks. As of July 23, 2026, the top tier looks like this, followed by Claude Opus 4.8, GPT-5.5 (xhigh), Claude Sonnet 5, and GLM-5.2.

| Model | Lab | Release | Index (max setting) | API price ($/M, in/out) |
|---|
| Claude Fable 5 | Anthropic | Jun 9, 2026 | 60 | $10 / $50 |
| GPT-5.6 Sol | OpenAI | Jun 26 preview → Jul 9 GA | 59 | $5 / $30 |
| Kimi K3 | Moonshot AI | Jul 16 | 57 (#4 of 189 models) | $3 / $15 |
A lot of character hides inside those three points. Read only the ranking and it is "two US leaders plus a Chinese chaser" — but Kimi K3 is an open-weight model, and no open-weight model has ever reached #4 on the frontier composite before. On the Japanese AI news show AI QUEST, AI researcher Shota Imai introduced K3 with the line that its developer is "number one in the world in AI research" — a reference to the citation record of Moonshot founder Yang Zhilin. His story is the heart of part two.
The same week, notably, Google delayed Gemini 3.5 Pro, while "Gemini 4 teaser" posts made the rounds on social media. The show read that as a sign of urgency, and I agree. This post stays focused on the three models that have actually shipped.
Choosing a model on the composite index alone is, to my mind, like running job interviews on test scores alone.
When I joined my CVC, the investment team had no unified diligence quality standard — every deal was examined on a different set of axes. Only after we fixed the evaluation axes first and templated them did deals become comparable at all. Model evaluation works exactly the same way. Unless you fix "what does this benchmark measure" up front, you drift wherever each lab's cherry-picked release notes push you. So let me fix three axes: (1) reproduction of real occupational work (GDPval-AA v2), (2) long-horizon agentic execution (AA-Briefcase), (3) frontend implementation (Frontend Code Arena).

| Benchmark | What it measures | 1st | 2nd | 3rd |
|---|
| GDPval-AA v2 | Real tasks across 44 occupations, 9 industries | Fable 5 Max (1,815) | GPT-5.6 Sol Max (1,747.8) | Kimi K3 (1,687) |
| AA-Briefcase | Long-horizon knowledge work | Fable 5 Max (1,587) | Kimi K3 (1,527) | GPT-5.6 Sol Max (1,495) |
| Frontend Code Arena | Head-to-head frontend implementation | Kimi K3 (1,679) | Fable 5 (1,631) | GPT-5.6 Sol (1,618) |
What surprised me most in this table is K3 beating GPT-5.6 Sol on AA-Briefcase. Long-horizon agents are the territory OpenAI has cultivated hardest, and I did not expect an open-weight model to take second place there. K3's first place in Frontend Code Arena matches what Tom's Hardware reported as "beating Fable 5." Though — let me be honest here — a 50-point gap in an arena-style eval feels, in my hands, less like "clearly won" and more like "pulled even."
The individual benchmarks are worth listing too. On SWE-Bench Pro, Fable 5 leads at 80.3% — the metric Anthropic positions as closest to production software work, with the gap over its own Opus 4.8 widening as tasks get longer and more complex. On Terminal-Bench 2.1, OpenAI reports 88.8% for Sol and 91.9% for Sol Ultra, above the 88.0% posted by both Claude Mythos 5 and GPT-5.5. A different winner on every benchmark — that is the actual state of July 2026.
After the scores come the internals and the price tags. To me this is the most interesting part of the generational shift.

Kimi K3 is the largest open-weight MoE (Mixture of Experts) ever shipped, at 2.8 trillion total parameters. It activates only 16 of 896 experts per token and leans on KDA and Attention Residuals for compute efficiency, with a 1M-token context window and native multimodality. Note that Moonshot has not published an active-parameter count — "2.8T runs on every token" is not how to read it.
Claude Fable 5 is the first model in the new "Mythos class" that Anthropic created above the Opus line. It shares identical weights with the restricted-access Claude Mythos 5, adding safety classifiers that reroute high-risk queries to Opus 4.8 (triggering in under 5% of sessions on average). Context is 1M tokens; output caps at 128k.
GPT-5.6 ships as a three-tier family: flagship Sol, half-price Terra, and the fastest, cheapest Luna — a "pricing staircase" designed in from day one.
| Item | Kimi K3 | Claude Fable 5 | GPT-5.6 Sol / Terra / Luna |
|---|
| Architecture | 2.8T MoE (896 experts / 16 active) + KDA | Mythos class (Mythos 5 weights + safety classifiers) | Undisclosed, three tiers |
| Context / output | 1M / undisclosed | 1M / 128k | Not officially confirmed |
| Input price ($/M) | $3 ($0.30 on cache hit) | $10 | $5 / $2.5 / $1 |
| Output price ($/M) | $15 | $50 | $30 / $15 / $6 |
| Weights | Open (staged release) | Closed | Closed |
Divide by intelligence and the picture flips. Per index point, Fable 5's crown (60) costs roughly 2× Sol (59) and about 3.3× K3 (57) on output price. — No, let me correct myself: that division is itself too crude. For a long-horizon agent saturating a 1M context, input costs diverge more than 30× between K3's $0.30 cache-hit rate and Fable 5's $10. Yet a few percentage points of task success rate can invert effective cost once retries are counted. Unit price only means something next to the question "which benchmark's winner looks like my workload."

One more thing, as The Decoder points out: K3's $3/$15 also announces the end of the "super-cheap Chinese AI" era. Through the K2.6 generation, Chinese labs competed on prices an order of magnitude lower; K3, having pulled level on quality, has moved its price up to the US floor. The era of choosing Chinese open models on cheapness alone ended this quarter, in my view.
My conclusion first: I stopped giving everything to one model.
Here is the allocation policy for ValuScope's agent fleet as of July 2026. As I wrote in No. 38, each agent demands different properties, so the allocation follows the benchmark axes.
| Task | Assignment | Rationale |
|---|
| Long-horizon diligence research agent | Claude Fable 5 | #1 on AA-Briefcase, 80.3% SWE-Bench Pro |
| Frontend implementation & prototypes | Kimi K3 | #1 on Frontend Code Arena (1,679) |
| General drafts, summaries, classification | GPT-5.6 Sol / Terra | Index 59 at ~1/3 the price, tier-tunable |
| High-volume batch processing | GPT-5.6 Luna / K3 cached | The $1/$6 floor tier and $0.30 cache hits |
The routing itself needs only a thin function. Through OpenRouter all three expose the same interface, so the minimal implementation is just this:
// model-router.ts: route tasks across the July 2026 frontier trio via OpenRouter
const MODELS = {
frontend: 'moonshotai/kimi-k3', // Frontend Code Arena #1 (1,679)
longAgent: 'anthropic/claude-fable-5', // AA-Briefcase #1 (1,587)
general: 'openai/gpt-5.6-sol', // Index 59 at ~1/3 of Fable 5 cost
bulk: 'openai/gpt-5.6-luna', // $1/$6 for high-volume drafts
} as const
type Task = keyof typeof MODELS
export async function complete(task: Task, prompt: string) {
const res = await fetch('https://openrouter.ai/api/v1/chat/completions', {
method: 'POST',
headers: {
Authorization: `Bearer ${process.env.OPENROUTER_API_KEY}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({
model: MODELS[task],
messages: [{ role: 'user', content: prompt }],
// cap output tokens to keep cost predictable ($15-$50 per 1M output)
max_tokens: 4096,
}),
})
if (!res.ok) throw new Error(`router: ${res.status} for ${MODELS[task]}`)
const json = await res.json()
return json.choices[0].message.content as string
}
Before you run it, here are the traps. I counted them — all four sit in the gap between benchmark numbers and production.
- Benchmark optimization suspicion: TechTimes flagged a "benchmark problem" in its GPT-5.6 review. The top three are heavily optimized against public benchmarks, and 1–3 point gaps barely correlate with felt differences in real work. Do not skip a small eval on your own tasks.
- K3's weights arrive in stages: K3 is reported as open-weight, but as of July 23 the full weight release trails the API launch. If your plan assumes self-hosting, confirm the release timing and license terms first.
- Fable 5's classifier rerouting: Under 5% of sessions, but swaps to Opus 4.8 happen automatically. If your use case needs eval reproducibility, always log and check the response's model ID. I missed this and spent half a day chasing phantom score variance.
- Data residency: If I name one case where this routing does not fit, it is regulated industries handling confidential data. K3's API runs primarily on Chinese-affiliated hosting, which internal policy may rule out — and self-hosting a 2.8T MoE presumes a multi-node GPU cluster, which is no realistic answer for individuals or small teams.
One table to close.
| Angle | Current winner |
|---|
| Composite index | Claude Fable 5 (60) |
| Cost performance | GPT-5.6 Sol (59 at ~1/3 the price) |
| Frontend implementation | Kimi K3 (#1 in the Arena) |
| Long-horizon agents | Fable 5, with Kimi K3 a close second |
Having finished writing, I am honestly not confident this routing will still hold in three months — the turnover is that fast. But the procedure of "fix the evaluation axes before looking at the numbers" survives any number of leaderboard rewrites.
Part two (No. 40) digs into the founding stories of Moonshot AI, Anthropic, and OpenAI, and the research lineages of their founders. Why does the first author of Transformer-XL obsess over 1M-token context? A model's design philosophy is legible in its founder's research history.
Next Issue Ideas
- Idea 1: The Realistic Cost of Self-Hosting Kimi K3 — Inference Infrastructure for a 2.8T MoE vs the API — Once weights land, compute the GPU cluster configuration, monthly cost, and throughput of running K3 yourself, down to the break-even point against API usage.
- Idea 2: A Three-Way Agent CLI Harness Match — Claude Code vs Codex CLI vs Kimi CLI on the Same Issues — Measure the harness-side differences, not the model, on identical issues in an identical repository: tool-call counts, wall time, success rates.
- Idea 3: Building a Private 30-Question Mini-Eval So Benchmarks Cannot Fool You — Against the problem that public benchmark gaps fail to correlate with felt quality, a step-by-step, implementation-included guide to carving an evaluation set from your own workload.
The figures and facts in this post are anchored to the following sources readers can verify (as of July 23, 2026).
This article is for informational purposes only and does not constitute investment advice or a recommendation of any specific service, stock, or contract structure. Benchmark scores and prices are published values as of the time of writing and are subject to change. Generative AI was used for parts of research, translation, and proofreading, with final review by ZYL0. See the disclaimer for details.