Should we add Thinking Machines Lab now that Inkling is out?

#47 · closed · 1 comments

View on GitHub ↗

hammer

## Context #14 evaluated Thinking Machines Lab as Tier 1 / "Watch closely — pedigree and funding are extraordinary but no model yet." That condition resolved on **July 15, 2026**: they shipped **Inkling**, a from-scratch open-weights flagship. This issue is the add/don't-add evaluation. ## The lab **Thinking Machines Lab** ([thinkingmachines.ai](https://thinkingmachines.ai), HF [thinkingmachines](https://huggingface.co/thinkingmachines), GitHub [thinking-machines-lab](https://github.com/thinking-machines-lab)) — founded February 2025 by **Mira Murati** (ex-OpenAI CTO), with **John Schulman** (co-founder, ex-OpenAI RL lead) and **Barret Zoph** (CTO). ~140–169 employees. Raised the largest seed in history — **$2B at $12B post** (July 2025; NVIDIA, Accel, ServiceNow, Cisco, AMD, Jane Street — [TechCrunch](https://techcrunch.com/2025/07/15/mira-muratis-thinking-machines-lab-is-worth-12b-in-seed-round/)). A November 2025 raise at **$50–60B collapsed by January 2026** ([Bloomberg](https://www.bloomberg.com/news/articles/2025-11-13/murati-s-thinking-machines-in-funding-talks-at-50-billion-value)), accompanied by researcher departures back to OpenAI. Products before Inkling: **Tinker** (fine-tuning API, Oct 2025; tinker-cookbook is 3.7k★) and the influential Connectionism research blog (batch-invariant ops 1k★, LoRA, on-policy distillation, modular manifolds). ## The trigger: Inkling ([announcement](https://thinkingmachines.ai/news/introducing-inkling/), July 15) - **975B total / 41B active MoE**, from scratch: 256 routed experts per MoE layer (6 active + shared experts), sliding-window:global attention interleaved 5:1, **1M context** - **Natively multimodal in one stack** — 45T tokens of text, image, audio, video; images as 40×40 pixel patches, audio as dMel spectrograms, *no separate encoders* - Post-training: SFT bootstrapped from synthetic data off open-weights models (explicitly incl. Kimi K2.5), then **>30M RL rollouts**; **controllable thinking effort** (comparable scores at ~⅓ the tokens; AA measures 25K avg output tokens/task vs 37–43K for DeepSeek v4 Pro / Kimi K2.6 / GLM-5.2) - Benchmarks (effort 0.99): HLE 46.0 w/ tools, AIME 97.1, GPQA 87.2, SWE-Bench 77.6, Terminal-Bench 63.8 - **AAII 41 (AA v4.1, verified live on [the model page](https://artificialanalysis.ai/models/inkling))** — [AA's writeup](https://artificialanalysis.ai/articles/thinking-machines-has-released-inkling-the-new-leading-u-s-open-weights-model) crowns it **the leading U.S. open-weights model**: above Nemotron 3 Ultra (38), Gemma 4 31B (29), gpt-oss-120b (24); GDPval-AA Elo 1238 above Kimi K2.6 (1190) - **Apache 2.0**, weights on HF (+ NVFP4 quant), served by Together/Fireworks/Modal/Databricks/Baseten, fine-tunable on Tinker day one - Companion model: **Inkling-Small** (276B total / 12B active), same post-training stack ## Case for adding 1. **The #14 watch condition is met, decisively** — not a preview or a derivative, but a from-scratch ~1T-parameter natively-multimodal flagship with a coherent thesis (open weights + customization via Tinker, against one-size-fits-all frontier APIs). 2. **Day-one AA coverage with a top-tier score** — unlike our recent adds (Xiaohongshu #44, SOOFI #45, KRAFTON #46, all AA-absent), Inkling arrives with a verified AAII 41 v4.1 and the "leading U.S. open weights model" title. The index is arguably incomplete every day we don't track this. 3. **Pedigree, capital, and infrastructure** ($12B valuation, billions in NVIDIA/Google compute commitments) make continued output likely despite the turbulence. 4. **Implementation is trivial**: `region: usa` exists, AA links exist, no UI changes needed. ## Case against / caveats 1. One model family, one day old — but the Tinker product line and research blog predate it, so this isn't a single-artifact lab. 2. **Organizational turbulence is real**: the collapsed $50–60B raise, the January 2026 researcher exodus back to OpenAI. Worth recording as news, not a reason to skip. 3. Their own positioning concedes Inkling is "not the strongest overall model" — it optimizes for customization economics. No bearing on trackability. I don't see a serious case against. This is the most clear-cut add since xAI. ## How we'd track it (sketch) - `data/labs/thinking-machines.yaml` — name "Thinking Machines", `type: private`, `region: usa`, `founded: "2025-02"`, `valuation: {amount: "$12B", type: post-money, date: "2025-07"}`; news: seed round (Jul 2025), collapsed raise (Jan 2026), Inkling launch (Jul 15, 2026) - Outputs: `inkling` (flagship model — `parameters: 975B`, `active_parameters: 41B`, `intelligence_index: 41` / `"AA v4.1"`, Apache 2.0, variants: Inkling-Small 276B/12B, NVFP4; TR: none yet — announcement post only, flag for the tech-report backfill list), `tinker` (product/tool, Oct 2025), and optionally a `connectionism` research-blog entry for the batch-invariant-ops / LoRA / on-policy-distillation line - People: Mira Murati, John Schulman, Barret Zoph, Lilian Weng - Top-level AAII anchors to the AA `inkling` page per our anchoring rule (the score is measured at xhigh effort — note it in the entry) ## Open questions 1. File the Connectionism research posts as tracked outputs or leave them as description color? (Lean: one grouped entry — the blog is unusually load-bearing for their reputation.) 2. Inkling-Small as a variant of `inkling` (lean) or its own entry? 🤖 Research and drafting with agent assistance.

Comments

hammer

Shipped in f8698a6 and deployed to labindex.ai: lab profile ([/labs/thinking-machines](https://labindex.ai/labs/thinking-machines)) + 3 outputs (inkling flagship w/ AAII 41 v4.1, tinker, connectionism), 5 people, 4 news items. Two deltas from the sketch, found during implementation: (1) leadership updated — Zoph/Metz/Schoenholz returned to OpenAI in Jan 2026, Soumith Chintala is CTO (Zoph recorded as Former); (2) OpenRouter links omitted — Inkling isn't in their model list yet. Next-sweep backfill: tech report, OpenRouter, AAOI, Inkling-Small HF weights.