TL;DR

Thinking Machines Lab released full weights for its first foundation model, Inkling, under Apache 2.0 on July 15, 2026, with deployment-tool support available at launch. The lab concedes that Inkling is not the strongest available model, making its open-first release strategy, deployment economics and ownership model the larger development.

Thinking Machines Lab, founded by former OpenAI chief technology officer Mira Murati, released the full weights for its first foundation model, Inkling, on July 15 under Apache 2.0 before offering a closed API. The release gives organizations direct access to a large, natively multimodal model, although the lab acknowledges that it does not lead the field on every benchmark.

Inkling is a 975-billion-parameter Mixture-of-Experts model that activates 41 billion parameters for each token. According to the lab’s materials, it has a 1-million-token context window and was pretrained on 45 trillion tokens spanning text, images, audio and video. It accepts text, image and audio inputs and produces text.

The lab published BF16 and NVFP4 checkpoints on Hugging Face, accompanied by day-one support for transformers, vLLM, SGLang and llama.cpp, among other tools. Apache 2.0 generally permits modification and commercial use, giving developers more control over deployment and fine-tuning than a closed, hosted API provides.

The vendor reports strong results on AIME 2026, GPQA Diamond, MCP Atlas and VoiceBench, but weaker performance than selected rivals on several coding, agentic and reasoning tests. The supplied figures are vendor-published benchmarks, some involving a pre-release checkpoint, and have not yet been independently replicated. Thinking Machines Lab also says Inkling is not the strongest model currently available, whether open or closed.

At a glance
reportWhen: released July 15, 2026; benchmark and l…
The developmentThinking Machines Lab released Inkling’s full model weights before a closed API, pairing an Apache 2.0 license with immediate support across major deployment frameworks.
AI Dispatch · Reality Check · 16 July 2026

The weights came first: what Inkling actually signals

Mira Murati’s lab shipped its first foundation model — and the model isn’t the story. The order of operations is: full weights, Apache 2.0, day one, before any closed API. Plus a rare concession — the lab says it’s not the strongest model available, open or closed.

975B / 41B
total / active · MoE
1M
context window
45T
pretrain tokens
T · I · A
text · image · audio in
Apache 2.0
the licence*
Licence over leaderboard — what’s actually open
Model weightsBF16 + NVFP4 checkpoints on Hugging Face — download, modify, commercialize, keep
Apache 2.0 licenceconfirmed on the model card & HF repo — the real thing, not a source-available lookalike
Day-0 toolingtransformers · vLLM · SGLang · llama.cpp · TokenSpeed · Unsloth
Training data / pipelinenot published — open weights ≠ open source. Industry norm, but say it plainly
Separate use policy?reported: a Model Acceptable Use Policy over parameters & modified versions, barring surveillance, deception & fully automated decisions affecting rights
Unverified — check the model card yourself. If it reads as reported, Apache 2.0 isn’t the whole legal picture, and for ISR / geospatial / public-safety builders that clause is a go/no-go, not a footnote.
▲ Where it’s strong
  • AIME 2026 97.1%
  • GPQA Diamond 87.2%
  • MCP Atlas (Nemotron 44.7%) 74.1%
  • VoiceBench · open-weight audio frontier 91.4%
  • FORTRESS adversarial · best open 78.0%
  • ForecastBench · calibration 61.1
▼ Where it’s behind
  • HLE text-only (GLM-5.2 40.1%) 29.7%
  • SWE-bench Pro (GLM-5.2 62.1%) 54.3%
  • Terminal-Bench 2.1 (GLM-5.2 82.7%) 63.8%
  • SWE-bench Verified (Fable 5 95.0%) 77.6%
  • Design Arena · 2nd open, behind GLM-5.2 ~10th
◆ The dial nobody’s talking about — controllable thinking effort

A 0.2 → 0.99 effort setting trades reasoning tokens against cost & latency, so you get a curve, not a point. On Terminal-Bench 2.1 it reportedly matches Nemotron 3 Ultra at ~⅓ the tokens. Peak score is a vanity metric when you serve millions of calls; the cost curve is what ships. (Bonus: its chain of thought compressed on its own during RL — nobody rewarded it; efficiency did.)

0.2 · fast & cheap 0.99 · max effort
⚑ The China question — & the irony

Pitched as the Western alternative to Chinese open weights (censorship-resistance training is the differentiator). But GLM-5.2 still wins on agentic/reasoning and Kimi K2.6 often on multimodal: best American open model, second in the open field. The irony — post-training was bootstrapped on synthetic data from Kimi K2.5.

⚠ Open weights you probably can’t run

BF16 needs ≥2 TB aggregate VRAM (8× B300 / 16× H200). NVFP4 still needs ≥600 GB. Not a workstation model — a 512 GB fleet falls just short. “Open” ≠ “runnable.” Mitigations: 1-bit GGUFs (~74% acc.), hosted eval routes, and Inkling-Small (12B active) — the release local-first builders actually want.

The take

Open weights used to be a consolation prize. Inkling is a strategic open release — Apache 2.0, natively multimodal, honestly marketed, published complete on day one, optimized for deployment rather than headlines (the model isn’t the product; the fine-tuning platform is). It doesn’t need to win every benchmark for that to matter. The frontier is learning that owning the base beats renting the API — arriving now from the inside. For the sovereignty buyer: ① a real Western hedge against being switched off · ② verify the use policy before you build · ③ check the VRAM, then benchmark vs GLM-5.2 & Kimi K2.6 on your task.

Sources: Thinking Machines Lab (announcement, model card, HF repo, 15 Jul 2026); Hugging Face; VentureBeat, TechCrunch, BenchLM, LinkLoot, XenoSpectrum, NewsCord; Nathan Lambert via X. Benchmarks are vendor-published (some via Artificial Analysis) & await independent replication; some reflect a pre-release checkpoint. The AUP is reported, not verified here.
thorstenmeyerai.com

Ownership Moves Ahead of Rankings

The order of release changes the commercial proposition. Instead of making customers depend first on a provider-controlled service, Thinking Machines Lab is offering weights that organizations can retain, modify and host in their own environments. That may appeal to governments, regulated businesses and developers seeking greater control over availability, data handling and model changes.

Inkling also puts more attention on cost per useful result rather than a single peak score. Its reported 0.2-to-0.99 thinking-effort control lets operators trade reasoning tokens against latency and expense. On Terminal-Bench 2.1, the company says Inkling can match Nemotron 3 Ultra while using roughly one-third as many tokens; independent testing is still needed.

AI Systems Performance Engineering: Optimizing Model Training and Inference Workloads with GPUs, CUDA, and PyTorch

AI Systems Performance Engineering: Optimizing Model Training and Inference Workloads with GPUs, CUDA, and PyTorch

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Open Access Still Has Limits

Open-weight releases have often followed commercial API launches or arrived as smaller alternatives to proprietary flagships. Inkling reverses that sequence by making complete flagship weights available on day one. Its Apache 2.0 license also differs from licenses that permit inspection but restrict modification or commercial deployment.

The release is not fully open source in the broadest sense. Thinking Machines Lab has not published the training dataset or full training pipeline. The model also requires substantial infrastructure: the reported BF16 configuration needs at least 2 terabytes of aggregate VRAM, while NVFP4 still needs about 600 gigabytes. That places the flagship beyond ordinary workstations and many smaller server fleets.

A preview of Inkling-Small, with 276 billion total and 12 billion active parameters, accompanied the flagship announcement. The lab says the smaller model matches or exceeds the larger version on several tests, but its full weights will follow after testing.

“Inkling is not the strongest model available today, closed or open.”

— Thinking Machines Lab’s launch announcement

Visual Studio Code Guide for Beginners: Master Programming, Debugging, GitHub Integration, Extensions, AI Tools, Terminal, Deployment, and Professional Development Workflows from Scratch

Visual Studio Code Guide for Beginners: Master Programming, Debugging, GitHub Integration, Extensions, AI Tools, Terminal, Deployment, and Professional Development Workflows from Scratch

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

License Scope and Scores Need Checking

Several parts of the release remain unresolved. A separate Model Acceptable Use Policy has been reported as applying to the parameters and modified versions, with restrictions involving surveillance, deception and automated decisions affecting rights. Its precise interaction with Apache 2.0 has not been verified in the supplied material, so prospective users will need to examine the current repository documents before deployment.

Inkling’s standing against GLM-5.2, Kimi K2.6 and proprietary systems also remains unsettled. The cited results cover different tasks and testing conditions, and some were reported through outside benchmark services. Claims about efficiency, multimodal performance and adversarial robustness require independent replication on production workloads.

Fine-Tuning AI: Customizing Large Language Models

Fine-Tuning AI: Customizing Large Language Models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Independent Tests and Smaller Weights Follow

Developers and enterprise evaluators are expected to test Inkling on their own workloads, verify the governing use terms and measure infrastructure costs against hosted alternatives. Attention will also turn to the release of Inkling-Small’s full weights, which could make the lab’s open-first strategy accessible to a wider group of operators.

Multimodal AI Systems Engineering: Building Production Vision-Language Models, Document AI, and Cross-Modal Retrieval Pipelines (Production AI Engineering Series)

Multimodal AI Systems Engineering: Building Production Vision-Language Models, Document AI, and Cross-Modal Retrieval Pipelines (Production AI Engineering Series)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main development in the Inkling release?

Thinking Machines Lab published Inkling’s full weights before a closed API, under Apache 2.0 and with immediate deployment-tool support. The release order makes model ownership and operational control the central development.

Is Inkling the highest-performing AI model?

No. The lab says Inkling is not the strongest model, open or closed. Vendor figures show competitive results on some reasoning, audio and calibration tests, while other models lead several coding and agentic benchmarks.

Can developers run Inkling on a workstation?

Most cannot run the flagship locally. Reported requirements are at least 2 terabytes of aggregate VRAM for BF16 or about 600 gigabytes for NVFP4, putting it in data-center territory.

Does Apache 2.0 make Inkling fully open source?

Not by itself. The weights are available for inspection and modification, but the training data and full pipeline have not been published. A reported separate use policy also needs legal and operational review.

What could Inkling change in the AI market?

If the release model proves sustainable, it could push competition toward deployable weights, predictable operating costs and customer control, rather than closed API access alone. Its influence will depend on real-world performance and adoption.

Source: Thorsten Meyer AI

You May Also Like

Molecular Assemblers: The Next Industrial Revolution?

Nearing the dawn of a new era, molecular assemblers promise revolutionary manufacturing breakthroughs, but challenges remain that could change everything.

Access to frontier AI will soon be limited by economic and security constraints

Recent developments indicate that access to advanced AI models will soon be limited due to security, economic, and geopolitical factors, affecting global AI deployment.

The Electric Scooter Upgrade That Changes Daily Commuting

Harness the power of electric scooter upgrades that can revolutionize your daily commute—discover how these changes can make a difference.

The Unreasonable Difficulty Of Time Series Forecasting

Experts highlight the significant difficulties in reliably predicting time series data, raising concerns over current forecasting methods’ effectiveness.