AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Thinking Machines Lab released full weights for its first foundation model, Inkling, under Apache 2.0 on July 15, 2026, with deployment-tool support available at launch. The lab concedes that Inkling is not the strongest available model, making its open-first release strategy, deployment economics and ownership model the larger development.

Thinking Machines Lab, founded by former OpenAI chief technology officer Mira Murati, released the full weights for its first foundation model, Inkling, on July 15 under Apache 2.0 before offering a closed API. The release gives organizations direct access to a large, natively multimodal model, although the lab acknowledges that it does not lead the field on every benchmark.

Inkling is a 975-billion-parameter Mixture-of-Experts model that activates 41 billion parameters for each token. According to the lab’s materials, it has a 1-million-token context window and was pretrained on 45 trillion tokens spanning text, images, audio and video. It accepts text, image and audio inputs and produces text.

The lab published BF16 and NVFP4 checkpoints on Hugging Face, accompanied by day-one support for transformers, vLLM, SGLang and llama.cpp, among other tools. Apache 2.0 generally permits modification and commercial use, giving developers more control over deployment and fine-tuning than a closed, hosted API provides.

The vendor reports strong results on AIME 2026, GPQA Diamond, MCP Atlas and VoiceBench, but weaker performance than selected rivals on several coding, agentic and reasoning tests. The supplied figures are vendor-published benchmarks, some involving a pre-release checkpoint, and have not yet been independently replicated. Thinking Machines Lab also says Inkling is not the strongest model currently available, whether open or closed.

At a glance
reportWhen: released July 15, 2026; benchmark and l…
The developmentThinking Machines Lab released Inkling’s full model weights before a closed API, pairing an Apache 2.0 license with immediate support across major deployment frameworks.

Ownership Moves Ahead of Rankings

The order of release changes the commercial proposition. Instead of making customers depend first on a provider-controlled service, Thinking Machines Lab is offering weights that organizations can retain, modify and host in their own environments. That may appeal to governments, regulated businesses and developers seeking greater control over availability, data handling and model changes.

Inkling also puts more attention on cost per useful result rather than a single peak score. Its reported 0.2-to-0.99 thinking-effort control lets operators trade reasoning tokens against latency and expense. On Terminal-Bench 2.1, the company says Inkling can match Nemotron 3 Ultra while using roughly one-third as many tokens; independent testing is still needed.

Amazon

AI model weights management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Open Access Still Has Limits

Open-weight releases have often followed commercial API launches or arrived as smaller alternatives to proprietary flagships. Inkling reverses that sequence by making complete flagship weights available on day one. Its Apache 2.0 license also differs from licenses that permit inspection but restrict modification or commercial deployment.

The release is not fully open source in the broadest sense. Thinking Machines Lab has not published the training dataset or full training pipeline. The model also requires substantial infrastructure: the reported BF16 configuration needs at least 2 terabytes of aggregate VRAM, while NVFP4 still needs about 600 gigabytes. That places the flagship beyond ordinary workstations and many smaller server fleets.

A preview of Inkling-Small, with 276 billion total and 12 billion active parameters, accompanied the flagship announcement. The lab says the smaller model matches or exceeds the larger version on several tests, but its full weights will follow after testing.

“Inkling is not the strongest model available today, closed or open.”

— Thinking Machines Lab’s launch announcement

Amazon

high-performance GPU for AI training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

License Scope and Scores Need Checking

Several parts of the release remain unresolved. A separate Model Acceptable Use Policy has been reported as applying to the parameters and modified versions, with restrictions involving surveillance, deception and automated decisions affecting rights. Its precise interaction with Apache 2.0 has not been verified in the supplied material, so prospective users will need to examine the current repository documents before deployment.

Inkling’s standing against GLM-5.2, Kimi K2.6 and proprietary systems also remains unsettled. The cited results cover different tasks and testing conditions, and some were reported through outside benchmark services. Claims about efficiency, multimodal performance and adversarial robustness require independent replication on production workloads.

Amazon

machine learning deployment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Independent Tests and Smaller Weights Follow

Developers and enterprise evaluators are expected to test Inkling on their own workloads, verify the governing use terms and measure infrastructure costs against hosted alternatives. Attention will also turn to the release of Inkling-Small’s full weights, which could make the lab’s open-first strategy accessible to a wider group of operators.

Amazon

multimodal AI model hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main development in the Inkling release?

Thinking Machines Lab published Inkling’s full weights before a closed API, under Apache 2.0 and with immediate deployment-tool support. The release order makes model ownership and operational control the central development.

Is Inkling the highest-performing AI model?

No. The lab says Inkling is not the strongest model, open or closed. Vendor figures show competitive results on some reasoning, audio and calibration tests, while other models lead several coding and agentic benchmarks.

Can developers run Inkling on a workstation?

Most cannot run the flagship locally. Reported requirements are at least 2 terabytes of aggregate VRAM for BF16 or about 600 gigabytes for NVFP4, putting it in data-center territory.

Does Apache 2.0 make Inkling fully open source?

Not by itself. The weights are available for inspection and modification, but the training data and full pipeline have not been published. A reported separate use policy also needs legal and operational review.

What could Inkling change in the AI market?

If the release model proves sustainable, it could push competition toward deployable weights, predictable operating costs and customer control, rather than closed API access alone. Its influence will depend on real-world performance and adoption.

Source: Thorsten Meyer AI

You May Also Like

The CNC Buying Guide Makers Wish They Had First

Learn essential tips to choose the right CNC machine, but beware—missing these steps can lead to costly mistakes and safety risks that you can’t ignore.

Industrial Exoskeletons Turn Workers Into Superheroes

Forces that transform workers into superheroes, industrial exoskeletons enhance strength and endurance—discover how these innovations are changing the future of work.

The Real Cost Of A Local-Inference Rig In 2026

Examining the hardware costs and constraints of running large language models locally in 2026, including VRAM limits and value strategies.

I think Anthropic and OpenAI have found product-market fit

Both Anthropic and OpenAI appear to have found product-market fit with enterprise coding and general-purpose AI tools, signaling a shift in revenue and growth prospects.