Thinking Machines Lab spent roughly a year and a half building Inkling after Murati founded the company following her departure from OpenAI in late 2024. The release was not a soft beta — the weights landed on Hugging Face available for immediate download, alongside a developer fine-tuning tool called Tinker. The company said it does not plan to monetize the model directly; Tinker is the product.
Sources: TechCrunch reported on the release and Thinking Machines' enterprise strategy on July 15. Axios covered the launch the same day, including specs and Murati's reasoning for going open-weight. The Register's July 16 story framed the decision as a direct contrast to OpenAI's approach. The official launch post is on Thinking Machines' own site. See TechCrunch on Inkling, Axios on the launch, The Register's take, and the official Thinking Machines announcement.
What Inkling actually is
Inkling is a mixture-of-experts (MoE) system with 975 billion total parameters, though only about 41 billion are active on any given inference. That architecture keeps per-token costs competitive with much smaller dense models while preserving raw capacity in the same tier as the latest frontier systems. It was trained on 45 trillion tokens of text, images, audio, and video — making it natively multimodal rather than having modalities retrofitted after pretraining.
The open-weight release means developers can download the model weights, run Inkling on their own hardware, and modify it freely without asking Thinking Machines for permission or routing requests through an external API. That is a meaningful practical difference from the approach taken by OpenAI, Anthropic, and Google with their flagship models.
Why open-weight matters right now
The major closed labs have moved toward tighter access controls over the past twelve months. Limited previews, government-coordinated rollouts, and opaque safety reviews have become routine for frontier model releases. Inkling is a direct bet that a meaningful segment of enterprise and developer demand is not served by that model — and that the costs of compute and R&D can be recovered by helping companies fine-tune and host their own versions instead.
Fine-tuning as the actual product
Thinking Machines was explicit that Tinker, its fine-tuning developer platform, is how it plans to generate revenue. Enterprise teams can adapt Inkling to their own data, terminology, and internal workflows without routing every prompt through an external vendor. For regulated industries — finance, healthcare, legal — that level of control matters more than raw benchmark scores or the name on the model card.
The compute calculation
An MoE architecture with 41 billion active parameters keeps inference costs far below what a dense 975B model would require. That makes self-hosting financially realistic for large enterprises, which was not true of earlier open-weight models at anything approaching frontier scale. The tradeoff is that routing between experts adds complexity — how well Inkling handles real-world edge cases relative to closed models will become clearer as the developer community runs it at scale over the coming weeks.
What to watch next
DeepSeek V4 is expected to graduate from preview to stable release on July 24 — also an open-weight model — making next week potentially the most active stretch for open-weight AI since Meta released Llama 3. Moonshot AI's Kimi K3 topped coding benchmarks within hours of its own launch earlier this week. The pattern is worth noting: the open-weight camp has had the most momentum in a cycle of news otherwise dominated by closed-lab delays and restricted access.
For Thinking Machines specifically, the key question is whether Inkling's real-world coding and reasoning performance matches what its architecture promises. That answer will come from the developer community running it in production, not from internal evaluations — and it will arrive fast.
Bottom line
Inkling is the most significant open-weight AI release of the year so far, from one of the most credible figures in the field. It is a direct argument that enterprises want models they can own and adapt, not just rent — and that the person who knew best how closed frontier models get built decided to build something different. Whether the bet pays off depends on how Tinker lands with enterprise developers, but the release itself changes what "frontier open-weight" means and is worth tracking regardless of which models you currently use.
LiveCue
Useful AI still has to work in the meeting you have today
Try LiveCue on Mac for real-time meeting prompts, follow-up questions, and private on-screen help.