The biggest AI labs are still fighting over whose chatbot is smarter, while Thinking Machines just released a model that can actually see and hear.
The Summary
- Thinking Machines released Inkling, a 975-billion-parameter open source model trained specifically for video and audio understanding
- First public model after 18 months building AI infrastructure out of view, positioning against OpenAI and Anthropic's generalist approach
- Open source release signals a strategic bet that specialized models beat jack-of-all-trades systems
The Signal
Thinking Machines just shipped its first receipts. After a year and a half operating in stealth mode, the lab dropped Inkling, a 975-billion-parameter model built exclusively for multimodal understanding. Not text-to-text with vision bolted on. Not a chatbot that can sort of handle images if you ask nicely. A model trained from the ground up to process video and audio as native inputs.
The parameter count alone puts it in heavyweight territory. For context, GPT-4 is estimated around 1.7 trillion parameters. Inkling lands just below that, but the architecture matters more than the size. Where the big labs chase general intelligence, Thinking Machines is betting against one-size-fits-all AI.
"Specialized models beat generalists when the task demands deep sensory understanding."
The open source release is the real signal. Anthropic and OpenAI guard their weights like nuclear codes. Thinking Machines is handing theirs over. That's either confidence or desperation, and the technical specs suggest the former. A lab doesn't spend 18 months building infrastructure in silence just to dump a mediocre model on GitHub for goodwill points.
Here's what matters for anyone building in the agent space:
- Video understanding unlocks physical world automation at scale
- Audio processing without text intermediation means faster, more natural agent interactions
- Open weights mean you can fine-tune for your specific domain without begging for API access
The timing is sharp. Every AI lab is racing toward agents that can operate autonomously. But agents need to perceive their environment, and text-only models hit a wall fast. You can't deploy an agent in a warehouse, a hospital, or a self-driving car if it can't process visual and audio streams in real time. Inkling positions Thinking Machines to compete not on conversational polish, but on sensory grounding.
The company stayed quiet for 18 months while competitors fought for headlines. Now they're entering the arena with a model that solves a different problem than everyone else is solving. That's either strategic clarity or a gamble that the market pivots toward specialized multimodal tools. The open source angle hedges both ways. If developers adopt it, Thinking Machines becomes infrastructure. If not, they still proved they can ship at scale.
The Implication
Watch what gets built on top of Inkling in the next six months. If robotics companies, surveillance platforms, and accessibility tools start forking it, that's validation that specialized beats general for real-world deployment. The open source bet also forces the closed labs to justify their walled gardens harder. Why pay OpenAI's API rates when you can run a 975-billion-parameter video model on your own metal?
For anyone building agents that need to operate in physical space, this is the first credible alternative to stitching together a dozen APIs from labs that didn't design for your use case. Download the weights, fine-tune for your domain, ship. That's the promise. Whether Thinking Machines delivers on performance will show up in benchmarks soon. But the infrastructure play is already live.