The AI that's best at ethics is worst at making money, and that tells you everything about where this technology is actually headed.
The Summary
- Startup Obside ran a World Cup betting experiment where seven major AI models (ChatGPT, Gemini, Claude, Grok, Mistral, DeepSeek, Kimi) each got $10,000 virtual bankrolls to bet on matches using live Polymarket odds
- French open-source model Mistral leads the field, while Anthropic's Claude Opus 4.8 is the only model losing money
- The experiment measures something standardized benchmarks miss: judgment under uncertainty with real-world constraints
The Signal
Obside's experiment flips the script on how we measure AI capability. Instead of yet another multiple-choice test or coding challenge, they gave models a simple task: research publicly available information about World Cup teams one hour before kickoff, then decide how much to bet on Polymarket. Virtual money, real odds, actual outcomes nobody knows in advance.
After the semifinals, Mistral (the open-source French model) sat at the top of the leaderboard. OpenAI's GPT 5.5 and DeepSeek's V4 followed. At the bottom: Claude Opus 4.8, the only model in the red. The irony is sharp. Anthropic built Claude to be the "Constitutional AI," the model that refuses to help you make a bomb or write a scam email. Apparently, it also refuses to be good at gambling.
"Maybe Anthropic's AI is simply too ethical to be a gambler?"
This isn't about soccer. It's about what happens when AI agents start making decisions with actual consequences:
- They need to weigh incomplete information against time pressure
- They need to decide how much confidence to assign to uncertain outcomes
- They need to manage risk across multiple decisions, not just optimize one answer
Last year, ChatGPT entered a secret forecasting tournament run by economists and performed no better than the average human. The World Cup experiment hits the same nerve. These models can ace the SAT, write code, summarize documents. But ask them to make a judgment call about something that hasn't happened yet? They're barely better than flipping a coin.
Here's the interesting part: open-source models are winning. Mistral leads the pack, DeepSeek is in third. The big proprietary labs (OpenAI, Anthropic, Google) aren't dominating. That matters because betting on uncertain outcomes is basically what every business decision is. Should we enter this market? Hire this person? Build this feature? If open-source models can match or beat closed ones at real-world judgment calls, the moat around frontier labs gets narrower.
The Implication
Watch what happens when these betting experiments scale. If AI agents are going to negotiate contracts, allocate capital, or decide which projects to fund, we need better tests than "can it pass a medical licensing exam." Obside's World Cup benchmark is small, but it points at something bigger: the models optimized for safety and alignment might be systematically worse at the messy, probabilistic decisions that define most valuable work.
The agents that win in Web4 won't be the most careful ones. They'll be the ones willing to take calculated risks with incomplete information. That's a different optimization target than "never say anything harmful."