The models can write the paper, but they can't tell you which paper is worth writing.

The Summary

The Signal

GPT-6 Astra just became the highest-rated AI model on the Epoch Capabilities Index, but an OpenAI researcher is telling you what that score doesn't measure: the ability to know what's worth measuring. The model can execute on almost any research task you give it. It just can't tell you which tasks matter.

This "research taste" problem isn't about technical capability. It's about judgment. A researcher with taste sees a field and knows which questions will crack it open. They sense dead ends before wasting months on them. They recognize when a seemingly boring problem is actually the key that unlocks everything else.

"AI's struggle with research taste highlights the need for human intuition in guiding meaningful scientific exploration and innovation."

Here's what this means for the agent economy everyone's building toward. We're automating execution at scale, but the hardest part of knowledge work isn't execution. It's deciding what to execute on. An AI agent can write a thousand research proposals overnight. But it can't tell you which one will still matter in five years.

OpenAI's GPT-6 Astra launch emphasized safety and competitive urgency, positioning the model as a milestone in AI development. But the capabilities gap in specialized domains like software engineering suggests the model's general intelligence hasn't yet translated to domain mastery in every vertical.

Key gaps that remain:

  • Research prioritization and problem selection
  • Long-term scientific judgment and taste
  • Specialized task performance in fields like software engineering

The implication cuts both ways. For companies building autonomous research agents, you're building a tool that needs a pilot. The agent can do the work, but someone still needs to point it at the right work. For researchers worried about being replaced, your competitive advantage just became clearer. It's not your ability to execute research. It's your ability to know which research is worth executing.

The Implication

Watch how AI research labs respond to this taste gap. The companies that figure out how to train or fine-tune models on research judgment, not just research execution, will unlock the next capability tier. That might mean new training approaches using historical impact data, or hybrid systems where AI generates options and humans rank them by intuition.

For knowledge workers, this is your wedge. The work that requires judgment about what's worth doing is the work that stays human the longest. Double down on developing taste in your field. The person who can look at ten AI-generated strategies and instantly spot the one that'll work is more valuable than the person who can execute all ten.

Sources

Crypto Briefing