The benchmark that matters isn't writing code—it's whether an AI can read a breakthrough paper and rebuild the science from scratch.
The Summary
- Inherent, a UK AI lab founded by DeepMind alumni, released Faraday, an AI agent that replicates scientific research from papers, reportedly outperforming Anthropic and OpenAI models on research reproduction tasks
- The company positions Faraday not as an assistant but as a "teammate" that can independently execute multi-step scientific workflows
- This shifts AI capabilities from pattern matching to genuine knowledge work—the kind that advances fields rather than just summarizing them
The Signal
Inherent's claim is specific: Faraday can read a scientific paper and reproduce the experiments, code, and results without hand-holding. That's not ChatGPT summarizing abstracts. That's an agent parsing methodology sections, identifying dependencies, writing functional code, debugging when things break, and validating outputs against reported findings.
The benchmark they're using measures success by whether the agent can independently replicate results from machine learning research papers. According to Inherent, Faraday achieved higher replication rates than Anthropic's Claude and OpenAI's models on the same task set. The company didn't release exact numbers, but the framing matters: they're competing on scientific rigor, not conversational charm.
"Replication isn't about being right once—it's about understanding a system well enough to rebuild it under new conditions."
The DeepMind pedigree is doing work here. These aren't founders pitching vaporware from a hackathon. They spent years inside the lab that built AlphaFold and AlphaGo. They know what good science looks like, and they know the difference between a model that can pass tests and one that can generate new knowledge. Faraday is their bet that the latter is now buildable.
Here's why this matters beyond academic flex:
- Research replication is a bottleneck. Most published ML papers never get reproduced because it takes weeks of expert time.
- If an agent can do it reliably, labs can validate findings faster, spot errors earlier, and build on solid foundations instead of shaky ones.
- This is the infrastructure for accelerating science itself—not just automating grunt work.
The "teammate" framing is intentional. Inherent isn't selling a tool you prompt. They're selling an agent you assign tasks to and walk away from. That's the Web4 promise: agents that operate with autonomy, not just responsiveness. If Faraday can replicate a paper, it can probably extend it, test variations, and propose new experiments. That's when agents stop being assistants and start being collaborators.
The Implication
If agents can replicate research, they can generate it. The timeline from "AI that helps scientists" to "AI that is a scientist" just compressed. Labs racing to build AGI are also racing to build systems that don't need humans in the loop for knowledge creation. Inherent is making that explicit.
For researchers, this is a forcing function. The value shifts from executing experiments to designing them, from running code to asking better questions. For AI labs, it's a new competitive surface: not who has the best chatbot, but who has the best autonomous scientist. Watch how Anthropic and OpenAI respond.