The easiest way to measure an AI's real capability isn't to ask what it can do in theory — it's to hand it someone's actual job and watch what breaks.
The Summary
- Platformer writer spent six months testing if Claude Fable 5 could replace their editor Casey, following up on an earlier experiment automating their own role — full experiment details here
- The test revealed specific limits in AI's ability to handle editorial judgment, context retention across long timelines, and the soft skills that make management work
- Key insight: AI can handle defined tasks well, but struggles with the connective tissue that turns a collection of tasks into actual leadership
The Signal
The Platformer experiment gave Claude Fable 5 the full scope of an editor's responsibilities: story assignment, deadline management, feedback loops, judgment calls on what's newsworthy versus what's noise. The six-month test built on an earlier round where the writer tried automating their own job, creating a before-and-after view of how AI performs against two different roles in the same organization.
The pattern that emerged matches what's happening across knowledge work right now. AI can absolutely crush the atomic tasks. It can draft feedback. It can track deadlines. It can even make decent calls on whether a story idea has legs. What it couldn't do was hold the through-line across weeks of conversation, remember the writer's previous missteps and growth areas, or read the room when a story needed to be killed for reasons that wouldn't show up in any brief.
"AI can handle defined tasks well, but struggles with the connective tissue that turns a collection of tasks into actual leadership."
The breakdown points weren't random. They clustered around three failure modes:
- Context decay: Claude would nail the first edit but forget key context by the third revision of the same piece
- Political blindness: It couldn't sense when a story would create internal friction or external blowback beyond what the words literally said
- Motivation misreads: It treated every piece of feedback as equally important, missing when a writer needed encouragement versus hard critique
This mirrors the broader agent economy pattern we're seeing. The companies winning with AI aren't replacing humans wholesale — they're finding the specific loops where AI maintains context well enough to close them without human checkpoints. Customer service queries that resolve in one session. Code reviews where the full context fits in the PR. Data analysis where the question and answer live in the same conversation.
Where it falls apart is exactly where this editing experiment fell apart: work that requires memory across long timelines, reading between the lines, and adjusting approach based on relationship history. The agent that can draft your performance review can't actually manage you, because management is 30% task execution and 70% held context about who you are and what you need to grow.
The Implication
If you're leading a team, this experiment is a blueprint for where to test AI and where to stay hands-on. Hand your agents the repeatable loops: first-pass edits, deadline reminders, formatting cleanup, research compilation. Keep the context-heavy relationship work yourself until the models prove they can remember Thursday's conversation by Tuesday.
The real question isn't whether AI can replace your boss. It's which 40% of your boss's calendar AI can clear so they can actually do the job only they can do. That's the unlock — not full replacement, but freeing the humans to operate in the spaces where context and relationships still matter more than speed.