Five people used AI agents for a week and got exactly what you'd expect from version 0.1 software: occasional magic, frequent confusion, and one guy who still has to call his mom.

The Summary

  • Business Insider staffers tested Muse (Meta) and Instinct across five use cases, from finding library books to tracking K-pop merch to identifying mystery credit card charges
  • Instinct successfully researched bike parts, created travel itineraries, and recovered an $80 forgotten subscription — then failed spectacularly at a simple car fee payment
  • The real finding: personal AI agents work best as research assistants and task reminders, but break down when actual transactions or complex multi-step workflows are involved

The Signal

Business Insider's newsroom experiment with Muse and Instinct reveals the current state of personal AI agents: impressive parlor tricks wrapped around a fragile core. When editor Will Martin asked Instinct to recommend a torque wrench, it delivered four options with pros, cons, and purchase links. When the editor-in-chief fed a cryptic credit card charge to both ChatGPT and Muse, only Muse decoded it. These aren't trivial wins. They're proof that agents can actually save you the kind of mental overhead that accumulates into decision fatigue.

But here's where the honeymoon ends. Martin also asked Instinct to handle a car fee payment. It backfired in ways the article doesn't detail, but the pattern is clear. The moment an agent has to navigate actual commerce, deal with authentication, or complete a transaction that requires more than research, the reliability drops off a cliff.

"Instinct secured me a refund on an $80 subscription I forgot to cancel after the free trial."

The gap between "find me options" and "complete this transaction" is where most agent promises currently die. Instinct could identify the forgotten subscription and guide Martin through cancellation, but it couldn't autonomously execute the refund. That's not a bug. That's the entire unsolved problem of Web4: giving agents enough access to act without giving them enough access to wreck your life.

Key patterns from the five-person test:

  • Agents excel at research, comparison, and reminder tasks (low-stakes, high-value)
  • Agents struggle with purchases, payments, and anything requiring verified identity
  • The "whoa moments" cluster around information retrieval; the "yikes moments" cluster around action execution

The other detail worth flagging: use cases ranged from library book reservations to K-pop merchandise tracking to mother-calling reminders. That spread tells you something. We're still in the "people trying random stuff to see what sticks" phase. There's no established workflow yet. No consensus on what personal agents should actually do. Just a bunch of early adopters poking at the edges, hoping to stumble into utility.

The Implication

If you're building agent infrastructure or thinking about where to place bets, this experiment hands you a roadmap. The research layer works. The transaction layer doesn't. The companies that solve authenticated execution without requiring users to hand over the keys to their entire digital life will own the next phase.

For everyone else: use agents like interns, not employees. Let them research, compare, remind, and draft. Don't let them buy, pay, or send anything you can't undo. We're one software update away from these things being genuinely useful and about six updates away from trusting them with our credit cards.

Sources

Business Insider Tech