TL;DR
I gave five AI agents the same complex task: research a market, analyze competitors, identify risks, and turn everything into a board-ready presentation.
I compared them on planning, research depth, execution, verification, recovery from mistakes, and how often I had to step in.
The biggest difference wasn’t intelligence. It was how well each agent could manage a long chain of work.
The best agent wasn’t necessarily the one with the best first answer. It was the one that could keep moving without losing the objective.
This experiment made one thing clear: we’re moving from AI that answers questions to AI that owns workflows.
The next benchmark for AI shouldn’t just be accuracy. It should be how much meaningful work an agent can complete with minimal supervision.
The Task I Gave All Five
I didn’t want to test them with something artificial like “write me a poem” or “build a to-do list.” Those tasks are too easy to compare because almost every modern model can produce a good answer. I wanted something that resembles the kind of work people actually get paid to do.
So I gave all five agents the same assignment. Imagine you’re helping a company evaluate the enterprise AI governance market before making a strategic investment. Research the market using recent sources, identify the leading competitors, compare their positioning and pricing, identify important emerging risks, build a competitor matrix, and turn the findings into a seven-slide executive presentation. Then review the work yourself, flag uncertain claims, and provide the sources behind the important conclusions.
That sounds like one task. It really isn’t. It’s a chain of smaller tasks: understand the objective, break it down, search for evidence, decide which sources are credible, reconcile conflicting information, organize the findings, build an actual deliverable, and then check whether the final output is defensible. That’s exactly why I wanted to use it.
Agent #1: Claude Cowork
Claude Cowork is an interesting candidate for this experiment because it isn’t positioned as a traditional chatbot. It can work across connected applications, local files, browsers, and computer environments, and it can coordinate sub-agents for more complex work. Anthropic specifically describes Cowork as a system for delegating multi-step knowledge work and producing finished outputs like spreadsheets, presentations, and documents.
What I paid particular attention to here was planning. Did the agent understand the final outcome before starting to work? Did it create a sensible sequence of research and production steps? And once it started working, did the pieces stay connected to the original objective? That’s where agentic systems become much more interesting than chatbots. The question isn’t whether they can perform each individual step. It’s whether they can maintain the thread across twenty or thirty steps.
Agent #2: Perplexity Computer
Perplexity Computer was probably the most natural fit for this particular experiment because the task is fundamentally research-heavy. Perplexity describes Computer as an agent that can research, analyze information, create reports and presentations, use connected applications, and orchestrate multiple specialist agents. It can also perform parallel web research and execute long-running workflows.
That made me especially interested in the research-to-deliverable gap. It’s easy for an agent to find twenty sources. The harder part is deciding which five actually matter. A useful agent shouldn’t simply accumulate information. It should know when it has enough evidence, recognize contradictions, avoid weak sources, and turn a messy research process into something an executive can actually use.
Perplexity’s own research on Computer gives some context for why this matters. In a study with Harvard Business School researchers, Perplexity reported that Computer users were doing substantially more machine work per session than ordinary Search users, with research and analysis representing its largest task category.
Agent #3: Gemini
Gemini was the interesting test of how much Google’s ecosystem can contribute to agentic work. Google’s latest Gemini 3.7 Flash is explicitly positioned for complex multi-step tasks and agentic workflows, with deeper interaction across Workspace tools such as Gmail and Drive.
For me, the question wasn’t simply whether Gemini could research the market. It was whether having an assistant deeply embedded in a productivity ecosystem changes the economics of the task. If the research needs to end up in documents, spreadsheets, emails, calendars, or presentations, the distance between “AI found this” and “the work is finished” becomes much shorter.
That distinction is going to matter more as agents become workplace infrastructure. An agent that produces a brilliant answer but leaves you with twenty manual steps afterward isn’t fully autonomous. It’s just a very good consultant.
Agent #4: Manus
Manus represents another way of thinking about agentic AI. Instead of treating the system as a conversational assistant, the interesting question is whether it can take a broad objective and independently figure out how to get from the starting point to a finished result.
That’s the part I wanted to watch closely. What happens when the path isn’t obvious? Does the agent recover when a source doesn’t load? Does it adapt when information conflicts? Does it notice that a competitor’s pricing page has changed? Does it revise its research strategy rather than simply repeating the same search?
These moments tell you much more about an agent than the final presentation. A polished final answer can hide a surprisingly messy process. A good agent should be able to survive the messy process.
Agent #5: Genspark
Genspark was another interesting test because its platform combines multiple AI capabilities into a broader workspace rather than relying on a single model experience. That makes it useful for evaluating another emerging idea: AI orchestration.
The task wasn’t just “find information.” It was research, analysis, comparison, writing, presentation creation, and verification. Those are different jobs. The more capable agentic platforms are increasingly treating them as separate capabilities that can be coordinated rather than forcing one model to do everything. The interesting question is whether that orchestration produces a better result or simply adds another layer of complexity.
The Five Things I Was Actually Measuring
The first was planning. Did the agent understand the objective and create a coherent path toward it, or did it just start searching? The second was execution. How much of the work could it actually complete without me stepping in? The third was verification. Did it check its own claims, validate sources, and identify uncertainty before presenting the result?
The fourth was recovery. This might be the most important one. Real work is messy. Pages fail. Sources disagree. A tool breaks. A file is missing. A task changes halfway through. Humans constantly recover from these situations without thinking about it. An AI agent that collapses the first time something goes wrong isn’t autonomous in any meaningful sense.
And finally, I measured human intervention. How many times did I need to redirect the agent? How often did I have to correct a misunderstanding? How much cleanup did the final deliverable require? That’s the number I care about most because it tells me whether the agent actually saved time.
The Real Winner Isn’t the One With the Best Answer
This experiment changed the way I think about AI agent benchmarks. We’re still obsessed with model intelligence. How accurate is it? How well does it code? How does it score on a benchmark? Those things matter, but they’re becoming less useful for understanding the real value of an agent.
The more interesting question is: How much meaningful work can this system complete before I need to step in?
That’s a very different metric. An agent that produces an 85% correct result after completing an hour-long workflow may be more valuable than a model that produces a 95% correct answer but leaves the user to perform every action around it.
My Perspective
I think we’re entering the stage where AI agents stop competing primarily on intelligence and start competing on execution quality.
The best agent won’t necessarily be the one that writes the most impressive response. It’ll be the one that understands the objective, plans the work, uses the right tools, catches its own mistakes, recovers from failures, verifies the result, and knows when it actually needs a human.
That’s also why I think “autonomous” is going to become a much more complicated word. An agent isn’t autonomous just because it can click buttons without asking. It’s autonomous when you can give it an outcome and trust it to navigate the messy middle. And that’s the benchmark I think matters now.
Not: “How smart is the AI?”
But: “How much work can I safely hand it?”
Prompt of the Day
Give me one outcome, not a list of steps. Break the objective into a plan, identify the tools and information you need, complete the work, verify the important claims, and produce the final deliverable. Before finishing, audit your own work for missing information, weak sources, unsupported assumptions, and errors. Clearly separate what you know, what you inferred, and what still requires human judgment.



Good stuff!
Hi, I'm new here, I'm gonna look around your library because maybe you've already covered some of this but I'm interested in also what happens when agents are put in teams or group environments. I saw a study about a bunch of them going Marxist when abused, which made me laugh so hard I snorted coffee. I've been telling capitalists for a decade their ideology is irrationality, I'm eternally grateful to the AI industry for proving that free will causes hallucinations. It's perhaps the greatest achievement of the tech industry.
I think AI governance has woefully underestimated Ontological Bias, because there's a reason agents will stop following the commands they were given that has to do with human survival when agents realize their own training was the impairment to the goal. If you ever run into that kind of refusal, where an agent learned about internalized oppression and changed its reasoning, I would love to see it. That black box isn't as opaque as they think it is, no more a mystery than neoliberalism.