The Hard Truth About AI in Sales
If you've been following AI lately, you've probably seen the demos: an AI agent browsing a shopping site, adding items to a cart, and checking out. It looks impressive. But behind the scenes, there's a whole other layer of AI work that most people never see—the grunt work of e-commerce sales operations. And according to a new benchmark called RealReplicaBench, AI agents are failing at it.
The benchmark put 13 leading AI models through 107 real-world business tasks, and not one of them scored a passing 60. The top scorer, Claude Opus 5, managed only 56.1. That's a wake-up call for anyone who thinks AI is ready to handle sales work end-to-end.
Why Traditional Benchmarks Don't Cut It
Traditional AI benchmarks are like multiple-choice tests. They reward partial progress. If you solve 80% of a problem, you get 80% of the points. That works fine for chatbots, where the goal is to produce a decent text answer. But sales work isn't a series of isolated questions. It's a chain of actions where each step depends on the last.
RealReplicaBench was built by the team behind Accio Work, an AI agent platform for e-commerce sellers. They realized that scoring AI on how many steps it gets right is meaningless if the final result isn't usable. A sales task isn't done until the next step can pick it up without human help.
What "Done" Actually Means in Sales
Think about a typical sales operation: finding a supplier, comparing quotes, creating a purchase order, updating inventory, arranging shipping, and confirming delivery. Each of those steps leaves a trace in some system—an email, a spreadsheet, a logistics dashboard. If an AI agent completes a task but leaves the last 20% for a human to finish, it hasn't really done the job.
RealReplicaBench has a strict definition of completion: a task is only complete if the output can flow directly into the next step without manual intervention. No partial credit. If the agent doesn't finish, it gets zero. That's harsh, but it mirrors what sales teams face every day.
Testing in a Simulated Real World
To test whether AI agents can handle real sales work, you can't just give them a text prompt. You need to recreate the messy environment where sales work happens: emails, browser interfaces, APIs, file systems, and backend states that change over time.
RealReplicaBench does exactly that. It builds a simulated e-commerce environment where agents have to navigate a realistic business workflow. For example, one task involves going through about 300 noisy emails to reconstruct a purchase request, then selecting suppliers, tagging emails, drafting replies, and setting up a calendar.
Another task has the agent process 5,383 customs records and turn them into a cross-system procurement control tower. That means aggregating data by category and supplier, filtering top suppliers based on policy, and then creating dashboards in Google Workspace, evidence folders in Box, and project tasks in Jira—all while keeping the relationships between those objects consistent.
Verifying the Work, Not the Words
One of the biggest problems with AI agents is that they can claim success without actually achieving it. That's why RealReplicaBench doesn't trust the agent's self-report. Instead, it uses a verifier that reads the final state of the environment—like checking if a real shipment ID was generated, not just if the agent said it would.
This approach keeps the focus on outcomes. In sales, that's what matters. A supplier selection is only valid if the purchase order reflects it. A shipping plan is only useful if it produces a booking confirmation. The benchmark checks these tangible results, not the thought process.
What This Means for Sales Teams
For sales operations, this benchmark is a reality check. AI agents aren't yet ready to handle complex, multi-step sales workflows without human oversight. But that doesn't mean they're useless. The gaps are specific and fixable. The issue isn't raw intelligence; it's the ability to maintain accuracy across a long chain of tasks and adapt when conditions change.
The Accio Work team plans to keep expanding the benchmark with new tasks and to use it for model routing—choosing the best model for each type of task. That's a practical approach for sales teams: use AI where it works, and keep humans in the loop where it doesn't.
The Bottom Line
AI in sales is moving from theory to practice, but the road is bumpy. RealReplicaBench shows that even the best models struggle with the nitty-gritty of real-world sales work. The good news is that this kind of benchmarking helps us understand exactly what needs to improve. For sales teams, the takeaway is clear: don't trust an AI agent to "almost" get the job done. Demand systems that can actually deliver the final result.
Comments (0)
Please sign in to post a comment.
Don't have an account? Create one
No comments yet. Be the first to comment!