Evaluating AI Agents: A production blueprint with Strands and AgentCore | Amazon Web Services
This post was co-written with Motorway and the AWS Prototyping and AI Customer Engineering (PACE) team.
Motorway, a UK-based online car marketplace, runs a daily auction where up to 8,000 dealers bid on up to 2,500 vehicles. Motorway worked with AWS Prototyping and AI Customer Engineering (PACE) to build an AI-powered dealer stock search agent that transforms how dealers find vehicles, replacing hours of manual filtering with natural language queries.
The challenge
The agent gives a confident-sounding response, but how do you prove it works reliably with real money on the line?
- Tool selection errors cause wrong search results, eroding dealer trust.
- Semantic search misinterpretations return irrelevant results. A query like “Petrol, Hybrid and electric cars up to 5 years old” requires the agent to correctly parse multiple constraints.
- Context drift in multi-turn conversations loses dealer refinements.
- Non-deterministic outputs make single-trial testing unreliable.
The solution
Together, Motorway and AWS built...
Copyright of this story solely belongs to amazon.com. To see the full text click HERE