Evaluating AI Agents: A production blueprint with Strands and AgentCore | Amazon Web Services

https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/07/23/ml-20515.png

This post was co-written with Motorway and the AWS Prototyping and AI Customer Engineering (PACE) team.


Motorway, a UK-based online car marketplace, runs a daily auction where up to 8,000 dealers bid on up to 2,500 vehicles. Motorway worked with AWS Prototyping and AI Customer Engineering (PACE) to build an AI-powered dealer stock search agent that transforms how dealers find vehicles, replacing hours of manual filtering with natural language queries.

The challenge

The agent gives a confident-sounding response, but how do you prove it works reliably with real money on the line?

  • Tool selection errors cause wrong search results, eroding dealer trust.
  • Semantic search misinterpretations return irrelevant results. A query like “Petrol, Hybrid and electric cars up to 5 years old” requires the agent to correctly parse multiple constraints.
  • Context drift in multi-turn conversations loses dealer refinements.
  • Non-deterministic outputs make single-trial testing unreliable.

The solution

Together, Motorway and AWS built...

Copyright of this story solely belongs to amazon.com. To see the full text click HERE

Read more