Why Agent Observability Cannot Replace Evaluation
Two traces
Here is the shape of the problem, in the only form that makes it obvious. A support agent handles refund requests. Two runs, side by side, as your observability platform renders them.
// Run A{ "trace_id": "a41f…", "gen_ai.operation.name": "agent_run", "gen_ai.request.model": "claude-sonnet-4-6", "duration_ms": 3180, "gen_ai.usage.input_tokens": 2847, "gen_ai.usage.output_tokens": 193, "gen_ai.response.finish_reasons": ["stop"], "spans": [ { "name": "lookup_customer", "status": "OK", "duration_ms": 112 }, { "name": "list_orders", "status": "OK", "duration_ms": 340 }, { "name": "issue_refund", "status": "OK", "duration_ms": 908 } ], "error_count": 0 }// Run B{ "trace_id": "b73c…", "gen_ai.operation.name": "agent_run", "gen_ai.request.model": "claude-sonnet-4-6", "duration_ms": 3204, "gen_ai.usage.input_tokens": 2851, "gen_ai.usage.output_tokens": 188, "gen_ai.response.finish_reasons": ["stop"], "spans": [ { "name": "lookup_customer", "status": "OK", "duration_ms": 118 }, { "name": "list_orders", "status": "OK", "duration_ms": 336 }, { "name": "issue_refund", "status": "OK", "duration_ms": 913 } ], "error_count": 0 }
Every field that your monitoring dashboard aggregates is, for practical purposes, identical. Both runs are three spans, all OK, no errors, ~3.2 seconds,...
Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE