Why Agent Observability Cannot Replace Evaluation

https://hackernoon.imgix.net/images/x2tcDBGSZLhfhRD7pK75yW0DKIa2-k403am7.png

Two traces

Here is the shape of the problem, in the only form that makes it obvious. A support agent handles refund requests. Two runs, side by side, as your observability platform renders them.

// Run A{ "trace_id": "a41f…", "gen_ai.operation.name": "agent_run", "gen_ai.request.model": "claude-sonnet-4-6", "duration_ms": 3180, "gen_ai.usage.input_tokens": 2847, "gen_ai.usage.output_tokens": 193, "gen_ai.response.finish_reasons": ["stop"], "spans": [ { "name": "lookup_customer", "status": "OK", "duration_ms": 112 }, { "name": "list_orders", "status": "OK", "duration_ms": 340 }, { "name": "issue_refund", "status": "OK", "duration_ms": 908 } ], "error_count": 0 }// Run B{ "trace_id": "b73c…", "gen_ai.operation.name": "agent_run", "gen_ai.request.model": "claude-sonnet-4-6", "duration_ms": 3204, "gen_ai.usage.input_tokens": 2851, "gen_ai.usage.output_tokens": 188, "gen_ai.response.finish_reasons": ["stop"], "spans": [ { "name": "lookup_customer", "status": "OK", "duration_ms": 118 }, { "name": "list_orders", "status": "OK", "duration_ms": 336 }, { "name": "issue_refund", "status": "OK", "duration_ms": 913 } ], "error_count": 0 }

Every field that your monitoring dashboard aggregates is, for practical purposes, identical. Both runs are three spans, all OK, no errors, ~3.2 seconds,...

Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE

Read more