Evaluation Studio
Built the evaluation surface to run golden test cases and measure reliability, groundedness, retrieval quality, latency, and escalation accuracy over time.Run golden test cases and inspect pass/fail, groundedness, retrieval hits, and escalation correctness.
Evaluation Studio
Built this deterministic eval runner so behavior can be verified repeatedly without relying only on LLM-as-judge scoring.Deterministic eval suite validates retrieval hits, citation presence, escalation accuracy, groundedness, and latency.
No eval run yet
Run the suite to produce case-level results and summary metrics.