Happy-path bias
Scripted flows miss live language, partial tools, and messy state.
AI agent testing
Go beyond scripted happy paths. Test tool use, state, recovery, and escalation on real traces—then turn reviewed failures into approved examples you can re-measure.
Simulated suites catch some regressions. Without approved replacements and production-grounded scenarios, testing does not improve the next similar run.
Scripted flows miss live language, partial tools, and messy state.
Checking the final string hides unsafe trajectories.
A failed test without an approved rewrite cannot teach runtime behavior.
Findings stay in tickets instead of becoming eval-ready experiences.
We combine scenario design, production-trace replay, human-approved corrections, and before/after measurement on one agent workflow.
Promote recurring live failures into versioned test scenarios.
Score tools, state, recovery, and escalation—not only final text.
Capture approved replacements and re-test with retrieval where appropriate.
Buyers need a release gate and a correction path—not another dashboard of red scores.
Lock prompts, tools, and expected trajectories so releases compare fairly.
Offline replay, controlled simulation, and targeted live review for high-risk clusters.
Define when a failed test becomes an approved experience for evaluation or retrieval.
Deepen trajectory scoring with AI agent evaluationConnect tests to LLM evaluation services and release gates
Turn production failures into versioned tests and approved fixes.
Map current tests against recurring production failures.
Tools, state, recovery, escalation, and safety checks.
When a failed test becomes an approved experience record.
Version prompts, tools, and expected trajectories for release comparison.
Start from the downloadable agent evaluation rubric CSV.
Reusable resource
Use the CSV as a testing scorecard for outcome, tools, state, recovery, escalation, safety, cost, and latency.
Illustrative template—not client data. Adapt it to your policies, privacy controls, and approval process.
Map current tests vs production failures and define trajectory assertions.
Versioned scenarios + approved corrections for one workflow.
Baseline vs changed agent comparison and a release recommendation.
A failed test is incomplete until there is an approved replacement and a re-measurement plan.
This page describes less than three’s production-trace review method: reconstruct failures, capture human-approved corrections, store them outside model weights, and measure whether scoped retrieval improves a defined workflow.
Primary references used for terminology and risk framing:
We combine scenario design, production-trace replay, human-approved corrections, and before/after measurement on one agent workflow.
No. Approved corrections stay outside model weights and can be retrieved at runtime. Fine-tuning can still complement the system for stable, common behavior.
A sample of production traces or conversation exports for one workflow, plus a working definition of success for that workflow.
We compare a fixed baseline against a changed agent on the same workflow using task success, edits, escalations, latency, and cost where available.
Bring your current test suite and a sample of production failures. We will propose one Agent Improvement Sprint.