Unlabeled failures
Errors without an ideal response cannot train or retrieve well.
Production traces → golden data
Transform real failures into normalized, reviewable records with an approved replacement, provenance, policy scope, and controls for evaluation or retrieval.
Production history shows outcomes. Without the chosen rewrite, critique, and approval status, traces cannot teach the next request.
Errors without an ideal response cannot train or retrieve well.
Made-up examples miss your tools, policies, and edge cases.
CSV dumps lack versioning, scope, and runtime retrieval design.
Pushing every fix into weights slows iteration and hides provenance.
Connect traces → human review → approved experiences → runtime retrieval → measured outcome.
OpenTelemetry preferred; conversation or observability exports work for a first sprint.
Internal reviewers, trained annotators, or both—your rubric, your approvals.
Each record keeps context, failed output, rewrite, critique, and status.
The artifact is a governed experience record, not a row of text. The fields below make it possible to audit retrieval and avoid applying one customer’s exception to another workflow.
Trace ID, user intent, relevant tools, policy state, and the failed output or trajectory.
The replacement response or action path, reviewer critique, approval status, and any escalation note.
Workflow, tenant, policy, confidence, and date metadata that limits where an experience can be used.
See the commercial golden-datasets service and its dataset design approachFollow the practical build process step by step
Raw traces become gold only after normalization, review, and approval.
OpenTelemetry spans, conversation exports, ticket threads, and observability dumps with stable identifiers.
Map request, tools, intermediate state, final output, outcome, and version fields into a reviewable record.
Use the trace-readiness checklist to reject incomplete, unsafe, or unresolvable candidates.
Raw trace → normalized candidate → human-approved gold → scoped retrieval or evaluation set.
Missing tool results, conflicting state, or retired policies are quarantined—not forced into gold.
Reusable resource
A print-ready checklist for required fields, privacy handling, eligibility, and rejection criteria.
Illustrative template—not client data. Adapt it to your policies, privacy controls, and approval process.
Coverage estimate and failure clustering from your trace sample.
Human-corrected golden set for one workflow plus retrieval prototype.
Baseline vs corrected comparison and next-step recommendation.
Every example points back to a real failure your users hit—so retrieval stays relevant and auditable.
This page describes less than three’s production-trace review method: reconstruct failures, capture human-approved corrections, store them outside model weights, and measure whether scoped retrieval improves a defined workflow.
Primary references used for terminology and risk framing:
Connect traces → human review → approved experiences → runtime retrieval → measured outcome.
No. Approved corrections stay outside model weights and can be retrieved at runtime. Fine-tuning can still complement the system for stable, common behavior.
A sample of production traces or conversation exports for one workflow, plus a working definition of success for that workflow.
We compare a fixed baseline against a changed agent on the same workflow using task success, edits, escalations, latency, and cost where available.
Book a call with a sample of production traces—we will map repeated failures and one sprint path.