Golden datasets

Build golden datasets for AI agents from production failures.

Turn real agent failures into governed, human-approved examples for evaluation and runtime retrieval—without hiding corrections inside model weights.

Most golden datasets never leave the lab.

Benchmarks and synthetic labels can score a model. They rarely capture the exceptions your humans already fix in production.

Generic benchmarks

Public eval sets miss your policies, tools, and escalation paths.

Annotation without context

Labels without the full trace cannot teach the next similar request.

Fine-tuning as the only fix

Training runs slow the correction cycle and bury individual examples.

Lost human edits

Support tickets and reviewer rewrites disappear instead of becoming reusable experiences.

What less than three builds

A golden dataset that is also an approved experience library: inspectable, scoped, versioned, and removable.

Production trace reconstruction

Requests, tool calls, intermediate steps, and final responses from OpenTelemetry or exports.

Human-reviewed corrections

Domain experts rewrite responses and trajectories with critiques tied to your rubric.

Runtime retrieval ready

Approved examples enter context as dynamic few-shots for similar future requests.

Design the dataset before you label it

A useful golden dataset is not a transcript archive. It has a narrow workflow boundary, a sampling rule, explicit acceptance criteria, and provenance for every approved example.

Sampling plan

Start with recurring, costly intents rather than a random export. Preserve difficult edge cases and known failure modes.

Annotation protocol

Define what reviewers can change, when they escalate, and how an approved answer is distinguished from a plausible one.

Provenance and versions

Keep the original trace, correction, reviewer, policy scope, and version together so each record can be audited or retired.

What a golden-dataset engagement delivers

Buyers receive governed records—not a raw dump. Formats and ownership stay inspectable.

Exact deliverables

Approved JSON/CSV records, annotation rubric, coverage map for one workflow, and a retrieval prototype for that same workflow.

Filled record example

Each record keeps source trace ID, observed failure, approved replacement, critique, scope, version, and approval status.

Coverage and sizing

Start with recurring, costly intents. Prefer a coherent hard set over a large random export.

Annotation QA

Reviewers resolve disagreements against a policy owner. Plausible answers without approval stay out of retrieval.

Security and PII

Minimize fields before review, mask sensitive values, and keep originals in the approved system of record.

Reusable resource

Download an illustrative golden-dataset template

Inspect the fields for source trace, observed failure, approved correction, critique, scope, version, and approval status.

Illustrative template—not client data. Adapt it to your policies, privacy controls, and approval process.

Who this is for

  • Customer support and voice agents with repeating failure modes
  • Healthcare admin, claims, and compliance review workflows
  • Teams that already edit agent output and need those fixes to compound
  • Operators who need measurable before/after—not another dashboard

Agent Improvement Sprint deliverables

A — Audit

Trace audit, failure patterns, and an annotation rubric tied to your success definition.

B — Dataset

Human-corrected responses/trajectories plus a retrieval prototype for one workflow.

C — Proof

Baseline comparison, quality review, cost/latency notes, and a rollout recommendation.

Controls that keep retrieval safe

Irrelevant or stale examples can hurt performance. Approval gates, metadata filters, and thresholds come before rollout.

Limitations

  • Retrieved examples can hurt performance if scope filters are weak or stale records remain active.
  • Golden sets drift as policies, tools, and products change; retirement rules are required.
  • Production traces can contain personal data and must be minimized before review.
  • Human reviewers disagree; unresolved cases should not become reusable gold.

Methodology by LTTAI LTD

This page describes less than three’s production-trace review method: reconstruct failures, capture human-approved corrections, store them outside model weights, and measure whether scoped retrieval improves a defined workflow.

Last updated: 2026-08-11

Primary references used for terminology and risk framing:

Common questions

What does golden datasets include?

A golden dataset that is also an approved experience library: inspectable, scoped, versioned, and removable.

Do you fine-tune our model?

No. Approved corrections stay outside model weights and can be retrieved at runtime. Fine-tuning can still complement the system for stable, common behavior.

What do we need to start?

A sample of production traces or conversation exports for one workflow, plus a working definition of success for that workflow.

How is success measured?

We compare a fixed baseline against a changed agent on the same workflow using task success, edits, escalations, latency, and cost where available.

Download the sample—or book a trace review.

Use the illustrative template now. For a scoped engagement, bring a slice of production traces and we will estimate experience coverage.