Training latency
By the time a fine-tune ships, the failure cluster and policies may have already changed.
Improve agents without fine-tuning
When the same failures keep returning, you do not need another training run first. Capture human-approved replacements from traces, retrieve them for similar requests, and prove lift on one scoped workflow.
Training embeds behavior slowly and opaquely. Repeated production exceptions need a faster, inspectable correction loop.
By the time a fine-tune ships, the failure cluster and policies may have already changed.
Individual approved answers disappear inside weights and cannot be audited or retired cleanly.
Stuffing every exception into the system prompt weakens what already works.
A red grade without an approved replacement does not change the next similar run.
We reconstruct failures, capture approved corrections outside model weights, retrieve scoped examples at runtime, and re-measure before wider rollout.
Start from real failures with tools, state, and outcomes—not synthetic guesses.
Keep rewrites, critiques, scope, version, and approval status inspectable and removable.
Compare task success, edits, escalations, latency, and cost on one workflow.
Fine-tuning still helps for stable, common behavior. For expensive repeating exceptions, a governed experience library is usually the faster path to proof.
Support, voice, claims, and compliance paths where humans already rewrite agent output.
Regulated or high-risk flows that must show which example influenced a decision.
Teams that need measured lift this sprint—not after the next training cycle.
Store corrections as golden datasets for AI agentsBuild the library from production traces
Prove whether approved retrieval beats waiting on a training cycle.
Which failure clusters belong in retrieval now vs a later fine-tune.
Approved corrections with scope, version, critique, and retirement rules.
Scoped few-shot injection for one workflow with filters and thresholds.
Task success, edits, escalations, latency, and cost on a fixed baseline.
Approval gates, negative retrieval tests, and expand/stop criteria.
Reusable resource
Inspect fields for source trace, failure, approved correction, critique, scope, version, and approval status.
Illustrative template—not client data. Adapt it to your policies, privacy controls, and approval process.
Cluster repeating failures and decide which belong in retrieval vs a later train.
Human-approved experience library + retrieval prototype for one workflow.
Baseline vs corrected comparison and a clear rollout recommendation.
We do not promise improvement before examining the workflow. The first engagement exists to produce evidence—and to decide whether fine-tuning is still needed later.
This page describes less than three’s production-trace review method: reconstruct failures, capture human-approved corrections, store them outside model weights, and measure whether scoped retrieval improves a defined workflow.
Primary references used for terminology and risk framing:
We reconstruct failures, capture approved corrections outside model weights, retrieve scoped examples at runtime, and re-measure before wider rollout.
No. Approved corrections stay outside model weights and can be retrieved at runtime. Fine-tuning can still complement the system for stable, common behavior.
A sample of production traces or conversation exports for one workflow, plus a working definition of success for that workflow.
We compare a fixed baseline against a changed agent on the same workflow using task success, edits, escalations, latency, and cost where available.
Send a sample of production traces. We will map which failures can improve through approved retrieval first.