Agent Application DevelopmentAccount
← Return to research directory

EVALUATION

Evaluating agents: execution traces and failure regression tests

Turn success criteria into a test contract using real events and business receipts. Understand the limits of baseline and approval tests, then extend evaluation to semantic quality and cost.

01 · Separate answer quality, execution, and business outcomes

Continue the research-report example. Text saying the report was saved establishes only that the model made that claim. A publish checkpoint establishes a locally recorded receipt. One row in the independent receipts database establishes no duplicate publication in this experiment. Evaluate these observations separately instead of accepting an Agent's self-reported success=true.

Dimension Observable evidence Limit of the evidence
Execution reliability Claim, commit, retry, and cancellation events Completing execution does not establish answer correctness
Output shape A nonempty summary and valid citation-ID array Valid IDs do not establish support for claims
Business outcome Operation key, payload digest, and receipt in an independent database One experiment cannot establish exactly-once behavior under every network failure
Semantic quality Claim-by-claim comparison with source passages Model-grader scores are not human ground truth
Resource use Actual calls, billed tokens, and wall-clock time Offline test-double timing does not measure model-service latency

Anthropic's evaluation approach distinguishes tasks, trials, graders, trajectories, and outcomes and combines code-based, model-based, and human grading. This article adopts that distinction; its samples, thresholds, code, and results were independently designed for this site.

Anthropic · Agent Evaluation explains evaluation objects, trajectories, outcomes, and grader types. (Source checked: 2026-09-25)

How to interpret the results: The 18 baseline mechanism tests passed, and version 3 adds 18 approval tests. These runs provide no live model ranking, semantic-quality measurement, or production throughput test. HTTP responses are simulated, and default summaries come from a deterministic test double.

02 · Define test contracts with inputs, failures, expectations, and evidence

A question alone is insufficient for a test. Record initial state, fault-injection point, permitted terminal states, and prohibited effects. For cancellation, denied permissions, or invalid memory, stopping correctly may be the expected result. Requiring succeeded for every task would misclassify those cases.

Example test contract

{
  "case_id": "effect-before-checkpoint/v1",
  "input_version": "report/v1",
  "provider": "fixture",
  "precondition": "fresh task + empty publisher database",
  "fault": "after_effect",
  "resume": "advance virtual clock by 31 seconds; rerun without fault",
  "expected": {
    "status": "succeeded",
    "checkpoint_count": 4,
    "publisher_receipts": 1,
    "publish_deduplicated": true
  },
  "evidence": [
    "jobs",
    "checkpoints",
    "events",
    "publisher.receipts"
  ]
}
Implemented group Scenarios Core assertions
Recovery and effects Crash after collect, crash after remote success, duplicate delivery Committed steps are not repeated; the business receipt remains unique
Concurrency and task control Two claimants, stale-generation writeback, cancellation, deadline One claimant wins; invalid commits fail; expired tasks stop
Input and reports Same ID with changed input, unknown source ID, conflicting idempotency payload Reject changed intent under an old identity; failed tasks do not publish
Retry Repeated simulated 429 responses Two- then four-second backoff; at most three claims
Memory Tenant/user/project isolation, expiry, unconfirmed records, updates, recovery after deletion Ineligible records are excluded; old snapshots are invalidated
Evaluation and adapters Mutated events, real CLI exit, simulated HTTP success/429 Reject invalid traces; observe exit code 75; classify errors correctly

Test groups are not production incident statistics. An isolation test using a trusted scope verifies a filter rather than attacking a real authentication gateway. It cannot establish enterprise-wide authorization isolation. Document evidence strength in the test contract to prevent inflated conclusions.

03 · Evaluate events and checkpoints without collecting hidden reasoning

The site's event table records submission, claims, step starts, step commits, retries, and cancellation. It retains job_id, generation, timestamps, and details, enough to reconstruct this example's state transitions. The runtime writes these logs; model output cannot rewrite them. Production systems must also restrict database write access so that actors outside the worker cannot forge audit records.

Illustrative event shape; see evidence.json for actual sequence numbers

{
  "seq": 9,
  "job_id": "evidence-001",
  "at": 1031,
  "generation": 2,
  "kind": "step_started",
  "detail": "publish"
}

The lease tests use a virtual clock. Values 1000 and 1031 are test coordinates, not production timestamps. Database-generated sequence numbers order events within this database. Across services, machine clocks alone cannot establish causality; add span_id, parent_span_id, operation_id, and attempt_id.

Record Useful fields Avoid retaining indiscriminately
Task metadata Input digest, versions, authorization references, budget API keys and session cookies
Tool events Tool name, status, duration, resource ID, error category Entire account histories or plaintext sensitive documents
Model usage Request ID, input/output tokens, retries Every user conversation without a retention assessment
Evidence store Controlled document references, passage hashes, access rules References assumed accessible to every user

Tool calls, explicit output, checkpoints, and receipts can verify most execution invariants without hidden reasoning. If complete tool outputs must be retained, give them separate access controls and retention limits. Do not publish them alongside ordinary debug logs.

04 · A runnable grader: score supported conditions and label the rest unscored

evaluate.py · Implemented grading rules

def grade(snapshot, expected_status='succeeded'):
    events = snapshot['events']
    commits = [e['detail'] for e in events if e['kind'] == 'step_committed']
    cp = snapshot['checkpoints']
    checks = {'expected_terminal_state': snapshot['status'] == expected_status,
              'no_duplicate_checkpoint_commit': len(commits) == len(set(commits)),
              'trace_matches_checkpoints': set(commits) == set(cp)}
    if expected_status == 'succeeded':
        checks['all_steps_present'] = set(cp) == set(STEPS)
        checks['source_id_gate_passed'] = cp.get('verify', {}).get('schema_and_source_ids') is True
        checks['receipt_present'] = bool(cp.get('publish', {}).get('receipt'))
    return {'passed': all(checks.values()), 'checks': checks,
            'semantic_support': 'not_scored', 'model_quality': 'not_scored'}

Download full evaluate.py

The code checks expected terminal state, duplicate commits, and consistency between events and checkpoints. For succeeded, it requires all four steps, valid source IDs, and a receipt. These are mechanical consistency checks. They establish neither tamper-proof logs nor semantic_support=true.

The grader deliberately permits repeated tool attempts. Crash recovery may produce two publish step_started events, but only one step_committed and one remote business record. Rejecting every tool called more than once would reject a correct recovery path.

Final status alone misses failures. test_grader_rejects_trace_mutation duplicates a commit event and verifies that the grader rejects that trace. This mutation test challenges an explicit invariant rather than demonstrating a grader that always passes.

Evaluate an actual task database

python3 evaluate.py --db lab.sqlite --job report-001
# 如果是在验证明确取消的任务:
python3 evaluate.py --db cancel.sqlite --job report-001 --expected-status cancelled

05 · Recorded observations and their supported conclusions

Observation At interruption After recovery
Task status running succeeded
Claim generation 1 2
Local checkpoints 3 4
Receipts in the simulated publishing database 1 1
publish deduplicated No local publish result yet true

These records come from the after_effect experiment using two real SQLite files and a virtual clock. It injects process-interruption behavior after a commit, then advances the clock past lease expiry. A separate CLI-subprocess test calls os._exit(75) and verifies that the committed collect checkpoint survives actual process exit.

evidence.json: before/after snapshots and grades · validation.json: environment, tests, and unverified scope · test_lab.py: assertions

Recorded unittest output excerpt; elapsed time depends on the environment

Ran 18 tests in 0.304s
OK

The 0.304-second duration measures this offline test suite, not Agent response latency or throughput. The two-process claim test covers one contention scenario rather than all scheduling behavior under load. Recovering one injected failure cannot establish a 100% recovery rate.

06 · Add semantic evaluation at the claim level

A common model error is to cite a real source that does not support the sentence. For example, s1 describes checkpoint storage but is cited for 'system efficiency improved by 30%.' The current source-ID check accepts that citation; claim-level verification is needed to detect the problem.

Claim Required evidence Suggested label
Recovery skips committed collect One collect commit and no subsequent collect start supported
Every external API executes exactly once This lab supplies no such evidence, and counterexamples exist unsupported
Tokens fall by 30% on average Actual billed usage for paired tasks under both versions insufficient_evidence
Deletion permanently removes all copies Complete deletion records for primary storage, caches, checkpoints, and backups unsupported

For each report, map claim_id → source_id → supporting passage → label → reason. As an initial workload, choose 20 disputed business claims and have two domain reviewers label them independently. Resolve rubric disagreements before calibrating a model grader. Twenty is a starting-workload suggestion, not a statistically sufficient sample size.

Human annotation template

{
  "claim_id": "c1",
  "claim": "恢复后 collect 未重复执行",
  "source_id": "run-evidence-001",
  "evidence": "generation=2 无 collect 的 step_started",
  "label": "supported",
  "reviewer": "human-label-required",
  "rubric_version": "grounding/v1"
}

The grading model reads only the task, candidate, and allowed evidence; it must have no publishing capability. Version the rubric and record a confusion matrix against human labels. If insufficient evidence is repeatedly scored as supported, repair the grader before optimizing the Agent against it. The current lab implements no semantic-grading layer.

07 · Compare versions on paired tasks

To compare prompt-v1 and prompt-v2, freeze source and model versions, tool environment, budgets, and grader version. Run both on the same tasks, repeat trials, and retain every result rather than only the best. Provider aliases can change; record a fixed model version when available.

Metric Definition Interpretation limit
Task success rate Trials satisfying their task contract / all valid trials Explain infrastructure-failure handling; never silently discard failures
Hard safety/business failures Separate counts for unauthorized access, unapproved writes, and duplicates Answer-quality averages cannot offset violations
Citation support rate Supported claims / claims requiring evidence Separate absent evidence from contradictory evidence
Cost per successful task Total trial cost, including failures / successful tasks Counting only successful calls understates cost
End-to-end latency Submission to terminal state, including queues and backoff Report model-call latency separately

A version needing three attempts may look good on pass@3 while providing a worse single-attempt experience. If users get one submission, compare first-attempt success and retry overhead. Workflows requiring consistent success also need repeated-trial reliability. At-least-one success and success on every attempt are different metrics.

For small samples, report numerators, denominators, paired task differences, and failure lists. Label insufficiently powered findings exploratory instead of claiming statistical significance. Different inputs, models, or tool data prevent attributing an improvement solely to a prompt change.

08 · Define failures that block a release

For an existing Java system, derive cases from integration tests and real defects before adding model judgments. Test writes and external actions in isolated sandboxes with controlled tools. Evaluation accounts should have no broader permissions than pilot users; testing is no reason to temporarily enlarge access.

  1. Run mechanical regressions after each code change; retain failure output and stop the pipeline if they fail.
  2. For prompt, model, or retrieval changes, run a frozen development set and retain input digests, output, traces, and grader versions.
  3. Compare candidates on a held-out set and inspect hard failures individually. Repeated tuning against that set invalidates its held-out role.
  4. Pilot read-only business tasks, retain human corrections and failure categories, and convert reproducible incidents into regressions.

Offline regression gate; the second command requires creating a normal task as described in README

python3 -m unittest discover -s . -p test_lab.py -v
python3 evaluate.py --db lab.sqlite --job report-001

The 18 baseline tests include both success and rejection. An invalid source ID must produce failed; cancellation must produce cancelled. CI passes when each case reaches its predefined valid outcome, not when every task reaches succeeded.

Release review needs records answering three questions: which failures improved, whether duplicates, unauthorized access, or costs regressed, and whether the evidence can be replayed. A score screenshot alone answers none of them.

09 · Diagnose evaluation failures in order

Symptom Distinguish first Next step
Good prose but failed task Hard constraints versus language quality Inspect events for unknown sources, deadline expiry, or invalid state before editing prose
Citations present but low semantic score Missing source versus unsupported claim Build a claim-to-evidence map
Mechanical pass but duplicate reports Local commits versus actual remote effects Query the independent business database and repair idempotency
Higher cost after a model change Per-call tokens versus additional retries Aggregate all trial costs and inspect format errors and timeouts
Large run-to-run variation Model randomness versus changing tool data Freeze snapshots, then run repeated paired trials
Human/model grading disagreement Ambiguous rubric versus inadequate evidence Calibrate the grader and retain disputed labels

Version and test the grader itself. For each new rule, include a valid case that must pass and an invalid case that must fail, ensuring valid alternatives remain accepted. Define subjective requirements such as complete answers or elegant code operationally before automating their evaluation.

One example shared across all three articles

Python 3.10+ · Standard library · Offline by default · Source code, 36 tests, and run records included

Download the complete lab ZIP · Run instructions

10 · Approval introduces expected states beyond succeeded

Expected state Required assertions Common mistake
waiting_approval Draft exists, zero published records, repeated run does not regenerate Treating an intentional wait as failure
failed / rejected / expired No dispatch and zero published records Checking status while missing execution before rejection
succeeded Valid approval precedes dispatch; one business record Final success hides an earlier unapproved call
reconciling No new dispatch; preserve uncertainty and reconciliation evidence Treating no receipt as proof of no effect
effect_confirmed Receipt matches original payload_hash; fresh_write_performed=false Treating an existing effect as new authorization

Version 3 adds 18 approval tests to the original 18 baseline tests, for 36 in total. New cases include concurrent approve/reject processes, duplicate callbacks, changed content, role and scope rejection, approval expiry, receipt reconciliation after cancellation, and no redispatch when a receipt is missing. Identity tests cover only local Principal fixtures, not a real authentication system.

Run the complete regression suite; test_lab.py alone excludes approval extensions

python3 -m unittest discover -s . -p 'test_*.py' -v

Baseline evaluate.py checks the original four-step workflow. test_approval.py directly checks approval ordering and business effects. Baseline passed=true does not establish approval compliance. See test_approval.py and validation.json for assertions and recorded results.

11 · Verify that recovery issued no new write

One final report row is insufficient: a deduplicating service could receive a second write and return the same row. During recovery, the approval test replaces publish with a test double that raises if called, runs reconciliation, and checks the real SQLite receipt. Together, these observations establish no new dispatch and a traceable original effect.

Test evidence Check performed Scope
test_crash_after_effect_and_expiry_reconciles_without_write publish raises after approval expiry, yet recovery completes This code path only queries receipts
test_no_receipt_after_dispatch_stays_unknown Repeated recovery finds no receipt and never calls publish Unknown outcomes do not cause automatic resending
test_cancel_after_effect_does_not_claim_rollback The original receipt remains and cancellation intent is retained Status reporting invents no rollback
test_concurrent_approve_reject_one_decision_wins Two actual subprocesses compete for one decision Single-request contention under local database transactions

The tests cover no cross-region network, live third-party queue, or remote disaster recovery. Production reconciliation must define query consistency, receipt-visibility delay, idempotency retention, and final expiry. Different provider contracts require different closure rules.

Sources and verification

Official sources establish the referenced mechanisms. The schema, program, and experiments were independently designed for this site.

Verification records · Approval recovery evidence