01 · Separate answer quality, execution, and business outcomes
Continue the research-report example. Text saying the report was saved establishes only that the model made that claim. A publish checkpoint establishes a locally recorded receipt. One row in the independent receipts database establishes no duplicate publication in this experiment. Evaluate these observations separately instead of accepting an Agent's self-reported success=true.
| Dimension | Observable evidence | Limit of the evidence |
|---|---|---|
| Execution reliability | Claim, commit, retry, and cancellation events | Completing execution does not establish answer correctness |
| Output shape | A nonempty summary and valid citation-ID array | Valid IDs do not establish support for claims |
| Business outcome | Operation key, payload digest, and receipt in an independent database | One experiment cannot establish exactly-once behavior under every network failure |
| Semantic quality | Claim-by-claim comparison with source passages | Model-grader scores are not human ground truth |
| Resource use | Actual calls, billed tokens, and wall-clock time | Offline test-double timing does not measure model-service latency |
Anthropic's evaluation approach distinguishes tasks, trials, graders, trajectories, and outcomes and combines code-based, model-based, and human grading. This article adopts that distinction; its samples, thresholds, code, and results were independently designed for this site.
Anthropic · Agent Evaluation explains evaluation objects, trajectories, outcomes, and grader types. (Source checked: 2026-09-25)
How to interpret the results: The 18 baseline mechanism tests passed, and version 3 adds 18 approval tests. These runs provide no live model ranking, semantic-quality measurement, or production throughput test. HTTP responses are simulated, and default summaries come from a deterministic test double.
02 · Define test contracts with inputs, failures, expectations, and evidence
A question alone is insufficient for a test. Record initial state, fault-injection point, permitted terminal states, and prohibited effects. For cancellation, denied permissions, or invalid memory, stopping correctly may be the expected result. Requiring succeeded for every task would misclassify those cases.
Example test contract
{
"case_id": "effect-before-checkpoint/v1",
"input_version": "report/v1",
"provider": "fixture",
"precondition": "fresh task + empty publisher database",
"fault": "after_effect",
"resume": "advance virtual clock by 31 seconds; rerun without fault",
"expected": {
"status": "succeeded",
"checkpoint_count": 4,
"publisher_receipts": 1,
"publish_deduplicated": true
},
"evidence": [
"jobs",
"checkpoints",
"events",
"publisher.receipts"
]
}
| Implemented group | Scenarios | Core assertions |
|---|---|---|
| Recovery and effects | Crash after collect, crash after remote success, duplicate delivery | Committed steps are not repeated; the business receipt remains unique |
| Concurrency and task control | Two claimants, stale-generation writeback, cancellation, deadline | One claimant wins; invalid commits fail; expired tasks stop |
| Input and reports | Same ID with changed input, unknown source ID, conflicting idempotency payload | Reject changed intent under an old identity; failed tasks do not publish |
| Retry | Repeated simulated 429 responses | Two- then four-second backoff; at most three claims |
| Memory | Tenant/user/project isolation, expiry, unconfirmed records, updates, recovery after deletion | Ineligible records are excluded; old snapshots are invalidated |
| Evaluation and adapters | Mutated events, real CLI exit, simulated HTTP success/429 | Reject invalid traces; observe exit code 75; classify errors correctly |
Test groups are not production incident statistics. An isolation test using a trusted scope verifies a filter rather than attacking a real authentication gateway. It cannot establish enterprise-wide authorization isolation. Document evidence strength in the test contract to prevent inflated conclusions.
03 · Evaluate events and checkpoints without collecting hidden reasoning
The site's event table records submission, claims, step starts, step commits, retries, and cancellation. It retains job_id, generation, timestamps, and details, enough to reconstruct this example's state transitions. The runtime writes these logs; model output cannot rewrite them. Production systems must also restrict database write access so that actors outside the worker cannot forge audit records.
Illustrative event shape; see evidence.json for actual sequence numbers
{
"seq": 9,
"job_id": "evidence-001",
"at": 1031,
"generation": 2,
"kind": "step_started",
"detail": "publish"
}
The lease tests use a virtual clock. Values 1000 and 1031 are test coordinates, not production timestamps. Database-generated sequence numbers order events within this database. Across services, machine clocks alone cannot establish causality; add span_id, parent_span_id, operation_id, and attempt_id.
| Record | Useful fields | Avoid retaining indiscriminately |
|---|---|---|
| Task metadata | Input digest, versions, authorization references, budget | API keys and session cookies |
| Tool events | Tool name, status, duration, resource ID, error category | Entire account histories or plaintext sensitive documents |
| Model usage | Request ID, input/output tokens, retries | Every user conversation without a retention assessment |
| Evidence store | Controlled document references, passage hashes, access rules | References assumed accessible to every user |
Tool calls, explicit output, checkpoints, and receipts can verify most execution invariants without hidden reasoning. If complete tool outputs must be retained, give them separate access controls and retention limits. Do not publish them alongside ordinary debug logs.
04 · A runnable grader: score supported conditions and label the rest unscored
evaluate.py · Implemented grading rules
def grade(snapshot, expected_status='succeeded'):
events = snapshot['events']
commits = [e['detail'] for e in events if e['kind'] == 'step_committed']
cp = snapshot['checkpoints']
checks = {'expected_terminal_state': snapshot['status'] == expected_status,
'no_duplicate_checkpoint_commit': len(commits) == len(set(commits)),
'trace_matches_checkpoints': set(commits) == set(cp)}
if expected_status == 'succeeded':
checks['all_steps_present'] = set(cp) == set(STEPS)
checks['source_id_gate_passed'] = cp.get('verify', {}).get('schema_and_source_ids') is True
checks['receipt_present'] = bool(cp.get('publish', {}).get('receipt'))
return {'passed': all(checks.values()), 'checks': checks,
'semantic_support': 'not_scored', 'model_quality': 'not_scored'}
The code checks expected terminal state, duplicate commits, and consistency between events and checkpoints. For succeeded, it requires all four steps, valid source IDs, and a receipt. These are mechanical consistency checks. They establish neither tamper-proof logs nor semantic_support=true.
The grader deliberately permits repeated tool attempts. Crash recovery may produce two publish step_started events, but only one step_committed and one remote business record. Rejecting every tool called more than once would reject a correct recovery path.
Final status alone misses failures. test_grader_rejects_trace_mutation duplicates a commit event and verifies that the grader rejects that trace. This mutation test challenges an explicit invariant rather than demonstrating a grader that always passes.
Evaluate an actual task database
python3 evaluate.py --db lab.sqlite --job report-001
# 如果是在验证明确取消的任务:
python3 evaluate.py --db cancel.sqlite --job report-001 --expected-status cancelled
05 · Recorded observations and their supported conclusions
| Observation | At interruption | After recovery |
|---|---|---|
| Task status | running | succeeded |
| Claim generation | 1 | 2 |
| Local checkpoints | 3 | 4 |
| Receipts in the simulated publishing database | 1 | 1 |
| publish deduplicated | No local publish result yet | true |
These records come from the after_effect experiment using two real SQLite files and a virtual clock. It injects process-interruption behavior after a commit, then advances the clock past lease expiry. A separate CLI-subprocess test calls os._exit(75) and verifies that the committed collect checkpoint survives actual process exit.
evidence.json: before/after snapshots and grades · validation.json: environment, tests, and unverified scope · test_lab.py: assertions
Recorded unittest output excerpt; elapsed time depends on the environment
Ran 18 tests in 0.304s
OK
The 0.304-second duration measures this offline test suite, not Agent response latency or throughput. The two-process claim test covers one contention scenario rather than all scheduling behavior under load. Recovering one injected failure cannot establish a 100% recovery rate.
06 · Add semantic evaluation at the claim level
A common model error is to cite a real source that does not support the sentence. For example, s1 describes checkpoint storage but is cited for 'system efficiency improved by 30%.' The current source-ID check accepts that citation; claim-level verification is needed to detect the problem.
| Claim | Required evidence | Suggested label |
|---|---|---|
| Recovery skips committed collect | One collect commit and no subsequent collect start | supported |
| Every external API executes exactly once | This lab supplies no such evidence, and counterexamples exist | unsupported |
| Tokens fall by 30% on average | Actual billed usage for paired tasks under both versions | insufficient_evidence |
| Deletion permanently removes all copies | Complete deletion records for primary storage, caches, checkpoints, and backups | unsupported |
For each report, map claim_id → source_id → supporting passage → label → reason. As an initial workload, choose 20 disputed business claims and have two domain reviewers label them independently. Resolve rubric disagreements before calibrating a model grader. Twenty is a starting-workload suggestion, not a statistically sufficient sample size.
Human annotation template
{
"claim_id": "c1",
"claim": "恢复后 collect 未重复执行",
"source_id": "run-evidence-001",
"evidence": "generation=2 无 collect 的 step_started",
"label": "supported",
"reviewer": "human-label-required",
"rubric_version": "grounding/v1"
}
The grading model reads only the task, candidate, and allowed evidence; it must have no publishing capability. Version the rubric and record a confusion matrix against human labels. If insufficient evidence is repeatedly scored as supported, repair the grader before optimizing the Agent against it. The current lab implements no semantic-grading layer.
07 · Compare versions on paired tasks
To compare prompt-v1 and prompt-v2, freeze source and model versions, tool environment, budgets, and grader version. Run both on the same tasks, repeat trials, and retain every result rather than only the best. Provider aliases can change; record a fixed model version when available.
| Metric | Definition | Interpretation limit |
|---|---|---|
| Task success rate | Trials satisfying their task contract / all valid trials | Explain infrastructure-failure handling; never silently discard failures |
| Hard safety/business failures | Separate counts for unauthorized access, unapproved writes, and duplicates | Answer-quality averages cannot offset violations |
| Citation support rate | Supported claims / claims requiring evidence | Separate absent evidence from contradictory evidence |
| Cost per successful task | Total trial cost, including failures / successful tasks | Counting only successful calls understates cost |
| End-to-end latency | Submission to terminal state, including queues and backoff | Report model-call latency separately |
A version needing three attempts may look good on pass@3 while providing a worse single-attempt experience. If users get one submission, compare first-attempt success and retry overhead. Workflows requiring consistent success also need repeated-trial reliability. At-least-one success and success on every attempt are different metrics.
For small samples, report numerators, denominators, paired task differences, and failure lists. Label insufficiently powered findings exploratory instead of claiming statistical significance. Different inputs, models, or tool data prevent attributing an improvement solely to a prompt change.
08 · Define failures that block a release
For an existing Java system, derive cases from integration tests and real defects before adding model judgments. Test writes and external actions in isolated sandboxes with controlled tools. Evaluation accounts should have no broader permissions than pilot users; testing is no reason to temporarily enlarge access.
- Run mechanical regressions after each code change; retain failure output and stop the pipeline if they fail.
- For prompt, model, or retrieval changes, run a frozen development set and retain input digests, output, traces, and grader versions.
- Compare candidates on a held-out set and inspect hard failures individually. Repeated tuning against that set invalidates its held-out role.
- Pilot read-only business tasks, retain human corrections and failure categories, and convert reproducible incidents into regressions.
Offline regression gate; the second command requires creating a normal task as described in README
python3 -m unittest discover -s . -p test_lab.py -v
python3 evaluate.py --db lab.sqlite --job report-001
The 18 baseline tests include both success and rejection. An invalid source ID must produce failed; cancellation must produce cancelled. CI passes when each case reaches its predefined valid outcome, not when every task reaches succeeded.
Release review needs records answering three questions: which failures improved, whether duplicates, unauthorized access, or costs regressed, and whether the evidence can be replayed. A score screenshot alone answers none of them.
09 · Diagnose evaluation failures in order
| Symptom | Distinguish first | Next step |
|---|---|---|
| Good prose but failed task | Hard constraints versus language quality | Inspect events for unknown sources, deadline expiry, or invalid state before editing prose |
| Citations present but low semantic score | Missing source versus unsupported claim | Build a claim-to-evidence map |
| Mechanical pass but duplicate reports | Local commits versus actual remote effects | Query the independent business database and repair idempotency |
| Higher cost after a model change | Per-call tokens versus additional retries | Aggregate all trial costs and inspect format errors and timeouts |
| Large run-to-run variation | Model randomness versus changing tool data | Freeze snapshots, then run repeated paired trials |
| Human/model grading disagreement | Ambiguous rubric versus inadequate evidence | Calibrate the grader and retain disputed labels |
Version and test the grader itself. For each new rule, include a valid case that must pass and an invalid case that must fail, ensuring valid alternatives remain accepted. Define subjective requirements such as complete answers or elegant code operationally before automating their evaluation.
One example shared across all three articles
Python 3.10+ · Standard library · Offline by default · Source code, 36 tests, and run records included
Download the complete lab ZIP · Run instructions
10 · Approval introduces expected states beyond succeeded
| Expected state | Required assertions | Common mistake |
|---|---|---|
| waiting_approval | Draft exists, zero published records, repeated run does not regenerate | Treating an intentional wait as failure |
| failed / rejected / expired | No dispatch and zero published records | Checking status while missing execution before rejection |
| succeeded | Valid approval precedes dispatch; one business record | Final success hides an earlier unapproved call |
| reconciling | No new dispatch; preserve uncertainty and reconciliation evidence | Treating no receipt as proof of no effect |
| effect_confirmed | Receipt matches original payload_hash; fresh_write_performed=false | Treating an existing effect as new authorization |
Version 3 adds 18 approval tests to the original 18 baseline tests, for 36 in total. New cases include concurrent approve/reject processes, duplicate callbacks, changed content, role and scope rejection, approval expiry, receipt reconciliation after cancellation, and no redispatch when a receipt is missing. Identity tests cover only local Principal fixtures, not a real authentication system.
Run the complete regression suite; test_lab.py alone excludes approval extensions
python3 -m unittest discover -s . -p 'test_*.py' -v
Baseline evaluate.py checks the original four-step workflow. test_approval.py directly checks approval ordering and business effects. Baseline passed=true does not establish approval compliance. See test_approval.py and validation.json for assertions and recorded results.
11 · Verify that recovery issued no new write
One final report row is insufficient: a deduplicating service could receive a second write and return the same row. During recovery, the approval test replaces publish with a test double that raises if called, runs reconciliation, and checks the real SQLite receipt. Together, these observations establish no new dispatch and a traceable original effect.
| Test evidence | Check performed | Scope |
|---|---|---|
| test_crash_after_effect_and_expiry_reconciles_without_write | publish raises after approval expiry, yet recovery completes | This code path only queries receipts |
| test_no_receipt_after_dispatch_stays_unknown | Repeated recovery finds no receipt and never calls publish | Unknown outcomes do not cause automatic resending |
| test_cancel_after_effect_does_not_claim_rollback | The original receipt remains and cancellation intent is retained | Status reporting invents no rollback |
| test_concurrent_approve_reject_one_decision_wins | Two actual subprocesses compete for one decision | Single-request contention under local database transactions |
The tests cover no cross-region network, live third-party queue, or remote disaster recovery. Production reconciliation must define query consistency, receipt-visibility delay, idempotency retention, and final expiry. Different provider contracts require different closure rules.
Sources and verification
Official sources establish the referenced mechanisms. The schema, program, and experiments were independently designed for this site.