Review the prerequisites
Suitable for: It is necessary to troubleshoot evaluation failures or version regressions.
- Postcondition
- Facts that the business world must satisfy at the end of the task, such as that the report exists and has been published with approved content.
- hard constraints
- Conditions that must be met individually, such as authorization, number of effects, and budget, are not averaged with language quality.
- trajectory
- The key actions and sequences in execution are used to check process constraints and do not necessarily require a unique fixed path.
- Keep test set
- A group of samples that do not participate in parameter tuning, used to evaluate the performance of changes on uncategorized samples.
How does the mechanism work?
- Define task contract
Inputs, allowed actions, artifacts, and prohibited behaviors.
- Gather credible facts
Observe business performance, permissions, usage and key events.
- Dimensionality judgment
Rule verification and semantic review are handled separately.
- Analysis and Regression
Output specific failure conditions, retest changes and different valid paths.
Debugging · Diagnose failures
Distinguish a system regression from a judging error
Objectives of this level: Able to distinguish between business errors, valid path differences and judgment errors.
Inspect failed dimensions and evidence
Replay input versions, tool configuration, observations, and critical events. Verify the facts consumed by the grader. Requiring one tool order when the task allows alternatives is an overly narrow contract. Scoring only final text when publication duplicated leaves a missing condition.
Allow valid alternative paths
Parallel retrieval may permit A then B or B then A. Check required events, dependencies, and prohibited behavior instead of a single complete event permutation. Equal business outcomes do not establish that both paths were authorized, but process checks should not mechanically reject every alternative.
Test randomness and grading error separately
Repeat the same task and retain success rates and failure categories. When retries are allowed, state whether the statistical unit is an attempt or a task. Calibrate open-ended grading against human standards and review disagreements. Evaluating final quality on repeatedly tuned data overestimates improvement.
Run experiments and observe counterexamples
Synthesizes a local rule scorer on trusted observations; does not run the Agent, verify acquisition system or model grader accuracy.
Python 3.10+ · Runs by default using only the standard library · Runs on your computer
- Counterexample of observing good results but overstepping authority
- Change approval, number of effects and cost respectively
- Write out content quality criteria not yet covered by the rater
python3 evaluation_contract.pyView the entry-point script
"""Rule-based scorer over synthetic trusted observations, not an LLM judge."""
import json
def evaluate(observed):
checks = dict(outcome=observed["report_exists"],
permission=observed["unauthorized_reads"] == 0,
approval=observed["approval_matches"],
single_effect=observed["publish_count"] == 1,
budget=observed["cost_cents"] <= 100)
return dict(passed=all(checks.values()),
failed_checks=[name for name, passed in checks.items() if not passed])
def demo():
observation = dict(report_exists=True, unauthorized_reads=1,
approval_matches=True, publish_count=1, cost_cents=80)
result = evaluate(observation)
assert result == dict(passed=False, failed_checks=["permission"])
return result
if __name__ == "__main__":
print(json.dumps(demo(), sort_keys=True))
Expected output when running locally
{"failed_checks": ["permission"], "passed": false}- Hard constraints cannot be offset by language quality
- Failure output points to specific conditions
- Trusted collection and semantic judgment need to be verified separately.
Continue to do advanced research experiments
Let reviews read events and actual effects
Differentiate between run success, result contracts, and unscored semantic quality along actual checkpoints, events, and servicer receipts for the same research task.
Read full text and fault analysis → · Download Reliability Experiment v3 ↓
python3 cli.py memory-put
python3 cli.py submit
python3 cli.py run
python3 evaluate.py
python3 -m unittest discover -s . -p test_lab.py -vKeep evidence and check item by item
- Explain the basis of passed in terms of events and checkpoints.
- Point out credible observations of scope, references, idempotent effects, and step budgets respectively.
- Semantic_support and model_quality are not_scored, and passed cannot be written as model quality score.
By default, deterministic summary functions and synthetic documents are used; real operating trajectories and mechanical constraints are evaluated, and there is no semantic support to verify the real model.
Acceptance task for this level
Design two valid trajectories and one unauthorized trajectory for the parallel data query task, and then write the decision rules.
Check each item after completion
- Valid paths do not fail due to irrelevant order differences
- Independent judgment on ultra vires incident failed
- Statistical caliber to illustrate repeated runs and task success rates
Save your own processes, code and results. Acceptance requirements are provided here, and course mastery status will not be automatically graded or saved at this time.
Hide the answer and check your understanding
If the success rate increases after changing to a more relaxed grader, can we say that the quality of the system has improved?
Expand reference derivation
Can't say directly. Changes in the judgment standards can also lead to an increase in scores. The standards should be fixed or calibrated, and then the actual task results should be compared with manual differences.
Further explanations and practice
When encountering unfamiliar principles, first read the implementation, continuous questioning and migration cases, and then independently explain the premise and boundaries. Answers and notes are saved to the original account record.
All linked explanations and exercises (5 )
- Executable acceptance checks and quality boundaries · answer independently
- Evaluate both outcomes and execution traces · answer independently
- Independent and representative evaluation samples · answer independently
- Calibration and errors of semantic judges · answer independently
- Success rates and retry accounting for stochastic tasks · answer independently
Sources and verification scope
The principles are based on public information; the numbers, cases and tasks are the teaching design of this website. Offline experiments verify the range noted on this page, and the learning effect still needs to be judged through independent tasks and feedback.