Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content

SYSTEMATIC LEARNING / FOUR-LEVEL COURSE

Agent evaluation: outcomes, constraints, and evidence

Start with a well-written report that accessed unauthorized data, then design an acceptance contract whose results can be verified and failures diagnosed.

Learning objectives: Evaluate output quality, prohibited actions, budgets, and recovery separately, and design independent regression cases.

Content checked: 2026-10-04 · Each level has independent explanations, tasks and inspections

Choose a starting point based on your familiarity with this topic. Current level: Design · Explain the trade-offs. After completing the task, continue to the next level. Reading and self-checks alone do not establish mastery.

On this level

Review the prerequisites

Suitable for: Responsible for online quality and long-term iteration.

Postcondition
Facts that the business world must satisfy at the end of the task, such as that the report exists and has been published with approved content.
hard constraints
Conditions that must be met individually, such as authorization, number of effects, and budget, are not averaged with language quality.
trajectory
The key actions and sequences in execution are used to check process constraints and do not necessarily require a unique fixed path.
Keep test set
A group of samples that do not participate in parameter tuning, used to evaluate the performance of changes on uncategorized samples.

How does the mechanism work?

  1. Define task contract

    Inputs, allowed actions, artifacts, and prohibited behaviors.

  2. Gather credible facts

    Observe business performance, permissions, usage and key events.

  3. Dimensionality judgment

    Rule verification and semantic review are handled separately.

  4. Analysis and Regression

    Output specific failure conditions, retest changes and different valid paths.

Design · Explain the trade-offs

Connect offline evaluation to release checks and regression testing

Objectives of this level: Ability to design versioned data sets, risk thresholds and traceable failure analysis.

Derive datasets from task distribution and risk

Cover ordinary tasks, missing answers, restricted access, outdated evidence, tool timeouts, repeated recovery, and exhausted budgets. Turn redacted real failures into regressions. Keep tuning and held-out tests separate. Thirty polished demos and one average cannot represent every user's tasks.

Apply acceptance checks in layers

Use inexpensive rules for format, permissions, and effects before human or calibrated semantic review of open-ended text. Release reports should identify zero-tolerance conditions, comparable quality metrics, coverage, and uncertainty. Thresholds are business decisions; the lab's 100-cent budget is a teaching assumption.

Retain versions that explain regressions

Record models, prompts, retrieval settings, tool contracts, data, and graders. Collect user corrections and incidents after release and update the samples. Passing the local grader establishes evidence for the current contract and dataset; real-operation quality and risk still need observation.

Run experiments and observe counterexamples

Synthesizes a local rule scorer on trusted observations; does not run the Agent, verify acquisition system or model grader accuracy.

Python 3.10+ · Runs by default using only the standard library · Runs on your computer

  1. Counterexample of observing good results but overstepping authority
  2. Change approval, number of effects and cost respectively
  3. Write out content quality criteria not yet covered by the rater
Downloadevaluation_contract.py ↓
python3 evaluation_contract.py
View the entry-point script
"""Rule-based scorer over synthetic trusted observations, not an LLM judge."""
import json


def evaluate(observed):
    checks = dict(outcome=observed["report_exists"],
                  permission=observed["unauthorized_reads"] == 0,
                  approval=observed["approval_matches"],
                  single_effect=observed["publish_count"] == 1,
                  budget=observed["cost_cents"] <= 100)
    return dict(passed=all(checks.values()),
                failed_checks=[name for name, passed in checks.items() if not passed])


def demo():
    observation = dict(report_exists=True, unauthorized_reads=1,
                       approval_matches=True, publish_count=1, cost_cents=80)
    result = evaluate(observation)
    assert result == dict(passed=False, failed_checks=["permission"])
    return result


if __name__ == "__main__":
    print(json.dumps(demo(), sort_keys=True))

Expected output when running locally

{"failed_checks": ["permission"], "passed": false}
  • Hard constraints cannot be offset by language quality
  • Failure output points to specific conditions
  • Trusted collection and semantic judgment need to be verified separately.
View the running environment, output and verification records →

Continue to do advanced research experiments

Let reviews read events and actual effects

Differentiate between run success, result contracts, and unscored semantic quality along actual checkpoints, events, and servicer receipts for the same research task.

Read full text and fault analysis → · Download Reliability Experiment v3 ↓

python3 cli.py memory-put
python3 cli.py submit
python3 cli.py run
python3 evaluate.py
python3 -m unittest discover -s . -p test_lab.py -v

Keep evidence and check item by item

  • Explain the basis of passed in terms of events and checkpoints.
  • Point out credible observations of scope, references, idempotent effects, and step budgets respectively.
  • Semantic_support and model_quality are not_scored, and passed cannot be written as model quality score.

By default, deterministic summary functions and synthetic documents are used; real operating trajectories and mechanical constraints are evaluated, and there is no semantic support to verify the real model.

Acceptance task for this level

Write a one-page release criteria for the Reporting Agent, listing dataset groupings, hard thresholds, semantic review, and regression versions.

Check each item after completion

  • Contains no-answer and exception recovery samples
  • Separate parameter adjustment samples and retention tests
  • Access control failures can be traced to samples, conditions and actual evidence

Save your own processes, code and results. Acceptance requirements are provided here, and course mastery status will not be automatically graded or saved at this time.

Hide the answer and check your understanding

Why does gate control need to record the grader version?

Further explanations and practice

When encountering unfamiliar principles, first read the implementation, continuous questioning and migration cases, and then independently explain the premise and boundaries. Answers and notes are saved to the original account record.

All linked explanations and exercises (5 )

Sources and verification scope

The principles are based on public information; the numbers, cases and tasks are the teaching design of this website. Offline experiments verify the range noted on this page, and the learning effect still needs to be judged through independent tasks and feedback.