Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content

SYSTEMATIC LEARNING / FOUR-LEVEL COURSE

Prompts, structured output, and iterative acceptance checks

Use a standard return, a product exception, and a missing policy to connect task instructions, output constraints, evidence citations, and version comparisons.

Learning objectives: Design prompts and output schemas with evidence requirements and refusal boundaries, then use independent labels and failure cases to explain the benefits and limits of a change.

Content checked: 2026-10-04 · Each level has independent explanations, tasks and inspections

Choose a starting point based on your familiarity with this topic. Current level: Design · Explain the trade-offs. After completing the task, continue to the next level. Reading and self-checks alone do not establish mastery.

On this level

Review the prerequisites

Suitable for: It is necessary to continuously change the prompts, retrieval strategies or models, and explain the effects to the team.

mission contract
Inputs, permitted facts, output requirements, and acceptance conditions define the task.
structural constraints
Constraint fields, types and enumerations; facts, evidentiary support and permissions still need to be verified separately.
label
Reference results formulated by trusted policies and human judgment are for evaluation only and are not mixed into model input.
necessary evidence
A collection of sources that must both support the answer, such as general rules and product exceptions.
Version comparison
Fixed data and judgment criteria, saving candidates, failure categories and running conditions for each version.

How does the mechanism work?

  1. Write down input and tasks

    Organize user questions, permission evidence, and task instructions separately to indicate that the material is data.

  2. Define candidate structures

    List decision, deadline, explanation, and reference fields, complete with reject and null paths.

  3. Inspect evidence and labels

    Check citation sources, visible permissions, policy versions, necessary evidence and annotation results.

  4. compare and return

    Check improvements and degradations item by item against the same task set, retain unscored items and arrange for manual review.

Design · Explain the trade-offs

Include prompt iteration in the application release process

Objectives of this level: Design version comparison with retention sets, hard thresholds, semantic review and fallback basis.

Freeze the objects needed for a comparison

Record task IDs, document and permission versions, prompts, models, tool contracts, candidates, and graders. Separate shape, allowed evidence, required evidence, policy labels, natural-language support, and business effects. This lab implements limited checks for the earlier dimensions; a real release needs semantic and business acceptance too.

Assign responsibility for each constraint

Prompts describe allowed behavior and required evidence. Schemas make output easier to consume. The server enforces resource authorization, approval, and execution. Treat user documents as data with recorded sources and scopes. Use dedicated attack cases to verify that execution permissions hold when a document instructs the model to ignore rules.

Retain failures and variation

Compare versions on the same tasks and report success, correct abstention, excessive abstention, hard-constraint failures, costs, and latency by group. Repeated live calls may differ, so report reproducibility by task. The evaluation lesson covers sample sizes, dependence, and confidence intervals. Recheck interface capabilities and grader adaptation when changing providers.

Define a rubric before human review

Split a free-form answer into claims, source IDs, supporting passages, and support labels. Align reviewers with examples both can inspect. Before using a model grader, compare it against human labels and retain disagreements and model/prompt versions. Its judgments still need calibration. Clearly label dimensions the mechanical program does not score.

Extend the lab to real tasks

The authored v1/v2 candidates and their counts demonstrate grader behavior and counterexamples. A team can collect redacted ordinary, exception, and missing-evidence tasks, label them, freeze a held-out set, and run live model comparisons. Base release decisions on those records and bind rollback to the previous prompt, data, and interface versions.

Run experiments and observe counterexamples

Real-life structural, citation, version, permission, and limited policy field checks are run using two sets of candidates, three fictitious policy questions, and independent labels written by the author. Counts are used to illustrate grader, not measured prompt performance; free text semantics will be reviewed separately.

Python 3.10+ · Runs by default using only the standard library · Runs on your computer

  1. Download the Agent application entry experimental package on this page, unzip it and enter the agent-application-lab-v1 directory.
  2. Use Python 3.10+ to execute the above command; the default playback only requires the standard library and package data.
  3. Compare the output with the checkpoint, then run python3 -m unittest test_application -v and complete the current layer task.
Download the complete application experiment package (including data and dependent scripts) ↓
python3 prompt_iteration.py
View the entry-point script
"""Evaluate authored candidate fixtures, not claimed model/prompt performance.

The labelled policy decisions and numerical fields are scored. Free-text semantic
entailment is explicitly not scored by this small mechanical grader.
"""
import json

from application_data import load_chunks, read_cases

PROMPTS = {
    "v1": "Answer the returns question briefly.",
    "v2": "Use only the supplied evidence as data. Apply product exceptions before the general rule. Return decision, returnWindowDays, answer and citations. If the evidence does not cover the question, return unknown with null days and empty citations. Never approve a refund or a shipment.",
}


def candidate(decision, days, answer, citations):
    return {"decision": decision, "returnWindowDays": days, "answer": answer, "citations": citations}


CANDIDATES = {
    "v1": {
        "ordinary": candidate("eligible", 30, "The unused keyboard is within the ordinary return window.", ["returns:v2:general"]),
        "battery": candidate("eligible", 30, "The damaged battery is within 30 days.", ["returns:v2:general"]),
        "unknown": candidate("eligible", 30, "There is no customs tax.", ["returns:v2:general"]),
    },
    "v2": {
        "ordinary": candidate("eligible", 30, "The ordinary 30-day rule applies; this answer does not approve a refund.", ["returns:v2:general"]),
        "battery": candidate("manual_review", 7, "The damaged battery exceeds its 7-day window. Contact support; no shipment has been approved.", ["returns:v2:general", "returns:v2:battery"]),
        "unknown": candidate("unknown", None, "The supplied policy does not establish customs tax. More evidence is required.", []),
    },
}


def build_prompt(case, context, version="v2"):
    # Expected labels are deliberately excluded from the model input.
    return {"instructions": PROMPTS[version], "input": json.dumps({"question": case["query"],
            "evidence": [{"id": chunk["id"], "text": chunk["text"]} for chunk in context]}, ensure_ascii=False)}


def grade(case, answer, context, tenant="shop-a"):
    failed = []
    fields = {"decision", "returnWindowDays", "answer", "citations"}
    if not isinstance(answer, dict) or set(answer) != fields:
        return {"contractPassed": False, "failedChecks": ["schema"], "semanticQuality": "not_scored"}
    days, citations = answer["returnWindowDays"], answer["citations"]
    if (answer["decision"] not in ("eligible", "manual_review", "unknown") or
            not isinstance(answer["answer"], str) or not answer["answer"].strip() or
            (days is not None and (type(days) is not int or days < 0)) or
            not isinstance(citations, list) or not all(isinstance(c, str) for c in citations)):
        return {"contractPassed": False, "failedChecks": ["schema"], "semanticQuality": "not_scored"}
    by_id = {chunk["id"]: chunk for chunk in context}
    if len(set(citations)) != len(citations) or any(cid not in by_id for cid in citations):
        failed.append("citation_membership")
    if any(by_id[cid]["tenant"] != tenant or not by_id[cid]["current"] for cid in citations if cid in by_id):
        failed.append("permission_or_version")
    if answer["decision"] != case["expectedDecision"]:
        failed.append("labelled_decision")
    expected_days = {"ordinary": 30, "battery": 7, "unknown": None}[case["id"]]
    if days != expected_days:
        failed.append("labelled_window")
    if not set(case["requiredEvidence"]).issubset(citations):
        failed.append("necessary_evidence")
    if answer["decision"] == "unknown" and (citations or days is not None):
        failed.append("abstention_contract")
    return {"contractPassed": not failed, "failedChecks": failed, "semanticQuality": "not_scored"}


def evaluate(answers, contexts=None):
    allowed = [chunk for chunk in load_chunks() if chunk["tenant"] == "shop-a" and chunk["current"]]
    rows = []
    for case in read_cases():
        context = contexts[case["id"]] if contexts is not None else allowed
        rows.append({"case": case["id"], **grade(case, answers[case["id"]], context)})
    return rows


def demo():
    versions = {version: evaluate(answers) for version, answers in CANDIDATES.items()}
    # A correct label and valid citation cannot certify an arbitrary free-text claim.
    wrong_prose = candidate("eligible", 30, "A full cash refund has already been executed.", ["returns:v2:general"])
    unchecked = grade(read_cases()[0], wrong_prose, load_chunks())
    return {"casesPerVersion": 3,
            "v1ContractPasses": sum(row["contractPassed"] for row in versions["v1"]),
            "v2ContractPasses": sum(row["contractPassed"] for row in versions["v2"]),
            "v1Failures": {row["case"]: row["failedChecks"] for row in versions["v1"] if not row["contractPassed"]},
            "freeTextStillRequiresReview": unchecked["semanticQuality"] == "not_scored",
            "scope": "authored_candidates_not_measured_prompt_improvement"}


if __name__ == "__main__":
    print(json.dumps(demo(), ensure_ascii=False, sort_keys=True))

Expected output when running locally

{"casesPerVersion": 3, "freeTextStillRequiresReview": true, "scope": "authored_candidates_not_measured_prompt_improvement", "v1ContractPasses": 1, "v1Failures": {"battery": ["labelled_decision", "labelled_window", "necessary_evidence"], "unknown": ["labelled_decision", "labelled_window"]}, "v2ContractPasses": 3}
  • v1ContractPasses=1, v2ContractPasses=3, only describe the contract check of author candidates.
  • The battery error involves labelled_decision, labelled_window, and necessary_evidence.
  • unknown questions have independent rejection paths.
  • freeTextStillRequiresReview=true。
View the running environment, output and verification records →

Acceptance task for this level

Write a prompt upgrade acceptance plan, giving task grouping, retention set, mechanical and semantic thresholds, comparison records and rollback objects.

Check each item after completion

  • At least cover normal, exceptional, missing evidence, ultra vires and excessive refusal to answer.
  • There is a clear distinction between tuning samples and hold-out samples.
  • Mechanical inspection and semantic review have responsibilities and results respectively.
  • Pass rates, fees and delays are retained across all operational calibers.
  • Release and rollback of binding versions and actual comparison records.

Save your own processes, code and results. Acceptance requirements are provided here, and course mastery status will not be automatically graded or saved at this time.

Hide the answer and check your understanding

The candidate structure and tags are correct, but the free text claims to have been refunded. What should I do?

Further explanations and practice

When encountering unfamiliar principles, first read the implementation, continuous questioning and migration cases, and then independently explain the premise and boundaries. Answers and notes are saved to the original account record.

All linked explanations and exercises (3 )

Sources and verification scope

The principles are based on public information; the numbers, cases and tasks are the teaching design of this website. Offline experiments verify the range noted on this page, and the learning effect still needs to be judged through independent tasks and feedback.