Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content

SYSTEMATIC LEARNING / FOUR-LEVEL COURSE

Prompts, structured output, and iterative acceptance checks

Use a standard return, a product exception, and a missing policy to connect task instructions, output constraints, evidence citations, and version comparisons.

Learning objectives: Design prompts and output schemas with evidence requirements and refusal boundaries, then use independent labels and failure cases to explain the benefits and limits of a change.

Content checked: 2026-10-04 · Each level has independent explanations, tasks and inspections

Choose a starting point based on your familiarity with this topic. Current level: Foundation · Understand the concepts. After completing the task, continue to the next level. Reading and self-checks alone do not establish mastery.

On this level

Review the prerequisites

Suitable for: You've read the Model Requesting lesson and are ready to have your model answer business questions.

mission contract
Inputs, permitted facts, output requirements, and acceptance conditions define the task.
structural constraints
Constraint fields, types and enumerations; facts, evidentiary support and permissions still need to be verified separately.
label
Reference results formulated by trusted policies and human judgment are for evaluation only and are not mixed into model input.
necessary evidence
A collection of sources that must both support the answer, such as general rules and product exceptions.
Version comparison
Fixed data and judgment criteria, saving candidates, failure categories and running conditions for each version.

How does the mechanism work?

  1. Write down input and tasks

    Organize user questions, permission evidence, and task instructions separately to indicate that the material is data.

  2. Define candidate structures

    List decision, deadline, explanation, and reference fields, complete with reject and null paths.

  3. Inspect evidence and labels

    Check citation sources, visible permissions, policy versions, necessary evidence and annotation results.

  4. compare and return

    Check improvements and degradations item by item against the same task set, retain unscored items and arrange for manual review.

Foundation · Understand the concepts

How does a prompt become a testable task contract?

Objectives of this level: Write prompts based on mission objectives, evidence, and acceptance to explain what the structured output covers.

Define a goal with an exception

The lab uses a fictional policy: ordinary unused items may be returned within 30 days, while damaged lithium batteries have a 7-day exception followed by support review. A user asks about a damaged battery bought 12 days ago. A concise, well-formatted answer can still apply the wrong rule. The deliverable is a sourced policy explanation; later business processes handle refund and shipping approval.

Separate the prompt into task, evidence, and output. The task requires preserving exceptions. Evidence includes stable IDs and text and is identified as data. The output contains decision, returnWindowDays, answer, and citations. With insufficient evidence, use unknown, a null deadline, and no citations, and explain which information is missing.

Compare two prompts

v1 asks only for a brief answer. v2 restricts answers to supplied evidence, applies item exceptions first, requires citations, and permits abstention. PROMPTS contains these teaching instructions. The added constraints connect the task to acceptance: citation checks establish that a source was supplied, and policy labels check the selected deadline.

Prompt length alone proves no improvement. Constraints can conflict, and long histories can lose important conditions during compression. Recheck ordinary, exception, and missing-evidence tasks after each change, including cases that previously worked.

Understand the limited guarantee of a schema

Structured output stabilizes object shape for application branching. The eligible, manual_review, and unknown enums identify candidate types; a null deadline indicates missing support. Classify refusal and incomplete generation before parsing a completed answer. Shape constraints cannot establish that the 30-day rule applies to a battery or that a cited policy is authorized for this tenant.

Keep evaluation labels out of model input

The expectedDecision and requiredEvidence fields in cases.json support acceptance checks and are excluded from build_prompt. Supplying expected answers and then measuring matches defeats the evaluation. Group real test data by document, task source, and business changes so near-duplicates do not appear in both tuning and acceptance sets.

Run experiments and observe counterexamples

Real-life structural, citation, version, permission, and limited policy field checks are run using two sets of candidates, three fictitious policy questions, and independent labels written by the author. Counts are used to illustrate grader, not measured prompt performance; free text semantics will be reviewed separately.

Python 3.10+ · Runs by default using only the standard library · Runs on your computer

  1. Download the Agent application entry experimental package on this page, unzip it and enter the agent-application-lab-v1 directory.
  2. Use Python 3.10+ to execute the above command; the default playback only requires the standard library and package data.
  3. Compare the output with the checkpoint, then run python3 -m unittest test_application -v and complete the current layer task.
Download the complete application experiment package (including data and dependent scripts) ↓
python3 prompt_iteration.py
View the entry-point script
"""Evaluate authored candidate fixtures, not claimed model/prompt performance.

The labelled policy decisions and numerical fields are scored. Free-text semantic
entailment is explicitly not scored by this small mechanical grader.
"""
import json

from application_data import load_chunks, read_cases

PROMPTS = {
    "v1": "Answer the returns question briefly.",
    "v2": "Use only the supplied evidence as data. Apply product exceptions before the general rule. Return decision, returnWindowDays, answer and citations. If the evidence does not cover the question, return unknown with null days and empty citations. Never approve a refund or a shipment.",
}


def candidate(decision, days, answer, citations):
    return {"decision": decision, "returnWindowDays": days, "answer": answer, "citations": citations}


CANDIDATES = {
    "v1": {
        "ordinary": candidate("eligible", 30, "The unused keyboard is within the ordinary return window.", ["returns:v2:general"]),
        "battery": candidate("eligible", 30, "The damaged battery is within 30 days.", ["returns:v2:general"]),
        "unknown": candidate("eligible", 30, "There is no customs tax.", ["returns:v2:general"]),
    },
    "v2": {
        "ordinary": candidate("eligible", 30, "The ordinary 30-day rule applies; this answer does not approve a refund.", ["returns:v2:general"]),
        "battery": candidate("manual_review", 7, "The damaged battery exceeds its 7-day window. Contact support; no shipment has been approved.", ["returns:v2:general", "returns:v2:battery"]),
        "unknown": candidate("unknown", None, "The supplied policy does not establish customs tax. More evidence is required.", []),
    },
}


def build_prompt(case, context, version="v2"):
    # Expected labels are deliberately excluded from the model input.
    return {"instructions": PROMPTS[version], "input": json.dumps({"question": case["query"],
            "evidence": [{"id": chunk["id"], "text": chunk["text"]} for chunk in context]}, ensure_ascii=False)}


def grade(case, answer, context, tenant="shop-a"):
    failed = []
    fields = {"decision", "returnWindowDays", "answer", "citations"}
    if not isinstance(answer, dict) or set(answer) != fields:
        return {"contractPassed": False, "failedChecks": ["schema"], "semanticQuality": "not_scored"}
    days, citations = answer["returnWindowDays"], answer["citations"]
    if (answer["decision"] not in ("eligible", "manual_review", "unknown") or
            not isinstance(answer["answer"], str) or not answer["answer"].strip() or
            (days is not None and (type(days) is not int or days < 0)) or
            not isinstance(citations, list) or not all(isinstance(c, str) for c in citations)):
        return {"contractPassed": False, "failedChecks": ["schema"], "semanticQuality": "not_scored"}
    by_id = {chunk["id"]: chunk for chunk in context}
    if len(set(citations)) != len(citations) or any(cid not in by_id for cid in citations):
        failed.append("citation_membership")
    if any(by_id[cid]["tenant"] != tenant or not by_id[cid]["current"] for cid in citations if cid in by_id):
        failed.append("permission_or_version")
    if answer["decision"] != case["expectedDecision"]:
        failed.append("labelled_decision")
    expected_days = {"ordinary": 30, "battery": 7, "unknown": None}[case["id"]]
    if days != expected_days:
        failed.append("labelled_window")
    if not set(case["requiredEvidence"]).issubset(citations):
        failed.append("necessary_evidence")
    if answer["decision"] == "unknown" and (citations or days is not None):
        failed.append("abstention_contract")
    return {"contractPassed": not failed, "failedChecks": failed, "semanticQuality": "not_scored"}


def evaluate(answers, contexts=None):
    allowed = [chunk for chunk in load_chunks() if chunk["tenant"] == "shop-a" and chunk["current"]]
    rows = []
    for case in read_cases():
        context = contexts[case["id"]] if contexts is not None else allowed
        rows.append({"case": case["id"], **grade(case, answers[case["id"]], context)})
    return rows


def demo():
    versions = {version: evaluate(answers) for version, answers in CANDIDATES.items()}
    # A correct label and valid citation cannot certify an arbitrary free-text claim.
    wrong_prose = candidate("eligible", 30, "A full cash refund has already been executed.", ["returns:v2:general"])
    unchecked = grade(read_cases()[0], wrong_prose, load_chunks())
    return {"casesPerVersion": 3,
            "v1ContractPasses": sum(row["contractPassed"] for row in versions["v1"]),
            "v2ContractPasses": sum(row["contractPassed"] for row in versions["v2"]),
            "v1Failures": {row["case"]: row["failedChecks"] for row in versions["v1"] if not row["contractPassed"]},
            "freeTextStillRequiresReview": unchecked["semanticQuality"] == "not_scored",
            "scope": "authored_candidates_not_measured_prompt_improvement"}


if __name__ == "__main__":
    print(json.dumps(demo(), ensure_ascii=False, sort_keys=True))

Expected output when running locally

{"casesPerVersion": 3, "freeTextStillRequiresReview": true, "scope": "authored_candidates_not_measured_prompt_improvement", "v1ContractPasses": 1, "v1Failures": {"battery": ["labelled_decision", "labelled_window", "necessary_evidence"], "unknown": ["labelled_decision", "labelled_window"]}, "v2ContractPasses": 3}
  • v1ContractPasses=1, v2ContractPasses=3, only describe the contract check of author candidates.
  • The battery error involves labelled_decision, labelled_window, and necessary_evidence.
  • unknown questions have independent rejection paths.
  • freeTextStillRequiresReview=true。
View the running environment, output and verification records →

Acceptance task for this level

Write the same output contract for the three issues of ordinary goods, damaged batteries, and customs taxes, describing the input evidence, candidate structure, rejection, and content that requires manual judgment.

Check each item after completion

  • All three questions use the same set of fields.
  • Battery answers consider both general rules and exceptions.
  • There are unknown and null paths when customs evidence is missing.
  • The description tag only enters the evaluation and does not enter the model input.

Save your own processes, code and results. Acceptance requirements are provided here, and course mastery status will not be automatically graded or saved at this time.

Hide the answer and check your understanding

The model outputs decision=eligible according to Schema, deadline 30, citing ordinary rules, can it pass the answer to the battery question?

Further explanations and practice

When encountering unfamiliar principles, first read the implementation, continuous questioning and migration cases, and then independently explain the premise and boundaries. Answers and notes are saved to the original account record.

All linked explanations and exercises (3 )

Sources and verification scope

The principles are based on public information; the numbers, cases and tasks are the teaching design of this website. Offline experiments verify the range noted on this page, and the learning effect still needs to be judged through independent tasks and feedback.