Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content

SYSTEMATIC LEARNING / FOUR-LEVEL COURSE

Prompts, structured output, and iterative acceptance checks

Use a standard return, a product exception, and a missing policy to connect task instructions, output constraints, evidence citations, and version comparisons.

Learning objectives: Design prompts and output schemas with evidence requirements and refusal boundaries, then use independent labels and failure cases to explain the benefits and limits of a change.

Content checked: 2026-10-04 · Each level has independent explanations, tasks and inspections

Choose a starting point based on your familiarity with this topic. Current level: Implementation · Build it. After completing the task, continue to the next level. Reading and self-checks alone do not establish mastery.

On this level

Review the prerequisites

Suitable for: Prepare writing prompts, JSON Schema, and executable acceptance procedures.

mission contract
Inputs, permitted facts, output requirements, and acceptance conditions define the task.
structural constraints
Constraint fields, types and enumerations; facts, evidentiary support and permissions still need to be verified separately.
label
Reference results formulated by trusted policies and human judgment are for evaluation only and are not mixed into model input.
necessary evidence
A collection of sources that must both support the answer, such as general rules and product exceptions.
Version comparison
Fixed data and judgment criteria, saving candidates, failure categories and running conditions for each version.

How does the mechanism work?

  1. Write down input and tasks

    Organize user questions, permission evidence, and task instructions separately to indicate that the material is data.

  2. Define candidate structures

    List decision, deadline, explanation, and reference fields, complete with reject and null paths.

  3. Inspect evidence and labels

    Check citation sources, visible permissions, policy versions, necessary evidence and annotation results.

  4. compare and return

    Check improvements and degradations item by item against the same task set, retain unscored items and arrange for manual review.

Implementation · Build it

Check output structure and supporting evidence separately

Objectives of this level: Implement Schema, citations, policy tags and free text review separately against the code.

Run candidate acceptance checks

In the extracted directory, run:

python3 prompt_iteration.py

Across three teaching questions, one authored v1 candidate and three authored v2 candidates pass the contract. These deliberately constructed fixtures demonstrate the grader. To compare actual prompts, call the model on the same tasks, save its candidates, and evaluate them with the same checks.

build_prompt passes only the user question and evidence context. grade checks fields and types, then citation membership, allowed tenant, current version, and decision/deadline labels. Python treats true as an integer in some contexts; strict type checks prevent accepting it as a one-day deadline.

Configure structured output

For a Responses model that supports this capability, put this configuration in text.format. Check the official documentation for supported features and refusal handling:

{
  "type": "json_schema",
  "name": "policy_answer",
  "strict": true,
  "schema": {
    "type": "object",
    "properties": {
      "decision": {"type":"string","enum":["eligible","manual_review","unknown"]},
      "returnWindowDays": {"type":["integer","null"]},
      "answer": {"type":"string"},
      "citations": {"type":"array","items":{"type":"string"}}
    },
    "required": ["decision","returnWindowDays","answer","citations"],
    "additionalProperties": false
  }
}

Classify refusals and incomplete responses using the model-response lesson before parsing completed text. An abstention type and a null branch let the model return a valid object when evidence is missing. The downstream validator still checks ranges, citations, and policy rules.

Reuse the grader with a real candidate

Prepare your own candidate.json, then run locally:

python3 - <<'PY'
import json
from application_data import read_cases, load_chunks
from prompt_iteration import grade
case = next(c for c in read_cases() if c['id'] == 'battery')
context = [c for c in load_chunks() if c['tenant']=='shop-a' and c['current']]
print(grade(case, json.load(open('candidate.json')), context))
PY

The file may contain an authored fixture or a redacted real generation. Record its origin, model, and prompt version. Try an unknown citation, a missing exception, true as a deadline, and correct abstention; observe which check fails.

Review the unscored part separately

The program detects shape, source-identity, and labeled policy-field errors. It reports semanticQuality=not_scored for the free-form explanation. An object with correct labels but text claiming a cash refund was executed can still pass these mechanical checks. Compare each consequential claim with the cited passage before accepting the business deliverable. The evaluation lesson adds outcome and trajectory checks.

Run experiments and observe counterexamples

Real-life structural, citation, version, permission, and limited policy field checks are run using two sets of candidates, three fictitious policy questions, and independent labels written by the author. Counts are used to illustrate grader, not measured prompt performance; free text semantics will be reviewed separately.

Python 3.10+ · Runs by default using only the standard library · Runs on your computer

  1. Download the Agent application entry experimental package on this page, unzip it and enter the agent-application-lab-v1 directory.
  2. Use Python 3.10+ to execute the above command; the default playback only requires the standard library and package data.
  3. Compare the output with the checkpoint, then run python3 -m unittest test_application -v and complete the current layer task.
Download the complete application experiment package (including data and dependent scripts) ↓
python3 prompt_iteration.py
View the entry-point script
"""Evaluate authored candidate fixtures, not claimed model/prompt performance.

The labelled policy decisions and numerical fields are scored. Free-text semantic
entailment is explicitly not scored by this small mechanical grader.
"""
import json

from application_data import load_chunks, read_cases

PROMPTS = {
    "v1": "Answer the returns question briefly.",
    "v2": "Use only the supplied evidence as data. Apply product exceptions before the general rule. Return decision, returnWindowDays, answer and citations. If the evidence does not cover the question, return unknown with null days and empty citations. Never approve a refund or a shipment.",
}


def candidate(decision, days, answer, citations):
    return {"decision": decision, "returnWindowDays": days, "answer": answer, "citations": citations}


CANDIDATES = {
    "v1": {
        "ordinary": candidate("eligible", 30, "The unused keyboard is within the ordinary return window.", ["returns:v2:general"]),
        "battery": candidate("eligible", 30, "The damaged battery is within 30 days.", ["returns:v2:general"]),
        "unknown": candidate("eligible", 30, "There is no customs tax.", ["returns:v2:general"]),
    },
    "v2": {
        "ordinary": candidate("eligible", 30, "The ordinary 30-day rule applies; this answer does not approve a refund.", ["returns:v2:general"]),
        "battery": candidate("manual_review", 7, "The damaged battery exceeds its 7-day window. Contact support; no shipment has been approved.", ["returns:v2:general", "returns:v2:battery"]),
        "unknown": candidate("unknown", None, "The supplied policy does not establish customs tax. More evidence is required.", []),
    },
}


def build_prompt(case, context, version="v2"):
    # Expected labels are deliberately excluded from the model input.
    return {"instructions": PROMPTS[version], "input": json.dumps({"question": case["query"],
            "evidence": [{"id": chunk["id"], "text": chunk["text"]} for chunk in context]}, ensure_ascii=False)}


def grade(case, answer, context, tenant="shop-a"):
    failed = []
    fields = {"decision", "returnWindowDays", "answer", "citations"}
    if not isinstance(answer, dict) or set(answer) != fields:
        return {"contractPassed": False, "failedChecks": ["schema"], "semanticQuality": "not_scored"}
    days, citations = answer["returnWindowDays"], answer["citations"]
    if (answer["decision"] not in ("eligible", "manual_review", "unknown") or
            not isinstance(answer["answer"], str) or not answer["answer"].strip() or
            (days is not None and (type(days) is not int or days < 0)) or
            not isinstance(citations, list) or not all(isinstance(c, str) for c in citations)):
        return {"contractPassed": False, "failedChecks": ["schema"], "semanticQuality": "not_scored"}
    by_id = {chunk["id"]: chunk for chunk in context}
    if len(set(citations)) != len(citations) or any(cid not in by_id for cid in citations):
        failed.append("citation_membership")
    if any(by_id[cid]["tenant"] != tenant or not by_id[cid]["current"] for cid in citations if cid in by_id):
        failed.append("permission_or_version")
    if answer["decision"] != case["expectedDecision"]:
        failed.append("labelled_decision")
    expected_days = {"ordinary": 30, "battery": 7, "unknown": None}[case["id"]]
    if days != expected_days:
        failed.append("labelled_window")
    if not set(case["requiredEvidence"]).issubset(citations):
        failed.append("necessary_evidence")
    if answer["decision"] == "unknown" and (citations or days is not None):
        failed.append("abstention_contract")
    return {"contractPassed": not failed, "failedChecks": failed, "semanticQuality": "not_scored"}


def evaluate(answers, contexts=None):
    allowed = [chunk for chunk in load_chunks() if chunk["tenant"] == "shop-a" and chunk["current"]]
    rows = []
    for case in read_cases():
        context = contexts[case["id"]] if contexts is not None else allowed
        rows.append({"case": case["id"], **grade(case, answers[case["id"]], context)})
    return rows


def demo():
    versions = {version: evaluate(answers) for version, answers in CANDIDATES.items()}
    # A correct label and valid citation cannot certify an arbitrary free-text claim.
    wrong_prose = candidate("eligible", 30, "A full cash refund has already been executed.", ["returns:v2:general"])
    unchecked = grade(read_cases()[0], wrong_prose, load_chunks())
    return {"casesPerVersion": 3,
            "v1ContractPasses": sum(row["contractPassed"] for row in versions["v1"]),
            "v2ContractPasses": sum(row["contractPassed"] for row in versions["v2"]),
            "v1Failures": {row["case"]: row["failedChecks"] for row in versions["v1"] if not row["contractPassed"]},
            "freeTextStillRequiresReview": unchecked["semanticQuality"] == "not_scored",
            "scope": "authored_candidates_not_measured_prompt_improvement"}


if __name__ == "__main__":
    print(json.dumps(demo(), ensure_ascii=False, sort_keys=True))

Expected output when running locally

{"casesPerVersion": 3, "freeTextStillRequiresReview": true, "scope": "authored_candidates_not_measured_prompt_improvement", "v1ContractPasses": 1, "v1Failures": {"battery": ["labelled_decision", "labelled_window", "necessary_evidence"], "unknown": ["labelled_decision", "labelled_window"]}, "v2ContractPasses": 3}
  • v1ContractPasses=1, v2ContractPasses=3, only describe the contract check of author candidates.
  • The battery error involves labelled_decision, labelled_window, and necessary_evidence.
  • unknown questions have independent rejection paths.
  • freeTextStillRequiresReview=true。
View the running environment, output and verification records →

Acceptance task for this level

Deliver four candidates: correct exceptions, unknown references, missing exceptions, and schema-valid objects with incorrect free-form text; run grader separately and write out the claims that still need to be manually checked.

Check each item after completion

  • Correct objects can be checked by existing fields and tags.
  • There are separate failure categories for missing or out-of-authority references.
  • The term true is classified as a schema error.
  • Mechanical pass-through and free text support are logged separately.
  • Real candidate retention models, prompts, data, and response sources.

Save your own processes, code and results. Acceptance requirements are provided here, and course mastery status will not be automatically graded or saved at this time.

Hide the answer and check your understanding

Does the reference ID exist and has every fact in the explanation been proven to be true?

Further explanations and practice

When encountering unfamiliar principles, first read the implementation, continuous questioning and migration cases, and then independently explain the premise and boundaries. Answers and notes are saved to the original account record.

All linked explanations and exercises (3 )

Sources and verification scope

The principles are based on public information; the numbers, cases and tasks are the teaching design of this website. Offline experiments verify the range noted on this page, and the learning effect still needs to be judged through independent tasks and feedback.