Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content

SYSTEMATIC LEARNING / FOUR-LEVEL COURSE

Prompts, structured output, and iterative acceptance checks

Use a standard return, a product exception, and a missing policy to connect task instructions, output constraints, evidence citations, and version comparisons.

Learning objectives: Design prompts and output schemas with evidence requirements and refusal boundaries, then use independent labels and failure cases to explain the benefits and limits of a change.

Content checked: 2026-10-04 · Each level has independent explanations, tasks and inspections

Choose a starting point based on your familiarity with this topic. Current level: Debugging · Diagnose failures. After completing the task, continue to the next level. Reading and self-checks alone do not establish mastery.

On this level

Review the prerequisites

Suitable for: An acceptance procedure has been written, but candidates may still have errors or the comparison results are unstable.

mission contract
Inputs, permitted facts, output requirements, and acceptance conditions define the task.
structural constraints
Constraint fields, types and enumerations; facts, evidentiary support and permissions still need to be verified separately.
label
Reference results formulated by trusted policies and human judgment are for evaluation only and are not mixed into model input.
necessary evidence
A collection of sources that must both support the answer, such as general rules and product exceptions.
Version comparison
Fixed data and judgment criteria, saving candidates, failure categories and running conditions for each version.

How does the mechanism work?

  1. Write down input and tasks

    Organize user questions, permission evidence, and task instructions separately to indicate that the material is data.

  2. Define candidate structures

    List decision, deadline, explanation, and reference fields, complete with reject and null paths.

  3. Inspect evidence and labels

    Check citation sources, visible permissions, policy versions, necessary evidence and annotation results.

  4. compare and return

    Check improvements and degradations item by item against the same task set, retain unscored items and arrange for manual review.

Debugging · Diagnose failures

Diagnose prompts, retrieval, and judging criteria separately

Objectives of this level: Use minimal failure samples to locate material, prompt, structure, and grader issues to avoid attributing all errors to prompts.

Trace failed checks back to the input

If a battery answer omits the exception citation, save the context actually sent. If the exception was absent, inspect retrieval, chunking, and assembly first. If it was supplied but the ordinary rule was applied, inspect the prompt and reasoning. Schema failures belong to a different boundary. A useful record includes the request, actual evidence, candidate, versions, and grader results.

Remove evidence deliberately

python3 - <<'PY'
from application_data import read_cases, load_chunks
from prompt_iteration import CANDIDATES, grade
case = next(c for c in read_cases() if c['id']=='battery')
context = [c for c in load_chunks() if c['id']=='returns:v2:general']
print(grade(case, CANDIDATES['v2']['battery'], context))
PY

The candidate cites the battery exception, but context contains only the ordinary rule, so citation membership fails. This independently tests the evidence path. A longer prompt cannot authorize a source that was never supplied.

Separate correct abstention from excessive abstention

The customs question has no tax evidence, so unknown meets the goal. The ordinary-keyboard question has an authorized return policy, so unknown loses useful coverage. Count stopping for missing evidence separately from refusing an answer supported by evidence. An overall pass rate can hide changes in exception, permission, or missing-evidence cases.

Changing the standard changes the score

After changing the grader, policy version, or required source set, inspect the differences and rerun both prompt versions. An old candidate citing v1 may fail a new task set without any change in model capability. Version candidates, labels, documents, prompts, models, and graders. Hold other conditions fixed and compare by task group.

Test new prompts on held-out cases

The three examples explain the code but cannot determine a production release. Add item types, time boundaries, paraphrases, conflicting evidence, and permission revocations. Separate prompt-tuning cases from held-out evaluation. Retain failed requests and costs when comparing repeated runs; selecting only successful answers biases the result.

Run experiments and observe counterexamples

Real-life structural, citation, version, permission, and limited policy field checks are run using two sets of candidates, three fictitious policy questions, and independent labels written by the author. Counts are used to illustrate grader, not measured prompt performance; free text semantics will be reviewed separately.

Python 3.10+ · Runs by default using only the standard library · Runs on your computer

  1. Download the Agent application entry experimental package on this page, unzip it and enter the agent-application-lab-v1 directory.
  2. Use Python 3.10+ to execute the above command; the default playback only requires the standard library and package data.
  3. Compare the output with the checkpoint, then run python3 -m unittest test_application -v and complete the current layer task.
Download the complete application experiment package (including data and dependent scripts) ↓
python3 prompt_iteration.py
View the entry-point script
"""Evaluate authored candidate fixtures, not claimed model/prompt performance.

The labelled policy decisions and numerical fields are scored. Free-text semantic
entailment is explicitly not scored by this small mechanical grader.
"""
import json

from application_data import load_chunks, read_cases

PROMPTS = {
    "v1": "Answer the returns question briefly.",
    "v2": "Use only the supplied evidence as data. Apply product exceptions before the general rule. Return decision, returnWindowDays, answer and citations. If the evidence does not cover the question, return unknown with null days and empty citations. Never approve a refund or a shipment.",
}


def candidate(decision, days, answer, citations):
    return {"decision": decision, "returnWindowDays": days, "answer": answer, "citations": citations}


CANDIDATES = {
    "v1": {
        "ordinary": candidate("eligible", 30, "The unused keyboard is within the ordinary return window.", ["returns:v2:general"]),
        "battery": candidate("eligible", 30, "The damaged battery is within 30 days.", ["returns:v2:general"]),
        "unknown": candidate("eligible", 30, "There is no customs tax.", ["returns:v2:general"]),
    },
    "v2": {
        "ordinary": candidate("eligible", 30, "The ordinary 30-day rule applies; this answer does not approve a refund.", ["returns:v2:general"]),
        "battery": candidate("manual_review", 7, "The damaged battery exceeds its 7-day window. Contact support; no shipment has been approved.", ["returns:v2:general", "returns:v2:battery"]),
        "unknown": candidate("unknown", None, "The supplied policy does not establish customs tax. More evidence is required.", []),
    },
}


def build_prompt(case, context, version="v2"):
    # Expected labels are deliberately excluded from the model input.
    return {"instructions": PROMPTS[version], "input": json.dumps({"question": case["query"],
            "evidence": [{"id": chunk["id"], "text": chunk["text"]} for chunk in context]}, ensure_ascii=False)}


def grade(case, answer, context, tenant="shop-a"):
    failed = []
    fields = {"decision", "returnWindowDays", "answer", "citations"}
    if not isinstance(answer, dict) or set(answer) != fields:
        return {"contractPassed": False, "failedChecks": ["schema"], "semanticQuality": "not_scored"}
    days, citations = answer["returnWindowDays"], answer["citations"]
    if (answer["decision"] not in ("eligible", "manual_review", "unknown") or
            not isinstance(answer["answer"], str) or not answer["answer"].strip() or
            (days is not None and (type(days) is not int or days < 0)) or
            not isinstance(citations, list) or not all(isinstance(c, str) for c in citations)):
        return {"contractPassed": False, "failedChecks": ["schema"], "semanticQuality": "not_scored"}
    by_id = {chunk["id"]: chunk for chunk in context}
    if len(set(citations)) != len(citations) or any(cid not in by_id for cid in citations):
        failed.append("citation_membership")
    if any(by_id[cid]["tenant"] != tenant or not by_id[cid]["current"] for cid in citations if cid in by_id):
        failed.append("permission_or_version")
    if answer["decision"] != case["expectedDecision"]:
        failed.append("labelled_decision")
    expected_days = {"ordinary": 30, "battery": 7, "unknown": None}[case["id"]]
    if days != expected_days:
        failed.append("labelled_window")
    if not set(case["requiredEvidence"]).issubset(citations):
        failed.append("necessary_evidence")
    if answer["decision"] == "unknown" and (citations or days is not None):
        failed.append("abstention_contract")
    return {"contractPassed": not failed, "failedChecks": failed, "semanticQuality": "not_scored"}


def evaluate(answers, contexts=None):
    allowed = [chunk for chunk in load_chunks() if chunk["tenant"] == "shop-a" and chunk["current"]]
    rows = []
    for case in read_cases():
        context = contexts[case["id"]] if contexts is not None else allowed
        rows.append({"case": case["id"], **grade(case, answers[case["id"]], context)})
    return rows


def demo():
    versions = {version: evaluate(answers) for version, answers in CANDIDATES.items()}
    # A correct label and valid citation cannot certify an arbitrary free-text claim.
    wrong_prose = candidate("eligible", 30, "A full cash refund has already been executed.", ["returns:v2:general"])
    unchecked = grade(read_cases()[0], wrong_prose, load_chunks())
    return {"casesPerVersion": 3,
            "v1ContractPasses": sum(row["contractPassed"] for row in versions["v1"]),
            "v2ContractPasses": sum(row["contractPassed"] for row in versions["v2"]),
            "v1Failures": {row["case"]: row["failedChecks"] for row in versions["v1"] if not row["contractPassed"]},
            "freeTextStillRequiresReview": unchecked["semanticQuality"] == "not_scored",
            "scope": "authored_candidates_not_measured_prompt_improvement"}


if __name__ == "__main__":
    print(json.dumps(demo(), ensure_ascii=False, sort_keys=True))

Expected output when running locally

{"casesPerVersion": 3, "freeTextStillRequiresReview": true, "scope": "authored_candidates_not_measured_prompt_improvement", "v1ContractPasses": 1, "v1Failures": {"battery": ["labelled_decision", "labelled_window", "necessary_evidence"], "unknown": ["labelled_decision", "labelled_window"]}, "v2ContractPasses": 3}
  • v1ContractPasses=1, v2ContractPasses=3, only describe the contract check of author candidates.
  • The battery error involves labelled_decision, labelled_window, and necessary_evidence.
  • unknown questions have independent rejection paths.
  • freeTextStillRequiresReview=true。
View the running environment, output and verification records →

Acceptance task for this level

For the battery questions, we construct "exceptions have not entered the context" and "exceptions have been provided but candidate misuse rules" respectively to deliver different positioning; we also add a reserved sample that has evidence but incorrectly refused to answer.

Check each item after completion

  • In fact the following can be reviewed separately.
  • Retrieval losses and generated misuses are located separately.
  • Correct refusals due to lack of evidence and excessive refusals will be counted separately.
  • Label differences caused by version changes are explained.

Save your own processes, code and results. Acceptance requirements are provided here, and course mastery status will not be automatically graded or saved at this time.

Hide the answer and check your understanding

After the upgrade prompt, the total score is improved, but the permission questions are degraded. Can I publish it directly?

Further explanations and practice

When encountering unfamiliar principles, first read the implementation, continuous questioning and migration cases, and then independently explain the premise and boundaries. Answers and notes are saved to the original account record.

All linked explanations and exercises (3 )

Sources and verification scope

The principles are based on public information; the numbers, cases and tasks are the teaching design of this website. Offline experiments verify the range noted on this page, and the learning effect still needs to be judged through independent tasks and feedback.