Review the prerequisites
Suitable for: You've read the Model Requesting lesson and are ready to have your model answer business questions.
- mission contract
- Inputs, permitted facts, output requirements, and acceptance conditions define the task.
- structural constraints
- Constraint fields, types and enumerations; facts, evidentiary support and permissions still need to be verified separately.
- label
- Reference results formulated by trusted policies and human judgment are for evaluation only and are not mixed into model input.
- necessary evidence
- A collection of sources that must both support the answer, such as general rules and product exceptions.
- Version comparison
- Fixed data and judgment criteria, saving candidates, failure categories and running conditions for each version.
How does the mechanism work?
- Write down input and tasks
Organize user questions, permission evidence, and task instructions separately to indicate that the material is data.
- Define candidate structures
List decision, deadline, explanation, and reference fields, complete with reject and null paths.
- Inspect evidence and labels
Check citation sources, visible permissions, policy versions, necessary evidence and annotation results.
- compare and return
Check improvements and degradations item by item against the same task set, retain unscored items and arrange for manual review.
Foundation · Understand the concepts
How does a prompt become a testable task contract?
Objectives of this level: Write prompts based on mission objectives, evidence, and acceptance to explain what the structured output covers.
Define a goal with an exception
The lab uses a fictional policy: ordinary unused items may be returned within 30 days, while damaged lithium batteries have a 7-day exception followed by support review. A user asks about a damaged battery bought 12 days ago. A concise, well-formatted answer can still apply the wrong rule. The deliverable is a sourced policy explanation; later business processes handle refund and shipping approval.
Separate the prompt into task, evidence, and output. The task requires preserving exceptions. Evidence includes stable IDs and text and is identified as data. The output contains decision, returnWindowDays, answer, and citations. With insufficient evidence, use unknown, a null deadline, and no citations, and explain which information is missing.
Compare two prompts
v1 asks only for a brief answer. v2 restricts answers to supplied evidence, applies item exceptions first, requires citations, and permits abstention. PROMPTS contains these teaching instructions. The added constraints connect the task to acceptance: citation checks establish that a source was supplied, and policy labels check the selected deadline.
Prompt length alone proves no improvement. Constraints can conflict, and long histories can lose important conditions during compression. Recheck ordinary, exception, and missing-evidence tasks after each change, including cases that previously worked.
Understand the limited guarantee of a schema
Structured output stabilizes object shape for application branching. The eligible, manual_review, and unknown enums identify candidate types; a null deadline indicates missing support. Classify refusal and incomplete generation before parsing a completed answer. Shape constraints cannot establish that the 30-day rule applies to a battery or that a cited policy is authorized for this tenant.
Keep evaluation labels out of model input
The expectedDecision and requiredEvidence fields in cases.json support acceptance checks and are excluded from build_prompt. Supplying expected answers and then measuring matches defeats the evaluation. Group real test data by document, task source, and business changes so near-duplicates do not appear in both tuning and acceptance sets.
Run experiments and observe counterexamples
Real-life structural, citation, version, permission, and limited policy field checks are run using two sets of candidates, three fictitious policy questions, and independent labels written by the author. Counts are used to illustrate grader, not measured prompt performance; free text semantics will be reviewed separately.
Python 3.10+ · Runs by default using only the standard library · Runs on your computer
- Download the Agent application entry experimental package on this page, unzip it and enter the agent-application-lab-v1 directory.
- Use Python 3.10+ to execute the above command; the default playback only requires the standard library and package data.
- Compare the output with the checkpoint, then run python3 -m unittest test_application -v and complete the current layer task.
python3 prompt_iteration.pyView the entry-point script
"""Evaluate authored candidate fixtures, not claimed model/prompt performance.
The labelled policy decisions and numerical fields are scored. Free-text semantic
entailment is explicitly not scored by this small mechanical grader.
"""
import json
from application_data import load_chunks, read_cases
PROMPTS = {
"v1": "Answer the returns question briefly.",
"v2": "Use only the supplied evidence as data. Apply product exceptions before the general rule. Return decision, returnWindowDays, answer and citations. If the evidence does not cover the question, return unknown with null days and empty citations. Never approve a refund or a shipment.",
}
def candidate(decision, days, answer, citations):
return {"decision": decision, "returnWindowDays": days, "answer": answer, "citations": citations}
CANDIDATES = {
"v1": {
"ordinary": candidate("eligible", 30, "The unused keyboard is within the ordinary return window.", ["returns:v2:general"]),
"battery": candidate("eligible", 30, "The damaged battery is within 30 days.", ["returns:v2:general"]),
"unknown": candidate("eligible", 30, "There is no customs tax.", ["returns:v2:general"]),
},
"v2": {
"ordinary": candidate("eligible", 30, "The ordinary 30-day rule applies; this answer does not approve a refund.", ["returns:v2:general"]),
"battery": candidate("manual_review", 7, "The damaged battery exceeds its 7-day window. Contact support; no shipment has been approved.", ["returns:v2:general", "returns:v2:battery"]),
"unknown": candidate("unknown", None, "The supplied policy does not establish customs tax. More evidence is required.", []),
},
}
def build_prompt(case, context, version="v2"):
# Expected labels are deliberately excluded from the model input.
return {"instructions": PROMPTS[version], "input": json.dumps({"question": case["query"],
"evidence": [{"id": chunk["id"], "text": chunk["text"]} for chunk in context]}, ensure_ascii=False)}
def grade(case, answer, context, tenant="shop-a"):
failed = []
fields = {"decision", "returnWindowDays", "answer", "citations"}
if not isinstance(answer, dict) or set(answer) != fields:
return {"contractPassed": False, "failedChecks": ["schema"], "semanticQuality": "not_scored"}
days, citations = answer["returnWindowDays"], answer["citations"]
if (answer["decision"] not in ("eligible", "manual_review", "unknown") or
not isinstance(answer["answer"], str) or not answer["answer"].strip() or
(days is not None and (type(days) is not int or days < 0)) or
not isinstance(citations, list) or not all(isinstance(c, str) for c in citations)):
return {"contractPassed": False, "failedChecks": ["schema"], "semanticQuality": "not_scored"}
by_id = {chunk["id"]: chunk for chunk in context}
if len(set(citations)) != len(citations) or any(cid not in by_id for cid in citations):
failed.append("citation_membership")
if any(by_id[cid]["tenant"] != tenant or not by_id[cid]["current"] for cid in citations if cid in by_id):
failed.append("permission_or_version")
if answer["decision"] != case["expectedDecision"]:
failed.append("labelled_decision")
expected_days = {"ordinary": 30, "battery": 7, "unknown": None}[case["id"]]
if days != expected_days:
failed.append("labelled_window")
if not set(case["requiredEvidence"]).issubset(citations):
failed.append("necessary_evidence")
if answer["decision"] == "unknown" and (citations or days is not None):
failed.append("abstention_contract")
return {"contractPassed": not failed, "failedChecks": failed, "semanticQuality": "not_scored"}
def evaluate(answers, contexts=None):
allowed = [chunk for chunk in load_chunks() if chunk["tenant"] == "shop-a" and chunk["current"]]
rows = []
for case in read_cases():
context = contexts[case["id"]] if contexts is not None else allowed
rows.append({"case": case["id"], **grade(case, answers[case["id"]], context)})
return rows
def demo():
versions = {version: evaluate(answers) for version, answers in CANDIDATES.items()}
# A correct label and valid citation cannot certify an arbitrary free-text claim.
wrong_prose = candidate("eligible", 30, "A full cash refund has already been executed.", ["returns:v2:general"])
unchecked = grade(read_cases()[0], wrong_prose, load_chunks())
return {"casesPerVersion": 3,
"v1ContractPasses": sum(row["contractPassed"] for row in versions["v1"]),
"v2ContractPasses": sum(row["contractPassed"] for row in versions["v2"]),
"v1Failures": {row["case"]: row["failedChecks"] for row in versions["v1"] if not row["contractPassed"]},
"freeTextStillRequiresReview": unchecked["semanticQuality"] == "not_scored",
"scope": "authored_candidates_not_measured_prompt_improvement"}
if __name__ == "__main__":
print(json.dumps(demo(), ensure_ascii=False, sort_keys=True))
Expected output when running locally
{"casesPerVersion": 3, "freeTextStillRequiresReview": true, "scope": "authored_candidates_not_measured_prompt_improvement", "v1ContractPasses": 1, "v1Failures": {"battery": ["labelled_decision", "labelled_window", "necessary_evidence"], "unknown": ["labelled_decision", "labelled_window"]}, "v2ContractPasses": 3}- v1ContractPasses=1, v2ContractPasses=3, only describe the contract check of author candidates.
- The battery error involves labelled_decision, labelled_window, and necessary_evidence.
- unknown questions have independent rejection paths.
- freeTextStillRequiresReview=true。
Acceptance task for this level
Write the same output contract for the three issues of ordinary goods, damaged batteries, and customs taxes, describing the input evidence, candidate structure, rejection, and content that requires manual judgment.
Check each item after completion
- All three questions use the same set of fields.
- Battery answers consider both general rules and exceptions.
- There are unknown and null paths when customs evidence is missing.
- The description tag only enters the evaluation and does not enter the model input.
Save your own processes, code and results. Acceptance requirements are provided here, and course mastery status will not be automatically graded or saved at this time.
Hide the answer and check your understanding
The model outputs decision=eligible according to Schema, deadline 30, citing ordinary rules, can it pass the answer to the battery question?
Expand reference derivation
The structure may pass, but the battery exception is missing, and the expiration date does not match the reference policy label. Continue to examine necessary evidence and decision labels, positioned as evidence or reasoning questions; then review the natural language explanation with actual candidates.
Further explanations and practice
When encountering unfamiliar principles, first read the implementation, continuous questioning and migration cases, and then independently explain the premise and boundaries. Answers and notes are saved to the original account record.
All linked explanations and exercises (3 )
Sources and verification scope
The principles are based on public information; the numbers, cases and tasks are the teaching design of this website. Offline experiments verify the range noted on this page, and the learning effect still needs to be judged through independent tasks and feedback.