Understand first
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Understand → Implement → Debug → Design
Differentiate structure, provenance, policy fields, and free text support with the same task set, retaining failures and comparing old and new versions.
Knowledge content check2026-10-04 · Check the source of the original question2026-10-04
It is recommended to understand first:
Your first model call and response contract →Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Prepare to write the principles into code
Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.
Reading implementation and trade-offs →Need to handle failures and changes in conditions
Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.
Continue to delve deeper into the problem →Need to design or review plans
Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.
Analyze engineering scenarios →LEARN · PRACTICE · REFLECT
First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.
Can be practiced directly. After logging in, answers, favorites, and notes will be saved to your account.
Log in and saveEach modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.
Core concept · Prompt iteration, evidence support, and independent acceptance checks
Preparatory concepts:Classify model responses, structured object, Evidence and reference labels
Prompts express task constraints, and structure outputs stabilize local shapes; business results are still subject to permitting evidence, independent labeling, semantic assertions, and enforcement fact-checking. Version improvements must be compared on fixed and independent tasks.
Type checks do not prove business facts. Prompts specify allowed sources, exception rules, and refusal conditions; schemas specify fields and enums; executors check authorization and ledgers. A battery exception to a 30-day return rule requires combining sources and applying the relevant exception.
Save the document and actual context. Check whether the exception entered the request, then inspect structure, citations, decisions, and deadlines. Complete evidence with a wrong decision points toward instructions or reasoning; missing evidence points toward retrieval or assembly. A claim that a refund occurred additionally requires approval and a business receipt.
Give expected labels only to the evaluator. Compare candidates on fixed tasks and criteria, preserving failures, valid refusals, excessive refusals, and ungraded cases. The worked example illustrates the procedure; actual results require model runs and an independent held-out set.
Define inputs, permitted evidence, candidate structure, and independent acceptance. Compare prompts on fixed normal, exceptional, and missing-evidence tasks. Schemas check shape; citations check sources and access; policy labels check decisions. Evaluate text and actual actions separately. Retain candidates, versions, error classes, and all run costs. Use untouched held-out tasks to judge improvement and regression.
The fictional return policy has the normal 30-day rule and a 7-day exception for damaged batteries. The prompt needs to describe allowed sources, exception application, output fields and missing evidence paths. Candidates include decisions, deadlines, explanations and citations. The current task delivers a policy explanation; refunds or shipping actions still require business approval and independent ledgers.
Structural checks for fields, types and enumerations, citations to check whether the source appears in the actual context and the current authorization and version, necessary evidence set to explain whether the answer retains general rules and exceptions, independent labeling to check marked decisions and deadlines. Every claim of a natural language explanation still needs to be supported by a source fragment. For output shape stability, please refer to Structured Outputs contract, and for action candidates and execution boundaries, please refer to Function calling.
Fixed tasks, documents, models, tools and graders, comparing old and new candidates and failure categories, with separate columns for correct rejections and excessive rejections. Tuning sets and hold-out sets are differentiated by source and business conditions; the cost, delays, and failures of real multiple runs are all accounted for. After changing the grader or policy label, re-run the two versions to explain the change in judgment definition.
The experiment on this site actually implements grader. One author-written v1 candidate and three author-written v2 candidates passed; these results only demonstrate the checking procedure. The real prompt effect requires actually calling the model and then saving the candidate. If the schema-valid object writes "Refund has been executed", the existing free text score shows not_scored, and semantics and business evidence must be provided. This boundary should be part of the release acceptance.
Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.
Level 1If the prompt adds "Answer based on materials only", why is independent acceptance still required?
The verification responsibilities that the program should bear are derived from the role of prompts.
Instructions express goals, but models can still misuse rules, omit exceptions, or make unfounded claims. The application first checks structure and allowed sources, then checks independent policy labels and necessary evidence, and finally reviews free text. For example, the battery issue only refers to the ordinary 30-day rule, and the JSON is legal but still does not meet the task contract. Only by retaining the real candidates and failure categories can we judge which section of the problem is solved by the prompt modification.
Level 1The reference IDs all exist, how to determine whether there are any missing product exceptions in the answer?
Change the acceptance conditions from the existence of a single source to the joint support of multiple sources.
Existence checks only confirm that the reference comes from an allowed collection. The necessary set of evidence for each question is first annotated by credible policy, and battery questions require both general rules and battery exceptions; candidates are then checked to see if they refer to them in actual context, and decisions and deadlines are compared. The free text also needs to be compared with the source fragments item by item to avoid misinterpretations despite citations.
Follow this answer further
Level 2Both necessary sources have entered the context, and the answer is still written for 30 days. Should I continue to revise the search?
Following the well-documented premise, continue to separate production errors from retrieval errors.
Save the actual request and candidate first, and confirm that the exceptions and applicable conditions are completely retained. If these are true, the recall has passed this check. The next step is to look at prompts, model inference and explanation support, and then use the same sample for regression after modification. Don't combine search metrics and answer correctness into a non-localizable score.
Follow this answer further
Level 3The decision and deadline fields are correct, but the text says the refund has been carried out, what more evidence is needed?
Proceed from correctness in policy fields to independent fact boundaries of external effects.
List refunds as independent business effect claims, and query the actual approval and refund ledgers, business operation IDs and receipts. Labels and references to policy explanations cannot infer the fact that refunds are implemented. The current teaching grader marks the free text as ungraded. The candidate should be subject to semantic and business review, and corresponding regression samples should be added.
Level 1More new tips from the same batch of samples pass through. How to avoid overfitting to the samples?
Put local improvements back into version comparisons with independent evidence.
Separate samples used for prompt tuning from frozen held-out samples and group them by document, task source, and business criteria to avoid near duplication across collections. Fixed model, data and grader, save all candidates and failures, group observations of normal, exception, lack of evidence and unauthorized actions. Repeat runs and costs of real models are also recorded to determine whether new versions meet predefined release thresholds.
First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.
Changing conditions:Common returns windows are migrated from the old version to the new version, and the old tasks still retain historical context.
Extended question:How to differentiate between changes in policy and changes in prompt quality when the total score of new tips changes?
Bind document, tag, candidate and grader versions, first checking which valid policy should be used by the task. Rerun both prompts with the same valid version and explain the differences. Old task recovery rechecks the current authorization and version, and the old policy results cannot be delivered just because the historical prompt has been passed.
The principles that remain unchanged:A comparison requires fixed, or explicitly explained, evidence and evaluation conditions.
Changing conditions:The decision and deadline match the label, but the explanation claims that a real refund has occurred.
Extended question:Can a grader's mechanical pass bring business to an end?
Treat refund execution as an independent claim, checking trusted approval objects, operation IDs and ledger receipts. Policy fields only support limited answer contracts, free text semantics and external actions still require their own acceptance. This failure is retained for regression, recording ranges that are not graded by existing graders.
The principles that remain unchanged:Each completion judgment relies on credible evidence to match it.
According to the official interface contract and the teaching program design that has been run on this site, business policies and candidates are fictitious samples. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.
After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.
Run prompt_iteration.py to construct incorrect free-form text in unknown references, missing exceptions, and schema-valid objects, delivering itemized acceptance and unscored explanations.