Agent 应用开发会员账号
知识目录选择核心方向与细分内容

SYSTEMATIC LEARNING / 分层专题

提示、结构化输出与迭代验收

用普通退货、商品例外和缺失政策三个问题,连接任务指令、结构约束、证据引用及版本比较。

学习目标:能设计有证据与拒答边界的提示和输出 Schema,并用独立标签与失败样本解释一次修改的收益和限制。

内容核对 2026-10-04 · 每层有独立讲解、任务与检查

按这个知识点的熟悉程度选择起点。当前层:进阶 · 定位故障。完成任务后可以继续下一层;阅读与自检不自动代表掌握。

本层学习目录

先补齐必要概念

适合:已写出验收程序,候选仍会出现错误或比较结果不稳定。

任务契约
输入、允许使用的事实、输出要求及验收条件共同定义任务。
结构约束
约束字段、类型与枚举;事实、证据支持和权限仍需分别验证。
标签
由可信政策和人工判断制定的参考结果,只供评测使用,不混入模型输入。
必要证据
支撑回答必须同时具备的来源集合,例如一般规则与商品例外。
版本比较
固定数据与判断标准,保存每个版本的候选、失败类别和运行条件。

原理怎样一步步成立?

  1. 写清输入与任务

    将用户问题、允许证据和任务指令分别组织,说明材料属于数据。

  2. 定义候选结构

    列出决定、期限、解释和引用字段,补足拒答与空值路径。

  3. 检查证据与标签

    核对引用来源、可见权限、政策版本、必要证据和标注结果。

  4. 比较并回归

    对照同一任务集逐项检查改善与退化,保留未评分项并安排人工复核。

进阶 · 定位故障

提示、检索和判断标准如何分开排错?

本层目标:用最小失败样本定位材料、提示、结构和裁判问题,避免把所有错误归给提示。

从失败检查项回看输入

电池候选少例外引用时,先保存本次实际上下文。如果例外从未进入请求,优先检查检索、切块与上下文组装;如果例外已提供而候选仍套用普通规则,再检查提示与推理。Schema 错误属于另一个边界。单次失败记录至少含请求、实际证据、候选、版本和 grader 结果。

直接制造证据丢失

python3 - <<'PY'
from application_data import read_cases, load_chunks
from prompt_iteration import CANDIDATES, grade
case = next(c for c in read_cases() if c['id']=='battery')
context = [c for c in load_chunks() if c['id']=='returns:v2:general']
print(grade(case, CANDIDATES['v2']['battery'], context))
PY

候选引用电池例外,但实际 context 只有普通规则,所以引用成员检查失败。这个实验使证据路径可以被独立验证;把提示写得更长不能让未提供的来源获得授权。

区分拒答和过度拒答

海关问题没有税费材料,unknown 符合目标;普通键盘问题已经有允许使用的退货政策,直接 unknown 则失去业务帮助。将“缺证据而正确停止”与“有证据却不回答”分开统计。不要靠一条整体通过率掩盖例外题、权限题或缺失证据题的变化。

标准变化也会改变分数

修改 grader、参考政策版本或必要来源集合后,先检查差异,重跑两个提示版本。旧候选引用 v1 政策可能在新题集失败,这并不自动说明模型能力退化。候选、标签、文档、提示、模型和裁判都需要版本;对照实验应固定其余条件,并按任务分组比较。

用保留样本检查新提示

已知的三道示例适合观察代码,规模不足以决定实际发布。为不同商品、时间边界、同义问题、证据冲突与权限撤销建立额外样本,区分用于改提示的集合和保留评测集合。观察重复运行差异时,也要保留失败请求和成本,避免只挑成功回答。

运行实验,观察反例

使用作者编写的两组候选、三道虚构政策问题和独立标签,实际运行结构、引用、版本、权限及有限政策字段检查。计数用于说明 grader,不是实测提示性能;自由文本语义另行复核。

Python 3.10+ · 默认运行只使用标准库 · 在你的电脑运行

  1. 下载本页的 Agent 应用入门实验包,解压后进入 agent-application-lab-v1 目录。
  2. 使用 Python 3.10+ 执行上方命令;默认回放只需标准库与包内数据。
  3. 对照输出与检查点,再运行 python3 -m unittest test_application -v,并完成当前层任务。
下载完整应用实验包(含数据与依赖脚本) ↓
python3 prompt_iteration.py
查看本入口脚本
"""Evaluate authored candidate fixtures, not claimed model/prompt performance.

The labelled policy decisions and numerical fields are scored. Free-text semantic
entailment is explicitly not scored by this small mechanical grader.
"""
import json

from application_data import load_chunks, read_cases

PROMPTS = {
    "v1": "Answer the returns question briefly.",
    "v2": "Use only the supplied evidence as data. Apply product exceptions before the general rule. Return decision, returnWindowDays, answer and citations. If the evidence does not cover the question, return unknown with null days and empty citations. Never approve a refund or a shipment.",
}


def candidate(decision, days, answer, citations):
    return {"decision": decision, "returnWindowDays": days, "answer": answer, "citations": citations}


CANDIDATES = {
    "v1": {
        "ordinary": candidate("eligible", 30, "The unused keyboard is within the ordinary return window.", ["returns:v2:general"]),
        "battery": candidate("eligible", 30, "The damaged battery is within 30 days.", ["returns:v2:general"]),
        "unknown": candidate("eligible", 30, "There is no customs tax.", ["returns:v2:general"]),
    },
    "v2": {
        "ordinary": candidate("eligible", 30, "The ordinary 30-day rule applies; this answer does not approve a refund.", ["returns:v2:general"]),
        "battery": candidate("manual_review", 7, "The damaged battery exceeds its 7-day window. Contact support; no shipment has been approved.", ["returns:v2:general", "returns:v2:battery"]),
        "unknown": candidate("unknown", None, "The supplied policy does not establish customs tax. More evidence is required.", []),
    },
}


def build_prompt(case, context, version="v2"):
    # Expected labels are deliberately excluded from the model input.
    return {"instructions": PROMPTS[version], "input": json.dumps({"question": case["query"],
            "evidence": [{"id": chunk["id"], "text": chunk["text"]} for chunk in context]}, ensure_ascii=False)}


def grade(case, answer, context, tenant="shop-a"):
    failed = []
    fields = {"decision", "returnWindowDays", "answer", "citations"}
    if not isinstance(answer, dict) or set(answer) != fields:
        return {"contractPassed": False, "failedChecks": ["schema"], "semanticQuality": "not_scored"}
    days, citations = answer["returnWindowDays"], answer["citations"]
    if (answer["decision"] not in ("eligible", "manual_review", "unknown") or
            not isinstance(answer["answer"], str) or not answer["answer"].strip() or
            (days is not None and (type(days) is not int or days < 0)) or
            not isinstance(citations, list) or not all(isinstance(c, str) for c in citations)):
        return {"contractPassed": False, "failedChecks": ["schema"], "semanticQuality": "not_scored"}
    by_id = {chunk["id"]: chunk for chunk in context}
    if len(set(citations)) != len(citations) or any(cid not in by_id for cid in citations):
        failed.append("citation_membership")
    if any(by_id[cid]["tenant"] != tenant or not by_id[cid]["current"] for cid in citations if cid in by_id):
        failed.append("permission_or_version")
    if answer["decision"] != case["expectedDecision"]:
        failed.append("labelled_decision")
    expected_days = {"ordinary": 30, "battery": 7, "unknown": None}[case["id"]]
    if days != expected_days:
        failed.append("labelled_window")
    if not set(case["requiredEvidence"]).issubset(citations):
        failed.append("necessary_evidence")
    if answer["decision"] == "unknown" and (citations or days is not None):
        failed.append("abstention_contract")
    return {"contractPassed": not failed, "failedChecks": failed, "semanticQuality": "not_scored"}


def evaluate(answers, contexts=None):
    allowed = [chunk for chunk in load_chunks() if chunk["tenant"] == "shop-a" and chunk["current"]]
    rows = []
    for case in read_cases():
        context = contexts[case["id"]] if contexts is not None else allowed
        rows.append({"case": case["id"], **grade(case, answers[case["id"]], context)})
    return rows


def demo():
    versions = {version: evaluate(answers) for version, answers in CANDIDATES.items()}
    # A correct label and valid citation cannot certify an arbitrary free-text claim.
    wrong_prose = candidate("eligible", 30, "A full cash refund has already been executed.", ["returns:v2:general"])
    unchecked = grade(read_cases()[0], wrong_prose, load_chunks())
    return {"casesPerVersion": 3,
            "v1ContractPasses": sum(row["contractPassed"] for row in versions["v1"]),
            "v2ContractPasses": sum(row["contractPassed"] for row in versions["v2"]),
            "v1Failures": {row["case"]: row["failedChecks"] for row in versions["v1"] if not row["contractPassed"]},
            "freeTextStillRequiresReview": unchecked["semanticQuality"] == "not_scored",
            "scope": "authored_candidates_not_measured_prompt_improvement"}


if __name__ == "__main__":
    print(json.dumps(demo(), ensure_ascii=False, sort_keys=True))

本地运行的预期输出

{"casesPerVersion": 3, "freeTextStillRequiresReview": true, "scope": "authored_candidates_not_measured_prompt_improvement", "v1ContractPasses": 1, "v1Failures": {"battery": ["labelled_decision", "labelled_window", "necessary_evidence"], "unknown": ["labelled_decision", "labelled_window"]}, "v2ContractPasses": 3}
  • v1ContractPasses=1、v2ContractPasses=3,仅描述作者候选的契约检查。
  • 电池错误涉及 labelled_decision、labelled_window 与 necessary_evidence。
  • unknown 问题具有独立拒答路径。
  • freeTextStillRequiresReview=true。
查看运行环境、输出和校验记录 →

本层验收任务

对电池问题分别构造“例外未进入上下文”和“例外已提供但候选误用规则”,交付不同定位;再增加一条有证据却错误拒答的保留样本。

完成后逐条核对

  • 实际上下文可单独复查。
  • 检索丢失与生成误用分别定位。
  • 缺证据正确拒答与过度拒答分开统计。
  • 版本变化导致的标签差异有说明。

保存自己的过程、代码与结果。这里提供验收要求,暂不自动评分或保存课程掌握状态。

收起答案,检查理解

升级提示后总分提高,但权限题退化,能否直接发布?

延伸原理与知识练习

遇到不熟悉的原理,先阅读实现、连续追问和迁移案例,再独立说明前提与边界。作答与笔记保存到原有账号记录。

本专题的全部关联解析与练习(3 道)

依据与验证范围

原理依据来自公开资料;数字、案例和任务是本站教学设计。离线实验验证本页注明的范围,学习效果仍需通过独立任务与反馈判断。