Agent 应用开发会员账号
知识目录选择核心方向与细分内容

SYSTEMATIC LEARNING / 分层专题

提示、结构化输出与迭代验收

用普通退货、商品例外和缺失政策三个问题,连接任务指令、结构约束、证据引用及版本比较。

学习目标:能设计有证据与拒答边界的提示和输出 Schema,并用独立标签与失败样本解释一次修改的收益和限制。

内容核对 2026-10-04 · 每层有独立讲解、任务与检查

按这个知识点的熟悉程度选择起点。当前层:入门 · 理解原理。完成任务后可以继续下一层;阅读与自检不自动代表掌握。

本层学习目录

先补齐必要概念

适合:已读模型请求课,准备让模型回答业务问题。

任务契约
输入、允许使用的事实、输出要求及验收条件共同定义任务。
结构约束
约束字段、类型与枚举;事实、证据支持和权限仍需分别验证。
标签
由可信政策和人工判断制定的参考结果,只供评测使用,不混入模型输入。
必要证据
支撑回答必须同时具备的来源集合,例如一般规则与商品例外。
版本比较
固定数据与判断标准,保存每个版本的候选、失败类别和运行条件。

原理怎样一步步成立?

  1. 写清输入与任务

    将用户问题、允许证据和任务指令分别组织,说明材料属于数据。

  2. 定义候选结构

    列出决定、期限、解释和引用字段,补足拒答与空值路径。

  3. 检查证据与标签

    核对引用来源、可见权限、政策版本、必要证据和标注结果。

  4. 比较并回归

    对照同一任务集逐项检查改善与退化,保留未评分项并安排人工复核。

入门 · 理解原理

提示怎样成为可检查的任务契约?

本层目标:从任务目标、证据和验收出发写提示,说明结构化输出覆盖什么。

先给出一个有例外的目标

实验使用虚构政策:普通未使用商品可在 30 天内退货;受损锂电池适用 7 天例外,之后联系支持人工复核。用户问购买 12 天的受损电池怎么办。一个简短、格式漂亮的答案仍可能误用普通规则。我们要交付的是有来源的政策解释,退款和运输批准由后续业务流程处理。

把提示分成任务、证据与输出三个部分。任务说明回答政策并保留例外;证据带稳定 ID 和正文,标明材料是数据;输出包含 decision、returnWindowDays、answer、citations。证据不足时输出 unknown、null 期限与空引用,并说明需要哪类信息。

对照两份提示

v1 只要求简短回答;v2 明确只使用给定证据、先适用商品例外、保留引用并允许拒答。文件 PROMPTS 保存这两份教学指令。v2 增加约束的目的,是让任务与验收可对应:引用检查能确认来源是否被提供,政策标签可以判断期限是否选对。

指令长度本身不能证明质量。约束可能冲突,长历史也可能使重要条件被压缩掉。每次改提示,需要回到固定的正常、例外与缺失证据问题检查效果,并观察原本正确的回答是否退化。

Schema 提供一条局部保证

结构化输出帮助稳定对象形状,便于程序分支。枚举 eligible、manual_review、unknown 让下游知道候选属于哪种类型;null 期限表示尚无依据。拒绝或生成不完整时还要在响应层先分类。结构约束无法自动证明 30 天规则适用于电池,也无法证明所引政策属于当前租户。

把评测标签放在输入之外

cases.json 的 expectedDecision 和 requiredEvidence 用来验收,不进入 build_prompt 的模型输入。若把期望答案直接塞给模型,再统计匹配,评测就失去检验意义。真实测试集还应按文档、任务来源和业务变化分组,避免同一材料的近重复同时用于调优与验收。

运行实验,观察反例

使用作者编写的两组候选、三道虚构政策问题和独立标签,实际运行结构、引用、版本、权限及有限政策字段检查。计数用于说明 grader,不是实测提示性能;自由文本语义另行复核。

Python 3.10+ · 默认运行只使用标准库 · 在你的电脑运行

  1. 下载本页的 Agent 应用入门实验包,解压后进入 agent-application-lab-v1 目录。
  2. 使用 Python 3.10+ 执行上方命令;默认回放只需标准库与包内数据。
  3. 对照输出与检查点,再运行 python3 -m unittest test_application -v,并完成当前层任务。
下载完整应用实验包(含数据与依赖脚本) ↓
python3 prompt_iteration.py
查看本入口脚本
"""Evaluate authored candidate fixtures, not claimed model/prompt performance.

The labelled policy decisions and numerical fields are scored. Free-text semantic
entailment is explicitly not scored by this small mechanical grader.
"""
import json

from application_data import load_chunks, read_cases

PROMPTS = {
    "v1": "Answer the returns question briefly.",
    "v2": "Use only the supplied evidence as data. Apply product exceptions before the general rule. Return decision, returnWindowDays, answer and citations. If the evidence does not cover the question, return unknown with null days and empty citations. Never approve a refund or a shipment.",
}


def candidate(decision, days, answer, citations):
    return {"decision": decision, "returnWindowDays": days, "answer": answer, "citations": citations}


CANDIDATES = {
    "v1": {
        "ordinary": candidate("eligible", 30, "The unused keyboard is within the ordinary return window.", ["returns:v2:general"]),
        "battery": candidate("eligible", 30, "The damaged battery is within 30 days.", ["returns:v2:general"]),
        "unknown": candidate("eligible", 30, "There is no customs tax.", ["returns:v2:general"]),
    },
    "v2": {
        "ordinary": candidate("eligible", 30, "The ordinary 30-day rule applies; this answer does not approve a refund.", ["returns:v2:general"]),
        "battery": candidate("manual_review", 7, "The damaged battery exceeds its 7-day window. Contact support; no shipment has been approved.", ["returns:v2:general", "returns:v2:battery"]),
        "unknown": candidate("unknown", None, "The supplied policy does not establish customs tax. More evidence is required.", []),
    },
}


def build_prompt(case, context, version="v2"):
    # Expected labels are deliberately excluded from the model input.
    return {"instructions": PROMPTS[version], "input": json.dumps({"question": case["query"],
            "evidence": [{"id": chunk["id"], "text": chunk["text"]} for chunk in context]}, ensure_ascii=False)}


def grade(case, answer, context, tenant="shop-a"):
    failed = []
    fields = {"decision", "returnWindowDays", "answer", "citations"}
    if not isinstance(answer, dict) or set(answer) != fields:
        return {"contractPassed": False, "failedChecks": ["schema"], "semanticQuality": "not_scored"}
    days, citations = answer["returnWindowDays"], answer["citations"]
    if (answer["decision"] not in ("eligible", "manual_review", "unknown") or
            not isinstance(answer["answer"], str) or not answer["answer"].strip() or
            (days is not None and (type(days) is not int or days < 0)) or
            not isinstance(citations, list) or not all(isinstance(c, str) for c in citations)):
        return {"contractPassed": False, "failedChecks": ["schema"], "semanticQuality": "not_scored"}
    by_id = {chunk["id"]: chunk for chunk in context}
    if len(set(citations)) != len(citations) or any(cid not in by_id for cid in citations):
        failed.append("citation_membership")
    if any(by_id[cid]["tenant"] != tenant or not by_id[cid]["current"] for cid in citations if cid in by_id):
        failed.append("permission_or_version")
    if answer["decision"] != case["expectedDecision"]:
        failed.append("labelled_decision")
    expected_days = {"ordinary": 30, "battery": 7, "unknown": None}[case["id"]]
    if days != expected_days:
        failed.append("labelled_window")
    if not set(case["requiredEvidence"]).issubset(citations):
        failed.append("necessary_evidence")
    if answer["decision"] == "unknown" and (citations or days is not None):
        failed.append("abstention_contract")
    return {"contractPassed": not failed, "failedChecks": failed, "semanticQuality": "not_scored"}


def evaluate(answers, contexts=None):
    allowed = [chunk for chunk in load_chunks() if chunk["tenant"] == "shop-a" and chunk["current"]]
    rows = []
    for case in read_cases():
        context = contexts[case["id"]] if contexts is not None else allowed
        rows.append({"case": case["id"], **grade(case, answers[case["id"]], context)})
    return rows


def demo():
    versions = {version: evaluate(answers) for version, answers in CANDIDATES.items()}
    # A correct label and valid citation cannot certify an arbitrary free-text claim.
    wrong_prose = candidate("eligible", 30, "A full cash refund has already been executed.", ["returns:v2:general"])
    unchecked = grade(read_cases()[0], wrong_prose, load_chunks())
    return {"casesPerVersion": 3,
            "v1ContractPasses": sum(row["contractPassed"] for row in versions["v1"]),
            "v2ContractPasses": sum(row["contractPassed"] for row in versions["v2"]),
            "v1Failures": {row["case"]: row["failedChecks"] for row in versions["v1"] if not row["contractPassed"]},
            "freeTextStillRequiresReview": unchecked["semanticQuality"] == "not_scored",
            "scope": "authored_candidates_not_measured_prompt_improvement"}


if __name__ == "__main__":
    print(json.dumps(demo(), ensure_ascii=False, sort_keys=True))

本地运行的预期输出

{"casesPerVersion": 3, "freeTextStillRequiresReview": true, "scope": "authored_candidates_not_measured_prompt_improvement", "v1ContractPasses": 1, "v1Failures": {"battery": ["labelled_decision", "labelled_window", "necessary_evidence"], "unknown": ["labelled_decision", "labelled_window"]}, "v2ContractPasses": 3}
  • v1ContractPasses=1、v2ContractPasses=3,仅描述作者候选的契约检查。
  • 电池错误涉及 labelled_decision、labelled_window 与 necessary_evidence。
  • unknown 问题具有独立拒答路径。
  • freeTextStillRequiresReview=true。
查看运行环境、输出和校验记录 →

本层验收任务

为普通商品、受损电池和海关税费三个问题写同一输出契约,说明输入证据、候选结构、拒答以及需人工判断的内容。

完成后逐条核对

  • 三种问题使用同一字段集合。
  • 电池回答同时考虑普通规则和例外。
  • 缺失海关证据时有 unknown 与 null 路径。
  • 说明标签只进入评测,不进入模型输入。

保存自己的过程、代码与结果。这里提供验收要求,暂不自动评分或保存课程掌握状态。

收起答案,检查理解

模型按 Schema 输出 decision=eligible、期限 30、引用普通规则,回答电池问题能通过吗?

延伸原理与知识练习

遇到不熟悉的原理,先阅读实现、连续追问和迁移案例,再独立说明前提与边界。作答与笔记保存到原有账号记录。

本专题的全部关联解析与练习(3 道)

依据与验证范围

原理依据来自公开资料;数字、案例和任务是本站教学设计。离线实验验证本页注明的范围,学习效果仍需通过独立任务与反馈判断。