先补齐必要概念
适合:已理解响应类型,准备实现一个最小应用。
- 模型请求
- 包含模型 ID、任务指令、用户输入及可用工具;凭据与真实用户身份由调用端管理。
- 候选动作
- 模型给出的工具名与参数。它表达建议,执行端还要核对格式、权限和业务条件。
- 调用关联
- call_id 把某次函数请求与它的返回结果关联;业务幂等键需由业务另外定义。
- 完成状态
- HTTP、生成过程和业务任务各有状态;只有业务验收能说明任务完成。
- 响应样本
- 本实验由作者编写的 Responses 形状样本,用于检查程序分支;真实 API 以当次返回和官方契约为准。
原理怎样一步步成立?
- 组装完整请求
任务指令说明允许行为,用户输入描述目标,工具定义提供局部参数契约。
- 分类接口结果
先检查 HTTP 结果,再检查生成状态和 output 的条目类型,保留拒绝与不完整原因。
- 执行受限动作
从可信会话绑定身份,独立校验参数、访问权限和调用预算;本实验只允许读取订单。
- 回传与核对
保留原始输出条目,把函数结果以对应 call_id 回传,最后对照工具回执核对候选答案。
初级 · 完成实现
组装请求,跑通一轮只读工具回传
本层目标:从输入、校验、权限、回执到答案核对完整走一遍,并知道真实 HTTP 请求接在哪里。
先运行可以逐步观察的程序
下载并解压本页实验包,然后执行:
python3 model_response_contract.py
python3 -m unittest test_application -v
正常路径先收到 read_order(o-1),程序查可信租户与用户,再返回包含状态的回执。下一轮候选答案必须与该回执一致,才得到 answer_ready。跨租户路径得到 permission_denied,工具结果为空;持续调用路径在两次允许调用后以 budget_exhausted 结束。
按四个函数读代码:build_request 组装请求;classify_response 分类结果;execute_read 做资源校验;run_loop 管理预算、回传和终态。ORDERS 是教学数据,PRINCIPAL 模拟可信登录态。换成真实业务时,从服务端会话取得身份并查询实际授权。
请求和响应怎样衔接
本轮工具候选保存为 function_call 条目。执行结果使用 function_call_output,带回同一个 call_id,output 是序列化的工具结果。下一轮输入同时保留原始响应条目与工具结果;使用 reasoning 模型时,相关输出条目也需按接口要求保留。应用自己的业务操作 ID 则用于账本、查询和幂等,单独管理。
run_loop 在并发候选执行前核对整批剩余调用额度。名字不在白名单、额外身份字段、未知订单或重复 call_id 都会被明确拒绝。写工具还需要审批、稳定业务键与结果未知核对,进入后续工具及恢复课程。
把接口示例接到真实响应
若已有自己的 API 账号,可按官方文档选择实际可用、支持 Responses 的模型,把环境中的 MODEL_ID 写入请求文件。下面只生成请求文件:
python3 - <<'PY' > request.json
import json, os
from model_response_contract import build_request
print(json.dumps(build_request('o-1', os.environ['MODEL_ID'])))
PY
官方接口请求的形式为:
curl https://api.openai.com/v1/responses \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
--data-binary @request.json --output response.json \
--write-out '%{http_code}\n'
凭据仅留在自己的环境,不放进网页、请求正文或提交的学习记录。该步骤调用真实服务时会产生实际用量;保存 HTTP 状态、脱敏响应、模型 ID、时间与 usage。网站验证记录覆盖离线分支,真实请求结果需要由当次调用另外记录。
真实响应先交给分类函数检查。若是工具请求,执行并回传结果后继续下一轮;若是拒绝或不完整,走独立终态。课堂分类器只支持示例中的非流式条目,接口新增能力要显式扩充适配与回归样本。
运行实验,观察反例
使用作者编写的非流式 Responses 形状样本、教学订单与可信身份模拟,验证请求组装、只读校验、call_id 回传、预算和答案核对。程序不调用在线模型或真实订单服务。
Python 3.10+ · 默认运行只使用标准库 · 在你的电脑运行
- 下载本页的 Agent 应用入门实验包,解压后进入 agent-application-lab-v1 目录。
- 使用 Python 3.10+ 执行上方命令;默认回放只需标准库与包内数据。
- 对照输出与检查点,再运行 python3 -m unittest test_application -v,并完成当前层任务。
python3 model_response_contract.py查看本入口脚本
"""A bounded read-only tool loop using authored Responses-shaped fixtures.
This file makes no API requests. The course separately documents the real API request.
The fixture response shape is deliberately limited to the non-streaming cases below.
"""
import json
ORDERS = {
"o-1": {"tenant": "shop-a", "user": "u-1", "status": "delivered"},
"o-2": {"tenant": "shop-b", "user": "u-2", "status": "processing"},
}
PRINCIPAL = {"tenant": "shop-a", "user": "u-1"}
TOOL = {
"type": "function", "name": "read_order", "description": "Read an authorized order status.",
"strict": True, "parameters": {"type": "object", "properties": {"order_id": {"type": "string"}},
"required": ["order_id"], "additionalProperties": False},
}
def build_request(order_id, model="reader-selected-model"):
return {"model": model, "input": [
{"role": "developer", "content": "Use read_order for order facts. Tool access is checked by the application. Return an object with order_id, status and source. Do not perform writes."},
{"role": "user", "content": "What is the status of order " + order_id + "?"},
], "tools": [TOOL], "max_output_tokens": 500}
def classify_response(response, http_status=200):
if http_status != 200:
return {"kind": "http_error", "retryable": http_status == 429 or 500 <= http_status < 600}
if not isinstance(response, dict):
return {"kind": "unsupported"}
status = response.get("status")
if status == "incomplete":
return {"kind": "incomplete", "reason": response.get("incomplete_details")}
if status in ("failed", "cancelled"):
return {"kind": status}
if status != "completed" or not isinstance(response.get("output"), list):
return {"kind": "unsupported"}
calls, texts = [], []
for item in response["output"]:
if not isinstance(item, dict):
return {"kind": "unsupported"}
if item.get("type") == "function_call":
if not all(isinstance(item.get(key), str) and item[key] for key in ("call_id", "name", "arguments")):
return {"kind": "invalid_tool_call"}
calls.append(item)
elif item.get("type") == "message":
if item.get("status") not in (None, "completed") or not isinstance(item.get("content"), list):
return {"kind": "unsupported"}
for content in item["content"]:
if not isinstance(content, dict):
return {"kind": "unsupported"}
if content.get("type") == "refusal":
return {"kind": "refused"}
if content.get("type") != "output_text" or not isinstance(content.get("text"), str):
return {"kind": "unsupported"}
texts.append(content["text"])
elif item.get("type") != "reasoning":
return {"kind": "unsupported"}
if calls:
return {"kind": "tool_requests", "calls": calls}
return {"kind": "text", "text": "".join(texts)} if texts else {"kind": "empty"}
def execute_read(call, principal):
if call["name"] != "read_order":
raise ValueError("unknown_tool")
try:
args = json.loads(call["arguments"])
except (ValueError, TypeError):
raise ValueError("invalid_arguments") from None
if not isinstance(args, dict) or set(args) != {"order_id"} or not isinstance(args["order_id"], str):
raise ValueError("invalid_arguments")
order = ORDERS.get(args["order_id"])
if not order or order["tenant"] != principal["tenant"] or order["user"] != principal["user"]:
# No protected resource fields are returned when access is denied.
raise ValueError("permission_denied")
return {"order_id": args["order_id"], "status": order["status"], "source": "read_order"}
def run_loop(responses, principal=None, max_turns=3, max_calls=2):
principal = principal or PRINCIPAL
trace, results, seen_ids, calls_used = [], [], set(), 0
transcript = build_request("o-1")["input"]
for turn, response in enumerate(responses, 1):
if turn > max_turns:
break
event = classify_response(response)
trace.append({"turn": turn, "kind": event["kind"]})
if event["kind"] == "tool_requests":
# Preserve output items, including reasoning items when present, before
# appending call-correlated tool results to the next request input.
transcript.extend(response["output"])
if calls_used + len(event["calls"]) > max_calls:
return {"status": "budget_exhausted", "trace": trace, "results": results, "calls": calls_used}
for call in event["calls"]:
if call["call_id"] in seen_ids:
return {"status": "duplicate_call_id", "trace": trace, "results": results, "calls": calls_used}
seen_ids.add(call["call_id"])
calls_used += 1
try:
result = execute_read(call, principal)
except ValueError as error:
return {"status": str(error), "trace": trace, "results": results, "calls": calls_used}
result_item = {"type": "function_call_output", "call_id": call["call_id"], "output": json.dumps(result, sort_keys=True)}
results.append(result_item)
transcript.append(result_item)
trace[-1]["nextInputTypes"] = [item.get("type", "message") for item in transcript]
elif event["kind"] == "text":
try:
candidate = json.loads(event["text"])
except ValueError:
candidate = None
confirmed = isinstance(candidate, dict) and set(candidate) == {"order_id", "status", "source"} and any(candidate == json.loads(result["output"]) for result in results)
return {"status": "answer_ready" if confirmed else "unverified_answer", "trace": trace, "results": results, "calls": calls_used}
else:
return {"status": event["kind"], "trace": trace, "results": results, "calls": calls_used}
return {"status": "budget_exhausted", "trace": trace, "results": results, "calls": calls_used}
def tool_response(order_id="o-1", call_id="call-1", name="read_order"):
return {"status": "completed", "output": [{"type": "function_call", "call_id": call_id, "name": name, "arguments": json.dumps({"order_id": order_id})}]}
def text_response(candidate):
return {"status": "completed", "output": [{"type": "message", "role": "assistant", "status": "completed", "content": [{"type": "output_text", "text": json.dumps(candidate)}]}]}
def demo():
normal = run_loop([tool_response(), text_response({"order_id": "o-1", "status": "delivered", "source": "read_order"})])
denied = run_loop([tool_response("o-2")])
loop = run_loop([tool_response(call_id="c-" + str(i)) for i in range(5)], max_turns=2)
refused = {"status": "completed", "output": [{"type": "message", "content": [{"type": "refusal", "refusal": "fixture refusal"}]}]}
return {"normal": normal["status"], "toolResultCallId": normal["results"][0]["call_id"],
"crossTenant": denied["status"], "deniedResults": len(denied["results"]),
"boundedLoop": loop["status"], "boundedCalls": loop["calls"],
"refusal": classify_response(refused)["kind"],
"incomplete": classify_response({"status": "incomplete", "incomplete_details": {"reason": "max_output_tokens"}})["kind"],
"scope": "authored_responses_fixtures_no_api"}
if __name__ == "__main__":
print(json.dumps(demo(), ensure_ascii=False, sort_keys=True))
本地运行的预期输出
{"boundedCalls": 2, "boundedLoop": "budget_exhausted", "crossTenant": "permission_denied", "deniedResults": 0, "incomplete": "incomplete", "normal": "answer_ready", "refusal": "refused", "scope": "authored_responses_fixtures_no_api", "toolResultCallId": "call-1"}- normal=answer_ready,最终对象与该次 read_order 回执一致。
- crossTenant=permission_denied 且 deniedResults=0。
- boundedLoop=budget_exhausted 且 boundedCalls=2。
- refusal 与 incomplete 使用独立处理路径。
本层验收任务
先复现离线正常、跨租户和预算耗尽三条路径,再组装 request.json。具备接口账号时另行运行并记录真实响应;使用离线分支也要交付完整请求、工具回执和核对过程。
完成后逐条核对
- 正常答案可以追到实际只读回执。
- 函数候选与返回条目使用同一 call_id。
- 下一轮输入保留原始输出及函数结果。
- 越权时没有受保护订单字段。
- 标注每份响应是作者样本还是实际接口返回。
保存自己的过程、代码与结果。这里提供验收要求,暂不自动评分或保存课程掌握状态。
收起答案,检查理解
只把工具结果放回下一轮输入,原始响应条目全部丢弃,会有什么问题?
展开参考推导
会失去函数请求与结果的完整关联,部分 reasoning 模型所需的原始输出也会缺失。按目标接口要求保留本轮输出,再追加 call_id 对应的函数结果,或使用其明确支持的响应延续机制。
延伸原理与知识练习
遇到不熟悉的原理,先阅读实现、连续追问和迁移案例,再独立说明前提与边界。作答与笔记保存到原有账号记录。
本专题的全部关联解析与练习(2 道)
依据与验证范围
原理依据来自公开资料;数字、案例和任务是本站教学设计。离线实验验证本页注明的范围,学习效果仍需通过独立任务与反馈判断。