Review the prerequisites
Suitable for: The normal chain has been run through and abnormal responses and wrong answers have been processed.
- Model request
- Contains model ID, task instructions, user input and available tools; credentials and real user identity are managed by the caller.
- Candidate action
- The tool name and parameters given by the model. It expresses recommendations, and the execution side also checks formats, permissions, and business conditions.
- Call correlation
- call_id associates a certain function request with its return result; the business idempotency key needs to be separately defined by the business.
- Completion status
- HTTP, generation process, and business tasks each have status; only business acceptance can indicate task completion.
- Response fixture
- In this experiment, the Responses shape sample written by the author is used to check the program branch; the real API is subject to the current return and the official contract.
How does the mechanism work?
- Assemble complete request
Task instructions describe allowed behaviors, user input describes goals, and tool definitions provide local parameter contracts.
- Classification interface results
Check the HTTP result first, then check the build status and output's entry type, retaining rejection and incomplete reasons.
- perform restricted actions
Bind identities from trusted sessions and independently verify parameters, access permissions and call budgets; this experiment only allows reading orders.
- Return and verification
Keep the original output items, return the function results with the corresponding call_id, and finally check the candidate answers against the tool receipt.
Debugging · Diagnose failures
Diagnose failures using statuses and receipts
Objectives of this level: Use clear failure samples to locate transfer, generation, parameters, authorization and answer check issues.
Locate the first failed boundary
Separate the investigation into HTTP transport, generation status, candidate arguments, resource authorization, tool receipts, and the final answer. Keep verifiable evidence for each stage. HTTP 429 indicates an interface limit. If a tool succeeded but the subsequent explanation was truncated, retain its successful receipt and adjust the explanation stage. These failures require different recovery actions.
Reproduce four small failures
Run this in the extracted lab directory:
python3 - <<'PY'
from model_response_contract import run_loop, tool_response, text_response
print(run_loop([tool_response(name='refund_order')])['status'])
print(run_loop([tool_response('o-2')])['status'])
print(run_loop([tool_response(), tool_response()])['status'])
print(run_loop([text_response({'order_id':'o-1','status':'delivered','source':'read_order'})])['status'])
PY
The cases demonstrate an unknown tool, denied permission, a duplicate call ID, and an answer without a receipt. The final answer happens to match the teaching order, but still returns unverified_answer: this run obtained no trusted lookup evidence.
Handle incomplete output explicitly
Set a tool response's status to incomplete. Even if its output appears to contain complete arguments, this lab stops without executing the candidate. Real streaming systems validate against complete call events. Record network interruption, output-limit exhaustion, and committed business actions separately. Before recovering, establish which effects occurred, then decide whether to regenerate or query an external receipt.
Investigate a 200 response that cannot be parsed
Check that the body came from the intended API and that its top-level status and output types match the adapter. HTML error pages, proxy responses, and similar JSON from another provider need separate handling. Do not delete characters, add closing brackets, or silently switch interfaces to manufacture success. Record fixed error categories and versions, and redact sensitive content under the applicable business rules.
Identify what can be retried
Transient generation failures may allow backoff within cost and time budgets. If a provider action or application write may already have occurred, reconcile the actual effect first. This lesson implements no remote retry protocol; the recovery lesson has a separate unknown-outcome lab. Debugging should establish confirmed facts, a safe retry target, and the next acceptance check.
Run experiments and observe counterexamples
Validate request assembly, read-only validation, call_id postbacks, budgeting, and answer checking using non-streaming Responses shape samples, tutorial orders, and trusted identity impersonation written by the author. The program does not call online models or real order services.
Python 3.10+ · Runs by default using only the standard library · Runs on your computer
- Download the Agent application entry experimental package on this page, unzip it and enter the agent-application-lab-v1 directory.
- Use Python 3.10+ to execute the above command; the default playback only requires the standard library and package data.
- Compare the output with the checkpoint, then run python3 -m unittest test_application -v and complete the current layer task.
python3 model_response_contract.pyView the entry-point script
"""A bounded read-only tool loop using authored Responses-shaped fixtures.
This file makes no API requests. The course separately documents the real API request.
The fixture response shape is deliberately limited to the non-streaming cases below.
"""
import json
ORDERS = {
"o-1": {"tenant": "shop-a", "user": "u-1", "status": "delivered"},
"o-2": {"tenant": "shop-b", "user": "u-2", "status": "processing"},
}
PRINCIPAL = {"tenant": "shop-a", "user": "u-1"}
TOOL = {
"type": "function", "name": "read_order", "description": "Read an authorized order status.",
"strict": True, "parameters": {"type": "object", "properties": {"order_id": {"type": "string"}},
"required": ["order_id"], "additionalProperties": False},
}
def build_request(order_id, model="reader-selected-model"):
return {"model": model, "input": [
{"role": "developer", "content": "Use read_order for order facts. Tool access is checked by the application. Return an object with order_id, status and source. Do not perform writes."},
{"role": "user", "content": "What is the status of order " + order_id + "?"},
], "tools": [TOOL], "max_output_tokens": 500}
def classify_response(response, http_status=200):
if http_status != 200:
return {"kind": "http_error", "retryable": http_status == 429 or 500 <= http_status < 600}
if not isinstance(response, dict):
return {"kind": "unsupported"}
status = response.get("status")
if status == "incomplete":
return {"kind": "incomplete", "reason": response.get("incomplete_details")}
if status in ("failed", "cancelled"):
return {"kind": status}
if status != "completed" or not isinstance(response.get("output"), list):
return {"kind": "unsupported"}
calls, texts = [], []
for item in response["output"]:
if not isinstance(item, dict):
return {"kind": "unsupported"}
if item.get("type") == "function_call":
if not all(isinstance(item.get(key), str) and item[key] for key in ("call_id", "name", "arguments")):
return {"kind": "invalid_tool_call"}
calls.append(item)
elif item.get("type") == "message":
if item.get("status") not in (None, "completed") or not isinstance(item.get("content"), list):
return {"kind": "unsupported"}
for content in item["content"]:
if not isinstance(content, dict):
return {"kind": "unsupported"}
if content.get("type") == "refusal":
return {"kind": "refused"}
if content.get("type") != "output_text" or not isinstance(content.get("text"), str):
return {"kind": "unsupported"}
texts.append(content["text"])
elif item.get("type") != "reasoning":
return {"kind": "unsupported"}
if calls:
return {"kind": "tool_requests", "calls": calls}
return {"kind": "text", "text": "".join(texts)} if texts else {"kind": "empty"}
def execute_read(call, principal):
if call["name"] != "read_order":
raise ValueError("unknown_tool")
try:
args = json.loads(call["arguments"])
except (ValueError, TypeError):
raise ValueError("invalid_arguments") from None
if not isinstance(args, dict) or set(args) != {"order_id"} or not isinstance(args["order_id"], str):
raise ValueError("invalid_arguments")
order = ORDERS.get(args["order_id"])
if not order or order["tenant"] != principal["tenant"] or order["user"] != principal["user"]:
# No protected resource fields are returned when access is denied.
raise ValueError("permission_denied")
return {"order_id": args["order_id"], "status": order["status"], "source": "read_order"}
def run_loop(responses, principal=None, max_turns=3, max_calls=2):
principal = principal or PRINCIPAL
trace, results, seen_ids, calls_used = [], [], set(), 0
transcript = build_request("o-1")["input"]
for turn, response in enumerate(responses, 1):
if turn > max_turns:
break
event = classify_response(response)
trace.append({"turn": turn, "kind": event["kind"]})
if event["kind"] == "tool_requests":
# Preserve output items, including reasoning items when present, before
# appending call-correlated tool results to the next request input.
transcript.extend(response["output"])
if calls_used + len(event["calls"]) > max_calls:
return {"status": "budget_exhausted", "trace": trace, "results": results, "calls": calls_used}
for call in event["calls"]:
if call["call_id"] in seen_ids:
return {"status": "duplicate_call_id", "trace": trace, "results": results, "calls": calls_used}
seen_ids.add(call["call_id"])
calls_used += 1
try:
result = execute_read(call, principal)
except ValueError as error:
return {"status": str(error), "trace": trace, "results": results, "calls": calls_used}
result_item = {"type": "function_call_output", "call_id": call["call_id"], "output": json.dumps(result, sort_keys=True)}
results.append(result_item)
transcript.append(result_item)
trace[-1]["nextInputTypes"] = [item.get("type", "message") for item in transcript]
elif event["kind"] == "text":
try:
candidate = json.loads(event["text"])
except ValueError:
candidate = None
confirmed = isinstance(candidate, dict) and set(candidate) == {"order_id", "status", "source"} and any(candidate == json.loads(result["output"]) for result in results)
return {"status": "answer_ready" if confirmed else "unverified_answer", "trace": trace, "results": results, "calls": calls_used}
else:
return {"status": event["kind"], "trace": trace, "results": results, "calls": calls_used}
return {"status": "budget_exhausted", "trace": trace, "results": results, "calls": calls_used}
def tool_response(order_id="o-1", call_id="call-1", name="read_order"):
return {"status": "completed", "output": [{"type": "function_call", "call_id": call_id, "name": name, "arguments": json.dumps({"order_id": order_id})}]}
def text_response(candidate):
return {"status": "completed", "output": [{"type": "message", "role": "assistant", "status": "completed", "content": [{"type": "output_text", "text": json.dumps(candidate)}]}]}
def demo():
normal = run_loop([tool_response(), text_response({"order_id": "o-1", "status": "delivered", "source": "read_order"})])
denied = run_loop([tool_response("o-2")])
loop = run_loop([tool_response(call_id="c-" + str(i)) for i in range(5)], max_turns=2)
refused = {"status": "completed", "output": [{"type": "message", "content": [{"type": "refusal", "refusal": "fixture refusal"}]}]}
return {"normal": normal["status"], "toolResultCallId": normal["results"][0]["call_id"],
"crossTenant": denied["status"], "deniedResults": len(denied["results"]),
"boundedLoop": loop["status"], "boundedCalls": loop["calls"],
"refusal": classify_response(refused)["kind"],
"incomplete": classify_response({"status": "incomplete", "incomplete_details": {"reason": "max_output_tokens"}})["kind"],
"scope": "authored_responses_fixtures_no_api"}
if __name__ == "__main__":
print(json.dumps(demo(), ensure_ascii=False, sort_keys=True))
Expected output when running locally
{"boundedCalls": 2, "boundedLoop": "budget_exhausted", "crossTenant": "permission_denied", "deniedResults": 0, "incomplete": "incomplete", "normal": "answer_ready", "refusal": "refused", "scope": "authored_responses_fixtures_no_api", "toolResultCallId": "call-1"}- normal=answer_ready, the final object is consistent with the read_order receipt.
- crossTenant=permission_denied and deniedResults=0.
- boundedLoop=budget_exhausted and boundedCalls=2.
- Refusal and incomplete use independent processing paths.
Acceptance task for this level
Deliver the minimum input, status, number of receipts, and location conclusions for the four faults; in addition, the model outputs incomplete recovery steps after the analysis tool succeeds.
Check each item after completion
- Each fault can be reproduced using functions within the package.
- Extra fields and override input fall under parameters and authorization respectively.
- When the tool receipt is missing, the text will not be passed directly.
- Subsequent build failures retain confirmed tool facts.
Save your own processes, code and results. Acceptance requirements are provided here, and course mastery status will not be automatically graded or saved at this time.
Hide the answer and check your understanding
After a model request times out, can the entire tool cycle be reissued uniformly?
Expand reference derivation
First determine which section it takes place in. Candidates that have not yet been executed can be re-verified; confirmed tool reads can be retained; write action timeouts may have been submitted and need to be checked with stable business keys and external receipts. Recurrence generation may also result in duplicate charges or different candidates.
Further explanations and practice
When encountering unfamiliar principles, first read the implementation, continuous questioning and migration cases, and then independently explain the premise and boundaries. Answers and notes are saved to the original account record.
All linked explanations and exercises (2 )
Sources and verification scope
The principles are based on public information; the numbers, cases and tasks are the teaching design of this website. Offline experiments verify the range noted on this page, and the learning effect still needs to be judged through independent tasks and feedback.