Understand first
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Understand → Implement → Debug → Design
String business tasks, model requests and tool calls into queryable links, and use quality regression to supplement operating indicators.
Knowledge content check2026-10-03 · Check the source of the original question2026-10-02
It is recommended to understand first:
Your first model call and response contract →Agent evaluation: outcomes, constraints, and evidence →Idempotency, unknown outcomes, and task recovery →Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Prepare to write the principles into code
Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.
View the code example →Need to handle failures and changes in conditions
Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.
Continue to delve deeper into the problem →Need to design or review plans
Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.
Analyze engineering scenarios →LEARN · PRACTICE · REFLECT
First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.
Can be practiced directly. After logging in, answers, favorites, and notes will be saved to your account.
Log in and saveEach modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.
Core concept · Task causality and resource attribution
Preparatory concepts:Distributed tracing, asynchronous queue, Business end state
Observability is about connecting a user task to real models, tools, and business effects. The timeline is used to explain waiting, and the resource account is used to explain costs. The two cannot be replaced by model self-statement or simple accumulation.
A report includes retrieval, model rounds, retries, and approval waits. Work can continue after an HTTP span ends. A stable run_id connects business records; traces express call relationships. Record each attempt’s stage, version, error class, and actual usage to distinguish model, queue, and approval delays.
Two fully overlapping two-second queries take about two seconds plus orchestration overhead, while both incur charges. Calculate elapsed time along the dependency graph’s critical path and sum deduplicated billing records. Do not add parent and child spans twice. Averages can hide a few extremely slow tasks.
Prefer event IDs, data versions, parameter structures, and redacted errors. Minimal reproduction payloads may need controlled storage with expiry. Explain reduced semantic reproducibility when text is omitted: a digest cannot reconstruct input. This attribution process is design guidance without production performance measurements.
Use run_id and trace_id to connect model, retrieval, tool, and approval spans. Record each attempt’s duration, state, tokens, version, and retry reason, separating request success from business success. Inspect slow stages and duplicate calls, then quality and costs by tenant and model version. Redact payloads, retain independent audits, and preserve critical write events despite sampling.
A user request generates a stable run_id; retries inherit the task ID but create a new attempt_id. The retrieval, model and tool spans are placed under the task span. The queue delivery carries the trace context, and the parent-child relationship is established after worker consumption. Each step records the start and end time, operation name, error type, model and prompt template version. OpenTelemetry's GenAI agent span convention provides semantics for agent, workflow, and tool operations; this convention is still in the Development state, and field mapping should be managed centrally and the adopted version should be fixed.
User-perceived delay includes queuing, execution and approval waiting. The time consuming of all sub-spans cannot be added up as the total time consuming, and concurrent tools will overlap. Cost statistics first read the supplier's usage, and then calculate the estimated amount based on the current price list. Cache hits and failed retries are marked separately. HTTP success only means that the request is transmitted successfully; "finding the correct account and giving a well-founded balance" is business success. Tool error rates are aggregated by tool, error type and tenant, and task quality is checked using fixed samples and business rules to avoid replacing acceptance with approval of the same model.
When troubleshooting, first find the slowest task, and then look at the critical path: whether the retrieval is slow, the model output is long, or the tool times out and triggers multiple retries. The repair hypothesis requires corresponding experiments. For example, after reducing irrelevant tool descriptions, check whether the correct tool selection is reduced in the same problem set; after compressing the context, check for reference omissions. Tracing of OpenAI Agents SDK covers model generation, tool invocation, handoff and other events; framework tracing is the starting point, and databases, queues and custom business nodes need to be connected.
General logging records resource identification, summaries, and redaction errors, and does not write full documents or keys by default. When it is necessary to save input and output, independent access rights, retention periods and deletion mechanisms are adopted. Aggregation indicators use low-cardinality fields and do not use run_id as a time series indicator label. Regular successful links can be sampled, failed and slow requests can be enhanced to retain; approval, permission determination and external writing enter independent persistence auditing. A tool timeout and a permission denial are injected during verification. The check link can locate the specific steps, and the denial will not be mistakenly counted as business success. At the same time, the prompt template release is included in the change event, so that a quality degradation can be associated with a specific version, instead of just seeing the average response time change.
The Python 3 standard library is ready to run. Only logical calls and retry statistics are demonstrated, and OpenTelemetry is not connected; tool_work_ms is the sum of workload, and cannot be regarded as task wall clock time consumption when concurrency exists.
import json
# 假设采集得到的三个 span;同一 tool_call_id 的不同 attempt 是重试。
spans = [
{"run_id": "r1", "tool_call_id": "t1", "attempt": 1, "status": "timeout", "duration_ms": 800},
{"run_id": "r1", "tool_call_id": "t1", "attempt": 2, "status": "ok", "duration_ms": 120},
{"run_id": "r1", "tool_call_id": "t2", "attempt": 1, "status": "ok", "duration_ms": 60},
]
result = {
"tool_attempts": len(spans),
"logical_tool_calls": len({s["tool_call_id"] for s in spans}),
"failed_attempts": sum(s["status"] != "ok" for s in spans),
"tool_work_ms": sum(s["duration_ms"] for s in spans),
}
print(json.dumps(result, ensure_ascii=False, sort_keys=True))
expected output
{"failed_attempts": 1, "logical_tool_calls": 2, "tool_attempts": 3, "tool_work_ms": 980}Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.
Level 1How does an asynchronous worker restore the trace context? What should I do if the context is missing?
After a task leaves an HTTP request, traceable associations still need to be maintained.
Pass the controlled trace context along with the queue message, create sub-spans or links according to semantics during consumption, and verify the format. If there is no context, you can create a new trace, still use the task ID to associate, and mark the broken link; do not guess the parent relationship based on the close time, and do not use the propagation header as a trusted user identity.
Follow this answer further
Level 2If a batch consumption in the queue is associated with multiple tasks, is it appropriate to use a parent span?
The parent question introduces an asynchronous context, and the child question adds multiple upstream batch processing relationships.
Multiple independent upstreams cannot all be the sole parent of the same span. You can create your own span for batch processing and use links to point to each message source; business events still retain their own run_id. Splitting sub-operations facilitates attribution, but do not throw away the true multi-source relationship in order to make up a tree.
Follow this answer further
Level 3If the same message is re-delivered, should the span_id consumed for the first time be reused?
The parent asks to preserve the causal relationship and continue to distinguish between message identity and execution attempt identity.
Different execution attempts should not be disguised as the same span. Create a new span for each consumption and associate it with the message ID, business key and attempt identifier. The target effect is deduplicated according to the business key. In this way, you can see the time and cost of repeated consumption without counting repeated attempts as multiple successful transactions.
Level 1Why can't the time consumption of concurrent spans be added directly?
After removing the components, explain the difference between statistical time-consuming and user waiting.
There is overlap in parallel operations, the parent span also includes waiting and orchestration, and direct summing will cause repeated calculations. Observe the critical path using time intervals and dependency diagrams, respectively displaying the wall clock time consumption and the cumulative workload of each component; costs are calculated based on actual billing events, which can be added up but must be deduplicated.
Level 1If I/O logging is turned off, how can I retain the minimum evidence of a reproducible problem?
Evidence is required for troubleshooting, but the scope of recording is subject to data minimization.
Preserves task and event IDs, input structures and summaries, versions, masked error codes, tool receipt status, and controlled data snapshot references. For parts that require the text to be reproduced, minimum samples are collected with authorization; if the text is not retained, the runtime link can only be reproduced, and the complete reproduction model answer cannot be claimed.
First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.
Changing conditions:Long periods of waiting for a person are added to machine execution.
Extended question:Does waiting for approval count as model delay?
Should not. Separate task end-to-end duration, active calculation duration, and approval wait, and retain pause and resume events; user wait indicators still include approval. Optimizing model speed cannot solve the backlog of approval queues and needs to be positioned by stage.
The principles that remain unchanged:The same task is connected by real causal events, and waiting and workload have independent definitions.
Changing conditions:Insufficient evidence of request-level charges.
Extended question:Can each tenant expense be accurately attributed?
First make an estimate based on available usage and public pricing, indicate the apportionment rules and differences, and then reconcile it with the bill. Shared caching, discounts, and retries keep estimates from actual payments; you can't force unknown bills into precise per-task costs.
The principles that remain unchanged:The cost conclusion shall not exceed the range supported by the collection and valuation evidence.
It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.
After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.
View verification records for independent examples
Design minimal Trace fields for multi-tool tasks to locate P95 latency sources.