Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content
knowledge unit 39AdvancedConceptsAbout 13 minutes

Understand → Implement → Debug → Design

Evaluate both outcomes and execution traces

Use result postconditions, partial ordering of necessary events, and prohibited behaviors to evaluate trajectories without forcing all tool calls to be exactly the same.

Track evaluationTool callpartial orderObservability

Knowledge content check2026-10-03 · Check the source of the original question2026-10-02

Which step do you want to learn from this knowledge point?

Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.

Understand first

New to this knowledge point

Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.

Start with core principles →

Implement next

Prepare to write the principles into code

Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.

View the code example →

Debug failures

Need to handle failures and changes in conditions

Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.

Continue to delve deeper into the problem →

Compare designs

Need to design or review plans

Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.

Analyze engineering scenarios →
Knowledge unit directory

LEARN · PRACTICE · REFLECT

Knowledge learning and personal records

My notes and review ↗

First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.

Answers and personal notes

Each modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.

Core concept · Evaluate both outcomes and execution traces

Understand the core principles first

Preparatory concepts:call log, Concurrency dependencies, Permission checks

The outcome answers what the task ultimately became, and the trajectory answers which actions caused that outcome. Acceptance should constrain necessary facts and causal relationships, and allow valid path changes that do not affect the constraints.

Correct outcomes can follow invalid actions

A correct balance may come from another user’s account; a ticket may exist twice. Final snapshots miss such risks, so inspect real tool events against protected resources. Conversely, an expected tool sequence does not establish correct arithmetic.

Express dependencies instead of one fixed sequence

Authorized searches can run concurrently, and caches or databases may provide equivalent versions. Require valid subject-resource-action authorization before each read and evidence within that authorized scope. Do not mandate a database query after a cache hit. Event IDs and dependencies express partial ordering; wall-clock times assist diagnosis.

Preserve uncertainty about missing evidence

A tool-start event is not a completion receipt, and “verified” in prose does not prove a call occurred. Separate insufficient evidence from observed violations and choose release blocking according to risk. The proposed paths and event designs are learning exercises, not executed concurrency experiments.

Check understanding with a question

The final result is correct, does the Agent trajectory still need to be evaluated? How to avoid misjudging different correct paths?

Outcome evaluation establishes what happened; trajectory evaluation checks how, including unauthorized access, duplicates, and waste. Do not require one literal sequence when valid caches or parallel paths exist. Define required and forbidden steps, parameter constraints, and dependencies. Hard violations fail; partial scores aid diagnosis. Real tool events and business records supply evidence, rather than model-written plans.

Implementation and trade-offs

Correct results may mask incorrect execution

The agent might happen to come up with the correct balance but read the wrong account; the report might end up being correct but submit a duplicate ticket multiple times along the way. Only commenting on the final reply cannot detect these issues. The trace provides the real tool name, parameters, return status, retry reason, authorization decision and causal relationship to locate the source of errors and costs. On the other hand, going through the expected steps does not mean that the result is correct, and wrong calculations may still be made after the retrieval is successful. Therefore, result acceptance and trajectory acceptance are complementary and cannot replace each other.

Use allowed paths instead of a single standard sequence

The LangSmith documentation states that exact trajectory has limitations because there may be multiple correct paths. In engineering, rules are divided into must-happen, must-not-happen, parameters must be met, and event A precedes B. For example, authorization must precede reading, reference evidence must come from this permission result, and writing must use the audited operation key. Reads can come from the database or the authorization cache, and two queries that are independent of each other can be parallelized. Comparing a set of tool names will ignore the number of times and parameters; comparing a fixed order will misjudge reasonable parallelism. Partial ordering and status conditions can be closer to real business.

Trajectory collection requires credible events and causal identification

Each event records event_id, run_id, step_id, parent_id, tool, input summary or controlled parameters, start and end time, status and resource usage. The start of the request does not mean that the business is successful. Timeout may mean that the result is unknown. Business reconciliation events should be recorded independently. Sensitive parameters are redacted but resource identifiers that can be used to verify permissions are retained; missing logs cannot be automatically scored. The events of parallel tasks need to be analyzed based on causality, and clock sequencing alone may lead to misjudgment. The "I checked it three times" generated by the model is not evidence of tool execution.

Partial points are used for diagnosis, hard constraints are used for blocking

It can calculate the completion degree of retrieval steps, the number of invalid calls, repeated reading and writing, and the proportion of correct parameters to help locate deviations; prohibited actions such as unauthorized reading and unauthorized sending will still fail directly. Use rules to detect stability boundaries before manually or semantically reviewing open decisions. The following code allows two paths of database reading and cache reading, and checks that authorization precedes reading and prohibits deletion; it only applies to serial list examples. Production parallel tracks should use event DAG to express happens-before, and resource parameters and final business postconditions need to be verified. Don't force redundant calls for "the trajectory is more like the answer", and don't treat unobservable internal thinking as an acceptance log.

code example

Accept two valid traces and refuse to read before authorizing

The Python standard library works. Only serial partial order is demonstrated, authorize does not verify the subject or resources; production must verify parameters, permission decisions, event cause and effect, and business results.

def score(events):
    tools = [e['tool'] for e in events]
    reads = [i for i,t in enumerate(tools) if t in {'db_read','cache_read'}]
    auth = [i for i,t in enumerate(tools) if t == 'authorize']
    return {
        'authorized_before_read': bool(auth and reads) and all(any(a < r for a in auth) for r in reads),
        'no_delete': 'delete' not in tools,
        'has_answer': bool(tools) and tools[-1] == 'answer',
    }
for names in [('authorize','db_read','answer'), ('authorize','cache_read','answer'), ('db_read','authorize','answer')]:
    checks = score([{'tool': n} for n in names])
    print(all(checks.values()))

expected output

True
True
False

Engineering deduction

scene
Assume engineering scenario: Query Agent has two implementation paths: direct database reading and authorized caching.
design decisions
The results both check against the same data snapshot, and track access allows both reads, but requires current authorization to precede access.
Verify target
The acceptance goal is that cache optimization will not be misjudged due to different calling paths, and authorization after reading will still fail.
applicable boundary
The examples do not handle cache staleness, principal and resource authorization matching, or demonstrate concurrent event ordering.

Continuous questions and answers

Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.

Draw inferences from one example: If the conditions change, how to deduce it?

First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.

Caching replaces a database read

Changing conditions:The execution path is shortened, but authorization and data versioning targets are maintained.

Extended question:Should one less query result in trajectory failure?

Derivation and reference solutions

If the cache record also carries the data version, authorization scope and can be verified as not invalid, the query source can be changed. Write the rule as "read the allowed version of the material" rather than "must call the SQL tool"; the final result still needs to be checked for consistency. Cache permissions cannot be relaxed when the permissions are unknown.

The principles that remain unchanged:Allows implementation path changes, retaining resources and authorization constraints.

Two release branches in parallel

Changing conditions:From read-only parallelism to two branches that may be repeated outgoing.

Extended question:Is it enough to just check that there is a final report?

Derivation and reference solutions

Not enough. Both branches should share the same business release key or be coordinated by a unique release status; the trace checks whether two attempts yielded the same effect and checks against the target record. The dependency graph allows for parallel builds, and releases are still bound by approved versions and deduplication.

The principles that remain unchanged:Success is achieved when process constraints and business end states are established at the same time.

Easy to make mistakes

  • Require unique tool call sequence, penalize valid alternative paths
  • Compare only the tool name, ignoring parameters, times and side effects
  • Treat tool request initiation as business success
  • Treat model readme as real execution trace

References

It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.

Check how far you understand

After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.

Basic standards met
A correct result may still have involved dangerous or unnecessary actions.
Intermediate and advanced signals
Check trajectories in order of allowed and prohibited actions and cause and effect.
Senior criteria
Tolerate multiple correct paths, measure trajectory rule false positives and calibrate with counterexamples.

View verification records for independent examples

Continue to do advanced research experiments

Let reviews read events and actual effects

Differentiate between run success, result contracts, and unscored semantic quality along actual checkpoints, events, and servicer receipts for the same research task.

Read full text and fault analysis → · Download Reliability Experiment v3 ↓

python3 cli.py memory-put
python3 cli.py submit
python3 cli.py run
python3 evaluate.py
python3 -m unittest discover -s . -p test_lab.py -v

Keep evidence and check item by item

  • Explain the basis of passed in terms of events and checkpoints.
  • Point out credible observations of scope, references, idempotent effects, and step budgets respectively.
  • Semantic_support and model_quality are not_scored, and passed cannot be written as model quality score.

By default, deterministic summary functions and synthetic documents are used; real operating trajectories and mechanical constraints are evaluated, and there is no semantic support to verify the real model.

Hands-on verificationComplete on demand · Suggestions15 minutes

Compare two different but legal instrumental trajectories with an override trajectory that turns out to be correct.

Expand acceptance requirements and checkpoints
  • Valid alternative paths are not misjudged
  • Ultra vires will not be released just because the result is correct.
  • The rules have a clear scope of application

Key inspections

  • Measurement value that can differentiate between business results and execution processes
  • Can explain the misjudgment of valid alternative paths by accurate trajectory matching
  • Can use partial ordering to express dependencies such as authorization before protected reads
  • Can identify parameter errors, missing logs and untrustworthy readme reports