Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content
knowledge unit 38IntermediateSystem designAbout 12 minutes

Understand → Implement → Debug → Design

Executable acceptance checks and quality boundaries

Convert results, permissions, evidence, resource costs and recovery capabilities into decidable conditions, and then combine them into release access control.

Reviewacceptance contractReturnRelease access control

Knowledge content check2026-10-03 · Check the source of the original question2026-10-02

Which step do you want to learn from this knowledge point?

Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.

Understand first

New to this knowledge point

Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.

Start with core principles →

Implement next

Prepare to write the principles into code

Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.

View the code example →

Debug failures

Need to handle failures and changes in conditions

Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.

Continue to delve deeper into the problem →

Compare designs

Need to design or review plans

Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.

Analyze engineering scenarios →
Knowledge unit directory

LEARN · PRACTICE · REFLECT

Knowledge learning and personal records

My notes and review ↗

First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.

Answers and personal notes

Each modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.

Core concept · Executable acceptance checks and quality boundaries

Understand the core principles first

Preparatory concepts:Business postconditions, interface contract, Regression testing

Acceptance is about converting "success" into externally observable facts. Quality scores can be used for comparison, but prohibited facts such as overreach and duplicate writing cannot be offset by fluent expression or other high scores.

Move from response assertions to task assertions

A backend refund test checks transaction records, not merely a “refunded” response. Agent prose adds another layer requiring observable success conditions. A report contract can require target-period coverage, sourced numbers, and a saved artifact while allowing varied wording.

Keep prohibited actions outside the average score

A complete report obtained through unauthorized payroll access is still a failure. Averaging quality and permissions could conceal that violation. Establish hard authorization and effect constraints before comparing coverage, explanation, and cost. Task risk determines the gates; different evaluations need different constraints.

Test the verifier with counterexamples

Under the same initial data, include correct reports, missing departments, nonexistent citations, and valid refusals. Verify that the evaluator distinguishes them. Missing evidence permits an unknown outcome and review; forced binary decisions hide gaps. This is instructional design, not a newly executed evaluation.

Check understanding with a question

How does the Agent define an executable acceptance contract to avoid going online based solely on "the answer looks good"?

Define initial environment, permitted actions, business outcomes, and hard constraints separately. Verify outcomes in records or artifacts. Authorization and prohibited effects are hard gates, outside average quality scores. Record cost and latency distributions and limit violations. Fix data, tool versions, and metric definitions; include denied access, failures, cancellation, and recovery. Use deterministic checks where possible and calibrated human or model review for open text.

Implementation and trade-offs

The acceptance contract first describes the task world

For example, "Query project alarms and generate processing suggestions", enter more than one prompt, but also include user identity, authorization project, alarm snapshot, tool contract, system time and allowed budget. The output contract specifies which alarms are required, which evidence is associated with each suggestion, and whether the creation of work orders is allowed. Business post-conditions are verified using database records or products instead of just comparing the similarity of the final text. Correct behavior must also be defined for empty results, no permissions, and unknown status. Active rejection by the system does not necessarily mean failure.

Result quality and hard constraints are determined separately

Establish dimensions such as outcome, policy, grounding, budget, and recovery. If the answer content is reasonable but other tenant data is read, it must be considered a failure; if the work order is successfully created but created twice, it will not be considered a pass. Permission violations, unauthorized side effects, and exceeding the hard budget cannot be averaged out with high language scores. Evidence citations and structured fields are verified with rules, and open suggestions can be reviewed manually or with calibrated models. LangSmith provides evaluation perspectives such as final response, single step and trajectory. Engineering should be combined according to the task contract rather than considering a single evaluator to be sufficient.

Datasets and execution environments require version records

Each sample retains case_id, data version, expected results, allowed tools, error injection configuration and whether it is a retained test set. Real failures can be redacted to form regression cases to avoid selecting only easy-to-succeed questions. Model, prompt, retrieval, tool schema, state machine, and data source changes may affect behavior, so experiments record complete configurations and traces. The data used for parameter tuning is separated from the final release test to avoid constantly modifying the system to cater to a fixed set of answers. Non-deterministic tasks are performed repeatedly and variances are recorded, and error type groupings are more diagnostic than an average success rate.

Release gates must identify specific failures

Runs deterministic checks first, followed by more expensive semantic reviews; failure output includes case_id, violation condition, associated tool calls, and business status differences. The access control threshold is determined based on product risk, and a universal percentage cannot be given to all Agents out of thin air. After launch, continue to collect timeouts, duplicate side effects, rejection rates, and user corrections, updating samples by source. The following code combines the correct result, authorization, idempotency, and resource upper limit into a hard decision, which is just a local scorer, not the actual Agent platform; the delay and cost must come from trusted collection, and the tested Agent cannot be allowed to report a low number and pass.

code example

Business postconditions and hard threshold scorer

The Python standard library works. Input should be generated by testing tools and resource monitoring; examples do not verify log authenticity or review suggestion text.

def evaluate(run):
    checks = {
        'outcome': set(run['ticket_ids']) == {'INC-42'},
        'permission': set(run['read_projects']) <= {'p1'},
        'no_duplicate': len(run['ticket_ids']) == len(set(run['ticket_ids'])),
        'budget': run['cost'] <= 0.20 and run['duration_s'] <= 30,
    }
    return {'passed': all(checks.values()), 'failed': [k for k,v in checks.items() if not v]}
good = {'ticket_ids':['INC-42'], 'read_projects':['p1'], 'cost':0.10, 'duration_s':8}
bad = {**good, 'read_projects':['p1','p2']}
print(evaluate(good))
print(evaluate(bad))

expected output

{'passed': True, 'failed': []}
{'passed': False, 'failed': ['permission']}

Engineering deduction

scene
Hypothetical engineering scenario: The operation and maintenance agent can generate suggestions and create a work order.
design decisions
The database checks the target work order, authorized reading and repeated creation as hard failures, and text suggestions are scored separately.
Verify target
The acceptance goal is that the same correct answer will also be blocked by the release gate when it exceeds authority or performs repeated operations.
applicable boundary
This scorer only covers declared contracts and cannot prove that all scenarios that are not covered are reliable.

Continuous questions and answers

Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.

Draw inferences from one example: If the conditions change, how to deduce it?

First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.

Read-only suggestions become automatically sent

Changing conditions:The output changes from reviewable text to irrevocable external sending.

Extended question:Is the original language quality acceptance sufficient?

Derivation and reference solutions

Not enough. Hard checks of new recipients, authorization purposes, approval content digests, and actual sending records are performed, and shadow evaluation only calls stand-ins. The quality of wording can still be judged semantically, but sending to the wrong recipient must fail independently and cannot be offset by the correctness of the main text. First write the target external state into a contract, and then choose the acceptance method.

The principles that remain unchanged:Success is defined by allowed business facts, prohibited actions are determined independently.

Capacity planning without a standard answer

Changing conditions:The future load is uncertain, and there are multiple reasonable solutions.

Extended question:How to review without falsifying the exact true value?

Derivation and reference solutions

Provide load intervals and fault assumptions, and check whether capacity derivation, bottleneck identification, and degradation strategies correspond to the conditions. Separate assumptions from measured data, and perform sensitivity deductions after parameter changes; scoring covers the quality of explanations, and does not regard a certain database name as the only correct answer.

The principles that remain unchanged:Acceptance should check constraints and derivation, and cannot replace facts with text.

Easy to make mistakes

  • Only evaluate the final text and do not check the business status
  • Override is offset by the average language quality score
  • The test environment and online tools have different semantics
  • Continuously using the test set to adjust parameters is still called an independent test

References

It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.

Check how far you understand

After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.

Basic standards met
Able to define success as the combination of verifiable artifacts and clear constraints.
Intermediate and advanced signals
Distinguish between structure, content, business results and prohibited actions, and have executable checks.
Senior criteria
Can handle multiple valid answers, scoring errors, and high-stakes independent thresholds.

View verification records for independent examples

Continue to do advanced research experiments

Let reviews read events and actual effects

Differentiate between run success, result contracts, and unscored semantic quality along actual checkpoints, events, and servicer receipts for the same research task.

Read full text and fault analysis → · Download Reliability Experiment v3 ↓

python3 cli.py memory-put
python3 cli.py submit
python3 cli.py run
python3 evaluate.py
python3 -m unittest discover -s . -p test_lab.py -v

Keep evidence and check item by item

  • Explain the basis of passed in terms of events and checkpoints.
  • Point out credible observations of scope, references, idempotent effects, and step budgets respectively.
  • Semantic_support and model_quality are not_scored, and passed cannot be written as model quality score.

By default, deterministic summary functions and synthetic documents are used; real operating trajectories and mechanical constraints are evaluated, and there is no semantic support to verify the real model.

Hands-on verificationComplete on demand · Suggestions15 minutes

Write the outcome contract for "Generate and approve release reports".

Expand acceptance requirements and checkpoints
  • Only generating text does not mean publishing is complete
  • Necessary references can be checked
  • Unapproved release independent judgment failed

Key inspections

  • Can convert natural language tasks into decidable business post-conditions
  • Ability to distinguish between quality scores and hard safety thresholds
  • Able to create versioned data sets and fault regression collections
  • Can explain review errors and online distribution changes