Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content
knowledge unit 42AdvancedConceptsAbout 18 minutes

Understand → Implement → Debug → Design

Success rates and retry accounting for stochastic tasks

Examine randomness, task-level statistics, pass@k, and stability.

pass@kStabilitystatistics

Knowledge content check2026-10-03 · Check the source of the original question2026-10-02

Which step do you want to learn from this knowledge point?

Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.

Understand first

New to this knowledge point

Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.

Start with core principles →

Implement next

Prepare to write the principles into code

Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.

Reading implementation and trade-offs →

Debug failures

Need to handle failures and changes in conditions

Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.

Continue to delve deeper into the problem →

Compare designs

Need to design or review plans

Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.

Analyze engineering scenarios →
Knowledge unit directory

LEARN · PRACTICE · REFLECT

Knowledge learning and personal records

My notes and review ↗

First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.

Answers and personal notes

Each modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.

Core concept · Success rates and retry accounting for stochastic tasks

Understand the core principles first

Preparatory concepts:proportion estimate, Repeat experiment, Task grouping

Reliable once, successful at least once many times, and successful every time are different questions. Metrics must be bound to the user's actual attempt policy and preserve differences between tasks and between repetitions of the same task.

Define the user’s attempt policy first

One successful email out of five attempts is inappropriate when only one send is allowed. Coding exploration can legitimately generate several candidates and select through tests. With accurate success detection, independent attempts, and fixed probability p, at-least-one success is 1−(1−p)^k and all-success is p^k. Without those assumptions, measure the actual policy instead of applying the formulas mechanically.

Separate tasks from repeated runs

Ten tasks repeated ten times produce 100 runs but only ten distinct tasks. Report per-task successes, strata, and aggregation rules. Repetitions share difficulty and environment; treating them as independent new tasks exaggerates generalization.

Report uncertainty and its assumptions

Small-sample proportions fluctuate. Independent Bernoulli samples with a common success probability can use intervals such as Wilson’s. Heterogeneous fixed task sets need an explanation before using binomial intervals. Analyze repeated runs by task, resampling whole tasks where appropriate. Zero temperature does not guarantee end-to-end determinism: tools, versions, and scheduling can vary. These formulas are conditional derivations, not new measurements.

Check understanding with a question

If you run the same question ten times and only succeed six times, how do you report and select the online version?

Report single-attempt success, repeated-run distributions, task counts, attempts, sampling, and cost. At-least-once and every-attempt success answer different questions. Compare versions on paired tasks and estimate uncertainty at task level. Match release criteria to actual allowed retries rather than presenting the best repeated result as single-attempt reliability.

Implementation and trade-offs

Clarify the denominator and delivery method

Six out of ten successes is an observation of the task under specified conditions and is not a proof of accuracy for all tasks. If the actual product is only run once, the single performance should be considered; if it is automatically tried multiple times, the costs, delays and side effects of all attempts must be included. Write actions with side effects do not always allow for multiple trials and errors to find one success.

Different indicators answer different questions

At least one success is suitable for evaluating the exploration potential of multiple candidates, while multiple successes are closer to the stability requirement. Don't confuse pass@k with the business goal of needing to be reliable every time. Under the assumption of independent and identically distributed trials, 1-(1-p)^k can be used to estimate at least one success, but in fact, failures of the same task are often related, and the estimate cannot be directly used as a measurement result.

Compare two versions

Fixed tool environment, data snapshots, allowed budgets and acceptance criteria, comparing results on a question-by-question basis. Easy and difficult tasks are observed separately to avoid one version only improving on a large number of easy questions. By repeatedly running related observations belonging to the same task, confidence estimates can be resampled by task grouping, rather than treating all attempts as independent new problems. Small sample improvements should clarify uncertainties.

Online threshold

Define key violations as independent blocking items, with thresholds for quality, cost and tail latency. If the success rate of the new version is slightly higher but the cost is doubled, you should consider whether the cost of each successful delivery and user waiting are acceptable. Canary release first and monitor real failures before deciding to expand traffic; don’t skip repeatable evaluations just because of a great demo.

Engineering deduction

scene
Interview hypothesis: Version A succeeds six times out of ten, version B succeeds seven times, but twice the amount of B is used.
design decisions
Expanded fixed-task integration pair comparison, reporting per-shot quality and total attempt cost.
Verify target
Able to judge whether improvements are stable and meet actual delivery requirements.
applicable boundary
Small sample sizes do not support firm conclusions and require reporting of uncertainties.

Continuous questions and answers

Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.

Draw inferences from one example: If the conditions change, how to deduce it?

First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.

Code suggestions allow test selection

Changing conditions:User accepts multiple candidates, external side effects occur after selection.

Extended question:How should I choose a version?

Derivation and reference solutions

Evaluate the entire candidate generation and test selection process, reporting true passes, false selections, and total time within budget; in addition, retain the quality of the first candidate. The benefits from multiple attempts must be deducted from testing and generation costs, and only the best patches cannot be displayed.

The principles that remain unchanged:Success rates are defined together with try and select policies.

Automatic refunds must not create duplicate effects

Changing conditions:Tasks have financial side effects, and at most one business can be refunded.

Extended question:Does six out of ten calls return success equal to 60%?

Derivation and reference solutions

First check whether the ten times are ten independent refund tasks or reissues of the same business key; the latter may have only one effect. Statistics on actual refunds and repeated violations of independent tasks, statistics on recovery and reconciliation capabilities for reissues, and requesting receipts cannot be used as business samples.

The principles that remain unchanged:The denominator is the clear task, and the business effect is separated from the number of calls.

Easy to make mistakes

  • Only keep the best once
  • Repeat attempts at any cost
  • Treat related duplicate samples as independent tasks

References

It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.

Check how far you understand

After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.

Basic standards met
Complete reporting on the number of successes, failures, and total attempt costs for repeated runs.
Intermediate and advanced signals
Can distinguish between once, multiple times at least once and multiple times successfully.
Senior criteria
Pairwise task comparisons, grouped uncertainty, and independent risk thresholds were used.

Continue to do advanced research experiments

Let reviews read events and actual effects

Differentiate between run success, result contracts, and unscored semantic quality along actual checkpoints, events, and servicer receipts for the same research task.

Read full text and fault analysis → · Download Reliability Experiment v3 ↓

python3 cli.py memory-put
python3 cli.py submit
python3 cli.py run
python3 evaluate.py
python3 -m unittest discover -s . -p test_lab.py -v

Keep evidence and check item by item

  • Explain the basis of passed in terms of events and checkpoints.
  • Point out credible observations of scope, references, idempotent effects, and step budgets respectively.
  • Semantic_support and model_quality are not_scored, and passed cannot be written as model quality score.

By default, deterministic summary functions and synthetic documents are used; real operating trajectories and mechanical constraints are evaluated, and there is no semantic support to verify the real model.

Hands-on verificationComplete on demand · Suggestions15 minutes

Explain how the results of A and B running three tasks each three times should be summarized, and which conclusions cannot be directly derived.

Expand acceptance requirements and checkpoints
  • Failed to delete without merit
  • Repeat count and budget transparency
  • Risk violations are reported individually

Key inspections

  • The metric denominator is consistent with the user retry strategy
  • Understand the correlation of repeated observations
  • Comprehensive decision-making based on quality cost risk