Understand first
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Understand → Implement → Debug → Design
Make the execution environment, permissions, status, evidence and completion judgment of the model into a project contract.
Knowledge content check2026-10-03 · Check the source of the original question2026-10-02
It is recommended to understand first:
Control and completion criteria in agent loops →Tool calls: structure, authorization, and business contracts →Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Prepare to write the principles into code
Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.
Reading implementation and trade-offs →Need to handle failures and changes in conditions
Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.
Continue to delve deeper into the problem →Need to design or review plans
Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.
Analyze engineering scenarios →LEARN · PRACTICE · REFLECT
First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.
Can be practiced directly. After logging in, answers, favorites, and notes will be saved to your account.
Log in and saveEach modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.
Core concept · Task contracts and independent acceptance checks
Preparatory concepts:access control, immutable version, Test acceptance
Harness turns goals into executable boundaries and inspectable artifacts. The model modifies the solution, not by modifying permissions or acceptance criteria to create a successful solution; recovery reconstructs the view of the work based on facts, rather than treating a summary of progress as the truth about the system.
Prompts guide interface use but provide no transaction, credential, or filesystem controls. If a gateway permits deleting acceptance files, “do not change the tests” remains a soft constraint. Separate task definitions, candidate artifacts, and acceptance procedures into permission domains. Runtime restrictions come from trusted task records.
“Fix login” needs permitted directories, an input baseline, expected behavior, and a validation environment. Contract changes need an explicit source and version, identifying which earlier evidence becomes stale. An independent verifier checks the current artifact. Model-written progress flags aid coordination without deciding success.
Progress notes describe the previous position; the repository and tool receipts show what exists now. Diagnose inconsistencies before proceeding instead of blindly repeating an old summary. Anthropic’s long-task articles motivate engineering records. Independent verification and permission separation here are additional design guidance, not automatic SDK guarantees.
A harness surrounds the model with tool access, permissions, budgets, persistence, environment preparation, event recording, and verification. Specify goals, constraints, allowed actions, artifacts, and executable acceptance checks. Reconstruct trusted context on recovery. Models propose plans and results; the harness determines permitted actions, evidence support, and completion.
The system can be divided into model adapters, tool gateways, state stores, policy controls, verifiers and event pipelines. The model adapter calls and responds uniformly, the tool gateway performs unified verification and authentication, the state store records the running facts, and the verifier checks the artifact. Prompt words describe the rules, but actual forbidden actions are rejected by the gateway. Anthropic's long-running engineering article uses progress notes, feature lists, and version history to help new contexts take over the work, which provides a practical direction: continuous work relies on readable project status, not just chat memory.
Inputs include task_id, goal, resource scope, baseline version, deadline, and budget. Allowed actions are limited by tools and resources. For example, specified branches can be modified, but production libraries cannot be operated. The artifact contract specifies the file path, format, source references, and required fields. The acceptance contract directly writes observable behaviors, such as "unlogged users can read public articles; the private management interface returns rejection; the build passes." The broad "experience is good and the content is rich" cannot replace acceptance. Models can suggest additional tests, but cannot remove existing failure criteria to create success.
At a minimum, the recovery package includes the most recently confirmed goals, current baseline version, completed subtasks, open issues, evidence locations, and next step candidates. The original log is retained for auditing, and the summary is used to reduce context costs. The two cannot replace each other. Recheck the repository status, artifact summary and authorization validity period before recovery; when the database records "Completed" but the files are missing, you should enter diagnosis. Progress only indicates the status of the work, with final success confirmed by an independent verifier. Manual approval saves the object version, approver, and scope instead of saving a global approved=true.
First, use a deterministic fake model to simulate exceptions: claiming success but no artifact, the artifact being modified externally, trying to access resources outside the scope, and the context being cleared and then restored. Every time an assertion is denied or diagnosed by contract, there must be reviewable evidence for success. Then use the real model to run typical tasks and record the completion rate, reasons for manual intervention, and costs. Harness is not about packaging Agent into a universal platform; it should start with a specific type of task, clarify which checks are sure to be executed, which judgments still require manual work, and then gradually expand.
Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.
Level 1How to prevent the Agent from modifying the acceptance criteria to pass itself?
When the model controls the candidate implementation, check again whether it also controls the grader.
Place the trusted acceptance definition in a non-writeable area of the model, or load and verify the digest from a fixed version. The editable tests in the working directory are only supplementary; the final check comes from the independent verifier, reading its exit status and evidence. Legitimate demand changes should generate a new contract version and explain the source of the change, and model deletion failure assertions cannot be allowed by default.
Follow this answer further
Level 2The independent test passes, but the Agent configures the service to be open only to the test account. Is it considered passed?
Independence solves the problem of tampering, and the next step is to test whether it truly represents the target.
This indicates that the existing acceptance coverage is insufficient. Add a representative account and configuration check that conforms to the actual scope of authority, and review whether it violates the original contract; if the original target requires ordinary users to be available, a single test account will not be successful enough. Protected acceptance to prevent question changes still cannot guarantee that the question itself is sufficient, and the behavioral conditions must be written in detail.
Follow this answer further
Level 3How to avoid constantly adding acceptance conditions so that the task can never be completed?
Strengthening acceptance also requires boundaries, otherwise the grader will be independent but unstable.
The scope of acceptance and key risks should be given before the start of construction. New inspections must indicate whether to verify the original requirements or expand the goals. The former completes the evidence, and the latter serves as a new task or an explicit contract change. Save the version and judgment basis so that both the model and the user can distinguish between omitted repairs and unlimited addition needs.
Level 1A constraint is lost after context compression. How to find out the details at runtime?
Summarization is lossy and hard limits cannot rely solely on model memory.
The amount, resource range and prohibited actions are saved in the server-side policy and re-verified before tool execution; the summary only helps model decision-making. When a constraint is missed, the gateway rejects and returns an understandable error, and the runtime rebuilds the context containing the constraint. For semantic constraints that cannot be judged by machines, the user's original words and sources are saved, and can be read back or manually reviewed before key decisions are made.
Level 1What should I do if the recovery record is inconsistent with the current repository version?
The persistent state may lag behind the actual environment, and recovery requires checking the facts first.
Pause operations that rely on older versions, compare baselines to current changes, and check documented artifacts and test evidence. Conflict-free results can be reused according to dependencies; relevant code or configuration changes will be re-verified, and permissions and approvals will also be re-confirmed. Do not overwrite the current repository directly and return to the old state. Protect external changes first and leave a record of the recovery decision.
First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.
Changing conditions:The test environment is not representative of live external services
Extended question:Does passing the local test prove that the task has been completed?
Distinguish between local contract verification and true integration verification. Use isolation credentials to verify necessary links, deliver code and clear unverified items when inaccessible, and cannot claim remote success. The transaction receipt, permission scope and environment identification returned by the third party serve as evidence, and the simulated response only proves the local processing logic.
The principles that remain unchanged:Every successful conclusion requires evidence consistent with the scope of its claim.
Changing conditions:Product quality requires human judgment
Extended question:How to maintain independent acceptance?
First define automatic checks such as size, format, and required content, and then allow authorized reviewers to review the fixed version according to specific usage goals. The approval record is bound to this version; the model can revise the draft and explain the trade-offs, but cannot write approved=true to itself. While there is a subjective component to quality scoring, the acceptance process still maintains clear sources and traceable results.
The principles that remain unchanged:Acceptance can be performed by humans, but the determination of source and artifact version must be independent of the self-assessment text.
It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.
After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.
Write the task contract and acceptance conditions for "Fix an API and deliver the patch".