Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content
knowledge unit 03AdvancedImplementationAbout 15 minutes

Understand → Implement → Debug → Design

Task contracts and independent acceptance checks

Make the execution environment, permissions, status, evidence and completion judgment of the model into a project contract.

Harnessmission contractAcceptanceexecution environment

Knowledge content check2026-10-03 · Check the source of the original question2026-10-02

Which step do you want to learn from this knowledge point?

Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.

Understand first

New to this knowledge point

Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.

Start with core principles →

Realize again

Prepare to write the principles into code

Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.

Reading implementation and trade-offs →

Will troubleshoot

Need to handle failures and changes in conditions

Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.

Continue to delve deeper into the problem →

Able to choose

Need to design or review plans

Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.

Analyze engineering scenarios →
Knowledge unit directory

LEARN · PRACTICE · REFLECT

Knowledge learning and personal records

My notes and review ↗

First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.

Answers and personal notes

Each modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.

Core concept · Task contracts and independent acceptance checks

Understand the core principles first

Preparatory concepts:access control, immutable version, Test acceptance

Harness turns goals into executable boundaries and inspectable artifacts. The model modifies the solution, not by modifying permissions or acceptance criteria to create a successful solution; recovery reconstructs the view of the work based on facts, rather than treating a summary of progress as the truth about the system.

Prompts cannot enforce every boundary

Prompts guide interface use but provide no transaction, credential, or filesystem controls. If a gateway permits deleting acceptance files, “do not change the tests” remains a soft constraint. Separate task definitions, candidate artifacts, and acceptance procedures into permission domains. Runtime restrictions come from trusted task records.

Turn vague goals into observable evidence

“Fix login” needs permitted directories, an input baseline, expected behavior, and a validation environment. Contract changes need an explicit source and version, identifying which earlier evidence becomes stale. An independent verifier checks the current artifact. Model-written progress flags aid coordination without deciding success.

Recovery reconciles the current facts

Progress notes describe the previous position; the repository and tool receipts show what exists now. Diagnose inconsistencies before proceeding instead of blindly repeating an old summary. Anthropic’s long-task articles motivate engineering records. Independent verification and permission separation here are additional design guidance, not automatic SDK guarantees.

Check understanding with a question

What responsibilities does Agent Harness have? How to design input and acceptance contracts for tasks?

A harness surrounds the model with tool access, permissions, budgets, persistence, environment preparation, event recording, and verification. Specify goals, constraints, allowed actions, artifacts, and executable acceptance checks. Reconstruct trusted context on recovery. Models propose plans and results; the harness determines permitted actions, evidence support, and completion.

Realization and trade-offs

Express harness responsibilities through interfaces

The system can be divided into model adapters, tool gateways, state stores, policy controls, verifiers and event pipelines. The model adapter calls and responds uniformly, the tool gateway performs unified verification and authentication, the state store records the running facts, and the verifier checks the artifact. Prompt words describe the rules, but actual forbidden actions are rejected by the gateway. Anthropic's long-running engineering article uses progress notes, feature lists, and version history to help new contexts take over the work, which provides a practical direction: continuous work relies on readable project status, not just chat memory.

Specify a concrete task contract

Inputs include task_id, goal, resource scope, baseline version, deadline, and budget. Allowed actions are limited by tools and resources. For example, specified branches can be modified, but production libraries cannot be operated. The artifact contract specifies the file path, format, source references, and required fields. The acceptance contract directly writes observable behaviors, such as "unlogged users can read public articles; the private management interface returns rejection; the build passes." The broad "experience is good and the content is rich" cannot replace acceptance. Models can suggest additional tests, but cannot remove existing failure criteria to create success.

Long task recovery requires fact layering

At a minimum, the recovery package includes the most recently confirmed goals, current baseline version, completed subtasks, open issues, evidence locations, and next step candidates. The original log is retained for auditing, and the summary is used to reduce context costs. The two cannot replace each other. Recheck the repository status, artifact summary and authorization validity period before recovery; when the database records "Completed" but the files are missing, you should enter diagnosis. Progress only indicates the status of the work, with final success confirmed by an independent verifier. Manual approval saves the object version, approver, and scope instead of saving a global approved=true.

Verify the contract

First, use a deterministic fake model to simulate exceptions: claiming success but no artifact, the artifact being modified externally, trying to access resources outside the scope, and the context being cleared and then restored. Every time an assertion is denied or diagnosed by contract, there must be reviewable evidence for success. Then use the real model to run typical tasks and record the completion rate, reasons for manual intervention, and costs. Harness is not about packaging Agent into a universal platform; it should start with a specific type of task, clarify which checks are sure to be executed, which judgments still require manual work, and then gradually expand.

Engineering deduction

scene
Hypothetical engineering scenario: The agent modifies the website across multiple rounds, and it is easy to forget the unfinished mobile behavior and announce the launch in advance.
design decisions
Establish a versioned task list and acceptance records; read the current code and progress during recovery; deployment is triggered by artifacts that pass acceptance.
Verify target
Expected behavior: The task remains unfinished when the mobile directory is not accepted, and the new context can continue to process this item.
applicable boundary
This is a general architecture example and does not claim to have unattended development capabilities; visual quality may still require human evaluation.

Continuous questions and answers

Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.

Draw inferences from one example: If the conditions change, how to deduce it?

First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.

Delivery relies on third-party APIs

Changing conditions:The test environment is not representative of live external services

Extended question:Does passing the local test prove that the task has been completed?

Derivation and reference solutions

Distinguish between local contract verification and true integration verification. Use isolation credentials to verify necessary links, deliver code and clear unverified items when inaccessible, and cannot claim remote success. The transaction receipt, permission scope and environment identification returned by the third party serve as evidence, and the simulated response only proves the local processing logic.

The principles that remain unchanged:Every successful conclusion requires evidence consistent with the scope of its claim.

Design draft without automatic criteria

Changing conditions:Product quality requires human judgment

Extended question:How to maintain independent acceptance?

Derivation and reference solutions

First define automatic checks such as size, format, and required content, and then allow authorized reviewers to review the fixed version according to specific usage goals. The approval record is bound to this version; the model can revise the draft and explain the trade-offs, but cannot write approved=true to itself. While there is a subjective component to quality scoring, the acceptance process still maintains clear sources and traceable results.

The principles that remain unchanged:Acceptance can be performed by humans, but the determination of source and artifact version must be independent of the self-assessment text.

Easy to make mistakes

  • Equivalent Harness to the Prompt template or model itself
  • Only natural language summaries are saved, evidence and versions are not saved
  • Treat model self-evaluation as final acceptance

References

It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.

Check how far you understand

After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.

Basic standards met
A model also needs a runtime, tools, and state management.
Intermediate and advanced signals
Ability to write inputs, permitted actions, budgets, artifacts and acceptance contracts.
Senior Signal
Make policy enforcement, isolation, recovery, and evidence collection verifiable boundaries.

Hands-on verificationComplete on demand · Suggestions15 minutes

Write the task contract and acceptance conditions for "Fix an API and deliver the patch".

Expand acceptance requirements and checkpoints
  • Acceptance is executable
  • Allow modification scope to be clear
  • Test evidence is linked to the actual artifact

Key inspections

  • Distinguish between prompts and runtime executable constraints
  • Can specify input, artifact, acceptance, and recovery contracts
  • Acceptance evidence cannot be replaced or deleted by the model itself