Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content

LEARNING PATH

Agent application learning paths

First use the seven-day introduction to establish concepts, then complete a set of runnable application experiments, and continue to study failure, recovery and system design.

20–30 minutes per day · Can be completed locally

Getting Started with Backend to Agent in 7 Days

For those of you who already have experience in back-end development and are new to Agent. 20–30 minutes per day, starting with three short questions and then completing a small task using a local simulation.

5 minutes of short answers, 5-10 minutes of reading, and 10-15 minutes of local experiments every day; for longer questions, only read the designated part of the day.

Paper and pencil, pseudocode, or a language you are familiar with are all acceptable. No model account, payment interface, or online execution is required. The seven-day output is an introductory experiment record, and whether it can be mastered still needs to be judged through independent answers and verification.

Day 1First run through the minimum execution closed loopThink of the Agent as a budget-constrained backend loop, distinguishing between model recommendations and runtime control.About 25 minutes

Answer independently first · 5 minutes

  1. What is the most obvious difference between Agent and ordinary interface calls?

    Answer with reference

    Ordinary interfaces are usually executed in fixed steps; the Agent can continue to call tools based on the next step suggestions given by the model. The runtime is still responsible for verification, execution and stopping.

  2. Can the runtime end when the model says "done"?

    Answer with reference

    Check the acceptance conditions of the task first. Model text can only be candidate results; there should be a clear end state when conditions, step limits, or unrecoverable errors are encountered.

  3. Why does the tool result still return the model?

    Answer with reference

    Let the model continue to make decisions based on real receipts. The response receipt should differentiate between success, failure and unknown results, and failure should not be packaged as completed.

Reading Correspondence · 5–10 minutes

Read the short answer and termination conditions of "Agent Loop" first; the essential LLM concepts focus only on tool invocation messages and structured output.

Local tasks · 10–15 minutes

Simulate a loop that queries order status: replace the model response with a fixed array and execute up to 3 steps.

  1. Write down the input orderId, output status/source, and the maximum number of steps, 3. Prepare a fake response of "Call Query Tool → Return Shipped".
  2. Let the query tool return {status: "shipped", source: "mock"}; record suggestions, tool receipts and end reasons at each step.
  3. Change the response to continuously calling the query tool and observe whether it stops after step 3; then change it to a non-existent tool name.

Check items one by one during acceptance

  • The normal path generates a query receipt, the output is marked source=mock, and the end reason is that the acceptance is met.
  • Repeatedly calling the path for up to 3 steps, the output budget is exhausted and the query cannot be claimed to be complete.
  • The unknown tool will not be executed, and there are errors in the log indicating that the tool does not exist and a clear end reason.

Pay attention to these fault symptoms:If the loop keeps running, declares success textually, or treats a non-existent tool as if it has been executed, it means that the runtime boundary is not implemented.

Day 2Add input boundaries to toolsFollow the familiar interface verification and check the format, business conditions and calling permissions separately.About 25 minutes

Answer independently first · 5 minutes

  1. If the JSON Schema verification passes, will the business request be valid?

    Answer with reference

    Not necessarily. Schema checks the structure and type; business conditions such as amount range and resource ownership also need to be verified on the server side.

  2. Can the userId provided by the model be directly used for permission determination?

    Answer with reference

    No. The identity of the caller comes from the trusted session, and the model parameters can only describe the request object, but cannot determine who they are or what permissions they have.

  3. What should be returned if verification fails?

    Answer with reference

    Return stable error codes and correctable information, such as INVALID_AMOUNT; do not execute the tool and do not report failure as success.

Reading Correspondence · 5–10 minutes

Read-only short answer, parameter verification sequence and basic compliance standards; complex authorization structure can be read later.

Local tasks · 10–15 minutes

Add validation to the local mock's "Create Refund Draft" tool and do not connect payments or real orders.

  1. Fixed session userId=u1; Prepare orders o1 (Vested to u1, Paid 100) and O2 (Vested to U2, Paid 100).
  2. Verify that orderId is a string and amount is a positive number; then verify that the order ownership and refund amount do not exceed the actual payment. Appends the local drafts array only after passing.
  3. Enter o1/20, o1/"20", o1/120, o2/20 in sequence, and record the return value and drafts length.

Check items one by one during acceptance

  • o1/20 Create a local draft, and the result is clearly marked that no refund has been performed.
  • The string amount, excess amount, and other people's orders are rejected respectively. After rejection, the length of drafts is still 1.
  • Modifying the incoming userId will not change the fixed session identity, and you will not be able to access o2.

Pay attention to these fault symptoms:A new draft is added to any rejected use case, or the model parameter covers the session identity, indicating that there is a gap in the verification sequence or authorization boundary.

Day 3Explain RAG with evidenceRecord the search, context and answers separately, and clarify the insufficient answers when encountering questions without information.About 30 minutes

Answer independently first · 5 minutes

  1. Is the model in RAG responsible for finding the document?

    Answer with reference

    The retriever finds candidate material, and the model generates an answer from the context actually supplied. Inspect retrieval and context first, then assess whether the answer is faithful to them.

  2. Is it enough to search for relevant titles?

    Answer with reference

    Not enough. Make sure the paragraph contains the facts you need to answer, and check that the paragraph enters the context of the model.

  3. What to do if there is no answer in the information?

    Answer with reference

    Return insufficient information and indicate which piece of evidence is missing. A seemingly reasonable company rule cannot be made up based on common sense.

Reading Correspondence · 5–10 minutes

Read-only search chain positioning method and short answer; today use keyword search to understand the link without installing the vector library.

Local tasks · 10–15 minutes

Use three pieces of local text to simulate a policy Q&A with citations, retaining evidence for each step.

  1. Prepare d1: "Apply for standard order refund within 7 days"; d2: "Member points are updated monthly"; d3: "Refund requires an order number". Set id for each item.
  2. Filter by the "refund" keyword, and then select d1 containing the number of days for "apply for a refund within a few days"; record the candidate id, actual context and answer reference.
  3. Test "whether membership refund is 30 days"; then deliberately remove d1 from the context and observe whether the answer turns into insufficient information.

Check items one by one during acceptance

  • The first answer says 'within 7 days,' cites d1, and has the corresponding original text in its context.
  • The information does not state an exception for members, and the second question cannot claim that "members have 30 days."
  • After removing d1, it is clearly reported that there is insufficient data, and the log can see "retrieved but no context passed in".

Pay attention to these fault symptoms:The citation id exists but the original text does not support the answer, or 7 days are still given after missing context, indicating that the generation is disconnected from the evidence.

Day 4Distinguish between task state and long-term memoryRecord which step of the current task is completed, and avoid treating task data as a memory shared by all users.About 25 minutes

Answer independently first · 5 minutes

  1. Are chat history, task status, and long-term memory the same thing?

    Answer with reference

    No. Chat history is a message; task status records the stages and intermediate results of this task; long-term memory stores filtered information that can be used in subsequent tasks.

  2. Should this request's order number be stored in memory shared by every task?

    Answer with reference

    Usually not required. It belongs to the current task status and will be processed according to the life cycle after the task is completed; long-term storage requires clear purpose and scope.

  3. Why do we still need to verify the scope when reading memory?

    Answer with reference

    Because similar content may come from other users or projects. Reads must be restricted by trusted organization, user or thread identities and cannot be matched by keywords only.

Reading Correspondence · 5–10 minutes

Read-only scoping and the short answer; the distinction between thread state and long-term memory is the focus today.

Local tasks · 10–15 minutes

Use two objects to simulate task status and user preferences, and check cross-task and cross-user reads.

  1. Create taskState[runId] and put step, orderId, and lastToolResult; create preferences[userId] and put only the language preference confirmed by the user.
  2. Let u1's run1 execute to query and save the result; when u1 starts run2, it only reads the language preference and does not reuse the order results of run1.
  3. Have u2 request to read u1's preferences, using fixed session authentication; then use the same run1 to read the state and continue to the next step.

Check items one by one during acceptance

  • run2 can get the language preference of u1, but it does not have the order number and tool results of run1.
  • u2 cannot read u1 preferences, and the user's identity is not determined by the parameters passed in by the model.
  • Resuming run1 can see the saved step and tool results and make it clear that they are only local experimental status.

Pay attention to these fault symptoms:New tasks inherit the results of old orders, all users share a preference object, or similar searches bypass identity restrictions, indicating that the scopes are mixed.

Day 5Simulation failure and recoveryKnow what retries, checkpoints, and idempotency solve respectively, and use receipts to handle "executed but no response received".About 30 minutes

Answer independently first · 5 minutes

  1. Does the timeout necessarily mean that the tool was not executed?

    Answer with reference

    Not necessarily. The request may have been executed, but the response was lost. The write operation should first check the receipt or use a stable idempotency key, and cannot directly replace the request with a new one and try again.

  2. Can checkpointing ensure that a write operation only happens once?

    Answer with reference

    No. External writes may not be in the same transaction as the local checkpoint and will be replayed during recovery. Operational idempotency and queryable receipts are also required.

  3. Should all errors be retried?

    Answer with reference

    It shouldn't be. Parameter and permission errors are corrected first, and transient errors are retried with limited time. Limit times and deadlines, and retain status in case of cancellation or budget exhaustion.

Reading Correspondence · 5–10 minutes

Read the short answer to retry and budget first; the checkpoint question only focuses on the boundary of "recovery may be repeated", leaving advanced content such as leases for the complete route.

Local tasks · 10–15 minutes

Simulate the situation where the local draft tool has been executed but the response is lost, and avoids repeated drafting after recovery.

  1. Use the local drafts of Day 2 and add receipts[operationId]; fix this business key run1:create-draft. Returns the same draftId when there is already a receipt.
  2. After first creating a draft and writing receipts, deliberately simulate a lost response and leave the task checkpoint pending.
  3. Resume execution with the same operationId and record the recovered receipt; test another mock that always fails, and try it a maximum of 2 times.

Check items one by one during acceptance

  • After the response is recovered, there is still only 1 drafts, and the recovery result returns to the original draftId.
  • Retries use the same business key, and the log can distinguish between the first execution and reading of existing receipts.
  • The total failure mock stops and marks failure after 2 times without claiming completion; indicating that the memory receipt will be lost after the process is restarted.

Pay attention to these fault symptoms:During recovery, repeated drafts, infinite retries, or memory objects are called persistence guarantees after changing the idempotency keys indicate that there are still omissions in the recovery plan.

Day 6Write an acceptance that can be checkedChange "Answer looks good" to a record of results, bounds, and execution that can be inspected step by step.About 25 minutes

Answer independently first · 5 minutes

  1. Does correct output format equal mission success?

    Answer with reference

    Not equal to. Also check facts, citations, permissions, and actual business receipts; legitimate JSON may also contain incorrect answers.

  2. Why include failed use cases in the review?

    Answer with reference

    The normal path cannot expose problems such as insufficient evidence, overstepping of authority, and duplication of recovery. The failure use case checks whether the system has stopped safely and reports the results truthfully.

  3. One day small sample passed, can you declare production available?

    Answer with reference

    No. It is only possible to report which use cases were tested, what results were observed, and the boundaries that were not covered. Production judgment also requires representative data and more complete verification.

Reading Correspondence · 5–10 minutes

Read-only results acceptance contract and basic standards; use manual checklist today and do not introduce LLM graders.

Local tasks · 10–15 minutes

Organize 4 acceptance use cases for the first five days of local experiments, recording actual results and uncovered items.

  1. Create a form: input, expected results, actual results, evidence, pass or fail. Four use cases are added: normal refund draft, other people's orders, insufficient data, and response loss recovery.
  2. Rerun or manually deduce existing mocks one by one, and save the number of drafts, original text quotes, error codes or recovery receipts. For use cases that have not been executed, write "not executed" and do not record them as passed.
  3. Deliberately remove the ownership verification of Day 2 and observe whether the unauthorized use case fails; review after restoring the verification and write down the untested database failures and real model performance.

Check items one by one during acceptance

  • Each test case has a clear expected result, and actual observations have supporting evidence. Record unexecuted cases separately from failures.
  • Removing the ownership check makes the unauthorized-access test fail; restoring it rejects draft orders for other users.
  • The result report only covers these 4 local mocks, and the use case coverage is clearly listed.

Pay attention to these fault symptoms:All use cases always pass, only check that the output has text, or mark unexecuted items as passed, indicating that the acceptance criteria cannot detect real problems.

Day 7Turn the experiment into a questionable projectUse real experimental records to explain goals, boundaries, failures and trade-offs, and distinguish between proven and unrealized.About 25 minutes

Answer independently first · 5 minutes

  1. Which part should you talk about first when introducing the Agent project in an interview?

    Answer with reference

    Let’s first talk about user tasks, input and output, and acceptance, and then talk about how models, tools, and states work together. This way every design choice has a clear purpose.

  2. Can these experiments be described as a production project without connecting a real model?

    Answer with reference

    No. It should be stated that it is an introductory experiment for local mocking. Which specific behaviors have been verified, and which ones involve the real model, persistence or online capabilities have not yet been verified.

  3. Need multi-agent or complex frameworks now?

    Answer with reference

    Start by choosing the smallest option based on status, recovery, and collaboration needs. This introductory experiment can use ordinary code loops; more complex solutions must have clear requirements and comparative evidence.

Reading Correspondence · 5–10 minutes

Read the short answers to the selection questions first; the project follow-up questions are only used to check your evidence and do not require completion of advanced production architecture.

Local tasks · 10–15 minutes

Introduce your Order Q&A and Refund Draft local experiment with one page of notes and 90 seconds of oral presentation.

  1. Draw Input → Loop → Tools/Retrieval → Status → Output, marking the locations of fixed sessions, budgets, and acceptance checks.
  2. Choose a normal and failure trajectory each, cite the evidence that was actually saved in the past few days; write down why you used mocks and code loops in the first place.
  3. Record or time the dictation for 90 seconds, and then answer "Has the tool timeout been executed?" "How to answer when there is no document?" "How to recover after restarting"; unverified content is listed as follow-up tasks.

Check items one by one during acceptance

  • The notebook includes goals, inputs and outputs, acceptance criteria, an architecture diagram, and two evidence-backed traces.
  • Orally can explain model and runtime boundaries, scoping, idempotency keys, and handling of insufficient data.
  • Clearly mark local mocks, memory status, and manual evaluation. Items to be implemented include at least persistence, real model access, and more evaluation use cases.

Pay attention to these fault symptoms:Telling only technical terms, without receipts or experimental records, or referring to mock results as production success rate, indicates that the project evidence is incomplete.

Connect the basics into one implementation

Be a return policy assistant who can accept acceptance

Use the same set of fictional orders, policy documents, and issues throughout response processing, prompt iteration, read-only tools, and RAGs. Each step preserves the actual output, and finally migrates the acceptance method to reliably perform experiments.

The default experiment does not require an account and uses the standard library and saved real local Embedding vectors; the generated responses are clearly labeled artificial fixtures. The course also provides real interface request examples, and you need to call and record the model results yourself. Offline passing does not mean that the model quality or production has been accepted online.

Download the complete application experiment v1 ↓ · Python 3.10+, unzip and run in the directory; optional reconstruction vector requires Python 3.10+, additional dependencies and model download.

01

Catch every model response

Distinguish between text, tool calls, rejections, incomplete and failed requests to avoid treating intermediate results as business completion.

python3 model_response_contract.py

Delivery in this step:Save the response classification, call_id binding and next round of input; the real interface experiment also saves the model name, parameters, request identification and actual response.

Item-by-item acceptance

  • Normal read-only query gets answer_ready
  • Cross-tenant request returns permission_denied, tool does not return protected order
  • Incomplete response stops, duplicate calls and batch calls are subject to a shared budget
02

Turn answer requests into inspectable contracts

Three questions are fixed to check the structure, business tags, references and behavior when unable to answer.

python3 prompt_iteration.py

Delivery in this step:Keep v1/v2 prompts, candidate answers, and failure reasons one by one; replace the candidate answers and use the same set of labels when accessing the model.

Item-by-item acceptance

  • Tags don't go into model input
  • Boolean true cannot pretend to be a legal number of days, and old versions and unauthorized references cannot pass
  • Fixture's v1/v2 passes are 1/3 and 3/3 respectively; semantic quality is still not_scored
03

Integrate order query into bounded tool loop

Submit the calls made by the model to the server for verification, and only allow the current user to read the orders of the current tenant.

python3 -m unittest test_application.ApplicationLabs.test_readonly_loop_requires_a_matching_tool_receipt test_application.ApplicationLabs.test_cross_tenant_request_has_no_resource_receipt test_application.ApplicationLabs.test_unknown_tool_is_not_executed -v

Delivery in this step:Track a request, authorization judgment, tool result and final answer in model_response_contract.py and draw their respective trust boundaries.

Item-by-item acceptance

  • Tool whitelist rejects unknown write operations
  • Tenant and user identities are bound by server-side fixtures and are not determined by model parameters.
  • The final text must be consistent with the successful tool receipt, and completion will not be declared if the receipt is missing.
04

From policy document to quoted answer

Actually run parsing, chunking, SQLite filtering, vector and keyword retrieval, RRF fusion, and then check answer citations.

python3 rag_pipeline.py
python3 application_project.py

Delivery in this step:Keep chunk ID, source and text summary, model version, two-way ranking, final context and acceptance results.

Item-by-item acceptance

  • Ordinary Returns vs. Battery Exception Requires Evidence to Enter Context
  • Old policy with another tenant's policy is filtered before retrieval
  • The unknown tax question is answered as unknown even if the document is retrieved; the vector must be reconstructed after changing the document or query.
05

Leave reproducible evidence of failures and rejections

Retrieval recall, output structure, authorization, and citations are reported separately by result contract, accounting for unscored semantic parts.

python3 -m unittest discover -v

Delivery in this step:Save test records and failure samples to explain top-k shortages, old vectors, override, no acknowledgment, and budget exhaustion respectively.

Item-by-item acceptance

  • A total of 32 tests passed in applied experiments and five basic experiments.
  • Missing exception evidence can be reproduced after reducing top-k
  • Error texts whose structures and labels pass still require manual semantic checking and are not summarized into model accuracy.
06

Migrate contracts to process recovery and memory revocation

Enter independent reliability research experiments to see how process exits, idempotent receipts, and memory failures affect recovery.

在对应课程的进阶实验中下载可靠性实验 v3,按 README 在独立目录运行。

Delivery in this step:Save events, checkpoints, and server receipts after first exit and recovery, and explain why an old draft had to be stopped.

Item-by-item acceptance

  • The process actually exits after collect. It recovers successfully after the lease expires and only submits collect once.
  • Exit after external success. Retry will only have one server effect.
  • Old drafts stop publishing when the memory is revoked; this independent experiment does not share status with the Return Assistant

Continue to go deeper

Complete project route

Learn in depth along the actual development sequence. At each stage, the knowledge principles and implementation are first understood, and then short answers and boundary experiments are used to check understanding.

01

Set up an execution closed loop

Can clearly draw the boundaries between models, states, tools, and runtime; explain why tool calls require verification, authorization, and receipts.

Hands-on verification

Implement a tool loop with a step budget. Invalid parameters, tool failures, and duplicate writes are simulated, and termination and compensation strategies are explained.

02

Access data and memory

It can locate errors along the retrieval chain, distinguish task status and cross-task memory, and implement permissions and revocation to the query link.

Hands-on verification

Create a dataset covering the same topic for different tenants. Verify that retrieval, caches, and historical context all respect permissions.

03

Verify long task reliability

Can illustrate the combination of checkpoints, leases, idempotency and cancellation; use result contracts and execution trajectories to jointly judge quality.

Hands-on verification

Inject downtime before and after tool requests. Document recovery paths, repeat requests, and external receipts, and use measurements to check whether you're accomplishing real business goals.

04

Complete the production system plan

Can explain link tracking, approval boundaries, framework selection, cost budget and system responsibilities; provide an enterprise architecture that can be accepted.

Hands-on verification

Design an enterprise research assistant: remove the request entrance, orchestration, tool gateway, retrieval, memory, evaluation and operation and maintenance, and explain the online and rollback conditions.