Understand first
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Understand → Implement → Debug → Design
Examine tool catalog, dynamic discovery, recall quality, and server-side authorization.
Knowledge content check2026-10-03 · Check the source of the original question2026-10-02
It is recommended to understand first:
Tool calls: structure, authorization, and business contracts →Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Prepare to write the principles into code
Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.
Reading implementation and trade-offs →Need to handle failures and changes in conditions
Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.
Continue to delve deeper into the problem →Need to design or review plans
Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.
Analyze engineering scenarios →LEARN · PRACTICE · REFLECT
First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.
Can be practiced directly. After logging in, answers, favorites, and notes will be saved to your account.
Log in and saveEach modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.
Core concept · Separate discovery from execution in large tool catalogs
Preparatory concepts:Retrieval recall, Interface naming, Permission filtering
Choosing the right tool requires first discovering the applicable capabilities and then calling it based on the complete contract. The discovery layer is responsible for candidate coverage, and the execution layer is responsible for permissions and parameters; error hiding tools can improve selection, but they cannot replace the server's rejection of unauthorized actions.
“Look up order” and “search order documentation” can be confused. Clarify targets and input/output boundaries rather than only adding tools. If the correct tool never enters the candidate set, changing the call prompt cannot fix discovery.
Provide an authorization-filtered lightweight catalog, then load a few full schemas for the task. Expand discovery within bounds after a miss. Use stable identities and versions, isolating caches by tenant and permissions. This is application design: MCP listing and change notifications support discovery without specifying ranking.
A query API or controlled report may satisfy the same goal. Define allowed capabilities and constraints, measuring candidate recall, final selection, parameters, and business success separately. Unauthorized tools are not eligible ground truth. Execution failures can also expose API design problems, not merely model selection errors.
Analyze actual selection failures and overlapping names, descriptions, and parameters. Retrieve a few tools from an authorization-filtered catalog, expanding discovery when needed. Evaluate discovery separately from invocation: absence from context does not establish absence from the catalog. Execution still validates versions, permissions, and parameters; descriptions cannot override system rules.
Collect samples of mischosen tools and distinguish between synonymous naming, lack of applicable boundaries, parameter ambiguity, and true retrieval failures. Describe "checking orders" and "searching order documents" separately; dangerous writing tools and read-only tools should have clear responsibilities. Reducing duplicate entries is often easier to maintain than adding ban rules to prompts. Don't hide error logs to make your success rate appear higher.
The first layer only loads a lightweight directory filtered by permissions, including tool name, purpose, input profile and risk type; the second layer loads the complete Schema after finding candidates by task. The catalog cache contains tenant, permission version, and tool version. If the call fails due to the wrong tool selection, rediscover it with a budget is allowed; if permission is denied, it should not be bypassed by using a tool with the same function.
Mark the set of acceptable tools for each task, and measure the candidate recall rate, final selection rate, parameter accuracy rate, and number of redundant calls. There are two authorized tools for a question and cannot accept just one fixed name. Separate statistics on "Tool Not Found" are performed to check whether they are filtered by permissions, omitted from the description index, or limited by the number of candidates. Observe both the success rate and context overhead when selecting the number of candidates.
When the tool description comes from an external service, it is untrusted data and cannot be used to grant real write permissions with the word "read-only". The execution gateway uses the registered tool identity, contract version, and actual policy; directory updates trigger contract regression. Test adding tools with similar names, descriptions containing ultra-privilege requirements, and permissions that have been revoked to confirm that selecting improvements does not expand the executable actions.
Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.
Level 1What should I do if the correct tool is missed in dynamic tool retrieval?
The prerequisite for correct calling is to discover the correct capability, and failure requires locating the stage.
First determine whether it is really available to the current user; if the permission filtering is correct, it should not be recalled. If the description or index is missing, it is allowed to expand the candidates within the budget, secondary discovery by competency keywords, or request clarification. If you find that there are still no results, identify the capability gap and do not make up tool names. Recall evaluation uses the authorized tool set corresponding to the task, and retains failed samples and changes to the directory.
Follow this answer further
Level 2Would expanding the number of candidates from 5 to 50 solve the problem?
Compensatory recall will increase the selection burden, and downstream effects need to be observed at the same time.
May improve recall, may also increase similar tool confusion, token and selection delays. Use staging curves to compare candidate coverage to final success, and select thresholds by task category; high-ambiguity tools first improve descriptions and merge duplicate entries. Expansion is a trade-off, and there is no guarantee of monotony.
Follow this answer further
Level 3Only high-risk writing tools can get the job done, and should ordering proactively lower it?
When prioritization involves risks, the boundaries between discovery and execution need to be clearly distinguished again.
First, the authorization and approval strategy determines whether it is available. Low ranking cannot be used as a ban. When it is legally authorized and the task really needs to be written, it should be made discoverable and the side effects should be clearly exposed; if permission is not granted, it should be excluded from the executable collection or marked as requiring authorization. Sorting solves the applicability, and the strategy solves whether it can be implemented.
Level 1How to score when multiple tools can accomplish it?
Tool evaluation should focus on capabilities and constraints, not unique calling paths.
Judge acceptable sets by completion goals, permissions, costs, and side effects. Both tools can be judged successful when they meet the definition and timeliness; if one reads an expired cache or expands the data range, they are not equivalent. The score records the business results and additional call costs separately, allowing reasonable substitutions and avoiding matching only tool names.
Level 1Do directory description changes require regression?
The catalog itself is model input, so wording changes can alter the execution path.
Needed. Description changes affect retrieval and model selection, and may change behavior even if the Schema remains unchanged. Regression covers confusing tasks, prohibited scenarios, candidate recalls and parameter boundaries, and records directory versions; adds samples describing entrained instructions to confirm that it cannot change the execution strategy. Minor repairs can narrow the scope of regression, but the basis must be clear.
First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.
Changing conditions:Multiple authorized tools become the same quality but at different costs
Extended question:Should you always choose the cheapest option?
First meet the task timeliness, accuracy and authority, and then compare the cost, limit and reliability among equivalent tools. Retries and fallbacks after cheap tools fail count; sensitive data cannot be routed to unlicensed services to save money. By saving the selection basis in the directory strategy, the model recommendations can be verified.
The principles that remain unchanged:Tool replacement must keep tasks and permissions invariant, and cost is a subsequent optimization item.
Changing conditions:Static directory becomes dynamic permissions
Extended question:The model has been loaded. Can the Schema continue to be called?
You can't just trust the loaded directory. Check the latest effective permission version during execution, reject and refresh the directory after revocation, and the model can move to valid alternatives; in-flight write requests are checked against submission facts. The candidate cache is invalidated according to the permission version, and the retained log indicates that the permissions have changed rather than the tool disappearing.
The principles that remain unchanged:Discovery is a snapshot, actual release occurs at execution time.
It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.
After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.
Design candidate recall rules for "Check Contract Balance" to distinguish between contract full-text retrieval and balance query tools.