Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content
knowledge unit 12IntermediateSystem designAbout 12 minutes

Understand → Implement → Debug → Design

Separate discovery from execution in large tool catalogs

Examine tool catalog, dynamic discovery, recall quality, and server-side authorization.

tool routingtool discoveryMCP

Knowledge content check2026-10-03 · Check the source of the original question2026-10-02

Which step do you want to learn from this knowledge point?

Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.

Understand first

New to this knowledge point

Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.

Start with core principles →

Realize again

Prepare to write the principles into code

Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.

Reading implementation and trade-offs →

Will troubleshoot

Need to handle failures and changes in conditions

Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.

Continue to delve deeper into the problem →

Able to choose

Need to design or review plans

Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.

Analyze engineering scenarios →
Knowledge unit directory

LEARN · PRACTICE · REFLECT

Knowledge learning and personal records

My notes and review ↗

First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.

Answers and personal notes

Each modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.

Core concept · Separate discovery from execution in large tool catalogs

Understand the core principles first

Preparatory concepts:Retrieval recall, Interface naming, Permission filtering

Choosing the right tool requires first discovering the applicable capabilities and then calling it based on the complete contract. The discovery layer is responsible for candidate coverage, and the execution layer is responsible for permissions and parameters; error hiding tools can improve selection, but they cannot replace the server's rejection of unauthorized actions.

Locate why selection failed

“Look up order” and “search order documentation” can be confused. Clarify targets and input/output boundaries rather than only adding tools. If the correct tool never enters the candidate set, changing the call prompt cannot fix discovery.

Choose discovery tradeoffs explicitly

Provide an authorization-filtered lightweight catalog, then load a few full schemas for the task. Expand discovery within bounds after a miss. Use stable identities and versions, isolating caches by tenant and permissions. This is application design: MCP listing and change notifications support discovery without specifying ranking.

Evaluate capability rather than one name

A query API or controlled report may satisfy the same goal. Define allowed capabilities and constraints, measuring candidate recall, final selection, parameters, and business success separately. Unauthorized tools are not eligible ground truth. Execution failures can also expose API design problems, not merely model selection errors.

Check understanding with a question

After receiving 200 tools, I always choose the wrong model. How would you narrow down your tool set?

Analyze actual selection failures and overlapping names, descriptions, and parameters. Retrieve a few tools from an authorization-filtered catalog, expanding discovery when needed. Evaluate discovery separately from invocation: absence from context does not establish absence from the catalog. Execution still validates versions, permissions, and parameters; descriptions cannot override system rules.

Realization and trade-offs

Check the tool itself first

Collect samples of mischosen tools and distinguish between synonymous naming, lack of applicable boundaries, parameter ambiguity, and true retrieval failures. Describe "checking orders" and "searching order documents" separately; dangerous writing tools and read-only tools should have clear responsibilities. Reducing duplicate entries is often easier to maintain than adding ban rules to prompts. Don't hide error logs to make your success rate appear higher.

Discover tools in two stages

The first layer only loads a lightweight directory filtered by permissions, including tool name, purpose, input profile and risk type; the second layer loads the complete Schema after finding candidates by task. The catalog cache contains tenant, permission version, and tool version. If the call fails due to the wrong tool selection, rediscover it with a budget is allowed; if permission is denied, it should not be bypassed by using a tool with the same function.

Evaluate discovery and execution separately

Mark the set of acceptable tools for each task, and measure the candidate recall rate, final selection rate, parameter accuracy rate, and number of redundant calls. There are two authorized tools for a question and cannot accept just one fixed name. Separate statistics on "Tool Not Found" are performed to check whether they are filtered by permissions, omitted from the description index, or limited by the number of candidates. Observe both the success rate and context overhead when selecting the number of candidates.

Prevent directories from becoming new risks

When the tool description comes from an external service, it is untrusted data and cannot be used to grant real write permissions with the word "read-only". The execution gateway uses the registered tool identity, contract version, and actual policy; directory updates trigger contract regression. Test adding tools with similar names, descriptions containing ultra-privilege requirements, and permissions that have been revoked to confirm that selecting improvements does not expand the executable actions.

Engineering deduction

scene
Interview hypothesis: Among the 200 enterprise tools, Agent often treats work order retrieval as customer information query.
design decisions
Organize the usage boundaries and then build a permission-filtered directory search.
Verify target
Candidate recall and selection rate can be measured separately, and the reasons why cannot be found can be explained.
applicable boundary
The tool number threshold should be determined through mission experiments and there is no uniform optimal value.

Continuous questions and answers

Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.

Draw inferences from one example: If the conditions change, how to deduce it?

First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.

Same ability, different fees

Changing conditions:Multiple authorized tools become the same quality but at different costs

Extended question:Should you always choose the cheapest option?

Derivation and reference solutions

First meet the task timeliness, accuracy and authority, and then compare the cost, limit and reliability among equivalent tools. Retries and fallbacks after cheap tools fail count; sensitive data cannot be routed to unlicensed services to save money. By saving the selection basis in the directory strategy, the model recommendations can be verified.

The principles that remain unchanged:Tool replacement must keep tasks and permissions invariant, and cost is a subsequent optimization item.

Revoke permissions during operation

Changing conditions:Static directory becomes dynamic permissions

Extended question:The model has been loaded. Can the Schema continue to be called?

Derivation and reference solutions

You can't just trust the loaded directory. Check the latest effective permission version during execution, reject and refresh the directory after revocation, and the model can move to valid alternatives; in-flight write requests are checked against submission facts. The candidate cache is invalidated according to the permission version, and the retained log indicates that the permissions have changed rather than the tool disappearing.

The principles that remain unchanged:Discovery is a snapshot, actual release occurs at execution time.

Easy to make mistakes

  • Throw in all your tools at once
  • Test-only model ultimately selects no-test candidate recall
  • Treat the security statement in the description as authorization

References

It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.

Check how far you understand

After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.

Basic standards met
Ability to improve tool naming, description, and parameter bounds using real mis-selection examples.
Intermediate and advanced signals
Two-stage discovery, set of acceptable tools, and failure classification are proposed.
Senior Signal
Can handle directory permission versions, contract drift and measurable candidate budgets.

Hands-on verificationComplete on demand · Suggestions15 minutes

Design candidate recall rules for "Check Contract Balance" to distinguish between contract full-text retrieval and balance query tools.

Expand acceptance requirements and checkpoints
  • Permissions filter first
  • Allow multiple authorized tools
  • Limited fallback when there are not enough candidates

Key inspections

  • Distinguish between tool recall and call success
  • Clarify two-phase discovery and failure fallback
  • Authorization is independently verified at execution time