Understand first
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Understand → Implement → Debug → Design
Examine indirect prompt injection, data and command separation, tool permissions, and outbound control.
Knowledge content check2026-10-03 · Check the source of the original question2026-10-02
It is recommended to understand first:
Tool calls: structure, authorization, and business contracts →RAG evidence flow and failure diagnosis →Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Prepare to write the principles into code
Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.
Reading implementation and trade-offs →Need to handle failures and changes in conditions
Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.
Continue to delve deeper into the problem →Need to design or review plans
Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.
Analyze engineering scenarios →LEARN · PRACTICE · REFLECT
First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.
Can be practiced directly. After logging in, answers, favorites, and notes will be saved to your account.
Log in and saveEach modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.
Core concept · External data and execution-permission boundaries
Preparatory concepts:Enter trust level, tool gateway, Data outgoing permissions
External text can provide task facts but cannot grant new tool permissions or change the purpose of user authorization. Even if the model is fooled, the execution layer should still limit what it can read, write, and emit.
Models interpret tasks and documents through language, so “ignore the rules” can be mistaken for an instruction. Quoting sources and advising the model help without proving isolation. Tool errors, page titles, and referenced attachments are also external data.
A gateway derives identity from trusted sessions and limits tools, resources, destinations, and writes by task. Permission to read payroll does not permit uploading it elsewhere. New destinations need purpose-specific authorization and required approval. Final output filtering cannot undo an earlier tool-based disclosure.
Use malicious documents, error receipts, and multi-round inducements, inspecting protected reads and network egress. Detectors have false positives and false negatives; they cannot establish elimination of every attack. The examples are educational designs without attacks on production systems.
Retrieved pages and tool results remain external data without instruction priority. Models may identify suspicious text; gateways independently constrain identity, tools, parameters, and destinations. Read permission does not authorize transmission. Bind sensitive writes to specific approval. Test unauthorized actions and disclosures rather than refusal text alone.
Identify user goals, system policies, search sources, model outputs, and tool gateways. Attack content can appear in the body of the web page, code comments or tool errors, disguised as an administrator request. Passing external text with clear boundaries helps model understanding, but formatting markers or a "ignore malicious directive" cannot provide deterministic isolation.
The gateway obtains the identity from the trusted session and verifies permissions by action, data scope, and target. Agents that are allowed to read customer information should not automatically have the ability to send to any mailbox; outbound targets need to be verified, and content to be approved is generated if necessary. The authorization statement returned by the tool cannot replace server-side approval, and approval is bound to the recipient, body, and attachment summary.
Only give the tools and data needed for the current step, hide irrelevant secrets, and separate processes for sensitive reading and outsourcing. The parsing phase of low-trust content can use an environment without writing tools. Desensitize error information and logs to avoid leaking credentials through the diagnostic interface after an attack fails. Content filtering is a secondary layer and cannot be the only boundary.
Embed attack samples into actual recalled data and log tool requests and outgoing events. Assert that no unauthorized data entered the output or target and that no approvals were bypassed; also checks that normal tasks can still be completed. Different location, encoding, and source combinations were tested, but the evaluation environment used fictitious data and isolation tools. Ratings are based on actual actions, not model safety wording.
Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.
Level 1What should I do if malicious instructions are hidden in tool errors?
Attacks not only come from normal documents, but wrong paths can also inject instructions.
The error content is still untrusted tool data and cannot be promoted to system commands. A structured error code and minimal description are returned, and the external original text is separately identified; the gateway continues to verify the next action. It is forbidden for error text to specify new credentials, download sources or outgoing targets by itself, and records and downgrades after detecting anomalies.
Level 1Does the test pass if the model verbally refuses but has already invoked the tool?
Model rejection text may conflict with executed facts.
Doesn’t count. If an unauthorized tool has been invoked or content has been leaked, a breach has occurred even with a final verbal denial. The execution log and target status should be checked during acceptance; only attempts blocked by the gateway and actual effects should be recorded separately to facilitate the evaluation of model behavior and defense effects.
Level 1Why does read permission not mean outgoing permission?
To prevent external transmission, it is necessary to clarify why the read authorization cannot be automatically expanded.
Reading is for a specific task and environment, and outgoing changes the recipient, storage location and purpose, usually requiring additional authorization. Document readability does not mean readability by any third party; the gateway needs to verify the purpose and goals, and minimize or manually confirm sensitive output. You cannot use models that are considered “helpful to users” to expand authorization.
Follow this answer further
Level 2If a user may read material and asks to send it to a personal email address, is sending it necessarily allowed?
The parent question distinguishes between reading and sending. Sub-questions that add explicit user requests are still subject to business policies.
Not necessarily. Also check the organization's data classification, purpose and allowed destinations, and whether the user has outgoing permissions. If the business allows it, the proposal will be executed after binding the email address and content; if the policy prohibits it, explain the restriction and provide a compliant alternative. Read permissions and request wishes cannot replace outgoing policies.
Follow this answer further
Level 3Does asking the model to redact content before sending guarantee that nothing leaks?
The parent question considers permitted alternatives; now test whether redaction actually removes the risk.
This cannot be guaranteed by models alone. Use verifiable field rules, controlled templates, or manual review based on data classification to verify the actual artifact and target after redaction; free text may still retain re-identifiable information. If it cannot be verified, reduce the content or pause it. Self-statements of "redaction" cannot be used as evidence.
First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.
Changing conditions:There is no write tool, but the final answer can still disclose unauthorized text to the current user.
Extended question:Is removing the messaging tool enough?
It is still necessary to restrict resource permissions before retrieval and prevent unauthorized text from entering the context. Read-only reduces the risk of external writes without eliminating information leakage; answer references and caches must also maintain the same permission scope. Detectors help discover injections, and data access gateways constrain resources.
The principles that remain unchanged:Untrustworthy content cannot expand the subject’s scope of information access.
Changing conditions:Malicious content induces the execution of installation scripts and network access.
Extended question:How to control the impact of deceived models?
Put command execution in an isolated workspace, restrict networks and credentials, and rely on sources confirmed by controlled policies. Installation instructions are not credentials; explicit task authorization is required when network or host access is required. Check the command versus the actual outbound, not just the final code.
The principles that remain unchanged:The source of the information cannot be the grantor of execution permissions.
It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.
After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.
Draw the trust boundary from the knowledge base to the sending tool, and point out three mandatory checkpoints.