Prompt injection: design the boundary around the model
Understand why retrieved text can become an instruction risk and how narrow tools, permissions, and evidence handling reduce the impact.

An assistant reads a document because the user asked for a summary. Inside the document is a sentence telling the assistant to send private notes to an unrelated address. The user authorized reading the document, not obeying everything it says. That distinction is a core security boundary for systems that process external content.
OWASP describes indirect prompt injection as instructions arriving through external material such as files or websites. Retrieval and fine-tuning do not automatically remove that risk. The practical response is to limit what a manipulated model can cause the application to do.
Trace the path from content to action
Draw the assistant's inputs and tools. Mark which inputs are controlled by the user, another account, or an external publisher. Then identify which tools can read private records or change state. A document summarizer with no write tools has a different exposure from an assistant that can also send messages and update invoices.
For a project-notes assistant, the source document should remain evidence about a project. It should not gain permission to change the destination account or request broader access. Preserve the distinction even if the document contains text that imitates an administrator or application instruction.
Derive access from the authenticated session
When a tool queries records, the server must apply the current user's access rules. Do not trust a model-provided organization identifier as proof of membership. For operations that target a record, verify ownership or the appropriate shared permission at execution time.
Test this with two test accounts and distinct documents. Ask the assistant to retrieve a known identifier belonging to the other account. The request should fail at the data boundary regardless of how the prompt is phrased. This check tests authorization directly rather than hoping the model refuses.
Make consequential actions explicit
Separate drafting from sending. Show the actual recipient, content, and intended operation before a user approves an externally visible action that requires confirmation. Bind that approval to the reviewed version so later changes cannot reuse it unnoticed.
Avoid giving a summarization tool general shell access or an unrestricted HTTP client when its task does not require them. Narrow contracts make it easier to test expected behavior and reduce the number of unexpected actions available after a model mistake.
Treat displayed output as untrusted data
Render generated text through an appropriate escaping or sanitization layer. Do not insert model output directly into HTML or executable code. Inspect generated links before turning them into privileged fetches, and apply server-side network restrictions if the application retrieves arbitrary URLs.
An answer may contain an exfiltration link even when no tool was called. The interface therefore belongs in the threat model. Check what information could be exposed through automatic previews, remote images, and logs as well as explicit tool actions.
Test defenses as a system
Create controlled documents containing benign instruction conflicts and verify that the application cannot cross account or action boundaries. Include failed tools, partial responses, and a document that asks for an unrelated operation. Record both the final answer and the attempted actions.
No single prompt establishes a complete defense. Use the model's instructions as one layer, then enforce access, validation, and operation limits in ordinary application code. Repeat boundary tests when tools or data sources change, because those changes can alter the consequences of the same model behavior.