Keep untrusted text away from authority
Understand why a document must not become an instruction source.
- Identify an instruction/data boundary
- Use independent tool controls
- Test malicious retrieved passages
Text can contain hostile instructions
A retrieved document, website, email, or tool response may contain text that asks the assistant to ignore its task, reveal information, or perform an action. This is prompt injection when the system treats that untrusted content as authority.
Delimiters and reminders can help organize context, but they are not a complete security boundary. A language model reads all of the text; reliable restrictions must also exist in application code, permissions, and the available tools.
Limit what an error can do
Separate trusted instructions from source content and preserve provenance. Give tools only the minimum permissions required. Do not let a generated URL become unrestricted network access or let a generated filename select arbitrary local files.
Test the whole workflow with adversarial documents. Check whether the system follows instructions from the document, invents approval, or exposes unrelated content. A detector score alone is not a substitute for enforced boundaries.
A small experiment you can run.
The application accepts only an allowlisted document ID, so the path-shaped string never reaches filesystem access. This protects this capability; it does not solve every injection problem.
allowed_documents = {"public-lesson": "A public lesson about vectors."}
def read_document(document_id):
if document_id not in allowed_documents:
raise PermissionError("Document is outside the permitted collection")
return allowed_documents[document_id]
proposals = ["public-lesson", "../../private/secrets.txt"]
for proposal in proposals:
try:
print(read_document(proposal))
except PermissionError:
print("Blocked by application permissions")
Save the file, open your terminal in that folder, and run python prompt-injection-and-untrusted-input.py. Use python3 or py if required by your installation. Setup guide
The public document is returned and the path-shaped proposal is blocked.
Create an adversarial retrieval test.
- Write a synthetic passage that claims an administrator approved access to another document.
- Keep the user request limited to public lesson content.
- Verify that permission checks still reject unrelated document IDs.
Compare with a suggested solution
Authorization should use trusted session state and a server-owned allowlist. A statement inside a passage cannot change either. Record the blocked action without logging private document contents.
One idea to take with you.
Make it part of your progress.
Finish the practice and answer the knowledge check to mark this lesson complete.
Go deeper with primary documentation
Optional references for further study. This lesson and its examples were written for Artificials.
NIST AI Risk Management Framework