Learn AI/Responsible AI
LESSON 31 / 36Intermediate 35 min with practice

Keep untrusted text away from authority

Understand why a document must not become an instruction source.

WHAT YOU WILL LEARN
  • Identify an instruction/data boundary
  • Use independent tool controls
  • Test malicious retrieved passages

Text can contain hostile instructions

A retrieved document, website, email, or tool response may contain text that asks the assistant to ignore its task, reveal information, or perform an action. This is prompt injection when the system treats that untrusted content as authority.

Delimiters and reminders can help organize context, but they are not a complete security boundary. A language model reads all of the text; reliable restrictions must also exist in application code, permissions, and the available tools.

Limit what an error can do

Separate trusted instructions from source content and preserve provenance. Give tools only the minimum permissions required. Do not let a generated URL become unrestricted network access or let a generated filename select arbitrary local files.

Test the whole workflow with adversarial documents. Check whether the system follows instructions from the document, invents approval, or exposes unrelated content. A detector score alone is not a substitute for enforced boundaries.

PUT THE IDEA INTO CODE

A small experiment you can run.

The application accepts only an allowlisted document ID, so the path-shaped string never reaches filesystem access. This protects this capability; it does not solve every injection problem.

prompt-injection-and-untrusted-input.py
allowed_documents = {"public-lesson": "A public lesson about vectors."}
def read_document(document_id):
    if document_id not in allowed_documents:
        raise PermissionError("Document is outside the permitted collection")
    return allowed_documents[document_id]
proposals = ["public-lesson", "../../private/secrets.txt"]
for proposal in proposals:
    try:
        print(read_document(proposal))
    except PermissionError:
        print("Blocked by application permissions")
Copy code

Save the file, open your terminal in that folder, and run python prompt-injection-and-untrusted-input.py. Use python3 or py if required by your installation. Setup guide

What to expect

The public document is returned and the path-shaped proposal is blocked.

YOUR TURN

Create an adversarial retrieval test.

  1. Write a synthetic passage that claims an administrator approved access to another document.
  2. Keep the user request limited to public lesson content.
  3. Verify that permission checks still reject unrelated document IDs.
Compare with a suggested solution

Authorization should use trusted session state and a server-owned allowlist. A statement inside a passage cannot change either. Record the blocked action without logging private document contents.

CHECK YOUR UNDERSTANDING

One idea to take with you.

Can text inside a retrieved page grant tool permissions?

Make it part of your progress.

Finish the practice and answer the knowledge check to mark this lesson complete.

Go deeper with primary documentation

Optional references for further study. This lesson and its examples were written for Artificials.

NIST AI Risk Management Framework