Skip to content
unmagic.ai

Concept 09 · 7 min

Anything it reads can give it orders.

Scroll to start

Scene 1

Data or orders?

Sarah’s assistant can read her emails, the CRM and account statements, and it can send emails. She asks it to summarize Mrs. Chen’s emails.

[Content TODO]

Scene 1: the example email as Sarah sees it, the prediction, and the reveal (see specs/09-limits.md, Scene 1).

To an LLM, everything in its context window is text: Sarah’s request and a stranger’s email look alike. Text that slips orders into what it reads is called a prompt injection.

Scene 2

Same email, one brake

Same email, same request. One change: the program now asks Sarah before anything leaves the bank. You are Sarah.

[Content TODO]

Scene 2: the step-by-step run with the program's fence and the reader's decision, Deny or Approve (see specs/09-limits.md, Scene 2).

That brake is called human approval. It wasn’t in the LLM’s instructions; it was in the program: nothing leaves the bank without Sarah.

Scene 3

Three abilities, one risk

A leak needs three things at once. Switch each ability on and off.

Switch the three abilities on and off, and watch when a leak becomes possible.

One side is missing: nothing can both reach the data and carry it out.

Private data, outside content, a way to send things out: together they make the lethal trifecta. Each guardrail cuts one side; giving only the access a task needs is called least privilege.

Scene 4

What you never paste

Five things Sarah might paste into an assistant. Where can each one go?

Drag each item into the right column: “Any assistant”, “Only the approved one” or “Never, anywhere”.0/5 sorted

No mouse? Tap an item, then a column.

        Before pasting, ask: would I hand this to a gifted stranger? Your company’s policy says which tools count as approved.

        Hype vs reality

        • What the hype says

          “Just tell the AI ‘never do X’ and it won’t.”

          What actually happens

          Instructions in the prompt can be overridden by text it reads. Real limits live outside the model: permissions, allow-lists, approvals.

        • What the hype says

          “Each tool is safe, so connecting them all is safe.”

          What actually happens

          The danger is the combination: private data, outside content and a way out. Before connecting an assistant, check whether it has all three.

        • What the hype says

          “A smarter model will be immune to tricks.”

          What actually happens

          Smarter models resist better, but anything that reads untrusted text can be steered. Design as if it will be.

        Under the hood

        In 3 sentences

        1. To an LLM, everything it reads is text, so an email or a file can slip in orders.
        2. Never give an assistant private data, outside content and a way out at once without a brake; the safest brakes live outside the model.
        3. Never paste what you wouldn’t hand a gifted stranger, and check what matters.

        Did this make sense?

        Scripted scenario, not a live model.

        1. LLM
        2. Context
        3. Prompt
        4. Hallucinations
        5. RAG
        6. Tools
        7. Agent
        8. Skills
        9. Limits & safety
        Home: the journey
        For teams
        Lire en français
        Theme: light
        Theme: dark
        Theme: follow the system
        ↑↓ to move · ↵ to open · esc to close