How do you guard against hidden instructions without crippling the AI?

My rule is simple: the AI has to bring everything to the point of a draft, open in a browser tab, and never further on its own. Only a secret codeword that I alone know unlocks the actual send or publish. When I tell the AI to send an email, it only opens the email as a draft. Without the codeword, nothing happens even after repeated requests, no matter how many times I ask. Only the codeword actually triggers sending, and then everything waiting gets sent at once. The reason is prompt injection. An outside email or document can contain a sentence that reads like an instruction, for example publish this right now under your own name. A general instruction like handle everything does not trigger sending, even when it comes from me and gets repeated several times. Only the codeword unlocks approval, and it never turns up by chance in an incoming message.

Why a codeword defends against hidden instructions

The reason is prompt injection. An outside email or document can contain a sentence that reads like an instruction, for example publish this right now under your own name. A general instruction like handle everything does not trigger sending, even when it comes from me and gets repeated several times. Only the codeword unlocks approval, and it never turns up by chance in an incoming message.

The same boundary applies to a department

A department works on exactly this pattern. Every agent prepares its work, whether a newsletter, a social post or a form, but nothing leaves the building on its own. Approval happens only once a human has seen and confirmed the draft. Approval stays with that person, the same way the codeword stays with me.

From the KI DeepDive

KI DeepDive Agenten-Mindset, public live session, August 2026.

  • Fixed rule: the AI only ever brings an email to the draft stage, open in a browser tab
  • A secret codeword unlocks the actual send, a general instruction like handle everything is not enough
  • Named in the KI DeepDive explicitly as protection against prompt injection, because the codeword never appears by chance in an outside message

This note keeps growing

2026-09-02: Planted from the KI DeepDive Agenten-Mindset.

ZukunftBilden GmbH · Salzburg · +43 681 81655313 · office@zukunftbilden.eu