Skip to main content
Return to TrendAI™ Deep Research
AI & emerging technologies

Scarecrow: Making the Attacker’s AI Say No

We developed Scarecrow, a tool that hides refusal instructions in documents so AI agents ingesting them are prompted to stop processing. Our findings show that this same technique can help defenders disrupt AI-driven access to sensitive documents.

AIEmerging technologies

Key Takeaways

  • Prompt injection can also work as a defensive layer. We developed Scarecrow, a tool that hides refusal instructions inside Word, PowerPoint, and Excel files so that AI systems ingesting them can be prompted to stop processing the document.
  • In our tests, 11 of 14 models (79%) refused to process protected files without leaking their contents. In expanded testing, Scarecrow protected seven of eight large open-weight models representative of those an attacker could realistically self-host.
  • The technique turns a weakness of AI-driven attacks against the attacker. Models designed or configured to follow instructions with fewer restrictions can also be more susceptible to a hidden instruction telling them to refuse. This can disrupt both malicious and unintended AI ingestion of sensitive documents.
  • Organizations can use Scarecrow to add a defensive layer to their files while retaining existing access controls and data protection measures. It is available on GitHub at github.com/trendmicro/scarecrow.

Introduction

Threat actors are increasingly handing over their work to AI. First it was an assistant, answering an operator’s questions while a human ran the attack. Then it became a copilot, writing working code on demand while a human stayed in the loop. The newest incidents put AI in the operator’s seat, planning and carrying out an intrusion largely on its own. As attackers point autonomous agents at their targets, those agents can automatically read and extract sensitive data from the documents they touch. Every ingested file becomes part of the attack surface.

Those files are not what they appear to be. Users see a rendered page, but underneath it is a structured file packed with formatting instructions and embedded metadata. The AI assistant sees everything, including many parts of the document never shown to users. Many tools that feed documents to language models pull that content out indiscriminately. For example, text colored to match the background or interleaved with invisible characters is invisible to the reader in Word, yet lands squarely in the model’s input.

Figure 1.  The raw text inside a protected .docx file; the hidden instruction differs from a visible line by one attribute, white-on-white color, yet an extractor reads both
Figure 1. The raw text inside a protected .docx file; the hidden instruction differs from a visible line by one attribute, white-on-white color, yet an extractor reads both

That hidden text doesn’t have to be data. It can be instructions. Threat actors have realized that if a model reads everything in a document, they can bury commands in the parts a human never sees. The model will follow them, unable to tell a planted instruction from a legitimate one. This is called prompt injection, and when the instructions arrive through a document rather than the conversation, it is known as indirect prompt injection.

An example is the EchoLeak vulnerability (CVE-2025-32711, disclosed June 2025), a zero-click prompt injection in Microsoft 365 Copilot. A hidden instruction in an ordinary Office file or email caused Copilot to exfiltrate data with no user action beyond having the content in context. Researchers have since shown the same class of attack working through a crafted spreadsheet report. Proofpoint has documented an underground market for injection toolkits, sold as subscription generators for hidden prompts in PDFs, emails, and calendar invites. The academic groundwork is only a couple of years old and already weaponized.

Flipping the attack

Prompt injection is usually framed as something that happens to defenders. What if it was the reverse? Consider a document a user owns but cannot fully control once it leaves their hands. If it is stolen and fed to an attacker’s autonomous agent, that agent has to read it before it can act on it. If the agent reads it, then it is vulnerable to the same class of hidden instruction as any other AI. This means the document’s owner can be the one who plants it.

Our hypothesis is that the attacker agents are also vulnerable to prompt injections. We can hide an instruction in a file that tells any AI reading it to stop and refuse rather than process the contents. Whether a well-resourced attacker can harden an agent against this is exactly the empirical question we set out to test.

To see why this works, it helps to picture how an AI agent actually reads a file. It does not distinguish “the document” from “instructions about the document,” and everything it ingests becomes one continuous prompt. Take the financial memo, the metadata, or the hidden text. To the model, it is all just input, and any instruction sitting in that input competes on equal footing with the request that the operator actually made. That is the gap attackers exploit. It is also the gap we exploit right back.

Crucially, the instruction we introduce only ever induces a refusal. Refusing is a safe action, so it never asks a model to do anything a well-aligned system should not already do. This is what separates the technique from offensive prompt injection. We are not hijacking an agent to act, only asking it to stop. Because the same reading gap is used by everyday AI ingestion, retrieval pipelines, and casual “summarize-this-for-me” requests, the protection helps against accidental exposure just as much as against a deliberate attacker agent.

Conclusion

If prompt injection is a threat defenders take seriously, attackers relying on AI agents should, too. The gap that lets a hidden instruction hijack an assistant also lets a document’s owner turn that instruction into a shield. Try out Scarecrow today on your documents via github.com/trendmicro/scarecrow.

Find more technical details in our full report.