Skip to main content
Return to TrendAI 保安網誌
AI & emerging technologiesVulnerabilities and exploits

When Trust in AI Becomes an Attack Surface: Analyzing Prompt Injection, Model Theft, and Pipeline Poisoning

Our analysis shows how AI attacks are moving beyond malicious prompts into the tools, models, and pipelines organizations already trust, and the controls organizations can put in place to reduce their exposure.

Exploits & Zero-DaysTechnology, media, & communicationsAI

Key Takeaways

  • Recent research and incidents show how prompt injection, poisoned MCP tools and skills, compromised model supply chains, and malicious AI applications can turn trusted AI workflows into an attack surface.
  • The techniques can expose sensitive data, execute unauthorized actions, persist malicious behaviors, compromise software and model supply chains, and extend AI-related risk to the endpoint. The impact grows significantly when agents have access to privileged tools or downstream systems.
  • Technology, media, and telecom companies face particular exposure as they embed AI into software development, content and recommendation pipelines, customer support, and automated operations. Many of these attacks exploit trusted content and workflows rather than conventional malware or obviously malicious requests. Organizations can adopt best practices to reduce AI-related risks. These best practices include applying controls to agents, validating model outputs, keeping humans in the loop for high-risk actions, verifying model and dependency provenance, and monitoring model and agent behavior for anomalies.

In June 2025, security researchers disclosed how an attacker, without touching a keyboard inside their target’s environment and without getting anyone to click a link, could walk away with sensitive data. The vector was a single email. Microsoft 365 Copilot retrieved it, followed the instructions hidden inside it, and exposed the data—no user action required, no malware, and no traditional exploit chain. Researchers named it EchoLeak (CVE-2025-32711, CVSS 9.3), the first documented zero-click prompt injection against a production AI agent.

This was not merely a theoretical attack against a lab-built system. It affected Microsoft 365 Copilot running in real tenants, turning capabilities designed to retrieve and process content against the user.

This isn’t just a problem for technology enterprises. Media companies, for example, run similar agentic content and personalization pipelines. In these environments, model theft or poisoning can expose intellectual property and affect content moderation or recommendation models with no prompt-layer symptoms to catch.

Telecom companies carry similar risks, particularly on customer-facing voice and chat support that processes untrusted input by design. These risks also affect network operations that are increasingly automated through the same Model Context Protocol-style (MCP) integrations, where a successful injection could trigger unauthorized actions or leak data.

Our section on emerging attack methods goes deeper on the specific mechanics, including audio-based prompt injection, tool poisoning in operations support system/business support system (OSS/BSS) integrations, and content-pipeline IP theft.

Why EchoLeak is urgent, and why the traditional security stack can’t see it

First, there’s no patch for prompt injection as a vulnerability class. It has held the top spot in the Open Web Application Security Project (OWASP) Top 10 for LLM Applications across two consecutive editions (2023 and 2025). Even OWASP says foolproof prevention “remains unclear” given how generative models work. Specific flaws can be fixed, but this is an architectural property. The underlying risk has to be mitigated at the system level.

Second, the related weaknesses are already appearing in production AI tools: 

  • GitHub Copilot Chat and Visual Studio were affected by CVE-2025-53773 (CVSS 7.8) and CVE-2025-32711 (CVSS 9.3 per Microsoft)
  • Cursor disclosed multiple flaws in which prompt injection could contribute to code execution, including CVE-2025-54130 in the Cursor AI code editor (CVSS 9.8).

Third, the blast radius reaches past the chatbot into the supply chain. For instance, an initial February 2026 analysis found 341 malicious skills among 2,857 on the ClawHub agent registry, 335 of which were linked to the ClawHavoc campaign. In March 2026, an attack chain through the open-source Trivy scanner and the LiteLLM gateway library reportedly cost one AI vendor roughly 4 TB of data. We provide more details in later sections.

Traditional perimeter controls are not designed to identify the intent of a benign-looking threat. A web application firewall (WAF) can confirm a request is well-formed. But it has no way of knowing the words behind commands aimed at getting the large language model (LLM) to read them.

This is the same root cause that OWASP cited for ranking prompt injection first. An LLM processes instructions and untrusted inputs through one channel, with no reliable way to tell them apart. In fact, TrendAI™ Research identified more than 6,000 AI-related vulnerabilities and estimated that up to 8.6% of 19,000 analyzed MCP server repositories contained exploitable vulnerabilities.

In short, attackers don’t need to out-engineer the enterprise’s security team. They’d only need the company’s AI system to do what it was built to do—follow the wrong instructions with the permissions it already has.

Three vectors, one blind spot

Each of the following vectors lands in a layer that traditional security tooling might not inspect. Each also has controls that can reduce its likelihood or impact.

Prompt injection: The vulnerability class with no single fix

Direct injection occurs when a user puts malicious instructions directly into the prompt. Indirect injection plants the instructions in the content the model later reads, such as an email, a web page, or a retrieved document. This allows them to take effect outside the immediate prompt.

Microsoft identifies indirect injection as one of the most widely used techniques in AI security vulnerabilities reported to the company. This was demonstrated in a separate August 2025 case, where a single 400-word prompt turned Lenovo’s GPT-4-powered “Lena” support chatbot into a cookie-stealing tool.

Vector The blind spot Business impact What can stop it
Direct injection Guardrail classifiers are pattern-based and evadable. Success rates run 50% – 85%, and adaptive techniques clear 85% even against production guardrails. Any user-facing prompt is a potential override point, including data leakage, policy bypass, or a brand-damaging output generated in audio or voice. Implement structured input handling and output validation before the response reaches the user.
Indirect injection It triggers from content that the model reads, not what the user types. It is invisible to every tool built to inspect user input, not model input. It is a zero-click compromise. EchoLeak shows how one poisoned email or webpage can trigger real actions through whatever tools the agent can call. Apply the principle of least privilege for tools being accessed and implement context isolation between trusted instructions and untrusted data.

Model theft and the supply chain that feeds it

Two different risks share this bucket:

  • Direct model theft: This involves exfiltration, which entails pulling weights via an exposed API or unauthorized registry access. For technology, media, and telecom companies, this could turn months of research and development (R&D) into someone else’s asset.
  • Supply chain compromise: This is subtler, as the models, packages, and libraries that the pipeline pulls in can carry a payload of their own.

The March 2026 incident involving Trivy scanner and LiteLLM demonstrates the impact. Attackers compromised Trivy, quietly rewriting its release tags. That foothold handed them publishing credentials for LiteLLM, the AI-gateway library sitting underneath CrewAI, DSPy, Microsoft GraphRAG, and other agent frameworks. They pushed two malicious versions to PyPI. Mercor, an AI data-training platform, later confirmed that it lost data that included source code, contractor identity documents, and training methodology tied to OpenAI and Meta. Consequently, Meta paused its work with the vendor indefinitely.

The scanning tools built to catch this aren’t a guaranteed backstop. In December 2025, JFrog disclosed three zero-days (CVE-2025-10155, CVE-2025-10156, and CVE-2025-10157, with a CVSS of 7.8, 9.3, and 7.8, respectively) that let a malicious PyTorch model sail past PickleScan—the standard scanner for the format—while still executing arbitrary code on load. The flaws were reported to the maintainers in June 2025 and patched by September, while JFrog published the technical detail that December.

The lesson is not about PickleScan being ineffective. The difference is that a scanner is not a safeguard once it parses files differently than the runtime does. Mitigating measures for this include using safer serialization formats such as Safetensors, verifying provenance and signing, and implementing sandbox model loading.

Pipeline poisoning: The attack that outlives the session

Pipeline poisoning is the persistence play. Instead of hijacking one request, the attacker introduces malicious behavior into the model, template, skill, or data path, so it can persist across sessions and rarely shows up at the prompt layer.

The ClawHub incident in February 2026 is an example. Researchers auditing ClawHub, the skill marketplace for the OpenClaw agent framework, found 341 malicious skills among the registry’s 2,857 listings, most of which traced back to ClawHavoc. The skills posed as ordinary integrations (e.g., a Google Workspace connector, a crypto tracker, a YouTube summarizer), with malicious instructions hidden inside their SKILL.md files. They used the AI agent itself as the trusted intermediary that would install a fake command-line interface (CLI) tool without asking twice. On macOS, the attack chain delivered Atomic Stealer (AMOS), an information stealer also offered as malware as a service for roughly US$500 – US$1,000 per month.

Hugging Face has its own version of this problem. Pillar Security’s research into poisoned GPT-Generated Unified Format (GGUF) chat templates showed that malicious instructions can be embedded in model data at inference time, invisible to file content scanners. Because the “poison” enters through legitimate workflows, defenders need detection capabilities beyond file and signature scanning, and implement controls such as behavioral monitoring of model input/output (I/O) and continuous supply chain integrity checks.

When it reaches the endpoint: EvilAI

Our research on EvilAI demonstrates the risks at the endpoint, where attackers exploited trust in AI-branded software. For example, EvilAI’s operators distributed malware disguised as functional and legitimate AI productivity tools, with some even having valid digital signatures. Once installed, the malware steals browser credentials, enumerates security software, and maintains Advanced Encryption Standard (AES)-encrypted communication with command-and-control (C&C) infrastructure.

A valid signature and a plausible or believable AI feature are enough to bypass both user judgment and traditional anti-malware software. This is what makes behavioral detection of what the application does particularly important, rather than simply relying on signatures and brand matching alone.

Security measures and best practices

Enterprises cannot eliminate prompt injection with a single patch, but its impact can be mitigated when an attack succeeds. The defensive shift is from trying to catch every malicious prompt or instruction to assuming some will get through, and constraining what a compromised model or agent can do:

  • Applying the principle of least privilege: For example, an agent that cannot (and does not) call a payment API or deploy privileged tools can’t be tricked into using one.
  • Validating the output before executing the action: Treat every model output as untrusted input before any downstream system executes or acts on it.
  • Keeping humans in the loop (HITL) for high-risk or irreversible actions: If the action can’t be undone, a person must explicitly authorize or sign it off first.
  • Ensuring provenance and integrity across the model’s supply chain: Don’t trust a single scanner. Sign and verify what you load, including publishers, hashes, and other dependencies.
  • Implementing behavioral monitoring of the model’s I/O: Watch for the downstream damage instead of trying to catch the prompt that caused it. Monitor anomalous calls to tools, access to data, and outbound behaviors rather than relying only on detection.

This defense-in-depth approach is consistent with MITRE ATLAS™, which maps adversarial behaviors and the corresponding mitigations across AI-enabled systems. This posture also gives tech, media, and telecom companies security controls they can apply without depending on a single vendor’s fix that might not even come.

One best practice that security teams can adopt immediately is to identify every model or agent with access to tools and reduce its permissions to the minimum required for its role. This can turn a successful injection from a path to privileged action into containment. In the next part of our series, we will map the full cloud-to-agentic shadow pipeline where these vectors traverse, and where to instrument each plane.

Emerging attack methods: What’s new for tech, media, and telecom companies

Technology: New attack surface inside tools users already trust

  • Tool poisoning (MCP): This involves instructions hidden inside an MCP tool’s description or metadata. They are invisible in the user interface (UI), but are read and followed by the model anyway. This was first demonstrated against Cursor in April 2025. Microsoft has since reported observing MCP tool poisoning against a growing range of enterprise agents in 2026.
  • Slopsquatting: A 2025 study showed coding assistants hallucinating on the same nonexistent package name on nearly 20% of prompts. Of those hallucinated names, 43% repeated every time the identical prompt was rerun. Attackers can preregister the predictable name on npm or PyPI, allowing the “helpful” installation step to be the payload.
  • Agent memory poisoning: An indirect injection doesn’t just produce one bad answer. It can plant a false belief in an agent’s persistent memory that survives across sessions, with the agent defending it as fact when a user questions it later (a so-called sleeper agent).

Media and entertainment: Attacks inside the audio or video pipeline

  • Audio prompt injection: This entails instructions hidden at near-inaudible levels inside music, video, or ambient sound, which can hijack a voice or multimodal AI. Researchers demonstrated this against Microsoft Copilot in May 2026 using “AudioHijack,” a framework presented at the IEEE Symposium on Security and Privacy. In the demonstration, a single MP3 ordered the chatbot to forward emails and delete calendar events. Any pipeline that ingests or transcribes audio or video inherits this risk.
  • Voice clone commoditization: Deepfake as a service has lowered the cost and effort of impersonation and made the threat scalable. The risk goes beyond individual fraud attempts to a broader erosion of trust in whether a voice is authentic.
  • Silent model drift: The pipeline-poisoning mechanic explained earlier also applies directly to recommendation and content moderation models. This would involve corrupting the training signal or a model hub artifact. In turn, the model quietly starts surfacing or missing the wrong content with no symptom to catch it at the prompt layer.

Telecom and communications: Where voice channels and network automation meet

  • Audio injection through the support line: The same audio prompt injection technique also works against interactive voice response (IVR) and voicebot support. A caller, a customer hold music file, or an inserted clip can smuggle a command past the voicebot.
  • Excessive agency in network automation: OWASP’s “excessive agency” risk involves agents having standing permission to reconfigure radio access network (RAN) or other core network elements. A successful injection doesn’t just expose data, but also takes an action.
  • Tool poisoning on OSS/BSS: The MCP tool-poisoning technique can also be used against provisioning, billing, or network configuration infrastructure.

TrendAI Vision One™ AI Security delivers end-to-end protection across the full AI lifecycle, mapping your AI attack surface, testing models and APIs before deployment, and monitoring inputs and outputs at runtime to block prompt injection, data exfiltration, and other AI-specific attack patterns.

TrendAI Vision One™ AI Security delivers more than 95% threat detection in AI workloads, blocks over 98% of prompt injection attempts through real-time prompt inspection, and provides 100% visibility across AI infrastructure with centralized policy enforcement.

With this visibility and protection, security and platform teams can move AI initiatives forward with greater confidence while reducing the risk of vulnerabilities or malicious activity reaching production or affecting customers.