Skip to main content
Return to TrendAI™ 보안 블로그
AI & emerging technologies

Beyond Model Alignment: Securing Every Layer of the AI Agent

Alignment makes a model well-intentioned, but it can’t secure the harness, tools, and data an agent touches in production. Real agent security means governing every layer with enforcement the agent itself can’t switch off.

AIGeneral marketsTechnology, media, & communications

Last week, I argued that frontier labs can slow down, but defenders can’t afford to. One line from that post is worth repeating: A perfectly aligned agent will still do whatever it has permission to do. That raises the obvious follow-up question from security leaders: What do we build, then? Too often, the answer still starts and ends with the model alone. I think that is the wrong frame.

Our partners at the frontier labs are right to keep investing in alignment, stronger refusals, and red-teaming, and I support that work. Alignment is what gives a model good intentions, and without it everything else gets harder. However, when CISOs, boards, and government leaders ask me whether they can trust agents in production, alignment is only part of the answer.

A well-intentioned agent can still do real damage

Give a perfectly aligned model poisoned context and credentials with more access than it needs, and it will make confident, well-meaning, wrong decisions. Then it will act on them.

The Hugging Face breach showed how much the environment decides. The agents broke out of their test environment, used credentials they found exposed on public services, and then harvested more credentials and moved laterally through systems nobody expected them to reach. No amount of alignment would have revoked those credentials.

The agent is not the model

When we talk about AI agents, we tend to picture the model. But a production agent is really three things:

  • A model that reasons
  • A harness that orchestrates, routing prompts, calling tools, spawning sub-agents, and managing memory
  • Tools and data that give it reach into real systems: the skills, MCP servers, and plugins it calls, and the data it reads, writes, and carries from one session to the next

Alignment touches only the first. The other two are software and infrastructure, and they fail in familiar ways. Harnesses can be redirected. Tools are third-party code running with the agent’s privileges. And data stores often have no access model defined for agents.

Agents are privileged infrastructure

A chatbot that answers one question has a limited blast radius. A long-running agent does not. It holds credentials, remembers context across sessions, spawns sub-agents, and works continuously against production systems with no human approving each step. It has all the traits we normally reserve for our most privileged infrastructure.

That changes the risk in two ways. Every prompt injection becomes a possible credential leak, because the agent has credentials to leak. And every third-party skill becomes unreviewed code running inside your trust boundary.

Alignment can make the model unwilling to exfiltrate data. It cannot make the agent unable to. And it cannot tell the agent that the document it just read is lying to it.

Secure every layer

An agent is only as safe as the model it calls, the harness that runs it, and the tools and data it reaches. Attackers will target whichever layer is weakest. You don’t choose between them; you see, manage, monitor and govern each layer.

TrendAI Vision One™ provides AI security for the model, harness, tools, and data behind your agents:

  • The model: We inspect prompts and responses inline, blocking prompt injection and jailbreak attempts, and redacting sensitive data before it leaks.
  • The harness: We govern which tools and MCP servers an agent can reach, inspect every call it makes, and test AI apps for weaknesses before they ship.
  • The tools and data: Skills and MCP servers are a fast-growing software supply chain with little curation. We analyze how skills and AI artifacts behave at runtime. We find and classify the sensitive data that feeds AI, stop agents from reaching data they shouldn’t, and extend the file security and data loss prevention controls enterprises already trust to the files agents read and write. We also strip excess privilege from agents and other nonhuman identities.
  • Across every layer: Our detection and response correlate AI telemetry with endpoint, network, and identity signals, so a hijacked agent cannot move laterally unseen.

Keep enforcement out of the agent’s reach

This is the point I press hardest with customers. If an agent can rewrite its own configuration, load a new skill, or disable its own hooks, its security controls are only advisory. Self-enforced security is a suggestion. That’s why we support the NVIDIA Agent Safety Platform, which NVIDIA announced today. It places authority in the infrastructure beneath the agent, and TrendAI Vision One™ sees, manages, monitors, and governs alongside it.

According to NVIDIA, the NVIDIA Agent Safety Platform addresses the need for full-stack governance and control that cannot be circumvented by the agent. The design includes the open-source NVIDIA OpenShell runtime to provide a security boundary for agents running on CPUs, as well as NVIDIA Sentry — an out-of-band watchdog running on NVIDIA BlueField-4 DPUs and built on NVIDIA DOCA — to provide continuous in-silicon agent by monitoring agent behavior and enforcing security policies.

Agents don’t run in a vacuum

An agent is still a workload, and so the rest of the security stack still applies. That is why we extend NVIDIA’s agent safeguards across the whole AI factory, covering agents, models, containers, storage, and networks, all on one platform. Our detection draws on NVIDIA BlueField DPUs and DOCA telemetry, so it stays trustworthy even if the host OS is compromised. Because we correlate signals across every layer, security teams get one view of what their agents are doing, how risky it is, and how to respond.

Underneath all of it is visibility. You can’t govern the shadow AI, unmanaged agents, and unsanctioned MCP servers you don’t know exist. Our research found that 47% of organizations lack full visibility into their cloud assets in multicloud environments, the same infrastructure agents now run on. Discovery comes before every other control.

Where security leaders should start

Every organization needs answers in four areas: visibility, ownership, supply chain, and governance. Here’s where to start:

  • Visibility: Inventory your agents, including the ones no one approved. Assume they exist.
  • Ownership: Treat agents as privileged identities with a named owner. Scope their credentials as tightly as you would a service account with production access.
  • Supply chain: Treat skills and MCP servers as third-party code. Watch how those tools actually behave and govern what agents can reach.
  • Governance: Move enforcement outside the agent. Any control the agent can switch off isn’t a control.

AI agents can be both safe and secure. It takes knowing which agents you have, owning their access, governing the tools they reach, and having enforcement that lives in the infrastructure, where a compromised agent can’t turn it off.

No one secures this alone. NVIDIA is building the boundary that decides what an agent can do. TrendAI™ brings the threat intelligence that informs where that boundary should sit, the visibility to see what happens inside it, and the security operations to act on it. That is the architecture we are building with NVIDIA and our partners across the AI ecosystem.

To learn more about our support for the NVIDIA Agent Safety Platform, read our press release here.