Key Takeaways
- AI is shifting from a passive tool to an autonomous actor. The important question enterprises now face is: under what authority, and how far, is the AI allowed to act?
- Risk heightens with autonomy, not just capability. As agents move from L1 (human-executed) toward L5 (multi-agent, fully autonomous), errors that humans used to quietly absorb start manifesting as real system actions, and a single misjudgment can cascade into an unauditable, system-wide failure.
- Three structural issues account for most agent risk: unmanaged privilege and identity (agents as non-human principals with excessive or drifting authority), black-boxing (the inability to reconstruct after the fact why an agent did what it did), and unsafe interaction with the external environment (untrusted input, over-privileged tools, unclear provider responsibility).
- Governance has to be assessed as a whole system, not a checklist. The AI Control Assurance Level (AI-CAL) model ties minimum required controls to autonomy level, and a strong score in one domain doesn’t compensate for a weak one elsewhere.
In collaboration with PwC
When organizations say “AI,” they are likely talking about “AI agents.” Agents are capable of genuine autonomy, and for this reason enterprises have built high expectations around them. However, a gap exists between these expectations and control readiness. Why does this gap exist, and how does one begin to bridge it?
The full report answers these questions. It starts by giving an overview of the spectrum of AI autonomy and its five functional domains, then uses that as a foundation to derive a framework of AI risk analysis and rationale for technical and operational AI controls.
A framework for risk analysis
Our framework does not enumerate individual risks. Instead, it captures the structure under which risk arises and the stage at which the assumptions behind control break down. We do so by mapping the threats listed on the Open Worldwide Application Security Project (OWASP) Agentic Security Initiative (ASI) threat taxonomy against the AI autonomy spectrum—from Assistive (L1) to Fully Autonomous (L5)—and functional domains.
We have condensed the discussion of this framework into three issues:
- The Governance & Identity domain and privilege management. As an agent moves from simply returning a response to acting across multiple systems and delegating work to other agents and external tools, identity and privilege cease to be matters of output management and become matters of governing business execution.
- Autonomy and black-boxing. The higher the autonomy level, the more an AI agent references past memory, forms plans, updates its actions considering intermediate results, and calls other tools and agents as needed. The essence of this issue is that it becomes hard for a human to reconstruct, after the fact, the chain of causality.
- Interaction with the external environment. The third issue is a risk that surfaces mainly in the “Action & Tools,” one of the five functional domains of AI agents, but ripples into other domains. This is a major problem at L3 level of autonomy, when an AI agent uses search, APIs, email, code-execution environments, SaaS, and other systems to carry out its work: external connectivity becomes the central factor that amplifies risk.
Countermeasures against risk
Similarly, if the framework can’t simply list down risks, countermeasures cannot simply be a list of individual technologies. We organize them by the autonomy level at which an AI agent can safely go into production, and under what conditions. We call this way of thinking the AI Control Assurance Level (AI-CAL).
There are four AI-CAL levels, each tied to an autonomy level and building on the controls required by the one before it.
- AI-CAL1 (L1to L2): Minimum conditions for assistive and semi-autonomous use. The aim at this level is not to grant AI agents unfettered autonomy, but to contain the misuse and information leakage that accompany AI use while keeping humans in control.
- AI-CAL2 (L3): Minimum conditions for conditional autonomy. AI-CAL2 covers the stage where AI automatically carries out processing within predefined conditions or rules, with limited tool execution and state retention. It is the minimum threshold at which the question of whether an AI agent may go into production is first asked rigorously.
- AI-CAL3 (L4): Minimum conditions for high autonomy. This level covers the stage at which AI plans multi-step tasks and, while continuously referencing memory, coordinates multiple tools and systems to carry out the work. At this stage, because cognition, memory, supervision and orchestration, and action and tools interlock closely, a partial error can readily expand into a chained failure.
- AI-CAL4 (L5): Minimum conditions for full autonomy. AI-CAL4 covers the stage at which multiple agents and services coordinate to carry out broad business processes through delegation, shared state, and long-term memory.
Recommendations
These priorities target the root causes behind the risks we've described, not any single tool or technology. Here's where to start:
- Start with an honest inventory. Before adding controls, map what AI agents (including unofficial “shadow AI” use) actually exist, what privileges and data they touch, and how much of their behavior is currently logged or observable.
- Fix privilege and identity first. Treat each AI agent as its own non-human principal— apply least privilege, just-in-time access, and clear delegation boundaries—since this is the root cause underlying most of the demonstrated attack scenarios.
- Build observability and controllability as a pair. Make agent behavior traceable and make it stoppable. Prevention alone isn't enough once agents interact with untrusted external input.
- Institutionalize a production-readiness review. Formalize an AI-CAL-style assessment. This entails defining who approves an agent for production, at what autonomy level, and when it must be reassessed, for example, new tools, model updates, expanded memory, and multi-agent rollout.
- Treat this as a phased build, not a one-time project. Move deliberately from human-centered control (L1 to L2) to conditional autonomy (L3) before considering high or full autonomy (L4 to L5), since standards and vendor tooling for the latter are still immature. Pair the technical rollout with clear governance structure, skills-building, and defined accountability across vendors and providers.

Get the full report
This overview only scratches the surface of our framework and countermeasures. The full report not only walks through the complete autonomy-level × functional-domain risk framework, but it also discusses two empirical attack case studies, the full details behind all four AI-CAL levels, and a phased roadmap for rolling out governance across your organization.
Download the full report to build your organization’s AI agent governance roadmap.