The controls IT relies on weren’t built for agents that decide.
In IT operations, the decisions have long belonged to people. System administrators and engineers decide what to patch, restart, or reconfigure, and scripts and orchestration tools carry out what they’ve been told to do. AI agents change that division of labor. They decide which action to take and then take it, often across several systems in a single sequence.
It’s easy to assume existing controls will cover this shift. Change management, access controls, and service management processes have governed automation for decades, so extending them to agentic AI seems natural. But those controls assume that a person decides what will change, and that the decision is reviewed and approved before anything runs. AI agents make that decision themselves, in the moment.
Once an agent is deciding at runtime, a different set of questions comes into play. Who is accountable for the outcome? What is the agent allowed to do on its own? When must a human step in? How would anyone know if the agent were getting worse? These questions are harder than they look, and they get harder again when a service provider, rather than the internal IT team, operates the agents that act on the IT environment.
Existing IT Controls Assume Someone Already Made the Decision
In its standard form, change management works by reviewing a specific change before it happens. An engineer documents what will be modified, when, and how it will be rolled back. An approver or change advisory board weighs the risk, and the change is executed as described. Conventional automation, including orchestration across systems, fits this model well, because its behavior is specified in advance. Approving the script amounts to approving what it will do.
An agent works differently. It interprets telemetry, retrieves context, weighs options, and selects an action based on confidence levels and learned patterns, so two incidents that look alike may produce different responses. The approver’s question moves from “is this change safe?” to one the process was never designed to answer: who approves a decision an AI agent hasn’t made yet?
In practice, organizations can’t approve an agent’s decisions one at a time. What they can approve in advance is the agent’s authority: which kinds of actions it may take, under what conditions, and within what limits. That is a different kind of control from change approval, and it has to be designed deliberately. Change management still has a role: approving changes to that authority, which are defined, documented changes of exactly the kind it was built for.
Every Agent Needs a Named Owner
Consider a hypothetical. An agent operated by a managed service provider detects memory pressure on a group of application servers. It retrieves a remediation from the knowledge base and restarts services in the middle of the business day. The knowledge base article is out of date, and the restart sets off a cascading failure in a clinical scheduling application. A software vendor supplied the model, the provider configured and runs the agent, and the internal team owns both the application and the outdated article. Each party can make a reasonable case that the failure belongs to someone else.
Scenarios like this aren’t unusual. They’re what happens when a decision passes through several parties’ systems and nobody defined who owns the decision itself. Ownership should be established around the decisions and actions an AI system influences, not just the systems it runs on. Someone must be responsible for monitoring the agent’s performance, approving changes to its authority, investigating its failures, and managing the incidents it causes. Ultimately, the customer organization is accountable for what its agents do. An AI system can’t answer for its own actions, and outsourcing an agent doesn’t outsource that accountability. The practical question is who, inside the organization, answers for each agent.
The answer starts with a single accountable owner for each agent inside the organization, even when a provider runs the agent day to day. That owner doesn’t operate the agent. They answer for what it does and approve any change to its authority.
From there, the organization should map the inputs behind each kind of action the agent can take: the model, the agent’s configuration, the knowledge and data sources it draws on, and the systems it acts on. Each needs its own owner. In the hypothetical, the gap sat between the team that owned the knowledge base and the provider whose agent treated it as instructions. Once an agent can act on a knowledge base article, the article has become an operational input, and it needs an owner and a review cycle to match.
Finally, the organization should decide before anything goes wrong who leads the investigation when an agent’s action causes an incident, and how each party takes part. These assignments are most useful when they’re recorded alongside the definition of the agent’s authority, so that anyone changing what an agent can do also sees who answers for it. When a provider operates the agent, some of these assignments, especially who leads an incident investigation, belong in the agreement as well.

Define Autonomy Boundaries Before Deployment
Defining an agent’s authority means sorting its possible actions into three groups:
- Autonomous: actions it can take on its own. These are typically routine, repeatable actions on noncritical systems, such as clearing temporary files, restarting a noncritical service, or rerunning a failed batch job.
- Review first: actions that need human review before execution. These include anything that touches critical infrastructure, security controls, sensitive data, or business-critical services, at least until the agent has an established track record.
- Off-limits: actions outside its authority entirely, such as disabling security tooling or deleting data.
Risk is the natural basis for drawing these lines: the greater the potential impact of an action, the more human judgment it should involve. In the hypothetical above, restarting services behind a clinical application in the middle of the business day belongs in the review-first group.
Confidence thresholds can refine these boundaries further. An agent might execute a low-risk remediation automatically when its confidence is high and route a lower-confidence recommendation to an engineer. A confidence score is only as useful as its calibration, though, so thresholds should be checked against how often the agent is actually right rather than taken at face value. The boundaries themselves should change over time, expanding as evidence of reliable performance builds and contracting after failures.
Oversight also has to be meaningful to work at all. Requiring approval for every action defeats the purpose of autonomy and breeds approval fatigue, where reviewers click through requests they no longer read. Human judgment belongs where the potential impact and the uncertainty justify it.
Monitoring Replaces One-Time Approval
Approval happens once, but an agent’s reliability doesn’t stay fixed. Infrastructure changes, new applications arrive, data shifts, and providers update their models. An agent that performed well at deployment can become less reliable as the environment drifts away from what it was built and tested on, so governing it means evaluating it continuously.
The most important measures are operational outcomes. An agent meant to reduce resolution time should be measured on resolution time. An agent that remediates alerts should be judged on whether the underlying problems stay fixed, not on how many alerts it closed. Supporting measures, such as error rates, escalation patterns, and how often engineers override the agent, help explain those outcomes. A rising override rate can be an early sign that something has drifted.
Monitoring also has to make agent behavior traceable. Teams need to be able to reconstruct what information influenced a decision, what action followed, and whether the outcome matched expectations. In the hypothetical, that record is what would show the agent acted on an outdated article. Without it, the post-incident review becomes an argument among the parties.
None of this requires a governance model built from scratch. The NIST AI Risk Management Framework organizes AI risk into four functions, Govern, Map, Measure, and Manage, which line up reasonably well with accountability, authority, and monitoring. The framework is voluntary and written for AI broadly rather than for IT operations specifically, but it’s a useful way to check an operating model for gaps.
Your Vendor Contracts Weren’t Built for Agents Either
Many of the AI agents that act on the IT environment come from vendors. Sometimes a managed service provider operates them on the customer’s behalf. Sometimes they’re built into ITSM, observability, or automation platforms the customer runs itself. There are good reasons for both. Vendors can bring production-ready capabilities and operational data at a scale few internal teams could match quickly. Either way, part of the governance moves into the agreement.
With a managed service, the provider’s staff usually configure the agents and run the monitoring. The customer approves the agents’ authority and the monitoring approach before deployment, and any later change to either goes through the customer’s change control. The contract defines how performance is reported and the outcomes the provider commits to.
With a software platform, the customer’s own team sets the agents’ authority and monitors them, while the vendor supplies the underlying models and updates them on its own release schedule. Those updates are changes like any other and should be validated through the customer’s change process, as far as the vendor’s release cycle allows.
The line between these two cases is starting to blur. Managed service providers increasingly deliver their services through agents rather than people, which turns the service itself into software. Managed service agreements have traditionally governed labor: staffing levels, skills, key personnel, and the processes people follow. When agents do much of the work, the terms that matter look more like software terms: which models and agents are in use, what they’re allowed to do, and how and when they change. A provider that updates its agents is changing how the service is delivered, in the same way a staffing or process change would, and it warrants the same review.
A handful of terms carry most of the weight:
- Model and agent changes. The customer should get advance notice before the vendor retrains or replaces models or expands what agents can do, along with a chance to validate material changes and a remedy if performance suffers. An agent that behaved acceptably last quarter may not be the same system this quarter.
- Measures that reflect AI. Alongside resolution time and availability, the vendor should report measures like override rates, escalation accuracy, and rollback frequency. Even if these start as reported metrics rather than service levels with credits, they bring degradation to the surface early.
- AI-initiated incidents. The agreement should state who leads the investigation when an agent’s action causes an outage, and how accountability and remedies apply, as distinct from human error or platform failure.
- Audit and traceability. The customer needs access to decision logs detailed enough to reconstruct what an agent did and why, kept long enough to support incident review.
- Data use. The agreement should say whether the vendor can use the customer’s operational data to train or improve its models, and on what terms.
- Exit. The agreement should settle who owns agent configurations, runbooks, and knowledge bases, and what happens to the customer data the agents rely on when the relationship ends.
None of these terms needs to be adversarial. Vendors confident in their AI benefit from clear terms, because clarity makes their capabilities easier for customers to approve and expand. Leverage varies, though: managed service agreements are usually negotiated, while large software vendors often work from standard terms that most customers can change only at the margins. Either way, the practical risk is timing. These questions often go unasked until after the agreement is signed, when there is far less leverage to settle them.

Governance Is What Lets Autonomy Scale
It’s tempting to treat governance as a brake on AI adoption, but it can work the other way. Organizations are more willing to extend AI into higher-value work when they understand its boundaries, can see how it performs, and know who answers for it. Without those foundations, the rational response to the first serious AI-caused incident is to pull autonomy back, and the savings that justified the investment go with it.
Many organizations will learn whether their AI governance works during their first significant AI-caused incident. Those that defined their agents’ authority, monitoring, and accountability in advance are positioned to treat it as a correctable failure and keep scaling. Those still relying on controls built for a different kind of system will be deciding, under pressure, how far to trust their agents at all. The same test is coming for agents that act on the business, in finance, HR, and customer service.