Skip to main content

Many government accountability systems were built around the basic assumption that people—not autonomous software—make consequential decisions.

Agentic AI is changing that assumption.

Before agencies decide how autonomous government should become, they should consider what authority they are prepared to delegate.

Agentic AI could help governments anticipate needs and orchestrate execution. And public expectations may already be shifting. In a recent Salesforce global survey on government services, more than 90% of respondents said they would use an AI agent to interact with government.1 The question is whether the public sector is ready to manage the accompanying risks.

Those risks may become tangible when viewed through the experience of a constituent. Imagine a mother logging in to check her son’s disability benefits, only to find that his coverage has been suspended with no explanation. She doesn’t know why it happened, how to appeal the decision, or how long it will take to restore the benefits her family depends on. She spends sleepless nights worrying about what comes next.

Most people never see the systems behind government services. But they feel the consequences when those systems fail.

When government uses AI to make important decisions, the stakes can be high. A flawed agentic response could suspend benefits, delay important services, or affect someone’s legal rights, and ultimately erode public trust in government. These actions may also need to withstand legal, regulatory, and public scrutiny months—or even years—later.

Without proper governance, agentic systems could create new challenges and compound existing deficiencies in government programs. This risk is especially acute because many government agencies still operate through fragmented systems, paper-based processes, and bureaucratic hierarchies designed for human decision-makers and are ill-suited to AI agents.

Agentic AI has the potential to transform how governments deliver services, but only if agencies also redesign their operating models and accountability frameworks.

As governments move from AI assistants that support human decision-making to AI agents that can act on their own, leaders face three fundamental questions.

  • What authority should an agent have?
  • What new risks may emerge when AI has agency?
  • What operating model allows agencies to govern those risks while improving mission delivery?

The promise of autonomous agents in the public sector

Government AI is evolving from copilots that assist human workers to agents that can plan, access data, coordinate with other systems, and act on their own (figure 1). That shift could help agencies improve service delivery in many areas.

AI agents could modernize permitting, licensing, grants, and procurement by coordinating reviews across offices, flagging missing documentation, and tracking compliance with statutory and regulatory requirements. Where approvals are delayed by system silos and manual handoffs, agents could shorten cycle times while improving consistency and transparency.

Agentic AI pilots in some states, for instance, review statutes and regulations to identify inconsistencies and opportunities for simplification. This offers an early look at how these systems could improve both government operations and citizen experiences.

AI agents can also help address behind-the-scenes gaps in benefits programs. In many states, eligibility determination and claims processing operate separately, creating gaps that monitoring agents could help bridge through real-time eligibility triggers and automated cross-checks.

AI agents could also scour policies, perform compliance checks, and monitor benefit usage to help agencies prevent fraud, waste, and abuse. Deloitte has found that using agents to pre-package relevant data points and behavioral patterns can shorten the time needed for human case review from 30 hours to just 4 hours.2

But as AI systems begin taking independent action, the risk equation changes. An agent that can initiate a workflow, update a record, trigger a payment, or influence an eligibility determination raises hard questions about authority, accountability, and control. The challenge is to confirm that systems capable of action do so in ways that remain lawful, transparent, and accountable.

Agentic controls to mitigate new public-sector risks

When an AI agent is embedded in a public program, citizens may have no practical alternative channel for service, raising the stakes for accessibility, fairness, and recourse. That makes accountability especially important: A public agency employing AI agents should be able to reconstruct events months or even years later to demonstrate that its decisions were lawful, fair, and consistent with policy.

Public agencies also operate within legal, procedural, and institutional constraints that make agentic technology harder to govern than in the private sector (see “Government-specific constraints”).

Government-specific constraints

  • Statutory mandates that limit discretion
  • Due process and appeal rights
  • Freedom of Information Act and records retention requirements
  • Procurement cycles
  • Legacy systems that shape what agents can do
  • Boundary risks to interagency data-sharing
  • Appropriations issues
  • Workforce rules and unions
  • Political accountability and reputational stakes
  • Public trust, both an outcome and a constraint

Autonomous AI introduces new risks when agents use tools, retain memory, interact with other agents, and create hard-to-audit chains of action (figure 3). For example, an eligibility agent might automatically suspend benefits for thousands of recipients after misinterpreting earnings data from another system, or an erroneous fraud classification in a tax system could spread across agencies to benefits and enforcement systems.

The right technical controls are only part of the solution. Those controls should be supported by strong governance and accountability mechanisms that ensure government AI agents are oversight-ready by design.

As discretion shifts to AI agents, accountability should remain with the human

In the public sector, an AI agent must remain bounded by law, visible to oversight, and contestable by affected individuals.

Human supervision is often seen as the default way to keep AI systems accountable. But agentic AI changes where discretion sits.3 In traditional systems, caseworkers exercise discretion; in agentic systems, that discretion is split across code, data, models, workflows, and people. And that shift creates new accountability risks (figure 4).

Government’s task is to make this redistributed discretion accountable. Rather than putting a human at every step, agencies should put accountable human judgment where it matters most in a workflow. That requires agencies to rethink how agents are supervised and build clear authority, oversight, and accountability into the system from the outset. As agencies develop their AI governance and accountability protocols, they should consider the following approaches.

Authority precedes autonomy

Before agencies decide how to supervise agents, they should determine what authority those agents can exercise. Should they only recommend actions, or can they initiate workflows, route cases, or execute transactions? Implementing an AI agent use policy in conjunction with a decision tree–based approach can take the guesswork out of the process. Clear limits on delegated authority are the foundation of effective governance.

Anchor AI in public-sector governance

Because the public sector globally already has some AI standards and policy guidelines in place—including, for example, the US Government Accountability Office’s AI Accountability Framework, the National Institute of Standards and Technology’s AI Risk Management Framework, Singapore’s Model AI Governance Framework for Agentic AI, and the United Arab Emirates’ AI Ethics—agencies can treat agentic AI governance as an extension of existing federal, state, and local expectations, rather than a separate compliance task.

Abu Dhabi’s TAMM 4.0: Building governance into AI-enabled public services

Abu Dhabi’s TAMM platform has more than 4 million users and brings more than 1,000 public services into a single user experience. The system can detect when a citizen may need to renew a document or update a registration and initiate the process automatically.

That shift is part of a broader, systemwide redesign led by Abu Dhabi’s Department of Government Enablement. AI systems are increasingly expected to meet data classification requirements, national cloud security standards, and auditing requirements for public infrastructure. Governance features—including validation agents, monitoring tools, and structured points of human review—are being integrated directly into AI-native platforms, especially in healthcare, infrastructure, and other high-risk sectors. Federal guidance sets expectations and reference points while giving operating entities flexibility in how they implement them.4

Use agents to strengthen oversight

As governments deploy more autonomous systems, monitoring should keep pace. This means first creating observability over all aspects of the agentic processes, including decisioning, prompting, resources used, token usage, and data accessed. Then specialized guardian agents can help human supervisors detect anomalies, monitor compliance with policies and regulations, enforce permissions and access controls, and identify unusual patterns warranting investigation.5

Guardian agents can also support auditing by reconstructing decision pathways, identifying high-risk cases for review, and monitoring workflows across systems and agencies.6 The goal is to give supervisors better visibility into complex agentic environments and help them intervene before small problems become larger failures.

A six-step playbook for oversight-ready AI agents

Once agencies have defined what authority AI agents can exercise, they need a way to translate those boundaries into day-to-day oversight. The following six-step playbook offers a practical approach for identifying agents, assessing their risks, setting operational controls, monitoring their behavior, responding to incidents, and addressing root causes.

1. Discovery: Know where the agents are operating

Create an AI agent registry. Track every deployed agent, its owner, mission use case, authority level, tools, data access points, and risk level. Like an asset inventory, the registry provides a basic control point for visibility, oversight, and accountability.

Classify agents by consequence. Agencies should consider mission impact, effects on rights and benefits, safety implications, privacy exposure, fiscal consequences, and public trust. Higher-consequence agents should face stronger accountability requirements, including explanation obligations, appeal pathways, remediation procedures, and human review for the most consequential decisions.

The risk taxonomy (figure 3) identifies what can go wrong; the consequence model (figure 5) can help agencies decide how much control each agent requires.

Once agencies understand the potential consequences of an agent’s actions, the next step is to evaluate whether the agent can operate safely and reliably within those boundaries.

2. Evaluation: Assess safety and reliability

Assess data and vendor risks. Determine the sensitivity of data that an agent can access and assess vendor practices for data handling, retention, model training, and customer data isolation. Know what you are buying through an AI Bill of Materials—a machine-readable inventory that documents every model, data set, dependency, and configuration used in an AI system. It provides standardized traceability for how an AI model was built, trained, validated, and deployed.

Test for reliability. Evaluate the model's tendency to generate inaccurate, unsupported, or biased information, as well as its resilience to prompt injection attacks via indirect prompts embedded in documents, websites, emails, or user-generated content.

3. Control: Set and enforce operational boundaries

Apply zero-trust principles to agents. Implement least privilege rules, scoped tool access, session limits, identity controls, and revocable permissions. A zero-trust approach can help agencies reduce the impact of mistakes, misuse, or compromise so agents act only within their authorized boundaries.

Assign clear accountability. Every agent should have a business owner, a technical owner, a data steward, a risk owner, and an escalation path. Clear ownership helps agencies identify who is responsible when issues arise and respond effectively.

Ukraine’s Diia.AI: An agentic model for a personalized, low-friction service experience

Diia.AI is a national AI assistant that allows citizens to access and complete public services through simple conversation, executing requests end-to-end. For example, a citizen can request an income certificate in their native language, specify the timeframe, and receive an official document within minutes.

Human oversight remains essential. The system follows a zero-trust model, enforces strict data protection, and does not grant AI direct access to personal data. Continuous feedback, adversarial testing, and a phased rollout—starting with the web portal before expanding to an app with more than 24 million users—have ensured reliability and resilience.7

4. Observation: Monitor agent behavior continuously

Monitor continuously. User behavior analytics and continuous anomaly detection can help identify previously unknown vulnerabilities or zero-day exploits in real time. Agencies can also use red teaming, scenario testing, bias testing, adversarial testing, and drift monitoring to identify emerging risks and changes in agent behavior.

Singapore offers one example. Its AI Guardian system comprises two complementary tools: Litmus, for safety testing and risk assessment, and Sentinel, for active protection against prompt injection, data leakage, and other AI vulnerabilities.8

5. Response: Prepare to intervene when problems arise

Execute targeted response protocols. Agencies should contain incidents by pausing compromised agents, revoking permissions, quarantining affected systems and cases, and reversing erroneous actions to limit harm.

Build leadership and supervision capability. Effective response also depends on people knowing how and when to intervene. Executives and managers should be prepared to redesign workflows, assign accountability, govern discretion, engage residents, and supervise multi-agent systems. Workers should also learn to supervise, challenge, override, audit, and improve agents—not simply prompt them.

The United Arab Emirates government has collaborated with Mohamed bin Zayed University of Artificial Intelligence to provide agentic AI training for 80,000 federal employees, including specialized learning tracks for leaders.9

6. Resolution: Fix root causes and improve

Identify and address root causes. Tighten permissions, update guardrails, and remediate exposed data to reduce the risk of similar incidents. Use lessons from the incident to strengthen controls and improve the agent over time. Restore affected systems to their intended state rather than the default baseline.

Building an accountable agentic future

Agentic AI can help public agencies anticipate needs, coordinate services faster, and deliver more personalized support. Realizing those benefits will depend on giving AI agents clearly defined authority and building accountability into how they operate. For citizens, the real test will be whether agencies can explain decisions, correct mistakes, and take responsibility when things go wrong.

Governments shouldn’t have to choose between speed and safety. By prioritizing governance, accountability, and risk management, they can deploy thoughtfully designed systems that improve mission delivery while maintaining the oversight and transparency needed to build public trust.

Continue the conversation

Meet the industry leaders

William D. Eggers

Executive director | Deloitte Center for Government Insights | Deloitte Services LLP
Deloitte United States

Joseph Conti

Managing director | Trustworthy AI Leader, Government and Public Services | Deloitte Consulting LLP
Deloitte United States

by

William D. Eggers

Deloitte United States

Joseph Conti

Deloitte United States

Amrita Datar

Deloitte Canada

ENDNOTES

  1. Salesforce, “Salesforce research: 90% of constituents ready for AI agents in public service,” Jan. 15, 2025.

  2. Interview with Amina Popowich, principal in enterprise operations and risk practice, Deloitte, March 11, 2026.

  3. See, for instance, Stephen Goldsmith and Junchung Yang, “AI and the transformation of accountability and discretion in urban governance,” Urban Governance 6, no. 2 (2026): pp. 159–171.

  4. INSEAD, “AI as public infrastructure: Lessons from the UAE for government transformation,” May 19, 2026; Innovation Library, “TAMM 4.0 shapes the future of digital government,” 2025.

  5. Deloitte and The Wall Street Journal, “Autonomous AI is rewriting the rules of risk management,” July 2026.

  6. Scott Holcomb, Clifford Doss and Amandeep Singh, “Taking a first step: Guardian agents for agentic AI applications,” Deloitte, accessed Sept. 24, 2026.

  7. World Economic Forum, “Making agentic AI work for government: A readiness framework,” April 24, 2026.

  8. AI Guardian, “AI Guardian: Safeguarding AI applications for Singapore’s public sector,” accessed Sept. 24, 2026.

  9. Emma Thompson, “UAE government partners with MBZUAI to train 80,000 federal staff in agentic AI,” EdTech Innovation Hub, May 26, 2026.

ACKNOWLEDGMENTS

Editorial (including production and copyediting): Kavita Majumdar, Aparna Prusty, Pubali Dey, and Anu Augustine

Design: Adamya Manshiva, Tushar Barman, and Natalie Pfaff

Cover image by: Adamya Manshiva and Tushar Barman

Knowledge services: Viswa Teja Vanapalli

COPYRIGHT