← Insights

Insights · OMEGACORTEX LLC

Human Oversight for AI Agents in Government: What OMB M-25-21 Asks For

By OMEGACORTEX LLC ·

Federal agencies are being encouraged to adopt AI faster. They're also being asked to show that the AI they deploy is tested, monitored and accountable. Both expectations come from the same document: OMB Memorandum M-25-21, "Accelerating Federal Use of AI through Innovation, Governance, and Public Trust," issued April 3, 2025 (M-25-21, PDF). It rescinded and replaced the earlier M-24-10.

If you build AI agents, meaning systems that don't just answer questions but take actions in other systems, this memo shapes what a federal customer will need from you. This article reads it from a builder's point of view. It's a plain-language summary, not legal advice. Each agency makes its own determinations, and the memo's text is what counts.

The short version

  • Most of the memo is about governance and adoption: Chief AI Officers, AI strategies, use case inventories and removing unnecessary barriers.
  • A specific set of minimum risk management practices applies to "high-impact" AI.
  • Whether a use case is high-impact depends on what the AI's output is used for, not on whether a human is involved.
  • For agentic systems, the practical consequence is to design for testing, monitoring, human intervention and an audit trail from day one, because retrofitting them later is expensive.

What counts as "high-impact" AI

The memo says AI is high-impact "when its output serves as a principal basis for decisions or actions that have a legal, material, binding, or significant effect on rights or safety."

Two details matter for agent builders.

First, a human in the loop doesn't automatically take a system out of scope. The memo states that "a high-impact determination is possible whether there is or is not human oversight for the decision or action." If a reviewer routinely approves whatever the agent recommends, the agent's output may still be the principal basis for the decision.

Second, the memo lists categories presumed to be high-impact (Section 6). Examples include safety-critical functions of critical infrastructure; medically relevant functions in healthcare; law-enforcement risk assessments; decisions on applications for critical federal services and benefits; and determinations about the terms or conditions of federal employment. An agency official can document that a specific use case in these categories doesn't meet the definition, but the starting presumption is high-impact.

Many useful agent deployments sit well outside these categories: drafting internal IT answers with sources, extracting fields from documents for a reviewer, summarizing a technical publication. Agencies still assess each use case, though, and a workflow can drift into high-impact territory as it gains authority. That's why the design questions below are worth answering early.

The seven minimum practices, and what each means for an agent

For high-impact use cases, M-25-21 requires agencies to implement minimum practices. Agencies had 365 days from issuance to document them, and they must safely stop using high-impact AI that doesn't comply. Here's each practice, with the engineering it implies for an agentic system.

1. Pre-deployment testing

Agencies must test before deployment and prepare risk mitigation plans that reflect expected real-world outcomes.

For agents: a test set built from realistic cases, including edge cases and adversarial inputs such as instructions hidden in documents, and re-run whenever the model, prompt or tools change.

2. AI impact assessment

Agencies must complete and periodically update an impact assessment covering, among other things, the intended purpose and expected benefit, the quality of the data, potential impacts, costs, independent review and risk acceptance.

For agents: document what the agent can access, what actions it can take and what happens when it's wrong. A clear list of tools and permissions makes this assessment far easier to write.

3. Ongoing monitoring

Agencies must monitor for performance and for adverse impacts over time.

For agents: log every input, retrieved source, tool call, output and approval. Track exception rates and reviewer overrides. Without a record of what the agent did, there's nothing to monitor.

4. Adequate human training and assessment

Operators must be trained on the specific system and how it's used.

For agents: reviewers need to know what the agent is good at, where it fails and what a low-confidence output looks like. A runbook is part of the deliverable, not an afterthought.

5. Human oversight, intervention and accountability

Agencies must ensure "human oversight, intervention, and accountability suitable for high-impact use cases" and, when practicable, a fail-safe that minimizes the risk of significant harm.

For agents: separate actions that need approval from actions the agent can take alone. Route low-confidence cases to a person. Provide a way to pause the agent and fall back to the manual process. Give every approval a named, accountable person.

6. Remedies or appeals

Affected individuals should have access to timely human review and a way to seek remedy.

For agents: if you can't reconstruct why a decision was made, you can't review it. Source-linked outputs and decision logs are what make an appeal workable.

7. Feedback from end users and the public

Agencies must provide ways for end users and the public to give feedback where available.

For agents: a simple channel to flag a wrong output, routed to the system's owner and fed back into the test set.

Pilots and waivers

The memo exempts pilot programs from the minimum practices if they are of limited scale and duration, the agency's Chief AI Officer has certified that the pilot may go forward (tracked centrally), participants can opt in or out where possible, and the minimum practices are applied where practicable. The CAIO can also issue written, risk-based waivers for specific requirements, which must be recertified annually.

For vendors, the lesson is practical. A pilot built without logging, testing or approval steps can't easily graduate into production. A pilot built with them already has most of the evidence an impact assessment needs.

A design checklist for agentic systems in government

  • Write down the decision or action the agent's output informs, and whether it could be a principal basis for a decision affecting rights or safety
  • Limit tools and permissions to what the workflow needs; enforce them in the downstream system, not just in the prompt
  • Define which actions need human approval, and by whom
  • Route low-confidence or out-of-scope cases to a person automatically
  • Log inputs, sources, tool calls, outputs and approvals
  • Build and maintain a realistic test set, re-run on every change
  • Provide a pause switch and a manual fallback
  • Deliver runbooks and reviewer training material
  • Provide a feedback channel tied to the system owner

These steps line up with the NIST AI Risk Management Framework's functions (Govern, Map, Measure, Manage), which many agencies already use to structure AI risk work (NIST AI RMF). For buying AI, OMB issued a companion memo the same day, M-25-22, "Driving Efficient Acquisition of Artificial Intelligence in Government" (M-25-22, PDF).

Where OMEGACORTEX stands

We build teams of AI agents where agents do the work and accountable people approve what matters. Our own development pipeline is designed so that agents work under explicit permissions, keep the source of every output, and a person is accountable for release decisions. That's the design pattern this article describes.

To be clear about what we don't claim: OMEGACORTEX LLC is a Washington self-certified micro business, registered and active in SAM.gov (UEI E822TKBBJGJ5, CAGE 23ST4). We don't claim prior federal agency performance, a CMMC certification, an agency authorization or an SPRS score. Contract-specific controls are a delivery approach that we tailor to the solicitation. Our federal page and capability statement have the details.

If your agency or prime team is weighing an agentic workflow and wants to design it for these practices from the start, Book a call · /gov/.


Sources