The date on the calendar
On April 3, 2025, OMB replaced the federal AI playbook with two memoranda: M-25-21, Accelerating Federal Use of AI through Innovation, Governance, and Public Trust, and M-25-22 on acquisition. Agency policies had to align with M-25-21 by May 6, 2026. Agencies must report their minimum practices for high-impact AI to OMB by September 22, 2026.
High-impact AI, in the memo's framing, is AI whose output materially affects rights, services, safety, or sensitive federal resources. If a system helps decide who receives a benefit, whose file gets flagged for review, or how a mission asset gets scheduled, it is in scope.
Most of what has been written about that deadline treats it as an administrative exercise. Designate a Chief AI Officer. Stand up a governance board. Publish the strategy. File the compliance plan. That work is real, and across the CFO Act agencies it is largely done. The part that is not done is the engineering, because the six minimum practices are not documentation requirements. They are requirements on the system itself.
The gap is technical, not procedural
In April 2026 the Government Accountability Office published its review of federal AI acquisitions, GAO-26-107859. It examined 13 AI acquisitions across the Department of Defense, the Department of Homeland Security, GSA, and the VA. Two findings matter to anyone building or buying one of these systems.
The first is scale. Federal agencies more than doubled their use of AI between 2023 and 2024. The number of systems that will eventually have to answer for their outputs is growing faster than the machinery for asking.
The second is capacity. Officials told GAO they had limited access to technical experts, data scientists among them, who could evaluate vendor proposals and assess how AI systems actually perform. GAO also found that none of the four agencies had a policy requiring them to collect lessons learned from AI acquisitions, so what one program discovers the hard way stays inside that program. All four concurred with the recommendation to fix it.
Read those together and the practical conclusion is uncomfortable but useful. A program office that cannot staff a data scientist cannot audit a model. It can, however, require a system that produces its own evidence. That shifts the burden from the buyer's expertise to the vendor's architecture, which is the only place it can realistically sit.
The six practices, as a specification
M-25-21 sets minimum risk management practices for high-impact AI: pre-deployment testing with a risk mitigation plan, a pre-deployment AI impact assessment, ongoing monitoring, adequate training and oversight, timely human review with an opportunity to appeal AI-enabled decisions, and the collection of feedback from users and the public.
As a governance checklist that reads like six documents. As a specification it is six things the software has to do.
- Pre-deployment testing implies a frozen evaluation set, a repeatable harness, and versioned results. If the testing lives in a notebook on one engineer's laptop, the practice leaves when the engineer does.
- Impact assessment implies the system knows which decisions it touches. Someone has to enumerate every decision path the model influences, which means those paths have to be enumerable in the code, not inferred from an architecture diagram drawn a year ago.
- Ongoing monitoring implies runtime instrumentation shipped with version one: input distributions, output distributions, fallback and refusal rates, latency, and an alert when any of them moves. Monitoring added in year two only sees year two.
- Training and oversight implies the interface teaches. A reviewer who cannot see what the model read cannot supervise what the model concluded, regardless of how many hours of training they sat through.
- Human review and appeal is the hardest requirement and the one most often deferred. An appeal only means something if a person can reconstruct why the system produced a specific output, for a specific case, on a specific date, under a specific model version. That is a data lineage and retention requirement, and retrofitting it is one of the most expensive things you can do to a running system.
- Feedback collection implies a path from a user marking an output wrong back into the evaluation set. Without that path, feedback is a mailbox.
None of this is exotic engineering. All of it is inexpensive at design time and punishing at retrofit time, which is precisely why it gets skipped during a pilot and then blocks the pilot from ever becoming a system.
Traceability is the load-bearing requirement
Five of the six practices collapse into a single architectural property: the system can explain, after the fact, exactly what it did and exactly what it used.
We build to that rule commercially, not only on government work. On XCreos, the AI underwriting platform we designed and built for commercial real estate, the design rule is absolute: no number appears anywhere in the platform without a traceable path back to the line in the source document it came from. AI does the reading. People stay able to check its work, one figure at a time.
That rule was adopted for a private-sector reason. An investment committee signs its name under a number, and a number nobody can defend is worse than no number at all. It turns out to be the same property a high-impact AI review is asking for, because the underlying question is identical: when this system is wrong, can a person see why, prove it, and correct it?
The other line we hold in both settings is where the model stops. Models read, extract, and research. Deterministic logic guards the numbers and rules that drive decisions. A language model is an extraordinary reader and an unreliable accountant. Where an output carries consequences, the arithmetic belongs in code you can test, and a conflict between two sources gets surfaced to a person rather than quietly averaged away.
Five questions that do not require a data scientist
GAO's finding about missing technical evaluators is the operative problem for most program offices. These five questions can be asked by a contracting officer or a program manager, and every answer is verifiable inside a live demo rather than in a proposal narrative.
- Show me one output and trace it to its inputs, including the model version and the retrieved source. If that takes more than a minute inside the product, it is not a feature. It is a favor someone does by hand.
- What happens when two sources disagree? Surfacing the conflict is a design decision. Silently resolving it is also a design decision, and it is the wrong one for anything high-impact.
- Where is the evaluation set, who owns it, and what happens to it when the model changes? A vendor who cannot show you the test set cannot tell you whether the next model version is better or worse.
- What is logged at runtime, for how long, and who can query it without opening an engineering ticket?
- When a user marks an output wrong, what physically happens next? Follow the answer all the way to the next release.
A vendor who built for high-impact use answers these from the product. A vendor who built a pilot answers them from the roadmap. The distinction is audible in about ten minutes, and it is the same distinction GAO is describing when it says agencies struggle to prove a system will behave reliably.
Where this starts
None of the above requires a larger AI budget. It requires deciding, before the build, which decisions the system is permitted to influence and what evidence it owes when it influences them. Written down early, that document is a specification. Written down after go-live, it is a finding.
That decision is the work of an assessment: two to four weeks, interviews with the people who own the decisions, and a plan that says what to build, what to buy, and what not to do at all. When the build follows, the same team that designed it operates it afterward, which is how the monitoring and feedback requirements stay true past launch instead of decaying into a dashboard nobody reads.
Megsoft holds active federal work at DoDEA and HHS, and is an SBA 8(a), Service-Veteran Owned, MDOT and VSBE certified small business. Our Government page has the contract vehicles and the capability statement.
