The Systemic Risks of AI-Driven Legacy System Transformation

0
2

Key Takeaways

  • Operation Epic Fury demonstrated that AI can generate and prioritize roughly 1,000 targets in the first 24 hours, sustaining a tempo more than double the 2003 Iraq invasion’s opening phase.
  • The Pentagon’s AI Acceleration Strategy removed many non‑statutory barriers but still lacks the policy, workforce, compute, data, evaluation, and authority‑to‑operate capacity needed for broad diffusion.
  • Missed deadlines for initial demonstrations of pace‑setting AI projects reveal a gap between rapid fielding and verified capability.
  • Generative (frontier) models hallucinate on 27‑85 % of short factual queries and can invent answers up to 94 % of the time, demanding new trust‑calibration, testing, and oversight mechanisms.
  • National Security Presidential Memorandum 11 pushes “adopt now, build the necessary systems later” while providing no clear office, funding, or methodology for assurance, increasing the risk of fielding untested AI.
  • Five core shortfalls—workforce, compute, data, evaluation, and authority to operate—compound each other; waiving processes alone cannot solve them.
  • Historical failures (Patriot fratricide in 2003, Israel’s Lavender system) show that automation bias and unverified outputs can lead to lethal mistakes when human review is rushed.
  • Congress authorized a testing sandbox, cross‑functional evaluation team, and governance subcommittee in the FY26 NDAA, yet none have been stood up despite expired deadlines.
  • Immediate actions recommended: designate workforce, compute, data, and evaluation as a single pace‑setting project; adopt provisional security accreditations for fast‑changing AI; fund independent testing infrastructure at universities and FFRDCs.
  • Without closing the capacity gap, the Pentagon’s AI adoption will remain confined to early‑adopter offices, jeopardizing reliability, safety, and long‑term strategic advantage.

Operation Epic Fury Shows AI at the Heart of Warfighting
The opening salvo of Operation Epic Fury made clear that “AI is now at the heart of American warfighting.” Using Claude through Palantir’s Maven Smart System platform, Central Command generated and prioritized roughly 1,000 targets in the first 24 hours, a tempo that “more than doubled the opening phase of the 2003 Iraq invasion.” Over the ensuing 38 days the campaign logged 13,000 total strikes, and Project Maven claims it can now push targeting to 5,000 in a single day. This speed underscores both the promise and the peril of deploying AI at scale without first verifying that the systems do what vendors claim.

AI Acceleration Strategy Removes Barriers but Leaves Gaps
Over the past 18 months the Pentagon cleared away non‑statutory barriers to AI adoption that were “long overdue for removal.” The AI Acceleration Strategy, released in January, stood up a Barrier Removal Board empowered to waive requirements across testing, contracting, and hiring, and mandated adoption of frontier models within 30 days of release. While these moves were necessary, they proved insufficient: “Building the department’s capacity to facilitate broad adoption is the second half of the acceleration effort.” Without parallel investments in policy, workforce, and infrastructure, adoption stays confined to a handful of early‑adopter offices.

Missed Demonstration Deadline Signals Implementation Risks
The strategy set July as the deadline for initial demonstrations across priority AI deployments—intended to prove that the strategy had unlocked real capability rather than merely accelerating systems the department cannot yet adopt at scale. However, “the department missed its own deadline for initial demonstrations. Nor has it made any statements about the status of the projects.” This lapse raises doubts about whether the pace‑setting projects can deliver verifiable, scalable AI or will stall once the overdue demonstrations finally arrive.

Generative AI Demands New Systems, Policies, and Investments
Unlike the narrow AI systems the Pentagon has fielded for years, generative models “do many tasks and change constantly,” rendering the traditional “test now, deploy forever” approach obsolete. Frontier model reliability varies enormously by task; on the HalluLens benchmark, “frontier models hallucinate on roughly 27 to 85 percent of short factual queries and invent answers about nonexistent entities up to 94 percent of the time.” This unreliability heightens automation bias, where operators may over‑trust articulate AI recommendations complete with precise coordinates and prioritized targets. Compressing evaluation to meet a 30‑day deployment mandate undercuts its purpose: a model whose failure modes remain unmapped is one operators learn not to trust, and systems operators don’t trust get bypassed in the field.

National Security Presidential Memo 11 Prioritizes Speed Over Assurance
National Security Presidential Memorandum 11, signed June 5, reinforced an “adopt now, build the necessary systems later” approach. The memo called for removal of “unnecessary barriers to rapid deployment” and listed assurance—ensuring AI is “reliable, robust, steerable, and controllable”—as one of its four pillars. Yet it “did not identify which office would be responsible for this function, nor did it direct funding for the work.” Instead, it paired the assurance emphasis with intense operational pressure by mandating deployment of new models within a month, effectively pushing operational deployment far ahead of independent testing. Two offices could theoretically assume the assurance role—the Chief Digital and AI Office’s Responsible AI Office and the Director of Operational Test and Evaluation—but both lack capacity: the former lost staff to a deferred‑resignation buyout and return‑to‑office mandate, while the latter saw its staff cut roughly in half and dropped nearly 100 programs from its oversight list without adapting methods for evaluating rapidly updating models.

The Five Pillars of the Capacity Problem
Frontier models update every few months, yet the Pentagon is not staffed, equipped with sufficient compute, or methodologically prepared to absorb them at that cadence. The Barrier Removal Board cannot waive its way out of a capacity problem. Five major shortfalls emerge:

  1. Workforce – insufficient machine‑learning engineers, acquisition officers versed in consumption‑based contracts, and operators trained to detect fluent‑sounding fabrications. Security clearances alone take the better part of a year before an engineer can start meaningful work.
  2. Compute – Pentagon‑owned facilities require six‑to‑eighteen‑month accreditation timelines, or more than two years for classified or complex systems, far slower than the weeks‑long update cycles of frontier models.
  3. Evaluation – reliance on vendor‑provided benchmarks via a “rent‑a‑bench” model leaves the department without independent methodologies; as of 2024, Maven correctly identified objects at 60 % accuracy versus 84 % for human analysts in the same 18th Airborne Corps evaluations, with no independent benchmarking to track improvement.
  4. Data – lack of a unified architecture prevents frontier models from accessing operational data; vendor contracts often fail to secure data‑rights, meaning the department funds systems it does not own.
  5. Authority to Operate – the traditional ATO process takes twelve to eighteen months, a cadence built for static systems, not for models that update every few weeks. Although Congress has written continuous authorization and cross‑service reciprocity into law, reciprocity remains aspirational: a mid‑2026 survey found none reporting timelines under six months, with “reciprocity” often still meaning a full redo of the old review.

These gaps will worsen if the Pentagon’s sole tool remains waiving processes, leaving the military unable to verify whether an AI solution performs as claimed—or does so safely.

What Failure Looks Like: Lessons from History
The department has already seen the consequences of fielding automated systems whose failure modes are unmapped. In 2003, Patriot air defense batteries shot down a British Tornado and a Navy F/A‑18 over Iraq, killing three aircrew; the batteries operated largely in automatic mode, conditioning operators to trust system outputs unconditionally. A Defense Science Board task force later found that when the system’s assumptions stopped holding, operators had no way to question what its sensors told them.

More recent parallels emerge from automation bias research. According to an investigation by +972 Magazine, Israel’s Lavender system flagged some 37,000 people in Gaza as suspected militants, with an estimated error rate of 10 percent, while human reviewers spent about 20 seconds deciding who to kill. When a system produces thousands of recommendations with fluent confidence, the human in the loop becomes a formality. Joint Publication 3‑60 calls for positive identification of a target, a collateral‑damage estimate weighed against military necessity, and legal review before a commander approves a strike—steps that only work if analysts have time to do them seriously.

During Operation Epic Fury, “No one, inside or outside the program, can say how many of the system’s object identifications during that campaign were correct.” A future failure could resemble a strike on a building the model confidently misidentified, approved by an operator trained to move fast, in a war where no evaluation record exists to establish whether the error was foreseeable. The department’s silence on current error rates, rather than a two‑year‑old accuracy figure, starkly reveals the capacity gap.

Congressional Action and Unmet Deadlines
Congress has attempted to close the gaps. The FY26 National Defense Authorization Act authorized three provisions: a testing sandbox, a cross‑functional evaluation team, and a governance subcommittee to assign clear accountability for AI oversight. None of these have been established, despite deadlines having come and gone. This inaction leaves the Pentagon reliant on vendor self‑assessment—a clear conflict of interest—and without the independent testing infrastructure needed to measure true performance against human baselines.

What the Pentagon Should Do Immediately
To keep pace‑setting projects from stalling once demonstrations finally occur, the Secretary should designate workforce, compute, data, and evaluation infrastructure as an official pace‑setting project under a single accountable leader drawn from the Responsible AI Office. Though assigning this mandate to an already depleted office poses immediate challenges, pace‑setting status provides standing access to the Barrier Removal Board and metrics that the Under Secretary for Research and Engineering reviews monthly. The project needs budget authority for its leader and output‑based metrics—engineers hired and cleared, compute delivered at each classification level, and data‑rights clauses executed in new contracts. Giving the project two years to show results creates a clear accountability checkpoint; if funds are not reallocated toward those hires, compute, and data rights by then, the Secretary should seek dedicated congressional funding.

Simultaneously, the Pentagon should pivot toward provisional security accreditations as the default for rapidly updating AI models. Provisional accreditation addresses cybersecurity risk, while capability evaluation—asking whether the system does what the vendor claims, under what conditions, and with what failure modes—remains a separate, independent check. Speeding up accreditation only makes sense if capability evaluation persists as a rigorous gate.

Finally, Congress should fund independent testing infrastructure for military AI housed in universities and federally funded research centers. Section 224 of the FY26 NDAA already directed the Department of War to establish a National Security and Defense AI Institute at a university. Such centers could measure target‑recognition accuracy against human analyst baselines on classified operational data, run adversarial tests for hallucination and prompt injection, and conduct human‑machine teaming experiments at operational tempo. Institutions like MIT Lincoln Laboratory and Carnegie Mellon’s Software Engineering Institute already perform classified evaluation work and could anchor this effort, with the sandbox and evaluation team Congress authorized serving as natural vehicles. Requiring regular reporting on results would create the transparency needed for commanders to field only what they trust.

Conclusion: Closing the Gap Is Essential for Broad, Lasting Adoption
Operation Epic Fury proved that the Pentagon can employ AI to prosecute a war at scale. Yet the initial pace‑setting projects, if they ever materialize, will only show whether the department can sustain that momentum across more systems. Neither speed nor early adoption guarantees that the military can verify system accuracy in combat or catch critical failures before they happen. Addressing the fivefold capacity shortfall—workforce, compute, data, evaluation, and authority to operate—is the prerequisite for driving broad diffusion and ensuring that AI adoption endures as a reliable, trustworthy component of American warfighting. Only by marrying rapid deployment with rigorous, independent assurance can the Pentagon harness AI’s potential without repeating the costly lessons of the past.

https://warontherocks.com/ai-and-the-risks-of-tearing-down-an-old-system/

SignUpSignUp form

LEAVE A REPLY

Please enter your comment!
Please enter your name here