A Guide to Government AI for Agency Leaders
Government AI is the use of artificial intelligence inside federal, state, local, and tribal agencies — where public records, controlled unclassified information, and classified data all remain under agency control. Unlike commercial deployments, the central constraint is not capability: it is sovereignty. Agency data, records subject to the Freedom of Information Act, and classified material must stay within boundaries the agency can prove, or the deployment does not happen at all.
The pressure to deploy is real and dated. Executive Order 14179, “Removing Barriers to American Leadership in Artificial Intelligence” (January 23, 2025), set the policy of “sustain[ing] and enhanc[ing] America’s global AI dominance” and directed development of an AI Action Plan within 180 days. That plan, “Winning the AI Race: America’s AI Action Plan,” was released by the White House on July 23, 2025, and identifies more than 90 federal policy actions. OMB Memorandum M-24-10, executed March 28, 2024, requires each CFO Act agency to develop an enterprise strategy for responsible AI use and to follow minimum practices when AI affects the rights or safety of the public. For agency leaders, the mandate is no longer whether to adopt AI — it is how to adopt it while keeping control of the data.
Why public-sector AI differs from commercial AI
Four structural differences shape every decision that follows.
Data classification, not just security. Federal information spans public records, Controlled Unclassified Information (CUI), and classified material. CUI is a defined program: established by Executive Order 13556 in November 2010 and implemented through 32 CFR Part 2002, with agency data reported annually. CUI carries marking and handling requirements that a vendor’s cloud region selection cannot satisfy on its own. Classified data sits above that, on networks such as SIPRNet, where only a closed, self-contained enclave is permitted.
FOIA-adjacent records obligations. Records created by an agency — including records produced with vendor assistance — can be subject to the Freedom of Information Act, 5 U.S.C. § 552, which entitles the public to access agency records absent an exemption. When a vendor holds agency-generated data, the agency must be able to retrieve it, produce it, and explain what happened to it. “The data lives in the vendor’s cloud” is not a defensible answer to a records officer, a FOIA request, or an audit.
Budget cycles, not annual renewals. The federal fiscal year runs from October 1 to September 30, and agency IT funding is largely discretionary, set by Congress annually. An AI deployment that costs $200,000 a year must be modeled as a multi-year appropriation request with justification, not as a card on file. Multi-year cost behavior, not the sticker price, is what the budget office evaluates.
Procurement is the architecture. Federal purchase decisions run through the Federal Acquisition Regulation and established vehicles. OMB Memoranda M-25-21 and M-25-22, referenced by GSA’s Buy AI program, direct agencies toward efficient, governed AI acquisition, and GSA now lists AI products and services as standard contracting options. In practice, this means the deployment model has to be specified, defensible, and vendor-verifiable before the RFP goes out — not negotiated after a demo impresses a stakeholder.
Deployment options by data sensitivity
The first architecture question is where the data may sit. The answer determines everything else — security baseline, certification path, cost model, and which vendors can even bid. The standard ladder looks like this:
| Data sensitivity | Deployment model | Certification / authorization path |
|---|---|---|
| Public / internal, low impact | Public cloud (IaaS/PaaS/SaaS), commercial | FedRAMP Low authorization |
| CUI / mission data, moderate impact | FedRAMP-certified public cloud, or government/private cloud | FedRAMP Moderate; agencies may leverage an existing agency ATO for an AI service |
| High-impact unclassified CUI, DoD environments | DoD-approved commercial government cloud | DoD Cloud Computing SRG Impact Level 5 (IL-5) |
| Classified (up to Secret) | On-prem or enclave data center inside the agency’s classified network | DoD Impact Level 6 (IL-6): a closed, self-contained enclave connected only to the classified network (e.g., SIPRNet); no commercial-cloud path |
| Highest sensitivity / disconnected requirements | Air-gapped on-prem deployment, no external network path | Agency-specific ATO with physical isolation; no external authorization body applies |
Three consequences follow from this table. First, FedRAMP is a gate, not a nice-to-have. FedRAMP is a GSA-operated federal program for the security authorization of cloud services, and it categorizes cloud offerings at Low, Moderate, and High impact levels based on FIPS 199. GSA’s procurement guidance is explicit that all cloud service providers used by the federal government must be FedRAMP-authorized or in the process of obtaining authorization — and it recommends reusing an existing agency authority to operate rather than starting from scratch. Second, the ladder is not a menu. CUI rules mean some datasets simply cannot go to a commercial region; DoD impact levels mean IL-5 and IL-6 workloads run under different security requirements guides and different vendor ecosystems. Third, the top of the ladder is on-prem by construction. Classified and air-gapped environments are closed, self-contained, and disconnected — which is why the sovereign AI discussion (see what is sovereign AI) and on-premise AI matter to government buyers long before they matter to commercial ones. For defense-specific environments, sovereign AI for defense covers the IL-5/IL-6 authorization path in detail.

The workloads agencies actually deploy
Federal AI is not speculative. The workloads that dominate agency roadmaps cluster in four areas:
Document processing and records management. Agencies hold massive backlogs of records, correspondence, and case files. AI-assisted document triage, classification, summarization, and FOIA request support — locating responsive records, identifying exemptions, generating summaries — is the most common first deployment, because the data is internal, the workflow is repetitive, and the upside is measured in staff hours recovered. The records obligations cut both ways: the same documents the AI helps process are the documents the agency must be able to produce under FOIA.
Citizen services. Public-facing intake, routing, and response — permit applications, benefit inquiries, 311-style services, multilingual help desks. The sensitivity here is lower (often public or PII-level), which makes cloud deployment viable, but the service-level and accessibility expectations are higher than in commercial settings because the agency is the only provider.
Predictive maintenance of public assets. Bridges, water systems, rail, and fleet vehicles generate sensor and inspection data that predicts failure before it happens. This is a classic on-prem or edge-adjacent workload: the data is local, the model benefits from staying near the asset, and the deployment avoids egress costs that would make continuous telemetry uneconomic.
Situational awareness and operations. Fusion of feeds for emergency management, infrastructure monitoring, and threat picture — the workloads that push agencies into the higher rungs of the sensitivity table, where air-gapped and classified deployment models apply.
A pattern worth noting: nearly all four workloads are regulated AI workloads in the practical sense — the deployment is governed by a specific set of requirements (CUI handling, ATO, audit logging) that the vendor must demonstrate, not merely claim.
The cost frame of TCO versus cloud egress
The commercial AI price frame is per-token or per-seat SaaS billing. The government frame is different, in four ways:
- Multi-year budget line, not annual renewal. A FY-cycle deployment is scored on three to five years of total cost of ownership — license, hosting, integration, security review, and operations. The budget office will model year two and year three pricing explicitly; usage-based bills that “can grow quickly without proper monitoring” (GSA’s own words) are a disqualifying risk if the envelope is not modeled.
- Egress is a real line item. Predictive-maintenance and situational-awareness workloads move data continuously. Routing that telemetry to a remote cloud for inference and back produces recurring egress charges that can dominate the license cost — a structural argument for processing at the edge or on-prem, where the data never leaves the facility.
- Authorization has a price, and reusing it saves money. A FedRAMP authorization and an ATO consume assessor time, staff time, and calendar weeks. GSA explicitly recommends leveraging existing authorities where available. An in-house platform that already carries the agency’s ATO avoids re-running the authorization for every new model or workflow.
- Procurement overhead is front-loaded. Security review, contract negotiation, and vehicle compliance consume staff months before the first dollar of subscription. Vendors who come with FedRAMP status, an ATO package, and pilot-ready environments shorten this phase materially — which is a legitimate part of the price comparison, not a tiebreaker.
How to procure AI in government
The RFP is where sovereignty gets decided. The practical checklist:
- Specify the deployment boundary in the statement of work. State explicitly where data is processed and stored: FedRAMP-authorized cloud, agency-operated data center, or air-gapped. “Vendor’s discretion” is not a requirement — it is a waiver of control.
- Require the authorization artifacts, not just the claim. FedRAMP authorization letter (and impact level), current ATO or provisional ATO path, and for DoD workloads the DoD Cloud Computing SRG impact level (IL-5 or IL-6). Verify against the FedRAMP Marketplace, which lists every FedRAMP-certified service.
- Demand data residency and portability language. The agency must be able to retrieve all data, models, and logs on contract end, and the vendor must not use agency data to train shared models without written authorization.
- Require audit and records support. Immutable audit logs for model inputs and outputs, plus a contractual commitment to support records production (including FOIA responses) for agency-generated outputs.
- Model the multi-year cost in the evaluation criteria. Weight three-year TCO, egress/egress-avoidance design, and usage-cap behavior explicitly in the scoring matrix rather than comparing sticker prices.
- Structure a pilot-to-production path. GSA’s guidance recommends evaluating solutions in testbeds, sandboxes, or pilot programs with a small user group before large-scale purchase. Specify the pilot’s exit criteria (accuracy threshold, ATO status, cost confirmation) so the pilot is a decision, not an extension.
- Engage the full governance bench before the RFP ships. GSA’s guidance names the coordination set: Chief Information Officer, Chief AI Officer, Chief Data Officer, Chief Information Security Officer, and Chief Privacy Officer. M-24-10 requires each CFO Act agency to develop an enterprise strategy for responsible AI use — the RFP should map to it.
Procurement is also where the market gap shows. Most incumbent AI vendors frame public-sector offerings around their cloud: the pitch is platform and models, and the deployment boundary is whatever the cloud allows. When an agency’s data cannot leave its own facility, that framing does not survive the requirements review. The sovereign framing — platform, model, and data all inside the agency’s control, with the authorization artifacts to prove it — is the frame that maps onto the actual sensitivity ladder. That is the core of the sovereign AI argument, and the one worth testing against a working system before an RFP is finalized. See the demo.
Last verified: 2026-09-05

