AI pilot compliance is a practical teaching problem because the failure point is often not the prototype itself. Federal reviews showed fast growth in agency AI activity, while documentation, policies, and inventory records did not always keep the same pace. In a July 29, 2025 review of 12 major federal agencies, GAO reported that AI use cases nearly doubled from 571 in 2023 to 1,110 in 2024, while generative AI cases rose from 32 to 282 GAO reported. For a kits-and-builds classroom, that gap is useful: students can build a review board simulation that tests whether a pilot has enough evidence to move toward approved use.
Why AI Pilot Compliance Fails After The Demo
AI Pilot Compliance Starts With Counts
The first technical issue is scale. A single prototype can be reviewed with informal notes, but hundreds of reported use cases require repeatable records. GAO’s 2025 finding on growth across 12 major agencies supports a clear lesson: an agency cannot evaluate risk, policy fit, or lifecycle status reliably if its inventory process is weak. In classroom terms, this means a model card or project poster is not enough. The review file needs a consistent identifier, a stated lifecycle stage, a data description, and a record of who owns the decision.
This is where a hands-on project kit can make the compliance gap visible. Give students several fictional AI project cards: a chatbot pilot, a document classifier, a scheduling assistant, and a data extraction tool. Some cards should include complete fields. Others should omit start dates, intended users, or lifecycle status. The task is not to judge whether AI is good or bad. The task is to decide which projects have enough evidence for review and which must return to the project owner for missing records.
Policy Lag Is A Technical Issue
GAO also found that many agencies struggled to keep use policies and guidance up to date with fast-moving generative AI technology. That is not only a legal or administrative issue. It affects system design. If a team does not know which data categories, review steps, or approval records apply, it may build a pilot that works during a demo but cannot be evaluated in the same way as a managed information system.
For students, the lesson is direct: compliance criteria should be treated like design constraints. A robot kit has voltage limits, current limits, and mechanical limits. A federal AI pilot has recordkeeping limits, data-use limits, and governance limits. If those constraints are absent from the build plan, the project may look finished while still lacking the evidence needed for approval. This framing keeps AI pilot compliance grounded in engineering practice rather than slogans.
Building A Classroom Review Kit
Kit Materials
A useful kit does not need live AI access. In fact, a paper-based review exercise may be safer for younger or introductory groups because the focus stays on evidence and decision records. The kit can include agency role cards, project owner cards, missing-field slips, lifecycle labels, and a review checklist. Teachers can connect this activity to broader readiness work, including AI adoption readiness lessons that examine governance, workforce gaps, data limits, and policy barriers.
- Use case cards with a project name, intended function, data type, and lifecycle stage.
- Evidence cards for pilot start date, owner, data boundary, user group, and review status.
- Policy cards that state which records must be present before a project can advance.
- Decision cards for approve, pause for missing evidence, or return for revision.
The kit should avoid asking students to make unsupported claims about model accuracy, bias, or security. Instead, it should ask whether the record contains enough information to start those reviews. That distinction matters. A missing lifecycle stage does not prove a system is unsafe. It proves the reviewer lacks a basic fact needed to place the project in the right review path.
Student Workflow
Divide the class into project teams and review teams. Project teams assemble evidence packets from the cards they receive. Review teams compare those packets against the checklist. A third group can act as the inventory office, tracking which projects are pilots, planned uses, or approved systems. This mirrors a common administrative challenge without requiring students to process real government data.
For cross-network classroom planning ideas, educators can also explore Natewin, a related site in the same network, while keeping the activity focused on the federal AI review evidence described here. The main build outcome is a defensible decision log. Students should be able to explain why a project advanced or why it was held for missing records.
Inventory Evidence Before Approval
Incomplete Records Change The Review
A December 12, 2023 GAO report covering 23 civilian Chief Financial Officers Act agencies found about 1,200 current and planned AI use cases, but only 5 of the 23 agencies provided fully comprehensive and accurate data; GAO cited missing elements such as pilot start dates, demographic data, or lifecycle stage GAO found. That finding is especially useful for a classroom build because the missing fields are concrete. Students can see that an inventory is not a ceremonial spreadsheet. It is the map reviewers use to understand what exists, what is planned, and what evidence is still absent.
AI pilot compliance becomes harder when basic fields are missing. A pilot start date can affect review timing. Lifecycle stage can affect whether a project is exploratory or moving toward operational use. Demographic data fields may be relevant when a system affects groups of people, depending on the use case and agency policy. The teaching point is not that every field has the same weight in every project. The point is that missing inventory fields prevent a consistent first review.
Policy Requirement Status Matters
The same 2023 GAO report stated that 10 of 23 agencies had implemented all selected AI policy requirements, 12 had implemented some but not all, and one agency was exempt. It also identified shortfalls connected to updated AI inventory guidance from OMB and missing occupational series for AI work. Those details are useful because they show that compliance gaps are not only project-level problems. They can also sit in the management layer that defines who tracks systems, how systems are classified, and what skills are assigned to AI work.
In a classroom kit, this can be modeled with two checklists. One checklist evaluates the individual pilot. The second evaluates the agency process. A student team may discover that its project packet is complete, while the mock agency still lacks a policy update or role assignment. That outcome teaches a careful lesson: a strong project file can still be slowed by an incomplete governance process.
Controls That Do Not Fit A Quick Demo

Data Boundaries And Audit Evidence
A quick demo often shows input and output. A compliance review asks for the surrounding evidence. What data entered the system? Who approved that data use? Was the project still a pilot, or had it started supporting daily operations? Which records show that reviewers can reconstruct the decision path? These questions do not require students to run an AI model. They require students to document a system boundary.
The most useful classroom diagram is a simple data flow: source data, processing step, AI service or model, user output, and stored records. Students can mark unknown points in red. If the data source is unknown, the pilot cannot be reviewed well. If stored records are undefined, later auditing becomes harder. If the user group is unclear, the review cannot identify who may be affected. This keeps AI pilot compliance tied to observable documentation rather than broad claims about innovation or risk.
Vendor And Model Records
Many pilots use tools or services that are not built fully inside the agency. A cautious classroom activity should not ask students to assess a real vendor. It can instead ask what records would be needed: the tool name, project owner, data category, intended function, and any policy limits the mock agency has set. The absence of these records is enough to trigger a return-for-revision decision in the exercise.
This approach also protects against overclaiming. Students are not asked to decide whether a model is explainable, secure, or fit for a mission based on a short output sample. They are asked whether the review packet contains the inputs needed for those later assessments. That is a better match for the federal findings, which point to gaps in inventories, guidance, and implementation status.
AI Pilot Compliance Gap In Federal Reviews
What The Kit Should Teach
The federal evidence supports a modest but valuable classroom conclusion: the gap between a pilot and an approved system is partly a documentation and governance gap. The 2025 GAO review showed rapid growth in reported AI and generative AI use cases. The 2023 GAO review showed that many agencies still had incomplete inventory data or partial implementation of selected AI policy requirements. Those facts make a strong case for teaching review evidence as part of technical education.
A well-designed kit should leave students with three habits. First, count and classify systems before judging them. Second, treat missing fields as technical blockers because they affect reviewability. Third, separate a working demonstration from an approved operational system. That distinction is the core of AI pilot compliance and a useful lesson for students who will later build, review, or manage automated systems in public-sector or regulated settings.