Research pillar · Procurement Innovation

Government Should Stop Doing All Its Own Buying

Government should keep every acquisition function that requires public authority and test different ways to perform the commercial support work underneath it. This is not privatization. It is a way to find out what works.

The question

For routine commercial purchases, which support model performs best: the agency itself, a government shared service, or a qualified commercial purchasing specialist, and where does AI improve the work without crossing into governmental decision-making?

Active · since 2026 · Procurement Innovation

Download the white paper

Research essay, Word document, August 2026

Abstract

This is a research essay on public acquisition, commercial buying and AI, written from a practitioner perspective based on years working with federal buyers, contractors, technology programs and acquisition professionals. It is dated August 2026.

The argument: government should keep every acquisition function that requires public authority and test different ways to perform the commercial support work underneath it. The pieces already exist. Agency acquisition teams, GSA and other government shared services, contractor acquisition support, commercial-first buying methods, and increasingly capable AI tools are all in use. What is missing is a disciplined comparison of those operating models on the same kinds of commercial purchases, under the same legal rules, with the same measures.

The proposal is to treat acquisition reform as research and development: set a hypothesis, establish a baseline, run parallel trials, keep inherently governmental decisions inside government, use AI to lower the cost of learning, and scale only what performs better.

Keywords: federal acquisition; procurement; group purchasing organizations; shared services; commercial buying; acquisition reform; artificial intelligence; experimentation.

The problem is ordinary commercial buying

I have worked around federal buyers, contractors, technology programs and acquisition professionals for years. I know why the rules exist. Public money is different from corporate money. Contracting officers carry authorities and responsibilities that should not be handed to a vendor just because a new operating model sounds efficient.

I also keep seeing government teams spend expensive human time on work that looks a lot like ordinary commercial purchasing: market research, price checking, supplier discovery, catalog work, product comparisons, should-cost analysis, renewal analysis and chasing down what another agency paid for substantially the same thing. That is the part I want to test.

The federal government spent more than $495 billion in fiscal 2024 on common products and services, according to GAO. Category management was created to make agencies buy more like one enterprise, and OMB has reported more than $111 billion in savings since that effort began. GSA is consolidating more common buying today; its OneGov program says it saved $1.1 billion in its first year by negotiating governmentwide technology agreements. That is GSA's own reported number and worth continuing to check, but the basic buying principle is hard to argue with: a large buyer should use its scale.

I am not arguing that government should outsource procurement. I am arguing that we should stop treating procurement as one indivisible activity. Some acquisition work is government authority. Some of it is commercial support. I want to test which operating model does the support work best, while government keeps the authority.

Authority against supportTwo columns. On the left, the acquisition functions that stay inside government under FAR Subpart 7.5. On the right, the commercial support work underneath those decisions, which is what the experiment tests.Stays governmentalDeciding what the government willVoting in source selectionObligating fundsAwarding the contractContracting-officer authorityFinal accountabilitySupport work under testMarket researchPrice checking and benchmarksSupplier discoveryCatalog and product comparisonShould-cost and renewal analysisFinding comparable government buys
Government keeps the left column in every model. Only the right column is under test.

A GPO is shared buying expertise

A group purchasing organization is not complicated. Several buyers use a common organization to gather pricing, negotiate terms, maintain supplier relationships and buy with more bargaining power than each buyer would have alone. Hospitals use them heavily. Large companies do similar things with centralized procurement, sourcing firms and managed purchasing services.

The customer still decides what it needs. A hospital can decide which surgical supplies meet its clinical requirements and still use a GPO to handle pricing and supplier relationships. A large company can set its cybersecurity standard and still use a centralized purchasing team to negotiate the software agreement.

That is the distinction I care about for government. I do not want a private purchasing company deciding what an agency needs. I want to know whether the agency has to repeat all the commercial homework underneath that decision.

Government already shares buying work. GSA's Multiple Award Schedule, GWACs, assisted acquisition, interagency acquisition, category management, OneGov and the Office of Centralized Acquisition Services all move work away from one-off agency purchasing. State and local governments use cooperative purchasing on the same logic: share procurement contracts so each public entity does not have to run the same procurement from scratch.

So this is not a foreign idea imported into federal acquisition. The question is whether some shared purchasing work now done by an agency or another federal office should also be tested against qualified commercial providers.

The FAR already draws the line

I would start every experiment with the legal line, not bolt it on later.

FAR Subpart 7.5 says contractors cannot perform inherently governmental functions. In acquisition, that includes decisions where a contractor would be exercising government authority: determining what the government will acquire, voting as a government member of a source-selection board, awarding contracts, or exercising contract authority reserved to government officials. I would keep those functions inside government. If an experiment pushes them outside, that is a bad experiment, not innovation.

The same rules recognize that contractors can support work close to those decisions. FAR 7.503 identifies acquisition-planning support, technical evaluation, assistance developing statements of work and contract-management support among activities contractors may perform with proper oversight. FAR 37.114 requires closer supervision when contractor support sits near government decision-making. DFARS 207.503 also allows contracted acquisition support while government employees retain the inherently governmental functions.

That gives a simple design rule: separate authority from support. Government keeps authority. Test different ways to provide the support.

Most of the pieces already exist

Government has tried shared buying, commercial acquisition and outside acquisition support before, and that history strengthens the case for a practical test rather than weakening it.

GSA is building a marketplace approach for acquisition shared services. Its Acquisition Shared Services and Solutions QSMO is meant to define standards, vet providers, monitor performance and consolidate demand. In plain English, GSA is already telling agencies that they do not each need to build every acquisition capability on their own.

The Interior Business Center has operated as a fee-for-service shared-services provider for more than 35 years and serves more than 150 federal organizations. Commerce has a Shared Services Procurement Office that provides full-lifecycle procurement to bureaus and offices that do not maintain their own procurement authority. Those examples settle one point: an agency does not have to do all of its own acquisition work. The FAR and DFARS settle another: contractors can legally perform substantial acquisition-support work as long as government keeps the decisions and authorities that belong to government.

DIU adds a third lesson. It often starts with a problem rather than a long government specification and relies on people who understand both government and commercial markets. GAO calls this dual fluency. That matters because a commercial purchasing specialist is useful only if it knows the commercial market well enough to tell government what is normal, what is expensive, which suppliers are credible and where a government-unique term is shrinking competition. Government still decides what it is willing to change.

George Mason University's Center for Government Contracting made a related point in 2026: buying commercial solutions well requires commercial behavior. A commercial product pushed through a heavily customized government process can stop looking commercial very quickly.

The missing test is between operating models

Take the same kind of commercial purchase and compare three ways of doing the support work under the same legal rules.

Model one is the agency's normal acquisition team. Model two is a government shared-service provider such as GSA or the Interior Business Center. Model three is a qualified private purchasing specialist with deep commercial market expertise. Government officials keep the same reserved decisions in all three models.

Then compare the results. Which team found the best suppliers? Which had the best price information? Which challenged unnecessary requirements? Which reduced cycle time? Which improved competition? Which gave small businesses a fair shot? Which created more work for government counsel or contracting officers? Which produced the lowest total taxpayer cost after counting the cost of the support itself?

I do not want to decide the winner before the test. In some categories GSA may be best. In others a strong agency office may win. A private specialist may be better at continuous price intelligence or supplier discovery. The useful answer may be a mix. That is a better acquisition-reform question than asking everyone to adopt the same new process.

Three operating models, one purchaseOne ordinary commercial category at the top, with three support models beneath it: the agency's own acquisition team, a government shared service, and a qualified commercial purchasing specialist. All three work under the same rules and the same measures.One commercial category, same rulesAgency acquisition teamthe current baselineGovernment shared serviceGSA, Interior Business Centerand similarCommercial specialistqualified private purchasingsupport
Same category, same legal boundary, same measures. Only the support model changes.

Treat acquisition reform like R and D

Government has spent years designing acquisition reform before anybody tested the operating assumptions underneath it. Reverse that order.

Write down a hypothesis. Choose a narrow commercial category. Define what stays governmental. Set the baseline from recent acquisitions. Run two or three approaches in parallel. Measure the result. Stop weak approaches early and expand the ones that work. That is how good product development works, and it is how research works. Government already funds experiments when it does not know the answer: FAR 35.016 recognizes broad agency announcements for research and experimentation where multiple technical approaches are expected. That is not necessarily the legal vehicle for an acquisition-support pilot, but it shows that experimentation is already a normal government idea when we admit we are still learning.

GAO has criticized acquisition reform efforts that create isolated workarounds without changing the underlying system, and has pointed to leading commercial organizations that use iterative cycles, continuous user engagement, testing and incremental delivery. Acquisition reform itself should use those habits.

GAO's warning about DIU is worth taking seriously too. DIU has experimented aggressively, but GAO found it still needed better measurable near-term goals to show whether the model was succeeding at scale. Experimentation without a baseline, target and decision rule becomes another program everybody calls promising.

Before starting a pilot, know what result would make you continue, what result would make you stop, and what result would tell you that you tested the wrong thing.

AI can make the experiments cheaper

AI changes the economics of these experiments, because you no longer need to build a large new acquisition-support organization just to test whether an idea is useful. Use AI first as temporary research equipment.

For a market-research test, have an AI system organize public contract history, vendor material, schedules, product catalogs and prior solicitations into a structured market map for a human acquisition team to review, then compare that with the normal market-research process. For price intelligence, use AI to normalize product names, configurations, license tiers and contract descriptions that do not line up neatly across procurement records. That does not make the resulting number automatically correct; it lets a human analyst spend time checking the important differences instead of matching thousands of rows.

For requirements, let AI compare a draft requirement with commercial product documentation, prior solicitations and market norms, and flag requirements that appear unusually specific, old, duplicative or expensive. The AI does not get to delete the requirement. It gives the program office and contracting team a question worth asking.

For supplier discovery, AI can help find credible small and nontraditional suppliers a team might miss because they do not speak fluent government contracting. GAO has identified market research as one potential federal use of AI in small-business contracting, while warning about inaccurate output, bias and data-security risks.

For experiment measurement, AI can keep the evidence organized: timelines, labor hours, suppliers considered, requirement changes, negotiated prices, protest activity and user feedback across the three operating models. The experiment should produce a record you can audit, not a stack of anecdotes. The key is that you can use AI for six months and throw the prototype away.

AI as research equipment, then a gateAI enters as temporary equipment for a single experiment, is judged against the human baseline, and only then, at a gate, becomes a candidate for permanent capability.Narrow task, small data scopeHeavy human reviewCompare with the human baselineNo material improvement: stopImprovement: decide who shouldOnly then: permanent capability
A successful prototype is not a mandate to build a permanent system.

Use AI as a test tool before making it infrastructure

Government technology programs often confuse a successful prototype with a mandate to build a permanent system, so make this distinction explicit.

In the first phase, use AI to lower the cost of learning: a secure commercial model, an agency-approved model, a retrieval tool over public procurement data, or a narrowly built analysis workflow. Keep the data scope small and the human review heavy. The goal is to find out whether the task benefits from AI at all. If the experiment shows no material improvement, stop. You have learned something cheaply.

If it works, then decide whether the long-term answer is an agency capability, a GSA shared service, a feature bought from a commercial provider, or simply a repeatable analyst workflow using approved AI tools. Government does not have to own the model or build another system.

GSA's current guidance for buying AI recommends starting with the agency need, using testbeds, sandboxes or pilots before large-scale purchases, protecting data and monitoring usage costs. That is good advice for AI acquisition and equally good advice for using AI to improve acquisition. OMB's M-25-22 pushes agencies to track AI performance, manage risk, avoid vendor lock-in, protect non-public government data and maintain documentation supporting transparency and ongoing evaluation, and calls for sharing AI acquisition lessons learned. GAO found in 2026 that agencies were not yet systematically collecting those lessons. Make the lesson log part of the experiment from day one.

Keep AI for the boring work only if it proves useful

Long-term AI support makes sense only where the experiment shows repeatable value.

AI can watch commercial markets continuously for price changes, new license models, product retirements, new entrants and contract terms. It can build price benchmarks from government purchasing history and flag proposed prices that are materially different from comparable buys. It can match a new requirement against existing contract vehicles before a team starts building another procurement, and detect when two offices are describing substantially the same need in different language.

It can maintain a searchable acquisition memory so a new contracting officer can find how a similar problem was handled two years ago without knowing which folder to search. It can watch supplier-performance signals, organize CPARS and delivery information, surface repeated issues, draft comparison tables, summarize market research, prepare negotiation questions and keep an audit trail of what data supported a recommendation.

AI does not need to be the contracting officer. It needs to make the contracting officer better informed. GSA is already building an AI-enabled Procurement Automation Ecosystem intended to use centralized federal data for automation, competition and smarter buying, which makes experimentation more important, not less: learn which AI-supported tasks actually improve results before automating weak processes at scale.

Separate the operating-model test from the AI test

Two different experiments are easy to mix together.

The first tests the operating model: agency team against government shared service against private purchasing specialist. If one group gets powerful AI and the others get spreadsheets, you have learned almost nothing about the operating model. Give all three equivalent access to the same approved AI capabilities, or keep AI out of all three.

The second tests AI itself. Take comparable acquisition-support tasks and compare human-only work with human-plus-AI work. Measure accuracy, time, labor effort, supplier coverage, quality of price analysis and the amount of rework caused by bad AI output.

This sounds obvious, but government pilots often change five things at once and then declare the whole package a success. Preserve a human-reviewed evidence set so the AI can be tested against known answers. Without that, 'the output looked pretty good' becomes the evaluation method, and that is not good enough for procurement work.

AI can make a bad process worse

AI makes it easy to produce more procurement paperwork faster. That is not the same as making procurement better.

A model can draft a seventy-page requirement from a ten-line problem statement. It can generate evaluation factors, market-research language and acquisition-plan prose in seconds. If the underlying process is unnecessarily complicated, AI helps manufacture the complication at industrial speed.

AI should challenge work before it expands it. Does this requirement exist because the mission needs it, or because it was copied from the last solicitation? Is this clause required for this purchase? Are we asking for a custom capability the commercial market already provides another way? Has another agency already bought this? Are we comparing like prices? What would happen to competition if we changed this term? Those are better questions for AI than 'write me a longer acquisition plan.'

Lines I would not cross with AI

AI does not make an inherently governmental decision. It does not decide what the government will acquire, vote in source selection, obligate funds, award the contract, exercise contracting-officer authority or own final accountability. FAR 7.5 still applies when the support tool happens to be AI.

Procurement-sensitive, source-selection, classified, CUI or proprietary supplier data does not go into an AI service unless the environment and contract clearly allow it. M-25-22 addresses protection of non-public government data and restrictions on vendors using agency inputs and outputs to train public or commercial models without consent.

No AI price benchmark is accepted without knowing what transactions, configurations and assumptions sit behind it. A confident answer built from mismatched SKUs is still wrong. A supplier-controlled AI tool does not get to quietly shape requirements toward that supplier's own products; the conflict-of-interest issue does not disappear because the recommendation came through a chatbot.

AI does not silently screen suppliers out. If a model is helping with supplier discovery, qualification or proposal analysis, humans should understand the criteria and check for bias, missing data and false negatives. GAO has warned about inaccurate outputs, data-security concerns and biased outcomes in potential AI uses for small-business contracting.

And keep an audit trail. If AI influenced a meaningful acquisition-support recommendation, you should be able to reconstruct what information it used, what it suggested, what the human accepted or rejected, and why.

Lines not to cross with AIA single boundary statement, with six specific prohibitions beneath it: no inherently governmental decision, no sensitive data into unapproved services, no unexplained price benchmark, no supplier-shaped requirements, no silent supplier screening, and always an audit trail.AI supports, it does not decideEvery line below holds regardless of howNo governmental decisionno acquiring, voting, obligating orNo sensitive dataunless the environment and contractNo unexplained benchmarkknow the transactions behind theNo supplier-shaped requirementconflict of interest survives theNo silent screeninghumans see the criteria and theAlways an audit trailwhat it used, suggested, and who
FAR 7.5 still applies when the support tool happens to be AI.

Government and commercial buyers make different mistakes

Government processes create mistakes that commercial purchasing teams can often correct quickly. Government can over-specify the solution before it understands the market; a commercial specialist can show what the market actually sells and what each government-unique requirement will cost. Government can repeat market research because teams cannot easily find or trust work another office already did; a shared purchasing operation can maintain that intelligence continuously. Government can negotiate without enough current price context because comparable buys are buried in inconsistent systems; commercial procurement teams live on benchmarks, and AI can make those benchmarks cheaper to maintain, though humans still have to validate that the comparisons are real. Government can make it hard for newer commercial suppliers to understand how to enter the market; an outside specialist can translate government needs into commercial market language.

Commercial purchasing has its own bad habits, and government controls can prevent them. A private buyer may favor speed over public competition, rely on preferred suppliers because they are convenient, accept opaque rebates or volume incentives, tolerate vendor lock-in because switching costs hit later, or discount socioeconomic policy, public records, protests and congressional scrutiny.

That is why government should not imitate private procurement wholesale. Borrow commercial strengths without discarding public safeguards.

A private GPO can create new problems

Healthcare GPOs are a useful warning. GAO found that the major healthcare GPOs it reviewed were largely funded through administrative fees paid by vendors, usually as a percentage of purchase price. Critics argued that this can create incentives that are not perfectly aligned with the buyer. I would not copy that structure into federal purchasing.

Government should pay the purchasing provider. Vendor-paid fees and rebates should be prohibited or fully disclosed and credited back. Government should own the pricing and transaction data produced by the work. Compensation should not simply rise when government spends more. Suppliers should not have to pay a private toll to reach federal customers. Providers should be recompeted, data should be portable and agencies should be able to change providers.

And there should be several providers. Replacing a government bottleneck with a private oligopoly does not solve much.

Start with a 180-day acquisition lab

Not a governmentwide transformation program. Six months of disciplined testing.

Days 1 to 30: choose two or three ordinary commercial categories and five to ten participating buying offices. Build the baseline from recent acquisitions. Write down the decisions that remain governmental. Pick the measures, the stop conditions and the data rules before anybody knows which model will win.

Days 31 to 60: test market research, supplier discovery, requirement comparison and price intelligence. These are good first experiments because the outputs are advice and data. Government still makes every procurement decision.

Days 61 to 120: use the best-performing support models on live acquisitions for tasks already permitted to be supported outside government. Keep contracting officers, program officials, counsel and source-selection authority in their normal roles. Run comparable work through agency and government shared-service teams.

Days 121 to 150: look beyond the headline measures. Did outside support save government labor or merely move work to oversight? Did suppliers understand who spoke for the government? Did small firms get better access? Did AI create rework because its outputs were wrong? Did better market information actually change a requirement or negotiation position?

Days 151 to 180: publish a short results report saying which work the agency did best, which work a government shared service did best, which work a private provider did best, where AI helped, where it hurt and which experiments should stop. Then run the next round.

The 180-day acquisition labFive phases across six months: choose categories and baselines, test support tasks, run live comparable work, look past the headline measures, then publish results and decide what stops and what scales.Days 1-30categories, baselines, boundariesDays 31-60test support tasksDays 61-120live comparable workDays 121-150look past the headlinesDays 151-180publish, stop or scale
Six months of disciplined testing, not a governmentwide transformation program.

Measure the whole cost

A cheaper unit price can hide an expensive process. A faster award can hide weak competition. A clever AI tool can hide hours of human cleanup.

Measure acquisition cycle time, government labor hours, outside-provider cost, AI and tool cost, prices paid, number and quality of suppliers considered, small-business participation, requirement changes, competition, protest activity, contract performance and user satisfaction.

Measure how much reusable knowledge the experiment leaves behind. Did government retain a clean market map, pricing history, supplier list and lessons learned, or did all of that intelligence stay with the provider? GAO's 2026 review of federal AI acquisitions found that selected agencies were not systematically collecting lessons learned even as AI use grew quickly. That is the habit to avoid.

If you cannot show the baseline and the after-state, do not call the pilot a success.

Measuring the whole costUnit price is one measure among many. The full set covers cycle time, government labor, provider and tool cost, prices paid, competition and small-business access, rework and protests, and the reusable knowledge left behind.Total taxpayer costCounted after the cost of the supportCycle timefrom need to awardGovernment laborhours spent, and on whatSupport and tool costprovider, AI, oversightPrices and competitionsuppliers considered, small-businessRework and protestsrequirement changes, contractKnowledge retainedmarket map, pricing, lessons learned
A cheaper unit price can hide an expensive process.

Put acquisition professionals on the hard work

This is not a head-count exercise. Selling it that way would miss the point and turn the acquisition workforce against the experiment for good reason.

Contracting officers and acquisition professionals should spend more time on requirements, competition strategy, negotiation, market structure, security, supply-chain risk, mission tradeoffs, source selection, contract authority and accountability, and less time manually rediscovering yesterday's price, searching for an old contract, normalizing product descriptions or copying standard language between documents.

A strong contracting officer with better market data, better commercial support and well-controlled AI is not a weaker government buyer. That is a stronger one.

Where this could be wrong

Private GPO-style support may not beat GSA. It may not beat a strong agency acquisition office. AI may help with market research and be useless for another task. Some categories may benefit from shared commercial expertise while mission-specific acquisitions stay close to the agency. That is fine. What we want is to learn where each model works.

Government already experiments with weapons, technology, medical treatments and operating methods. Acquisition is important enough to get the same treatment: write down the hypothesis, keep public authority inside government, test the parts that can safely be tested, use AI where it lowers the cost of learning, measure what happens and change only what the evidence supports.

If acquisition transformation cannot test which parts of acquisition government actually needs to perform, it risks becoming another round of process improvement around assumptions that were never tested.

What would have to be true before calling a pilot successful

These are the open questions the research agenda has to answer with evidence rather than impression.

Total cost

Did it lower total taxpayer cost after counting outside support, AI tools and government oversight?

Where the time went

Did contracting officers spend less time on routine work and more on judgment and negotiation?

Competition

Did competition improve, including access for small businesses and commercial entrants?

The line held

Did government keep every decision and authority that should remain governmental?

AI without new harm

Did AI improve speed or insight without creating material accuracy, security, bias or audit problems?

Knowledge retained

Did government retain reusable market, pricing, supplier and lessons-learned data?

Beat both baselines

Did the approach beat both the agency baseline and the best available government shared-service option?

Clarity to suppliers

Did suppliers understand who represented the government and who was only providing support?

Durability

Would the result still look good after a recompete, protest, difficult delivery or market change?

Author note

This is an opinion and research essay, not legal advice. The argument is intentionally framed as a testable operating-model proposal.

Any live experiment would need agency counsel, contracting leadership, security, data-governance, ethics and program officials to define the legal and operational boundaries before work begins.