Start With a Procurement Framework, Not an AI Policy

A city does not need one exotic procurement model in 2026; it needs a repeatable process for deciding whether artificial intelligence is justified, selecting suppliers, governing data, testing effects, measuring public value, and terminating systems that fail. The framework should cover everything from low-risk tools that draft routine documents to high-risk systems that influence benefits, housing, employment, inspections, policing, or zoning. This is important because purchasing a model is also purchasing an administrative decision system: software determines what evidence staff see, which applications receive attention, how residents are classified, and whether a human can meaningfully challenge an outcome.

Also worth reading: How should urban planners and municipal leaders develop effective AI procurement guidelines for local government contracts? · What are municipal AI procurement best practices for modern city governments? · What is a municipal algorithm audit framework and how should city planners implement it?

A workable municipal framework should therefore connect technical review with procurement law, civil rights, public-records rules, accessibility obligations, labor policy, information security, privacy, and the city’s budget process. Technical performance cannot compensate for an unlawful purpose, inaccessible service, undisclosed data transfer, or expansion of discretion without accountability. The framework should be risk-tiered rather than vendor-neutral in the abstract; an applicant-support tool that reduces filing errors, for example, should not face the same evidentiary and contestability requirements as a predictive system used to screen tenants for code enforcement.

For 2026, cities should treat the procurement framework as a governance program with published standards, assigned owners, documented decisions, and scheduled audits. Numeric thresholds should be adapted to local law, but cities can begin with a simple structure: consequential decisions about individual access to services or property receive enhanced review; internal tools that cannot materially affect rights receive standard controls; and experimental systems receive limited authority. A city that lacks mature AI expertise should initially restrict deployments to reversible, assistive uses rather than attempting to launch autonomous decision systems.

Define Municipal AI by Function, Not by Marketing Label

Many procurement agreements describe products as decision support, automation, predictive analytics, or generative AI, but these labels can obscure the same underlying administrative function. A system is municipal AI when software materially analyzes data, predicts or recommends an outcome, generates content used in a government decision, or triggers an action affecting residents, staff, or property. The city should define scope this way so that a foundation model, conventional business software, machine-learning model, rules engine, or outsourced human review process cannot escape scrutiny simply because the contract does not call it AI.

This functional definition also prevents cities from overstating their own risk classification. Not every algorithm is a high-risk system, and not every spreadsheet-assisted process requires a specialized review board. Nevertheless, sophistication is not a safe proxy for harm. A small model that denies a food-renewal application, a vendor’s claim that its tool merely “assists” staff, or an opaque rule inherited from a purchased platform may create serious rights and due-process concerns. Procurement documents should ask what decision is being made, what role the software plays, what human authority remains, and what happens if the output is wrong.

The city should also define prohibited and specially controlled uses before it begins collecting proposals. Categories warranting special attention include facial identification, biometric surveillance, predictive policing, automated benefit denials, tenant screening, immigration-related predictions, and systems that rank residents for enforcement or scarce services. These uses should not be prohibited automatically in every jurisdiction, but each should require a specific legal analysis, public purpose, necessity finding, alternatives assessment, and community consultation. The absence of an express ban should never be interpreted as municipal approval.

Use a Risk-Tiered Review with Clear Gates

A risk-tiered process is more useful than a single checklist because cities operate both convenience tools and systems that can affect a person’s liberty, housing, income, or physical safety. One workable structure uses three procurement tiers, with illustrative—not universal—thresholds that local counsel should reconcile with constitutional, statutory, civil-rights, labor, and procurement requirements.

TierTypical municipal useCore requirementsApproval and renewal
Tier 1: LimitedMeeting summaries, internal drafting, translation with human review, non-consequential searchSecurity and privacy baseline, vendor and data inventory, accuracy check, staff training, published noticeDepartment approval; annual owner review
Tier 2: OperationalApplicant guidance, case routing, inspection prioritization, planning analysis, staff schedulingData and model documentation, accessibility testing, bias and impact assessment, human appeal path, public reportingCross-functional review and limited pilot; annual independent review
Tier 3: ConsequentialBenefits eligibility, housing or zoning decisions, policing, employment, biometrics, other rights-intensive usesAdvance legal and civil-rights analysis, necessity and alternatives review, notice, contestability, heightened audit, public explanation of data and criteriaSenior approval or council authorization where required; fixed-term contract with renewal evidence
The numerical boundary between tiers can be based on the probability and severity of harm, the scale of affected residents, irreversibility, autonomy, and the sensitivity of the data. A tool affecting more than 10,000 residents deserves closer scrutiny than a small internal prototype, but scale alone is not decisive: a system that silently suspends shelter payments for 20 families can still be serious. Cities should also add a presumption that moving from one tier to another triggers a new review, rather than allowing a pilot to become permanent through repeated extensions.

Each stage should have a defined gate. Legal and data review should occur before contract negotiation; testing should occur before production; accessibility and impact findings should occur before expansion; and performance should be reviewed before renewal. A vendor should not be allowed to narrow the review by describing accuracy, security, or fairness commitments as proprietary. Core contract terms—including audit rights, documentation standards, data retention, and remedies for deficient performance—must be negotiable before award.

Build the Evaluation Around Public Value and Administrative Fairness

A city should not select a model primarily by asking how advanced it appears. The central question is whether AI improves a documented public-service problem in a way that ordinary process improvement, better staffing, procurement modernization, or policy redesign cannot achieve as safely. A useful procurement statement might require a reduction in application errors, shorter processing times, more consistent inspection coverage, or better access to multilingual services. It should also identify who benefits, who may be burdened, what baseline currently exists, and how the city will determine whether the result is meaningful.

Evaluation must combine several kinds of evidence. Technical measures should include error rates, calibration, uptime, latency, robustness to missing data, and performance across relevant languages and demographic groups. Administrative measures should include processing time, abandonment rates, staff workload, appeal rates, reversal rates, and the consistency of decisions. Public-value measures should examine whether residents can obtain timely service, understand decisions, correct inaccurate information, and avoid unreasonable barriers. A system that cuts average processing time by 40% while raising incorrect denials or complaints by 60% has not necessarily improved government.

The framework should require a comparison with credible alternatives. Cities should compare the proposed system with a simpler rules-based process, additional staffing, workflow redesign, shared data infrastructure, and less data-intensive methods. “Human in the loop” is itself a design choice and can hide unworkable review practices. If staff routinely approve outputs in seconds because a manager measures speed, the human safeguard is procedural theater. Contracts should therefore specify how much authority reviewers have, what information they see, how much time they have, how often they override the system, and what organizational incentives reward careful judgment.

Finally, metrics must be segmented enough to reveal harm. Citywide averages can conceal failure concentrated among tenants, applicants with disabilities, residents in particular neighborhoods, or people using a language with less training data. Where appropriate, the city should publish group-level error, delay, denial, appeal, and reversal rates, with privacy safeguards for small populations. Technical sophistication should never substitute for evidence that ordinary residents actually experience fairer and more effective public service.

Govern Data, Vendor Dependence, and Contract Lock-In from the Start

Data governance should be treated as a procurement obligation rather than an appendix negotiated after the product is selected. The solicitation should identify the legal authority for each dataset, its quality and provenance, permitted uses, retention periods, sharing restrictions, and whether the supplier may use municipal data to train a general or commercial model. Cities should prefer vendors that can separate service data from model training, document deletion on request, and support records needed for public accountability. Any transfer to a cloud provider, affiliated company, or overseas facility should be disclosed consistently with local records, privacy, and security law.

The city should also distinguish data provided for a specific service from data inferred, purchased, or created through the system. A procurement framework should address whether individual profiles are portable, how long inferred information remains in the vendor’s environment, and what residents can access or correct. Contract language should cover data breach notification within a fixed period, audit cooperation, subcontractors, disaster recovery, and secure deletion after termination. If these provisions are omitted, the city may discover that reproducing a service, migrating records, or testing an alternative supplier will take longer than developing the original tool.

Vendor dependence requires special attention. Cities often face unequal bargaining power, proprietary model architectures, changing pricing, and few qualified alternatives. Contracts should therefore include a term of no more than three years for higher-risk systems, with options exercised only after measurable performance is met. Initial terms of two or three years can be joined by two one-year extensions, but renewal should require a public record explaining results, unresolved risks, and why reprocurement is not preferable. The city should retain the right to require exportable data, documentation, and interoperability in standard formats, subject to security and legal requirements.

A model’s source code need not always become public, but a city should be able to inspect important system behavior and receive enough information to operate, audit, challenge, or replace the product. “Trade secret” claims should not defeat statutory transparency, public-records, discovery, or auditor access to relevant facts. The stronger the vendor’s control over the system, the more the contract should require portability, transition assistance, and penalties tied to failure to cooperate during migration.

Require Independent Testing, Procurement Pilots, and Community Participation

No vendor-controlled demonstration should serve as the city’s entire evaluation. A short pilot can determine whether a product works in one building or on one dataset, but it cannot establish safety across departments, neighborhoods, languages, and changing conditions. The pilot design should state its duration, sample size, intended users, excluded uses, decision authority, and stop conditions. As a general rule, a Tier 2 deployment should begin with at least 90 days of controlled operation, while a Tier 3 rights-intensive system should normally run for longer and under stronger independent oversight.

Testing should include adversarial, accessibility, security, privacy, and administrative scenarios rather than only historical accuracy on vendor-selected records. The city should test incomplete files, inconsistent addresses, changed legal rules, appeals, rare but serious errors, and interactions with people who lack reliable internet access or digital literacy. It should evaluate whether applicants can contest an error and whether staff disclose that AI was involved. For systems affecting people’s access to essential services, a decision cannot be considered valid merely because a salaried employee clicked “approve.”

Community participation is most useful when it occurs before requirements harden and again before renewal. Residents, disability advocates, tenant organizations, labor representatives, civil-rights groups, small businesses, and frontline workers can identify harms that a technical evaluation misses. Participation should include translated materials, accessible meetings, paid involvement where feasible, and direct feedback from groups historically excluded from the municipal process. The city must explain what evidence it received and how it changed—or did not change—the procurement; consultation without visible influence can create a false impression of consent.

An independent reviewer should have access to source data where lawful, not merely a prepared summary from the vendor. Reviewer independence can be strengthened by prohibiting the supplier from selecting every test case, supplying hidden evaluations, or rewriting an unfavorable conclusion. The final contract should permit a reasonable number of assessment cycles—for example, one validation, one pre-launch review, and one annual audit for higher-risk systems—without forcing the city to authorize unrestricted model updates.

Plan for Workforce Capacity, Accessibility, and Service Access

A municipal AI framework fails if it treats workforce development as optional training added after deployment. Procurement should specify who will operate, supervise, challenge, and maintain the system, and the city should budget for those roles before award. Frontline employees need authority to pause automated processes, access complete case information, explain outputs, and escalate suspected harm. They also need time and tools to exercise judgment; a requirement that every output receive “human review” means little if staffing levels assume the city cannot actually perform that review.

Cities should use the framework to identify skills gaps rather than rely on an abstract promise to “upskill” staff. Data governance, accessibility testing, language access, records management, model evaluation, procurement negotiation, and civil-rights enforcement may require different expertise. Smaller cities may share a regional evaluation office, use a joint procurement, or contract with a university or nonprofit, but responsibility should not be outsourced. Atlanta and other cities can offer useful models for central coordination because local departments often lack the capacity to evaluate sophisticated vendors independently.

Accessibility is part of operational quality, not a separate compliance checkbox. Cities should test interfaces used by residents with disabilities, assistive technologies, older devices, low bandwidth, and limited English proficiency. A system that improves staff efficiency but makes an application inaccessible can increase exclusion. Contract requirements should cover WCAG-oriented digital accessibility, alternative channels for essential services, plain-language notices, and the ability to receive human assistance without losing priority. Vendors should not reduce accessibility testing to a certificate issued for a different product or version.

Workforce data should also be protected. Monitoring employee keystrokes, screen activity, or performance through AI can introduce new surveillance and labor-relations concerns. Cities should limit collection to justified operational data, disclose what is monitored, prohibit undisclosed scoring, and establish review and appeal procedures. The goal is not to replace experienced public employees with opaque management software; it is to use technology in a way that strengthens public expertise and accountability.

Prevent Common Procurement Failures and Require an Exit Plan

The most common failure is beginning with a vendor demonstration instead of a public problem. City officials are then invited to accept whatever data, assumptions, and success metrics the supplier offers. Another frequent error is treating software deployment as an IT project, which causes legal, civil-rights, accessibility, labor, and records staff to enter the process after specifications and budget commitments are already fixed. Cities can avoid these failures by requiring an agency-neutral problem statement and cross-functional review before any solicitation is released.

A second set of mistakes concerns measurement. Technical accuracy can conceal unequal impact; a low complaint rate can reflect residents lacking an accessible appeal route; and faster decisions can result from pushing difficult cases downstream. Cities should not adopt arbitrary targets such as “95% accuracy” without defining the task, error cost, affected population, and test conditions. Nor should they equate vendor-reported performance with independent validation. Targets should be tied to service outcomes and thresholds for pausing use when material failures appear.

The framework must also anticipate system retirement. Every material AI contract should include an exit plan naming the owner, data to be returned or deleted, records to be retained, services to be maintained, residents to be notified, and systems to be switched back to if the AI is withdrawn. The city should not renew merely because automated service has become deeply embedded in operations. Lack of an alternative can be evidence of dependency, not proof that the product remains the best option. A tool that cannot be disabled within a defined period—for example, 30 days for an essential service—should face procurement review before extension.

Exits can be made safer through ordinary continuity planning. Procurement teams should keep validated manual or non-AI procedures available where feasible, test them during pilots, and require vendors to support transition. Sunsetting a harmful system is not a service failure; continuing it because replacement is inconvenient is. The city should report planned retirements and post-deployment reviews so that lessons travel across departments and future solicitations.

Act in 2026 with a 180-Day Foundation and a Two-Year Maturity Path

A city does not need to wait for national legislation, a comprehensive municipal code, or perfect technical expertise before acting. It can designate an accountable official, publish interim definitions and prohibited-use principles, inventory active and planned AI purchases, and require risk tiers in the next competitive solicitation. It should review existing vendor renewals first because a city with hundreds of decentralized tools may gain more public value by governing current purchases than by announcing a new innovation program.

During the first 180 days, the city can establish a cross-functional review group, approve a standard AI procurement addendum, require vendors to disclose system function, data use, subcontractors, and performance claims, and launch one limited pilot with an independent evaluation. It should identify at least three baseline service measures and at least three harm measures before pilot approval. A small-city template may use the same concepts as a large-city process even when external staffing and budget are limited; the city can share legal review, testing, or contract expertise regionally rather than pretend the work is costless.

Over the following two years, the city should move toward annual auditing, public reporting, accessibility standards, staff training, and periodic procurement consolidation. Performance reports should disclose the number of systems by tier, the number of pilots and production deployments, the most common decision types, unresolved complaints, data incidents, contract renewals, and systems retired. By 2028, a mature program should be able to show not only how many AI products it bought, but which public problems improved, which populations experienced additional burdens, and whether another supplier could compete for the work.

The decisive rule is that the city must be able to explain and stop an automated government decision. If no one can identify the responsible official, relevant data, review authority, success measure, appeal route, or termination date, the city is not ready to deploy. This approach may slow some purchases, but it reduces the greater risk of automating illegality, embedding vendor power, and discovering harm only after residents have lost time, money, opportunity, or trust.