Direct Answer: Treat Municipal AI Procurement Tiers as Risk Controls
Cities should organize AI purchases into at least four risk tiers: Tier 1 for low-risk tools, Tier 2 for tools handling operational but non-sensitive data, Tier 3 for systems influencing decisions about people or public resources, and Tier 4 for autonomous or legally sensitive uses. The tier should determine the evaluation method, contract controls, security review, pilot duration, approval authority, and post-deployment monitoring—not simply which features a vendor advertises. For example, a tool that drafts an internal meeting summary may fit Tier 1, while software that recommends which housing applications receive inspection priority may require Tier 3 review. A system allowed to make or execute decisions without meaningful human review may belong in Tier 4 and could be unsuitable for deployment even after ordinary procurement review.
Also worth reading: How Should Cities Control AI Purchasing Decisions in Municipal Procurement? · How Do Cities Buy AI Urban Planning Software Without Locking Themselves Into Risky Procurement? · How Should Cities Review AI Vendors Before Using Permit Review Software?
The distinction matters because conventional procurement often compares price, functionality, and vendor credentials, whereas AI can produce uncertain outputs whose behavior depends on prompts, data, user behavior, and model changes. A contract that adequately describes conventional software may not address biased recommendations, confidential inputs, model updates, data retention, generated-content errors, or third-party model providers. Municipal AI Procurement Tiers therefore give officials a repeatable way to match oversight to potential harm. They also help smaller vendors compete by supplying the same core evidence at every tier without undergoing an unnecessarily burdensome process for every low-risk tool.
No universal four-tier model is mandatory across all cities as of October 2, 2026. Each jurisdiction must interpret its constitution, statutes, public-records rules, civil-rights obligations, records schedules, security policies, and procurement thresholds. The framework is best understood as a proposed governance structure rather than a claim that existing cities use one standardized tier system. Its purpose is to produce documented decisions that can withstand audit, public scrutiny, and vendor challenge while preserving practical speed.
How Cities Can Define Four AI Procurement Risk Tiers
Tier 1 should cover narrow, reversible uses with no personal, regulated, privileged, or operational data. Examples include generating internal drafting templates, classifying publicly posted agenda items, or summarizing documents already approved for public release. Officials should require a short data-flow statement, ordinary security review, vendor terms prohibiting unauthorized training, user training, and a 30-to-90-day trial. Human approval should remain routine, and the department should be able to disable the tool without interrupting essential public services. A written inventory record and lightweight exit plan are still necessary because “low risk” does not mean “no risk.”
Tier 2 should cover tools that process internal operational information, moderate-scale records, or data with limited sensitivity. A permit-system assistant that searches internal guidance, or a customer-service tool that drafts responses from non-confidential reference material, may fit this tier. Review should add testing against representative inputs, role-based access controls, logging, breach-notification duties, and contractual limits on secondary use. A 60-to-120-day pilot is commonly more appropriate than an immediate citywide launch, especially where the vendor cannot explain error rates for the city’s language, geography, or population. Approval should rest with the department head and information-security officer rather than require a new procurement action for each prompt-level change.
Tier 3 should apply where AI contributes materially to decisions affecting residents, money, enforcement, access to services, or mission continuity. Examples include permit prioritization, inspection targeting, benefits triage, grant screening, or infrastructure maintenance recommendations. The procurement record should document training-data provenance where known, subgroup performance, human-review design, appeal routes, accessibility, logging, model-change notices, and what happens when the service is unavailable. Independent testing is warranted when decisions could disadvantage residents, while a fixed pilot of at least 120 days can help expose edge cases before full use. Contracts should also preserve procurement remedies for inaccurate outputs, unauthorized disclosure, and failure to correct material defects.
Tier 4 should cover uses involving autonomous action, sensitive personal or legally protected data, real-time safety control, or decisions for which human “review” would be nominal rather than effective. Some jurisdictions may prohibit this category outright unless a specific law authorizes it and a formally accountable official accepts residual risk. Cities should require algorithmic impact assessment, executive approval, cybersecurity and privacy review, legal analysis, accessible challenge mechanisms, contingency operations, and an independent audit after deployment. A 180-to-365-day pilot may be justified for high-impact systems, but time alone does not make an unacceptable use acceptable. Procurement should pause if the city cannot measure errors, override outputs, restore manual service, or explain responsibility for harm.
Why Risk Tiering Is Better Than a Single AI Review Process
A single review process either treats every AI product as ordinary software or treats every product as high-risk. The first approach can miss harms from probabilistic outputs and changing data practices. The second can delay beneficial tools, increase prices, and discourage smaller businesses that cannot afford enterprise compliance programs. Risk tiering offers a middle path: governance intensity rises with the consequence of error, scale, data sensitivity, reversibility, and degree of human control.
The framework also supports procurement teams evaluating competitive vendors on comparable evidence. Rather than asking five cloud platforms five unrelated questions, a city can ask each finalist to demonstrate access controls, retention settings, incident response, output logging, accessibility, model-change notification, and data portability. IBM’s discussion of watsonx Orchestrate illustrates the commercial direction toward platforms that connect procurement and other enterprise workflows, but such orchestration does not remove the city’s duty to assess accuracy and public accountability. Similarly, National League of Cities guidance on generative-AI ethics and governance emphasizes public values, accountability, inclusion, transparency, and privacy rather than technology deployment alone.
Tiering can improve competition without lowering standards. A small firm may lack a large security certification staff but still offer a well-documented Tier 1 tool; demanding the same costs as for an autonomous welfare system would favor only the largest vendors. Conversely, avoiding scrutiny merely because a product is sold as a “copilot” would create a loophole. Procurement language should describe the intended use, decision role, data, affected population, and override capacity—not rely on product labels such as assistant, copilot, analytics, or automation. The intended function and realistic operating practice determine the tier.
Several measurements should be recorded so oversight remains proportional. Cities can track percentage of outputs sampled for review, severity and frequency of errors, subgroup disparities, appeal outcomes, average response time, staff hours saved, service demand, vendor spending, security events, and model or feature changes. A useful rule is that no use should remain in its original tier after a material change in data sensitivity, scale, autonomy, user population, or consequence. For example, a Tier 1 meeting assistant that becomes connected to confidential bid evaluations should be reassessed immediately rather than treated as an ordinary configuration update.
| Feature | Tier 1: Low Risk | Tier 2: Moderate Risk | Tier 3: High Impact | Tier 4: Autonomous or Sensitive |
|---|---|---|---|---|
| Typical use | Drafting and public-information summaries | Internal operational search or customer-service drafting | Recommendations affecting residents, money, or access | Autonomous or real-time high-consequence action |
| Typical data | Public or non-sensitive | Internal or limited operational data | Personal, confidential, or protected records | Highly sensitive data or essential-system control |
| Baseline review | Data statement and security check | Security, privacy terms, testing, and logging | Impact assessment, subgroup testing, appeal design | Legal approval, executive acceptance, independent audit, contingency plan |
| Suggested pilot | 30–90 days | 60–120 days | At least 120 days | 180–365 days, or no deployment |
| Human role | Routine approval | Review material outputs before action | Substantive review with authority to override | Independent control; nominal review is insufficient |
| Reassessment trigger | New data, audience, scale, function, or autonomy | Material change in data or workflow | Expanded population or higher consequence | Any material change; heightened monitoring |
The first implementation step is to create a cross-functional team involving procurement, IT, cybersecurity, privacy, legal counsel, accessibility, records management, labor or human resources, and the responsible program office. Urban planning departments should include planners and frontline service staff because they understand where incomplete data or bad recommendations can affect permits, housing, transportation, capital projects, and public participation. A committee should not decide every purchase by itself; it should establish the tier rules and advise on disputed classifications. Department leaders must retain responsibility for whether a tool meets program needs and whether staff use it as intended.
Next, the city should issue an approved-use inventory and a request-for-proposal or evaluation template. Each entry should identify the vendor, model family if disclosed, purpose, user group, data categories, decision impact, external components, hosting location, retention period, training rights, logging method, human reviewer, appeal route, annual cost, and termination method. Vendors should be told which evidence is required at each tier before bids are scored. This reduces late-stage demands for affidavits and prevents subjective comparisons. It also allows the city to ask whether the same result can be achieved with rules-based software, a conventional database, licensed data, or a smaller approved model.
Pilot contracts should separate experimentation from automatic rollout. Cities can set a 90-day starting point, define baseline service measures, and require monthly review rather than relying on a demonstration conducted with idealized examples. For a planning application-screening tool, baseline measures might include review time, abandonment rate, error correction, staff workload, and disparities by neighborhood or applicant group. For a public-notice drafting tool, measures might include legal review time, accessibility defects, corrections, and staff satisfaction. Savings should be calculated against total operating cost, including subscriptions, usage fees, integration, training, monitoring, records retention, security review, and eventual migration.
Contracts need explicit provisions on city data ownership, prohibition on model training on city inputs unless specifically approved, defined retention and deletion periods, breach notice, subcontractor disclosure, audit access, exportability, model-change notification, accessibility, business continuity, indemnity where lawful, and termination assistance. A fixed term with renewal options is usually safer than an indefinite “pay as you go” arrangement because it creates regular opportunities to reassess need and performance. If a vendor claims it cannot provide model logs, training-data details, or deletion certification, the city should determine whether that limitation is acceptable before disclosing data or paying for integration.
Cost, Pricing, and Contract Structure Considerations
AI procurement costs vary too widely for a defensible citywide price. A department may pay roughly $20 to $200 per user per month for a packaged productivity tool, while departmental workflow products, cloud consumption, and custom systems may cost from several thousand to millions of dollars annually. These figures are planning ranges, not published universal prices; actual contracts depend on users, usage, model type, storage, integrations, support, security requirements, and implementation. More capable models or longer context windows can increase per-request charges, while custom fine-tuning, retrieval infrastructure, evaluation, and human review add costs that are not always visible in the headline subscription fee.
Cities should compare total cost over at least three years and include the cost of manual review. If an AI tool saves an employee two hours per week, the financial benefit should be tested rather than assumed, and saved time must actually be redirected or removed from the process. Conversely, a low-cost tool that creates legal review or appeals may be expensive in practice. Procurement scoring should therefore include expected error-handling cost and opportunity cost, not merely license price. A “free” service funded through usage data or dependent on undocumented API changes may not be inexpensive for a municipality.
Pilot budgets should fund independent evaluation where appropriate. For Tier 1 and many Tier 2 uses, the city might reserve a modest portion of the subscription budget for configuration, training, and a short evaluation. Tier 3 purchases may need dedicated funds for domain testing across languages, neighborhoods, disability-related use cases, and historical error analysis. Tier 4 systems may require redundant infrastructure, manual fallback operations, audit support, and long-term staffing even when the vendor charges only for each transaction. Public contracts may use ceiling prices or usage bands, but unlimited usage without a budget cap creates fiscal exposure; unlimited liability without evidence of insurance or financial capacity creates commercial risk instead.
Small contractors can compete by supplying concise security documentation, transparent interfaces, data-export options, and clear incident procedures. They should not be penalized merely for not offering every enterprise feature, because those features may be unnecessary at a lower tier. However, inability to meet mandatory controls—such as deletion, incident notification, access restrictions, or accessibility—can justify exclusion regardless of company size. The goal is not to make the lowest bidder win; it is to obtain the best defensible service at a transparent price.
Alternatives and Ways to Avoid Overbuying
The principal alternative is conventional rules-based automation. A permit-routing matrix, searchable database, form validator, or statistical dashboard may perform a task more predictably and cheaply than a generative model. Cities should establish when the task does not need linguistic generation before acquiring an AI product. Open-source models, hosted models, and commercial enterprise systems may also be compared, although open-source licensing does not automatically mean lower total cost or stronger security. The relevant question is who operates the system, who bears update and compliance duties, and whether the city can maintain continuity.
Another alternative is to buy a narrower product and add human oversight rather than procure a broad platform intended for every department. A city could use a document-retrieval tool for planning records without allowing it to alter zoning descriptions or generate final staff findings. This reduces integration requirements and narrows the potential impact. However, restricted-use language must be enforceable through permissions and workflow design, not merely a vendor promise that users “should not” rely on outputs for another purpose. Tools that write directly into case-management or permit systems create more risk than read-only tools even when they use the same underlying model.
Some cities may avoid autonomous AI altogether by issuing an internal moratorium on uses that lack adequate authority, data, or review capacity. That can be prudent, but a blanket prohibition can also divert staff to less transparent systems or prevent useful experimentation. A time-limited moratorium of perhaps 6 to 12 months, coupled with an inventory and tier-definition process, can allow the city to learn before making permanent law. New York’s experiments with public-sector technology and broader city responses to AI, as reported by Bloomberg Cities, show why structured experimentation matters, but individual city programs should not be treated as proof of safety elsewhere.
Common Procurement Mistakes and When Cities Should Pause or Act
One common mistake is tiering by vendor reputation. A well-known cloud provider can host an unsafe application, while a small developer can deliver a controlled and testable tool. Another is classifying a product by its marketing name: calling a system a “decision-support” tool does not reduce risk if staff effectively follow its recommendation because managers lack time or independent information. Cities should interview users, examine workflow screenshots, inspect the actual deployment plan, and test how outputs appear in ordinary operations. Shadow recommendations, automatic routing, and “human in the loop” language all require scrutiny.
Another mistake is treating the pilot as production. Pilot data can still contain sensitive information, and real residents may be affected even during a limited test. Agreement periods should be written into the pilot, personal data should be minimized, and escalation and deletion rules should operate from the first day. Cities also err by evaluating only successful demos rather than adversarial, ambiguous, multilingual, outdated, and incomplete records. Where a model has material effects on people, vendors should provide relevant performance measures, while the city should independently verify them where feasible.
A city should pause procurement when it cannot identify the vendor’s downstream processors, cannot secure deletion guarantees, cannot explain how model updates are tested, or cannot operate the service if the vendor exits. It should also pause when baseline human performance is unknown, making improvement impossible to prove, or when the tool could affect emergency decisions without a tested manual fallback. Escalation to a higher tier is warranted when usage expands from 50 to 50,000 users, when internal drafting becomes public-facing advice, or when a recommendation becomes a final administrative action. Thresholds should be set by consequence rather than arbitrary user counts alone.
Conversely, cities should not delay indefinitely when controls are proportionate. A Tier 1 drafting tool with public data, a 60-day trial, routine approval, and no operational impact can often be approved through an existing technology purchasing process. Waiting months for a bespoke legal review may expose staff to less accountable shadow tools. The correct timing question is whether the city knows what it is buying, has tested the intended use, can observe performance, and can stop safely. For higher tiers, the city should allow more time and demand stronger evidence because the cost of silent failure is greater.
A Defensible Adoption Standard
A sound municipal AI procurement system does not promise perfect outputs. It creates conditions in which officials state the purpose, identify risks, measure performance, preserve human authority, protect data, examine unequal effects, and accept or reject the residual risk. Tiering makes those conditions visible to elected officials, employees, residents, auditors, and bidders. It also creates a reason to revisit a purchase when vendors change models, workflows drift, or the community’s expectations change.
By October 2, 2026, cities should have an inventory of active and contemplated AI tools, a written tier definition, and a standard evaluation form. They should be able to answer which tools process personal or confidential data, which influence individual outcomes, which have been tested, which have fallback procedures, and what each costs annually. Within 90 days, a department can pilot a low-risk internal use; within 6 months, it can establish committee review and publish a reporting template; and within 12 months, it can audit higher-risk systems and retire tools that lack clear value. Those timelines are governance targets rather than statutory deadlines.
The best test is whether the city would be comfortable explaining its decision in public using records an auditor can verify. If the answer is yes, the tier need not be named after a particular framework. If the answer is no because nobody knows what data is used, how errors are detected, or who can stop the system, the city should not proceed. Municipal AI Procurement Tiers are therefore not a procurement obstacle; they are a method for buying capable technology without giving up public control.