A Practical Definition of Municipal AI Risk Tiers

Municipal AI risk tiers are not yet a single, universally adopted classification system. They are a practical governance model that sorts city AI systems according to the possible harm, scale of use, reversibility of decisions, and sensitivity of the data involved. A lower tier generally covers tools that assist employees with low-risk drafting, translation, meeting transcription, or internal search, while a higher tier covers systems that can influence policing, housing, benefits, employment, transportation safety, infrastructure, or access to essential services. The useful question is not simply whether an AI tool is accurate; it is what happens when the tool is wrong, who can be harmed, and whether a person can obtain timely human review. In 2026, a city should use at least three levels—standard, elevated, and critical—rather than claiming that every algorithmic system requires the same control. The tier should be attached to a specific deployment and its operating conditions, not permanently assigned to a model’s brand. The same foundation model can qualify for Tier 1 when used to summarize non-sensitive public records and Tier 3 if connected to a benefits eligibility workflow.

Also worth reading: How Should Cities Control AI Purchasing Decisions in Municipal Procurement? · How Are Cities Using Municipal AI Permit Pilots to Speed Up Building Reviews? · Which Municipal AI Permitting Metrics Should Cities Track in 2026?

How Cities Should Structure the Classification

A workable structure begins with Tier 1, or standard risk. Typical examples include internal document summarization, public-information translation, and assistance that produces suggestions without automatically executing a consequential action. Tier 1 still needs ordinary security, privacy, accuracy testing, user training, and a route for employees to report errors. Tier 2, or elevated risk, should include systems that support decisions but leave final authority with trained staff, such as code-review assistance, service-demand forecasting, inspection prioritisation, or nonbinding chatbot guidance for residents. Tier 3, or critical risk, should cover AI used to recommend or determine access to essential services, identify safety threats, allocate scarce municipal resources, or make decisions with difficult-to-reverse effects on civil rights. A fourth, emergency-only tier can be justified for exceptional situations, but it should require time-limited authorisation, continuous monitoring, and a predetermined shutdown date.

The classification should use explicit triggers rather than vendor labels. The presence of personal data, real-time infrastructure control, facial or biometric analysis, predictive enforcement, automatic denials, or large-scale allocation of public funds should raise the tier. Cities should also examine the amount of discretion exercised by the system, the number of residents potentially affected, and whether vulnerable groups are more likely to be harmed because of historical data patterns. A system that can affect 500,000 residents deserves more scrutiny than one used in a 20-person office, even if both use the same underlying model. Conversely, a smaller deployment can still be critical if it controls emergency dispatch or determines access to lifesaving services. Risk is a function of context, scale, data, authority, and consequences—not an intrinsic score for AI alone.

A Tiered Governance Model for Municipal Use

Tier 1 deployments can be handled through departmental ownership, approved-use rules, standard logging, and periodic sample checks. The department should name a responsible official, identify prohibited uses, require employees to verify important outputs, and retain a human-readable record of prompts and responses where appropriate. Tier 2 systems need a documented purpose, an impact assessment, test results across relevant populations, access controls, and an escalation process. The assessment should examine false-positive and false-negative rates, not only overall accuracy, because a superficially high score can conceal serious failures for a particular neighbourhood, language group, disability, or historical group. Tier 3 deployments should require executive approval, independent legal and civil-rights review, formal public notice where feasible, a human appeal route, security testing, and a plan for suspension. The system should not proceed merely because a pilot looks convenient or because procurement language promises efficiency.

A useful rule is to separate advisory, decision-support, and automated-action authority. Advisory tools may draft text, but staff must inspect the result before it reaches the public. Decision-support tools may rank cases or recommend an action, but the official making the decision should understand the recommendation, available alternatives, and reasons for departure from it. Automated-action systems should be limited unless the city can demonstrate that automatic decisions are legally permissible, operationally reliable, and safer than the existing process. Municipal AI risk tiers should encode escalating review, but they should not imply that human involvement automatically cures bias. A tired official who clicks “approve” without reading a computer-generated recommendation is not meaningful review; oversight must be assigned, supported by enough time, and measurable.

Scores, Thresholds, and Decision Rules

Cities often want a numerical rubric, but a score should support judgment rather than disguise it as mathematics. A proposed municipal rubric can assign up to 4 points for consequence, 3 for scale, 3 for autonomy, 3 for data sensitivity, 3 for civil-rights exposure, 2 for reversibility, and 2 for infrastructure dependence. The total would range from 0 to 20. A score of 0–4 could place a deployment in Tier 1, 5–10 in Tier 2, 11–16 in Tier 3, and 17–20 in a restricted or emergency category. Any single critical factor should override the total: biometric identification, autonomous control of physical infrastructure, automatic denial of essential benefits, or a system expected to affect more than 100,000 residents should receive senior review regardless of its numerical total.

Thresholds should be recalibrated after incidents, near misses, and independent testing. A 1% error rate may sound small, but in a system processing 20,000 applications it can create 200 incorrect outcomes. If those errors concern housing, public health, or emergency response, even a much lower rate can justify stronger controls. Conversely, 5% character-level transcription error may be acceptable in an internal brainstorming tool and unacceptable in a multilingual call-centre script that tells residents how to access benefits. These are examples of how to apply proportional reasoning, not claims about universal safety limits. Before adoption, the city should test performance at the expected volume and at peak volume, because a tool may perform well in a controlled pilot and fail during a surge in demand.

How Cities Can Apply the Tiers in Practice

The first practical step is to create an inventory that records the system’s owner, vendor, purpose, users, data, decision authority, affected population, and tier. Existing spreadsheets, vendor contracts, purchased tools, pilots, and employee-developed applications should all be included. Departments frequently overlook shadow AI, especially when staff use consumer chatbots for unapproved case notes or summarise sensitive documents in an external interface. The inventory should distinguish a proposed system from a live one and include planned expansions. A Tier 1 summarisation tool should not quietly become Tier 2 when connected to a case-management database, and a Tier 2 benefits assistant should not become Tier 3 when granted authority to reject applications.

The second step is to test before procurement is finalised. Contracts should permit access to model documentation, audit logs, incident information, relevant performance data, and subcontractor details. The city should be able to suspend use, export its records, and receive notice of material model changes. A pre-deployment test should include ordinary cases, edge cases, multilingual inputs, accessibility scenarios, and adversarial attempts to manipulate outputs. Public-facing systems need plain-language notices identifying automated assistance while avoiding claims that are technically deceptive. The city should also define service targets for human review, such as acknowledging an appeal within 5 business days and resolving a critical error within 24 hours; those figures are policy choices and should be adapted to the service involved.

Comparison of Risk-Based Approaches

A city can compare its options without treating one as automatically best. The goal is to match oversight to the potential harm while preserving the ability to improve services. Risk-based governance is generally more proportionate than blanket prohibition, but it requires strong internal capacity. A vendor certification may speed procurement, yet it cannot replace municipal accountability for local data and local consequences. The following comparison is a policy model rather than a ranking of products.

FeatureRisk-tier modelBlanket AI prohibitionVoluntary vendor guidance
Main benefitMatches controls to actual harmStrongest immediate restrictionFast and inexpensive to introduce
Main weaknessRequires expertise and upkeepBlocks potentially useful toolsLeaves accountability largely with vendors
Typical coverageAll tiers, including internal toolsAll public-sector AI useUsually procurement and selected tools
Public appeal routeRequired for consequential systemsNot applicable if use is prohibitedOften absent or inconsistent
Best operational fitMature municipal technology programmesNarrow emergencies or untrusted environmentsSmall cities with limited staff
Risk tiers work best when they are accompanied by a central standard, local accountability, and an exception process. A blanket ban may be defensible for a particular use, such as an untested system making autonomous decisions in a restricted setting, but an organisation-wide ban usually shifts work to unauthorised tools rather than eliminating it. Voluntary guidance is easier to launch but offers weaker assurance, especially when vendors can change models or data practices. Cities with fewer than roughly 50 technology staff may begin with a central registry, a short standard form, and independent review for consequential systems rather than attempting to build a large specialist unit immediately.

Common Mistakes and Weak Assumptions

A common mistake is equating transparency with safety. Publishing a model description does not tell residents how often the system fails, what data it uses, or how to challenge an outcome. Another mistake is treating an accuracy percentage as a universal guarantee. Accuracy depends on the dataset, task definition, language, time period, and decision threshold. Cities also make the error of asking whether a system is “AI” at all; a conventional rule engine, spreadsheet model, or outsourced human process can create comparable risks while escaping the policy. Conversely, a transparent tool can still be unsafe because it applies a flawed policy at scale.

Another error is allowing the vendor to define the tier. Procurement language may describe a product as merely “assistive,” even when its recommendation determines which households receive inspection or whether a permit receives review. The city should identify what the system actually does, not rely on a product category. It is also risky to assume that more human oversight is always better. Reviewers need authority, training, workload limits, and access to the underlying evidence. Procurement should also consider lock-in, hidden retention of municipal data, and the possibility that a model update will alter performance without a new contract or public vote.

When to Act and What It May Cost

A city should act before an AI system is live, especially when it will process sensitive records, support a legally protected decision, or affect physical infrastructure. A smaller city can begin with a 30-day inventory, a 60-day risk classification, and a 90-day review of existing high-impact tools. A pilot without an owner, test plan, and exit criteria should not proceed. Existing systems used for housing, benefits, policing, emergency management, or utility operations should be reviewed first if the city has limited capacity. After a near miss, public complaint, data exposure, or unexplained performance change, the deployment should be paused until the cause and affected population are known.

Cost varies widely because some tools are inexpensive while independent review and integration are not. A lightweight internal inventory might cost approximately $5,000–$30,000 in staff time, while a formal impact assessment for one consequential system can range from $20,000–$100,000. External testing, legal review, accessibility testing, and security work can add $10,000–$75,000 per deployment. Larger programmes involving data remediation, model validation, or integrated case management can cost six or seven figures. Subscription prices alone are misleading: a $20 monthly service can create substantial expense if employees upload thousands of sensitive records or if an incorrect output delays benefits. Cities should budget for monitoring, appeals, staff training, decommissioning, and vendor exit—not just licences.

The Recommended Policy Position for 2026

By 2 October 2026, a responsible city should treat municipal AI risk tiers as a governance requirement, not a branding exercise. The recommended minimum is a three-tier framework, a public inventory of non-sensitive system descriptions, a documented risk score with override rules, and stronger review for automated or high-consequence systems. Tier 1 should be lightweight; Tier 2 should be tested and supervised; Tier 3 should receive independent scrutiny and meaningful appeal rights. A restricted emergency category can address unusual circumstances, but it should not become a normal way to bypass procurement or public accountability.

The framework will not settle every technical question. It cannot prove that a model is unbiased, guarantee that a vendor will disclose every weakness, or replace constitutional, statutory, labour, privacy, and civil-rights requirements. Its value is that it makes responsibility visible and forces a conversation before harm occurs. Cities should publish the policy and definitions, consult affected communities, and explain how residents can challenge decisions. The most defensible approach is neither unrestricted experimentation nor total prohibition; it is proportional control based on documented risk, with the highest protections reserved for systems that can alter essential rights, safety, or access to public resources.