A Practical Definition of Municipal AI Risk Tiers

Cities do not yet have a universally adopted classification called Municipal AI Risk Tiers. The term instead describes a practical governance model in which an AI system is assigned a risk tier based on its intended use, the sensitivity of the data it processes, the authority it receives, and the severity of foreseeable harm. As of October 2026, there is no single citywide standard comparable to a building code or globally uniform classification. This means a municipal government must define its own tiers while remaining compatible with privacy law, sector regulation, cybersecurity requirements, procurement rules, and public-sector duties.

Also worth reading: How do urban planners conduct an AI algorithmic impact assessment for municipal systems? · How do automated building permit review systems accelerate municipal housing approvals? · How Should Cities Build a Municipal AI Governance Framework in 2026?

A sound system should not classify an AI product merely by its underlying model. A general-purpose chatbot supporting internal document search and a chatbot connected to a 311 service may use the same foundation model but require different controls because they differ in data access, operational authority, human review, and public impact. Risk should therefore be assessed at the level of the actual municipal use case, including integrations, users, decisions, data, deployment environment, and what happens when the model fails.

The central purpose is not to label AI as inherently safe or dangerous. It is to match governance intensity to potential harm. Lower-tier applications can usually operate with baseline privacy, security, and transparency controls. Higher-tier systems need stronger testing, independent review, restricted permissions, human approval, incident reporting, appeal mechanisms, and sometimes a formal prohibition. The model should also allow a deployment to move to a higher tier after an expansion, new data source, removal of human review, or change in decision authority.

A Four-Tier Model for Municipal AI Use

A workable municipal framework commonly uses four tiers: prohibited or Tier 0, low risk or Tier 1, elevated risk or Tier 2, and critical risk or Tier 3. These labels are recommendations, not an existing statutory scheme. Tier 0 covers uses that should not be authorized because they are unlawful, incompatible with fundamental rights, or difficult to control even with ordinary safeguards. Examples might include fully automated decisions producing permanent deprivation of essential services or uses designed to manipulate vulnerable groups in ways that exceed lawful authority.

Tier 1 covers low-impact uses such as public-information search, translation, meeting transcription, and draft internal summaries when outputs are reviewed before external use. Tier 2 covers systems that influence individual access, eligibility, inspection priority, service routing, or resource allocation without making a final legal decision. Tier 3 covers high-impact systems that can materially affect safety, essential services, civil rights, policing, employment, housing, benefits, healthcare, or large-scale infrastructure. The same vendor model may appear in several tiers depending on its integration and purpose.

FeatureTier 1: Low RiskTier 2: Elevated RiskTier 3: Critical Risk
Typical useTranslation, internal search, draftingCase triage, eligibility support, inspection prioritizationAutomated service denial, safety control, compulsory enforcement
Typical dataPublic or low-sensitivity operational dataPersonal, confidential, or regulated dataHighly sensitive data or control of essential infrastructure
Human controlRoutine review before consequential useHuman approval for each material decisionMandatory authority by law plus independent approval and monitoring
Core testsAccuracy, privacy, prompt-injection screeningBias testing, red-team exercises, appeal designIndependent assurance, continuity testing, security assessment, formal authorization
Review cycleAt least annuallyAt least quarterly and after material changesAt least monthly for high-impact operations plus continuous alerting
Failure responseCorrect and reissue outputPause affected workflow and examine affected peopleContain immediately, preserve evidence, notify responsible authorities, and provide remediation
This table is a starting design rather than a compliance certificate. A low-risk label cannot excuse weak cybersecurity, and a high-risk label is not a substitute for deciding whether deployment is permissible. Public notice, records access, and due process remain relevant even when an AI system performs only an administrative function.

How Risk Should Be Scored

The best tier assignment combines consequence, data sensitivity, autonomy, scale, reversibility, and exposure to manipulation. Consequence measures whether failure can create inconvenience, financial loss, discrimination, physical danger, loss of liberty, or interruption of an essential service. Data sensitivity considers whether the system handles public records, employee information, health details, financial records, precise location, credentials, legal communications, or information about children and other vulnerable residents.

Autonomy determines how much a human can genuinely intervene before harm occurs. A system that recommends a possible answer is different from one that silently closes a case, changes a priority, disconnects a service, or directs enforcement. Scale also matters: an error affecting one visitor is less serious than an error affecting thousands of residents or a citywide service. Reversibility should receive credit only if the action can truly be undone quickly and without disproportionate harm. A theoretically reversible termination decision may still be non-reversible if the person loses housing, income, or essential access while the decision is disputed.

A possible scoring rubric can assign 0 to 5 points for each of six factors, producing a maximum of 30 points. A system scoring 0–7 could enter Tier 1, 8–15 Tier 2, 16–24 Tier 3, and 25–30 Tier 4 or prohibited pending executive review. This numerical approach makes discussions more concrete but does not remove judgment. Any factor involving unlawful automated rights deprivation, access to critical infrastructure, or a substantial risk of physical harm should trigger senior review regardless of the total.

Several 2026 policy developments increase the need for such discipline without proving that every AI deployment requires the same treatment. New York City has shown increasing attention to AI procurement, workforce implications, student restrictions, and agent safety, while Dublin has adopted a responsible-AI strategy and South Korea’s AI Basic Act offers lessons for city governments. These developments point in the same direction: authority and context matter. They do not establish a globally uniform municipal tier system, and cities should not treat an executive order, vendor statement, or policy discussion as a complete risk analysis.

Who Should Assign and Approve Each Tier?

The accountable business owner should initiate the assessment, but that owner should not assign the lowest possible tier without independent challenge. A small AI steering group can classify routine Tier 1 uses. Elevated and critical uses should require review by privacy, cybersecurity, legal, procurement, accessibility, records management, labor, civil-rights, or sector specialists as applicable. The data protection impact assessment is particularly important where processing is likely to create significant effects for residents or where special-category data is involved.

Ownership must remain inside the department using the system. The model vendor, consultant, or central innovation office can supply technical evidence, but it cannot become the final decision-maker for whether residents receive benefits, inspections, protections, or remedies. If a city outsources a service, contractual language should identify the responsible agency, define audit access, require notice of model or data changes, and preserve the city’s ability to suspend the system. A contract saying the vendor supplies a “secure AI platform” is not a risk classification.

Some systems should never be delegated to vendors as governance decisions. A vendor may recommend a classification and controls, while a designated municipal official approves operation. For Tier 3 systems, approval should require a written purpose statement, architecture description, affected-population analysis, data map, test results, residual-risk statement, appeal process, shutdown plan, and expiration date. Authorization should expire after a defined period, such as 12 months, unless renewed. An emergency deployment may need a shorter period or temporary authority, but it should not become permanent through repeated renewals without evidence.

Minimum Controls by Risk Level

Every tier needs baseline controls, including a documented purpose, authorized users, access logging, retention limits, vendor inventory, privacy review, security patching, prompt-injection screening where relevant, employee training, and a way to report harmful output. Municipal documents should record whether the system uses public, confidential, personal, or specially protected information. They should also describe human responsibilities, including who reviews output and who can stop the system.

Tier 2 and Tier 3 deployments require stronger controls. Cities should test performance across relevant demographic groups and language groups, examine false-positive and false-negative rates, conduct adversarial testing, and compare the AI result with an appropriate non-AI baseline. An accuracy percentage alone is insufficient unless the test population and consequences are disclosed. A 95% accurate model that can wrongly deny 500 applicants may be less defensible than a 97% accurate model when the base rate, appeal cost, and alternative review process are properly considered.

Tier 3 systems should have independent validation, segregation of duties, restricted credentials, continuous monitoring, and a tested continuity mode. There should be a designated person or authority that can pause the service, a process for identifying affected decisions during a defined look-back period, and a remedy for residents. Authorities must also consider the “human in the loop” problem: having a nominal reviewer who sees only a confident answer and has too little time to challenge it is not meaningful human oversight.

Comparison between a lighter framework and a formal assurance process clarifies the trade-off. A lightweight model is cheaper and faster for a translation tool used by one department, but it is inadequate for a benefits eligibility engine used across hundreds of thousands of cases. A formal process consumes more staff time and may expose inconsistent systems to greater scrutiny. That cost is part of the system’s real operating expense and should be included before procurement, not discovered after a harmful error.

Governance choiceLightweight internal reviewFormal multi-party assurance
Best fitLow-impact, contained workflowsPersonal-data, civil-rights, safety, or essential-service uses
Indicative staff effort20–60 staff hours per deployment80–300 staff hours before launch
Decision speedDays or several weeksSeveral weeks to several months
AdvantageLow administrative burdenBetter challenge, documentation, and public accountability
LimitationCan miss hidden dependenciesCannot make a prohibited or unsafe use acceptable
Main requirementAccurate scope and annual reviewIndependent evidence, assigned controls, appeal path, and expiry date
## Procurement, Budgeting, and Cost Expectations

AI risk governance has no single market price. The figures vary by system integration, data preparation, model type, hardware, privacy requirements, evaluation depth, and whether the city already has staff and cloud contracts. A contained Tier 1 pilot using an existing approved model might cost roughly $10,000–$50,000 for configuration, testing, training, and a limited period of support. A Tier 2 workflow connected to case-management or permitting data may cost $75,000–$300,000. A Tier 3 system requiring legacy-system integration, high assurance, accessibility testing, and operational resilience can exceed $300,000 and may reach $1 million or more.

These are planning ranges, not official procurement figures. Token consumption, storage, security scanning, observability, and vendor support may be billed separately. The largest recurring costs often involve data cleaning, staff review, record retention, model monitoring, retraining where appropriate, and remediation after errors. A cheap API call may therefore produce an expensive public service if it creates thousands of appeals or staff interventions.

Budgets should state the total cost of ownership over at least the intended pilot and first operational year. Procurement should require a calculation of expected volume, concurrency, data transfer, retention, integration work, user support, and incident response. It should also clarify who owns prompts, logs, embeddings, fine-tuned weights, derived data, and audit records. Public bodies need contractual access to those records when necessary to investigate a decision, and deletion requirements should account for legal holds and unresolved appeals.

Cost pressure must not justify hiding risk behind an unverified vendor label. Certifications or general control frameworks may reduce duplicated work, but they do not prove that a specific municipal use is lawful, accurate, accessible, or fair. The city should examine the exact configuration and current technical evidence. A vendor’s claim that its system is “SOC 2 compliant,” “ISO certified,” or built on a “large language model” is marketing evidence, not a complete tier assignment.

Common Mistakes and Bad Governance Assumptions

A frequent mistake is treating model size as the principal risk metric. A large generative model can be used safely for internal summarization with restricted access, while a small rule-based system can create serious harm if it automatically suspends an essential service. Risk comes from context and consequence. Another mistake is assuming human review automatically solves the problem, particularly when reviewers cannot understand the recommendation, need to process too many cases, or are discouraged from disagreeing with the system.

Cities also make the error of assessing a prototype rather than the production service. New permissions, broader data, automatic execution, changed prompts, model updates, and new vendors can alter the risk after approval. A tiered policy should require reassessment after any material change and at least annually for active systems. Higher-risk services may need quarterly review because drift, security vulnerabilities, and organizational changes occur faster than an annual procurement cycle.

A third error is measuring only model accuracy and ignoring workflow outcomes. An AI system should be evaluated by the number of incorrect decisions, disparities, appeals upheld, service delays, security events, staff overrides, and residents who cannot obtain a remedy. Fourth, cities may announce AI usage without publishing enough meaningful information for public scrutiny. Notices should identify the system’s purpose, accountable department, types of data, role of human review, and complaint channel without disclosing sensitive security details.

Finally, some governments treat a ban as the only safe option or regard unrestricted experimentation as the best route to innovation. Blanket bans can prevent beneficial uses and divert public and academic work into less accountable environments. Unrestricted pilots can expose residents to harm. The better response is proportional governance: prohibit clearly unacceptable uses, pilot lower-risk uses under controlled conditions, require stronger evidence for higher impacts, and stop when controls fail.

When a City Should Pause, Re-Tier, or Act Immediately

A city should pause a system after credible evidence of material harm, widespread inconsistent performance, unauthorized access, manipulated outputs, loss of meaningful human oversight, or a vendor change that invalidates testing. It should also pause when a service depends on a single model provider that cannot meet security, records, or continuity requirements. For Tier 3 systems, containment should begin with the technical operator, followed by preservation of logs and prompt or decision records needed to identify affected people.

Not every error requires the same emergency response. A staff member receiving one poor translation can be corrected within the normal support process. A model that produces many incorrect eligibility notices may require reassessment of every decision made within a defined period. If residents were denied housing, income, healthcare, or another essential service, the city should consider notice, re-evaluation, appeal acceleration, and remediation. Deleting logs to limit reputational damage would be an additional governance failure.

Regulatory deadlines should also trigger action. Although the exact rules depend on the jurisdiction, an organization should inventory obligations early rather than waiting for an enforcement event. Public authorities subject to the EU AI Act may face risk-based obligations that apply differently depending on the system and deployment context. Privacy assessments may be required for certain high-risk processing, and automated decisions involving people can trigger transparency or rights requirements under applicable data-protection law. Municipal legal staff should map those obligations before assigning the internal tier.

The final decision should be time-bounded and evidence-based. A steering group may authorize a pilot for 90 days, require review after the first 30 and 60 days, and prohibit production automation during the trial. It may permit a Tier 2 tool for six months while measuring false positives, disparities, appeal outcomes, staff workload, and incidents. A Tier 3 system should begin with a narrow setting, a small population, restricted permissions, and no automatic adverse decision before any wider release.

Putting a Municipal AI Risk Framework Into Operation

A city can begin by creating a central inventory covering pilots, purchased tools, embedded products, and internally built systems. Hidden AI should be treated as a governance problem because vendors may add automated features without procurement teams recognizing them. The inventory should identify owners, vendors, models, data sources, users, decisions supported, autonomy level, public impact, and current approvals. The city can then apply a common scoring method and target the most consequential 20 percent of systems for deeper review.

Within the first 90 days, a municipality should issue an interim policy, designate accountable officials, define escalation thresholds, and require a basic assessment for every new system. By month six, it should publish the taxonomy, approve a small set of controlled pilots, and establish testing and incident procedures. By month 12, it should complete the highest-risk inventory, publish aggregate transparency information, audit a sample of lower-risk systems, and revise controls based on evidence. This timeline is a governance recommendation rather than a legal deadline.

Success should be judged through concrete indicators. Examples include 100 percent of active systems having an owner, at least 95 percent of new tools receiving a classification before deployment, all Tier 3 systems having tested shutdown procedures, and publication of major incidents within a policy deadline. Numeric targets should not create perverse incentives: reaching 100 percent compliance can encourage superficial reviews, and omitting a system may make the denominator look successful.

The defensible approach is therefore neither “AI first” nor “AI never.” Cities should make lawful purposes and accountable use conditions explicit, assign risk before deployment, and increase scrutiny as autonomy and public consequence grow. The strongest municipal framework will evolve after use, because technology and operations change. Regular review, transparent decisions, and a credible remedy when systems fail matter more than a polished label or an impressive demonstration.