# How Should Cities Buy AI Software Without Creating New Risks in 2026?

urbanplanadvisor.com · September 30, 2026

> The Direct Answer for Municipal AI Procurement A municipal AI procurement guide should treat AI purchases as governed technology acquisitions, not...

## The Direct Answer for Municipal AI Procurement

A municipal AI procurement guide should treat AI purchases as governed technology acquisitions, not experimental software decisions. Cities need a documented process covering business need, data authority, vendor evidence, human oversight, security testing, measurable acceptance criteria, and a practical exit plan before a contract is signed. As of October 1, 2026, the central issue is no longer whether public agencies will encounter AI, but whether their purchasing controls can keep pace with tools that now reach procurement, permitting, public works, schools, inspections, and administrative decisions. Recent examples include New York City schools facing pressure to pause some software purchases while guidance was still being finalized, and Honolulu’s planning office using an AI-assisted tax-style product to help applicants reduce mistakes. These cases show why a general policy must be connected directly to contract review and operational controls.

**Also worth reading:** [How can municipal governments practically integrate AI into urban planning workflows without creating policy chaos or technical debt?](https://urbanplanadvisor.com/knowledge/how_can_municipal_governments_practically_integrate_ai_into_urban_planning_workflows_without_creating_policy_chaos_or_technical_debt.php) · [How Should Cities Review AI Vendors Before Using Permit Review Software?](https://urbanplanadvisor.com/knowledge/how_should_cities_review_ai_vendors_before_using_permit_review_software.php) · [Which AI Planning Software Should Cities Compare in 2026?](https://urbanplanadvisor.com/knowledge/which_ai_planning_software_should_cities_compare_in_2026.php)

The best guide answers four practical questions. What public problem will the system solve, and what would happen if the city simply used a conventional process? What evidence shows that the proposed tool is accurate, secure, and appropriate for the intended decisions? Which official remains accountable when an automated recommendation affects a resident, contractor, or employee? Can the city inspect its operation, transfer its records, challenge vendor assertions, and leave without disruption? A useful threshold is not an exact model size or spending level; it is whether a procurement could materially affect rights, safety, money, or access to public services. Small purchases can meet that threshold when cameras, eligibility, or automated enforcement are involved, while a low-risk internal drafting tool may justify a lighter process.

No single framework fits every city. A municipality with limited technical capacity may begin with restrictions on sensitive uses, a standard intake form, and independent testing, then add a formal review board as volume grows. Larger cities may maintain enterprise agreements, specialist counsel, red-team exercises, and continuous monitoring across dozens of vendors. The correct answer is therefore a tiered municipal AI procurement guide: proportionate enough to remain usable, but strict enough that convenience never replaces public accountability. AI Urban Planner can support planning departments with that framework, but software selection should follow governance and evidence rather than promises about efficiency.

## What Makes AI Procurement Different From Ordinary Software?

Conventional software usually executes rules that administrators can inspect and test. AI systems infer patterns from data or prompts, so performance can change with user wording, neighborhood conditions, scanned-document quality, or new populations. A system may be excellent in a vendor demonstration yet behave differently when residents submit incomplete applications, contractors use unfamiliar formats, or rare cases trigger different outputs. The city must therefore test the actual workflow and representative edge cases rather than relying only on aggregate accuracy reported by the supplier. A claim of “90% accuracy” is not decision-ready without definitions for the dataset, task, population, error types, and human reviewer.

The data environment creates another difference. Municipal records may contain identity data, employee information, legal privileges, utility details, health information, inspection histories, or location records that residents reasonably expect the city to protect. Vendors may also use submitted information to train general models, retain prompts for support, or transfer information across subsidiaries. A city cannot accept a contractual promise that data is “secure” without determining where it is stored, who can access it, how long it is retained, whether it is used for training, and what obligations survive termination. Public bodies should ask whether sensitive records are necessary at all, especially where a lower-risk pilot could use synthetic, sampled, or de-identified data.

AI also changes the allocation of responsibility. Automating a repetitive step does not remove legal or policy responsibility from the agency that authorizes the result. For example, a tool that identifies likely permit errors should not silently reject an application, and an analytics system that ranks vendors should not secretly determine bid eligibility. Human review must be meaningful rather than ceremonial: the reviewer needs time, authority, relevant information, and training to disagree with the system. If city staff routinely approve every output, the “human in the loop” is nominal, not protective. Procurement documents should define which actions require review, how dissent is recorded, and how residents can contest outcomes.

Finally, AI contracts can be harder to exit than traditional licenses. Model updates may alter outputs, vendors may combine tools into enterprise platforms, and essential records may become dependent on proprietary APIs. The city should test export, deletion, continuity, and transition before purchase. A three-month pilot may be sensible, but a month is often too short to observe rare errors, while a year can consume budget without generating decision-grade evidence. The duration should reflect risk, frequency of use, and the time required to validate outcomes—not simply the vendor’s standard subscription term.

## A Risk-Tiered Municipal Review Model

A practical municipal AI procurement guide can use four procurement tiers. The lowest tier covers low-impact tools such as drafting internal memos, summarizing public documents, or suggesting meeting agendas, provided no sensitive personal data and no external decision is involved. A light review can require the department head to confirm the purpose, vendor privacy terms, staff training, and a record of intended use. This keeps routine work from consuming the same legal and technical resources as biometric surveillance or automated benefits decisions.

The middle tier should apply to tools that handle public records, recommend permit or case outcomes, score applications, or assist procurement analysis. Here, a cross-functional panel should include the operating department, procurement, legal counsel, privacy or information-security staff, and a representative familiar with affected residents. It should require a vendor questionnaire, data-flow diagram, performance test, error analysis, contract provisions, and defined human review. The team may permit a limited pilot of 60 to 180 days if the system will remain advisory and cannot directly determine eligibility, enforcement, payment, or safety.

The highest tier covers uses that may substantially affect rights or safety, including facial recognition, predictive policing, automated eligibility denial, consequential employee decisions, critical infrastructure, or intelligence analysis for enforcement. Such purchases should normally require written council authorization where local law demands it, independent testing, a public explanation of the evidence, and a prohibition on facial-recognition or biometric surveillance unless clearly authorized by applicable law. A city should not rely only on a vendor’s assurance that a use case is “responsible”; procurement officials should examine whether the use is necessary, legally authorized, and proportionate to the public objective.

| Feature | Light-Tier Internal Tool | Higher-Risk Operational Tool | Comparison: Why It Matters |
| --- | --- | --- | --- |
| Typical use | Drafting, summarizing, meeting support | Permit review, case triage, supplier analysis | Determines testing and review depth |
| Data | Public or non-sensitive information | Restricted records, contracts, or personal data | Changes privacy and security exposure |
| Decision authority | Staff edits every output | Advisory output may affect timing or routing | Requires meaningful human control |
| Pilot duration | 30–90 days may suffice | 60–180 days with documented edge-case testing | Rare failures need time to surface |
| Approval | Department head or delegated officer | Procurement, legal, IT, privacy, and executive review | Prevents one unit bypassing controls |
| Contract focus | Confidentiality, acceptable use, exit | Security, audit, IP, subcontractor, indemnity, and termination terms | Allocation of risk must match impact |

Tiering is not a substitute for judgment. A low-cost recruiting tool can create discrimination exposure, while a free internal writing tool may expose confidential records. Officials should classify systems by intended use and available data, then reassess when a pilot expands, a model is retrained, or the vendor is acquired by another company. Reclassification should be routine after incidents and at least annually for active high-risk systems, rather than waiting for contract renewal.

## Required Evidence Before a City Signs a Contract

Start with a written problem statement and a measurable public objective. “Improve efficiency” is too vague; “reduce avoidable permit correction cycles by 20% while preserving appeal rights” is testable. The baseline should be captured before deployment, including staff hours, error rates, processing time, resident satisfaction, and the cost of current rework. If reliable baseline data does not exist, the city can use a small manual study rather than assuming that vendor projections are facts. Benefits should include avoided errors and time, but also staff learning, service accessibility, and the cost of oversight.

The evaluation should separate several performance questions. Does the tool identify the intended issue, and how often does it miss it? Does it produce false alarms, and what operational burden do those create? Does performance differ by language, disability, neighborhood, document quality, or type of applicant? Can unauthorized users influence its recommendations? Can a resident understand how the result was reached and challenge it? For generative systems, city evaluators should also test hallucination, citation accuracy, prompt injection, leakage of confidential data, and refusal behavior. An overall score can hide serious failures in small but important groups.

Procurement teams should request evidence proportionate to the claim. A mature vendor may provide independently audited controls, penetration-test summaries, model cards, data lineage, and past deployment results, while disclosing that some details are confidential. Smaller suppliers may lack comparable documentation, which is not automatic rejection but raises the need for testing, contractual protections, financial review, and perhaps a smaller commitment. Independent validation is stronger than a polished demonstration. The city may also purchase a limited right to conduct security testing, but it must ensure that testing does not expose residents, critical systems, or law-enforcement information.

Cost claims need normalization. Compare the total cost over at least three years, including licenses, usage or inference charges, data preparation, integration, security review, staff training, oversight, evaluation, accessibility changes, and eventual migration. A quoted $20,000 annual license can become a six-figure commitment when a city needs cloud credits, consultants, duplicate systems, and specialist review. Conversely, an expensive platform may be justified if it replaces a costlier manual process, but only measured baselines support that conclusion. Vendors should disclose what happens when query volumes, model versions, or user counts change.

## Contract, Security, and Public Accountability Clauses

The contract should identify the city as the decision owner and prohibit undisclosed changes to models, training methods, subprocessors, or data locations. Vendors must provide a current data-flow diagram, role-based access controls, encryption requirements, incident-notification deadlines, audit evidence, and a process for residents’ records requests where applicable. A 24-hour notice requirement may be useful for suspected security events, while notifying the city within 24 hours does not mean the city must notify the public within 24 hours; that separate legal deadline depends on local law and the seriousness of the event. The agreement should not promise an impossible universal guarantee about emerging risks.

The city must decide who may use the system, whether prompts and outputs become public records, and how staff will avoid entering unnecessary confidential details. If a vendor claims it does not retain prompts, the city should verify the relevant product configuration and supporting evidence. Restrictions on training on city-provided data must be explicit, and the vendor should explain whether de-identified data remains subject to contractual controls. Subprocessors should be disclosed, and material changes should require notice or consent. These terms are particularly important when a department buys through a larger platform that bundles analytics, messaging, document generation, and location services.

Performance remedies should be operational. The contract can define minimum service levels, required accuracy on agreed test sets, maximum unacceptable error rates, response times, and credits or termination rights for repeated failure. The city should not set an unrealistic promise that no system will ever err. Instead, it should require prompt correction, notice, retesting, and a safe operating mode. For example, if a permit tool falls below an agreed threshold for correct issue identification, it might route all flagged applications to trained reviewers until the vendor remediates the defect. Repeated failure should create enforceable consequences, including termination without penalty if the vendor caused the problem.

Audit rights should cover the system, not only the contract. Depending on law and risk, the city may need records of model versions, changes, inputs, outputs, reviewer actions, override rates, complaints, incidents, and vendor testing. Public accountability also requires a plain-language explanation of the tool’s purpose, data categories, decision role, limitations, vendor, cost, and appeal route. Internal reports should measure whether the deployment produced its promised public value. If it did not, the city should modify or end it rather than allowing sunk cost to justify continued use.

## Practical Steps for a First-Time City Buyer

A first-time city buyer should form a small working group within the first 30 days and assign one executive sponsor. The group should include procurement, the operational owner, information technology, legal or privacy staff, and a frontline user. It should inventory AI already operating through existing SaaS tools, cloud services, cameras, and vendor pilots, because a city may be using AI without calling it that. The inventory should record the supplier, purpose, data, owner, annual cost, contract date, and whether the use can affect residents. This first step often matters more than selecting a prestigious model.

During days 30 to 60, departments should develop an intake form and classify proposed systems by impact. The form should ask for the public problem, baseline, intended user, decision consequence, data categories, external parties, model provider, hosting location, human oversight, and exit strategy. Procurement should issue model clauses, security minimums, and a standard vendor evidence request. Legal staff should align these documents with records law, public-ethics rules, procurement statutes, civil-rights obligations, labor requirements, and local surveillance restrictions. Templates are useful only when department leaders use them on every purchase, including low-cost renewals.

Between days 60 and 120, the city can run a controlled evaluation with representative, lawfully obtained data. A vendor demonstration may support this stage, but it should not count as independent testing. The test plan should include ordinary cases, edge cases, known failure modes, adversarial inputs, and groups that may be affected differently. For a planning or permitting tool, for example, that could mean varied site plans, languages, historical records, property types, and incomplete submissions. The city should record false positives, false negatives, latency, accessibility, staff workload, and security findings rather than merely asking whether users “liked” the product.

A pilot should begin only after the city knows who can pause it. The operational owner should set a stop condition, such as a serious security event, a material rights impact, or performance below the approved threshold. Staff should receive training on limitations, overreliance, secure use, escalation, and documentation. After 60 to 180 days, the department should issue a factual report and choose among continuing, changing, expanding, or ending the pilot. A negative result is useful when it prevents a harmful purchase. Cities should also publish appropriate public information and retain confidential details where disclosure could create a security or privacy risk.

## Costs, Alternatives, and Common Procurement Mistakes

AI pricing is too variable for a single credible public price. A city should expect to pay for software licensing or usage, implementation, integration, data preparation, evaluation, security, training, governance, and ongoing monitoring. Internal pilots may appear inexpensive, but staff time can be the largest cost. A vendor’s free tier can be appropriate for non-sensitive experimentation, yet it is rarely suitable as the permanent production environment for public records. Annual prices may range from thousands for a narrow departmental tool to hundreds of thousands or more for an enterprise platform, while usage-based systems can produce unpredictable bills if consumption grows without an agreed cap. These are planning ranges, not universal market prices.

Several alternatives deserve comparison. The city can improve forms, staffing, document standards, or workflow automation without AI. It can procure a fixed-rule rules engine, hire temporary staff for a measured backlog, or run a conventional software pilot first. It can also use a shared civic platform, require an open standard, or ask an existing vendor to provide an AI feature under the current contract. Each option has different costs: rules-based tools are easier to test but may not address unstructured documents; hiring can be slow and expose residents to turnover; shared platforms reduce duplication but can create vendor dependence. AI should be selected only when its additional capability is necessary and its risks are controllable.

| Alternative | Best Use | Cost Profile | Main Limitation |
| --- | --- | --- | --- |
| Process or form redesign | Common errors and poor handoffs | Usually modest, mostly staff time | May not solve unstructured-data problems |
| Conventional rules engine | Repeatable eligibility or routing rules | Licensing plus maintenance | Brittle when exceptions dominate |
| Commercial AI pilot | High-volume language, document, or image tasks | Subscription, usage, evaluation, oversight | Variable behavior and vendor dependence |
| Open or civic technology | Shared records and interoperable services | Development and long-term maintenance | Requires technical capacity and funding |
| No purchase | Low-volume or low-value use | Avoids new vendor and privacy exposure | Leaves the existing process unchanged |

Common mistakes begin with buying before defining the problem and with accepting a generic vendor proposal. Officials often treat a demonstration as proof, use an average accuracy figure that hides group differences, or assume human review protects the city when reviewers lack time or authority. They may also overlook public records, accessibility, language access, data retention, subcontractor access, and the cost of reproducing a vendor’s analysis. A pilot without a predeclared stop date can become an unapproved rollout, and a pilot without production exit terms can leave the city with an untested service. Finally, procurement may negotiate price while leaving security, audit, incident response, intellectual property, and termination rights vague.

## When Cities Should Act, Pilot, or Pause

A city should act when it has a defined public need, lawful data, accountable ownership, and enough capacity to evaluate the result. It should generally pilot before production when the system handles sensitive records, assists consequential decisions, uses new data, or has limited municipal evidence. Immediate deployment may be reasonable for a low-impact internal tool if the data is public, the output is reviewed, and the vendor has passed basic security and privacy checks. The relevant question is not “Is AI mature?” but “Is this use mature enough for this city and this population?”

A pause is warranted when staff cannot explain what the system does, the vendor will not disclose data use, the city cannot test representative cases, or no official can stop deployment. Procurement should also pause when guidance is unfinished, as illustrated by the reported call for New York City schools to delay software purchases while AI guidance was being finalized. A pause is not an automatic rejection; it can be a short governance step that resolves missing authority, data terms, or testing. Cities should document the reason, owner, and deadline so the delay does not become indefinite.

The public case for workforce development is also strong. Reports on cities getting AI right commonly connect implementation to training, not simply software acquisition. Frontline staff need to know when not to use a tool, how to verify an output, and how to report harm. A city that purchases first and trains later may increase throughput while quietly transferring risk to employees and residents. Budgets should therefore include at least initial instruction, scenario exercises, accessible user guidance, and refresher training. Training should be role-specific: permit reviewers need different instruction from elected officials or public-records staff.

By October 2026, a sound municipal AI procurement guide should be treated as a living operating manual. Review it after major legal changes, a serious incident, a vendor acquisition, a new use of the tool, or at least once a year for higher-risk systems. The measure of success is not the number of AI contracts signed. It is whether residents receive services more accurately and fairly, public funds are used efficiently, officials can explain and challenge automated recommendations, and the city can change course when evidence shows that the technology is not working as promised. That standard is demanding, but it is more defensible than awarding a contract to the fastest vendor.

## Quick answers

### What is the safest first AI purchase for a city?

A low-impact internal tool with non-sensitive data is usually the safest starting point, such as an assistant that summarizes public meeting materials for staff review. The city should still verify the vendor’s terms, prohibit unnecessary data entry, and require a person to check outputs before they are used externally.

### How long should a municipal AI pilot last?

A 60- to 180-day pilot is common for operational tools, although 30 to 90 days may be enough for a narrow internal experiment. Duration should reflect the frequency of use, risk, and number of edge cases the city needs to test, not simply the vendor’s requested rollout schedule.

### Should cities publish their AI vendor contracts?

At minimum, cities should publish a plain-language account of the system’s purpose, supplier, cost, data categories, decision role, oversight, and complaint route when public disclosure is lawful. Security-sensitive details and privileged material may require redaction, but commercial confidentiality should not automatically justify hiding basic information.

### Can a city rely on human review to make AI safe?

Only if reviewers have enough time, authority, information, and training to disagree with the system. A reviewer who must approve every output for speed is not a meaningful safeguard, especially when the tool handles eligibility, enforcement, safety, or other consequential decisions.

### What should happen if an AI procurement fails its test?

The operational owner should pause or restrict the system, preserve relevant records, and document the failure against the approved thresholds. The city may require vendor remediation, additional testing, compensation, or termination; a failed pilot should be treated as a governance decision rather than a reason to hide the result.

Canonical: https://urbanplanadvisor.com/knowledge/how_should_cities_buy_ai_software_without_creating_new_risks_in_2026.php
Markdown: https://urbanplanadvisor.com/knowledge/how_should_cities_buy_ai_software_without_creating_new_risks_in_2026.php/index.md
