A Practical Answer to the Urban AI Procurement Guide

Cities should procure artificial intelligence as governed operational software, not as an indistinguishable promise of efficiency. For urban planning, the defensible approach is to begin with a specific public-service problem, establish measurable baseline performance, require evidence from comparable deployments, and preserve human authority over decisions that affect zoning, housing, transportation, public safety, utilities, and environmental justice. A useful Urban AI Procurement Guide therefore asks five questions: What decision will the system improve? Who can be harmed by an error? What evidence supports the vendor’s claims? Can the city inspect, challenge, and exit the system? and Who remains accountable when something fails? As of 29 September 2026, no city should need a new AI system merely because a vendor labels it autonomous, agentic, predictive, or next generation. Those descriptions describe architecture, not public value.

Also worth reading: How Should Cities Buy AI Responsibly Without Locking In Costly Vendor or Surveillance Risks? · How Do Cities Build a Responsible AI Planning Workflow in 2026? · How Should Cities Evaluate AI Planning Tools for Safer, Faster Development Review?

Procurement should also distinguish between systems that produce information and systems that make or execute decisions. A model that prioritizes inspection requests can support planners, but a model that automatically recommends which neighborhoods receive inspections can create discriminatory exposure. A traffic-management system that estimates congestion differs from one that changes signal timing in real time. The first has a narrower operational footprint; the second can affect travel times, emergency response, emissions, and transit reliability within minutes. A responsible contract makes that distinction explicit and applies stronger controls to higher-risk uses. This is important because technical performance can conceal social harm: a system may optimize average travel time while worsening reliability for bus riders, pedestrians, or residents near busy intersections.

No universal citywide price exists. Planning-related AI purchases can range from a few thousand dollars for a narrowly scoped analytical pilot to tens of thousands for a commercial tool with implementation and security review, and from roughly $100,000 to several million dollars for a production platform requiring data integration, model monitoring, audit access, and ongoing public oversight. Subscription prices are only a small part of total cost. Cities should require a five-year total-cost disclosure covering acquisition, data preparation, integration, cybersecurity, legal review, model validation, staff training, vendor support, infrastructure, audits, contract exit, and records retention. A low bid is not economical if the city cannot operate it or if staff must continue making decisions without a functioning explanation.

Define the Public Problem Before Choosing AI

The first procurement document should be a service-needs statement, not a technical specification. It should identify the accountable office, affected residents, current process, known failure, service target, legal authority, and reason conventional tools are inadequate. For example, “improve urban planning” is not procurable. “Reduce the median review time for low-risk sidewalk permits while preserving the 95th-percentile correction rate and collecting data on disparate service impacts” is testable. Likewise, “predict flooding” should be replaced with a defined question about which alerts residents or emergency managers need, at what spatial resolution, how much lead time is operationally useful, and which critical facilities require priority.

Cities should compare AI with less novel alternatives before accepting it. These can include better data management, added staff, process redesign, rule-based optimization, conventional forecasting, shared data platforms, or a procurement of the underlying data rather than a predictive model. A mature vendor should welcome this comparison because it can expose whether AI is genuinely suited to the task. Vendors that cannot explain why a deterministic workflow fails to meet the need are likely to overstate their system’s necessity. This approach is consistent with the Federation of American Scientists’ emphasis on AI procurement guardrails: technology purchases in public institutions need safety criteria, evidence thresholds, and avenues for students and other vulnerable users to challenge harm.

Before a pilot, officials should set a baseline. Depending on the use case, that baseline may include 12 to 24 months of service data, documented case-processing times, false-positive and false-negative rates, geographic coverage, resident complaints, appeal outcomes, or implementation costs. If reliable baseline data do not exist, the city should improve measurement first. A pilot without a counterfactual is an expensive demonstration. A credible test should use a comparison site, a time-based split, or another design capable of separating model performance from seasonal conditions, policy changes, and unrelated investments. The procurement should require a minimum accuracy threshold relevant to the purpose, but accuracy alone is not sufficient; it must be reported across neighborhoods and relevant demographic groups wherever lawful and appropriate.

Match the Tool to the Decision and Its Risk

Not every planning problem needs machine learning, generative AI, or an agentic system. Some tasks are better handled by a geographic information system, optimization model, rules engine, or ordinary statistical model. Generative systems can draft planning narratives or summarize public comments, but their generated statements still require source verification and human review. Predictive models can estimate demand or identify potential hazards, but their outputs should be treated as decision support unless the city has deliberately authorized a limited automated action. An agent that files routine documents or schedules inspections introduces different risks from a forecasting model because it can initiate actions, interact with other software, and propagate errors across systems.

A four-level risk classification is practical. Level 1 covers internal, low-impact search or visualization; Level 2 covers recommendations that experienced staff may independently verify; Level 3 covers operational recommendations that can materially affect residents; and Level 4 covers automated decisions involving legal rights, safety-critical actions, or direct access to sensitive personal data. Higher levels should trigger privacy review, an impact assessment, independent testing, public documentation, appeal procedures, and stricter limits on vendor access. The classification should be based on actual functionality, including integrations and actions, rather than on the vendor’s product label.

The table below compares three broad procurement options. It is not a substitute for legal and technical review, because a product’s features and local data practices can change the risk assessment. Its purpose is to make the city’s choices explicit before a sales demonstration establishes momentum.

FeatureBuy a narrow off-the-shelf toolCommission a pilot and managed serviceBuild an owned planning capability
Best fitStandard visualization, document search, or limited analyticsA defined forecasting, inspection, or service-delivery problemRepeated strategic needs requiring strong local control
Typical first-year costAbout $5,000-$50,000About $50,000-$500,000+About $250,000 to several million dollars
Time to useful test1-3 months3-9 months12-36 months
Control of methods and parametersUsually limitedContractual and technical access should be negotiableMaximum operational control
Major riskHidden limitations and weak customizationDependence on vendor road map and data accessScarce staff, maintenance burden, and model drift
Exit conditionExport data and reproducible reports at contract endTransferable models, prompts, integrations, and audit recordsMaintain staff capacity and documented recovery paths
## Test Vendors With Public-Value Evidence

A request for proposals should test whether vendors can support their performance claims in the city’s context. Vendors should submit documented deployment results, sample contracts, security materials, model or system documentation, data-flow diagrams, incident history, and information about human review. References should be checked independently. A claim based on average accuracy across a national customer base does not establish performance under the city’s housing patterns, climate, language mix, infrastructure, or enforcement practices. Nor should a city accept only vendor-produced benchmarks; the contract should reserve the right to test a representative workload using locally approved data.

Procurement teams should ask whether performance degrades under changed conditions. A flood model trained on historical events may fail when cloud-cover patterns, drainage configurations, sensor failures, or extreme rainfall exceed the training distribution. A housing-demand tool may reproduce historical discrimination if past allocation data reflect unequal access. An inspection-prioritization model may improve average case resolution while directing intensive enforcement toward historically over-policed communities. The Federation of American Scientists and the Nature discussion identified in the research context both point toward procurement practices and measurement that must account for social consequences, not just technical sophistication.

The evaluation should also examine distributional effects. Cities should compare error rates, service times, benefits, burdens, and appeal reversals across relevant geographic and demographic groups. There is no ethical or legal basis for blindly ignoring a materially disproportionate outcome, but group-level analysis also has privacy and practical limits. Results should be reported with sample sizes, confidence intervals, missing-data rates, and uncertainty. A percentage based on 12 cases is not equivalent to one based on 12,000 cases, and a perfect score in a small test may be meaningless. Contracts should identify minimum data-quality thresholds, such as completeness, recency, geographic coverage, and label reliability, and state what happens when those thresholds are not met.

The scoring formula should give public value more weight than product polish. A defensible weighting may assign 25% to service effectiveness, 20% to equity and civil-rights risk, 15% to privacy and cybersecurity, 10% to transparency and auditability, 10% to interoperability and data portability, 10% to workforce and operational capacity, and 10% to total cost. Security may warrant a separate pass-or-fail gate rather than being diluted by a high design score. The exact percentages should reflect the use case, but vendors should not receive the largest evaluation shares for a visually impressive interface or use of proprietary technology.

Write Contractual Guardrails and Exit Rights

A procurement contract should treat monitoring as a continuing obligation. It should define prohibited uses, permitted data categories, retention and deletion schedules, role-based access, encryption, security-incident deadlines, subcontractor controls, restrictions on secondary use, and the city’s right to obtain audit evidence. It should also prohibit material model or product changes without notice, testing, and written acceptance. The city needs to know whether generated text was produced by retrieval from public records, entered by staff, supplied by residents, or inferred from sensitive data. Every relevant field should be distinguishable in an audit record.

Human review must be real rather than ceremonial. Reviewers need authority to reject a recommendation, sufficient time to examine supporting evidence, training that explains likely failure modes, and staffing that does not reward rubber-stamping system outputs. The contract should report agreement between reviewers and the system, overrides, reasons for overrides, and outcomes after overrides. If operators reject more than 50% of recommendations during the first three months, the city should pause expansion and determine whether the system is unsuitable, poorly integrated, or assigned to the wrong staff. If automated queues steadily grow because city staff cannot keep up, the vendor’s promise of efficiency has not been realized.

Exit rights deserve equal attention. At termination, the vendor should deliver the city’s data in a documented, commonly usable format and provide required documents, configurations, audit logs, and model artifacts consistent with intellectual-property and public-records obligations. A reasonable transition assistance period may be 90 to 180 days. Contracts should prohibit hostage pricing for data extraction or essential integration, and should state how the city can continue a critical service if the vendor fails. Although portability can be difficult for some proprietary services, the city should not accept a claim that its own operational data and required records are permanently inaccessible.

Implementation, Workforce Capacity, and Accountability

The winning bidder is not the final decision-maker. The city must assign an accountable program owner, technical owner, privacy or records contact, civil-rights reviewer, and frontline users. Those roles should be named in the implementation plan. A steering committee should meet at defined intervals—monthly during a pilot and quarterly during stable production, for example—and its minutes should record risks, incidents, performance changes, corrective actions, and decisions to scale, modify, or stop. For high-impact systems, an independent evaluator or city auditor should have access to relevant evidence.

Training is a budget line, not an optional launch activity. Planners should learn the limits of the data, how uncertainty is represented, how bias can enter a workflow, and how to document overrides. Procurement staff need to monitor subscriptions, data clauses, service levels, and renewal dates. IT staff need integration and incident-response skills. Communications teams need a plan for explaining defects or errors without concealing them. Residents and affected workers need a usable notice explaining when AI is involved, what information is used, how decisions are checked, and how to request human review or appeal an outcome.

The city should run a time-limited pilot before a broad rollout. A common duration is 8 to 16 weeks, with 3 to 6 months reserved for data preparation, security review, and baseline measurement. Expansion should occur in stages, such as one district or workflow before citywide deployment. Success should depend on predefined service, equity, reliability, and cost measures. Fiscal teams should not confuse a higher contract ceiling with economic benefit; the city should estimate avoided labor, faster service, fewer errors, reduced harm, and total operating expenses. Research cited by the Center on Data Innovation also emphasizes workforce upskilling, because institutional capacity determines whether cities can use AI effectively rather than merely purchase it.

Alternatives, Common Mistakes, and When to Act

The most common mistake is procurement driven by novelty. A vendor demonstration may show instant drafting, automated maps, or a persuasive chatbot, while the operational requirement remains untested. Another error is choosing a model before defining who can override it. Others include buying before cleaning data, treating pilot accuracy as a permanent fact, expanding automatically after a short test, comparing vendor results on incompatible benchmarks, and ignoring labor consequences. Staff may become measured by the system without receiving authority to question it. Residents may have no practical route to contest an adverse result. A technically functional deployment can therefore remain a public failure.

Cities should also resist the idea that public officials must independently build every model. A managed service can be appropriate when the need is bounded and the market has mature providers. Building is more defensible when the capability is central to repeated public duties, existing off-the-shelf tools cannot meet essential requirements, or procurement must retain stronger control over sensitive data and decisions. A third alternative is a cooperative arrangement with a university, nonprofit, neighboring city, or regional public agency, although the public authority must still define accountability, security, procurement rights, and records obligations. The San Jose nonprofit model mentioned in the research context illustrates institutional experimentation, but it would still need clear public-purpose, funding, governance, and liability arrangements.

Immediate action is appropriate when a documented service problem is recurring, current data support measurement, stakeholders agree on success criteria, and a human-controlled pilot can test whether AI outperforms simpler options. Waiting is usually wiser when the city lacks data, staff, or legal authority; when the system would make high-impact decisions without meaningful review; when expected benefits are too small to justify the lifecycle cost; or when another agency can provide the capability. A useful threshold is not “How advanced is the model?” but “How much verified public value is expected relative to risk and total cost?” Cities should set a 60% minimum confidence for a small operational pilot where feasible, require 95% availability only where an outage has a documented service-level need, and demand materially better-than-baseline performance before expansion.

A Durable Citywide Standard

A defensible urban AI procurement program should end with a repeatable governance record rather than simply a signed software contract. For each deployment, the city should retain a problem statement, risk classification, data inventory, vendor evidence, evaluation plan, equity analysis, security assessment, contract controls, staff roles, incident log, appeal mechanism, and exit plan. Managers should review the program at least annually, while more frequently reviewing systems exposed to material model changes, new data sources, or high-impact decisions. Budget planning should reserve renewal, monitoring, and decommissioning costs before a system becomes embedded in daily work.

The final question is not whether AI can support urban planning. It can already assist with supplier evaluation, flood detection, wildfire detection, permitting triage, infrastructure search, and analysis of complex spatial data. The harder question is whether a particular purchase improves a defined public service enough to justify its costs, risks, and institutional burden. Cities that answer with evidence, enforceable procurement terms, affected-community scrutiny, and a credible off-ramp can adopt AI without surrendering public accountability. Cities that answer with a sales pitch cannot.

The practical takeaway is straightforward: define the service, classify the risk, establish a baseline, test alternatives, require local evidence, contract for monitoring and exit, and scale only after measurable benefit. That process may produce fewer pilots than a technology-first strategy, but it is more likely to produce durable systems that planners and residents can trust, use, and correct.