The Direct Answer for Urban AI Purchasers
Urban authorities should not buy artificial intelligence on the strength of a vendor’s model size, demonstration, or promise of efficiency. The contract should define what the system may do, what data it may use, how decisions can be challenged, and what happens when the service fails. For planning departments, housing authorities, transit agencies, and city governments, the most defensible package normally includes data-use restrictions, training rights, audit access, performance metrics, security obligations, vendor-incident notice, model-change controls, human oversight, and a right to terminate or transition the service. These clauses matter because algorithmic errors can affect permits, inspections, service allocation, public safety, and residents’ access to essential services. A well-drafted AI procurement contract does not make the technology risk-free; it makes responsibility visible before deployment. The exact wording should be adapted to the jurisdiction, the sensitivity of the data, and whether the system is advisory, decision-support, or authorized to take operational action.
Also worth reading: How Should Cities Build an AI Procurement Checklist for Urban Planning in 2026? · How should urban planners and municipal leaders develop effective AI procurement guidelines for local government contracts? · What are the standard AI contract audit procedures for municipal planning and urban technology?
The legal position is developing rather than settled. Public debate in 2025 and 2026 has focused on government restrictions involving AI vendors, training-data rights, and controversial procurement practices. The U.S. Court of Appeals for the Federal Circuit rejected Anthropic’s challenge to a Pentagon procurement exclusion, while the Supreme Court was considering related litigation involving Anthropic and President Trump in 2025. Regardless of the outcome, the underlying procurement lesson remains: buyers need precise definitions, especially for terms such as “training,” “authorized use,” “bias,” “safety,” and “material model change.” The GSA’s revised federal AI clause discussions in 2025 also show how large-scale buyers are trying to standardize treatment of AI systems. Urban buyers should learn from that work without assuming that federal language automatically applies to local contracts.
Data Use, Training, and IP Rights Must Be Explicit
The first group of clauses should control what information enters the system and what the provider may do with it. A public buyer should prohibit the vendor from using agency records, prompts, maps, photographs, personally identifiable information, or derived outputs to train a general or customer-specific model unless the contract gives express written permission. It should also state whether de-identified or aggregated data may be retained, whether human reviewers can see prompts and outputs, and whether the vendor may use the buyer’s name, logo, or results in a case study. Silence is dangerous because many cloud AI services are priced on the assumption that the provider can process data to improve its platform. A clause saying “we protect your data” is too vague to answer who owns the information and what activities are excluded.
The contract should distinguish between customer data, public records, licensed data, telemetry, prompts, embeddings, and generated outputs. It should prohibit secondary use for advertising, product development unrelated to the service, model training, benchmarking against other customers, or sale to third parties. Urban planning systems may contain especially sensitive material, including parcel-level data, utility maps, environmental records, mobility traces, and information about protected classes or vulnerable residents. The data schedule should define deletion or return at termination, including backups and derived artifacts, and require certification of deletion. A practical threshold is to require a contractual event notice within 24 hours for a confirmed breach and within 48 hours for a suspected material incident, although the buyer can adjust these periods to national law and operational capacity.
| Feature | Strong buyer position | Weak or ambiguous position |
|---|---|---|
| Training | No provider training on agency data without express written consent | “Data may be used to improve services” |
| Retention | Return or securely delete inputs, logs, prompts, and derived data at exit | “Commercially reasonable retention” only |
| IP | Agency retains rights in records and approved deliverables; limited vendor license stated | “Vendor owns all outputs” without exceptions |
| Deletion | Deadline, backup treatment, and written certificate required | No certificate or maximum deletion period |
| Disclosure | Buyer’s name and results cannot be advertised without permission | Vendor may publish case studies at its discretion |
Performance Metrics and Human Oversight
AI Procurement Contract Clauses should contain measurable service obligations rather than subjective assurances such as “high accuracy.” The agreement should define the task, the population, the ground truth, the evaluation period, the failure cost, and the remedy. A planning department might measure address geocoding accuracy, duplicate-property detection, permit-risk classification precision, or the share of recommendations reviewed by a planner. It should not claim that an overall accuracy percentage proves suitability if the errors are concentrated in older neighborhoods, low-income districts, or residents with limited English proficiency. Metrics should be reported by subgroup where privacy and sample size allow, and the vendor should explain material changes to the benchmark.
The contract should also define the acceptable level of human oversight. A human being must be able to understand the recommendation, access relevant evidence, correct an error, and stop the system from taking action. The system should not treat a person’s unexplained confidence score as a substitute for substantive review. For decisions affecting housing, zoning, enforcement, benefits, or public safety, the agreement should identify the official accountable for the decision, the training required for reviewers, and the maximum percentage of cases that may proceed without review. There is no universal “human in the loop” percentage: a 10% review requirement may be inadequate for detention or emergency decisions but excessive for a low-risk search suggestion. The appropriate threshold depends on consequence, reversibility, and the reliability demonstrated in local testing.
Performance clauses should include audit rights, model and version logs, and a process for investigating adverse outcomes. The vendor should preserve records sufficient to reconstruct a decision, subject to privacy and security limits. It should report downtime, latency, false-positive rates, false-negative rates, override rates, appeal outcomes, and incidents involving inappropriate outputs. A service credit is useful but should not be the only remedy where a serious error continues. The buyer should be able to require correction, suspension, retraining, replacement of affected personnel, or termination without paying an early termination charge.
Security, Resilience, and Operational Dependencies
AI systems create dependencies that ordinary software clauses may not cover. The agreement should identify hosting locations, subprocessors, encryption standards, access controls, logging, vulnerability testing, penetration-test rights, and the vendor’s incident-response process. It should require notice of a material vulnerability within a defined period, such as 24 to 72 hours after confirmation, with earlier notice when a vulnerability is actively exploited. The buyer should be able to inspect independent audit reports, but confidentiality terms must not prevent the buyer from receiving information needed to assess risk. Security language should also address model theft, poisoned data, prompt injection, insecure tool use, and unauthorized retrieval from connected municipal systems.
Resilience is particularly important for urban services. The contract should state recovery-time and recovery-point objectives, expected uptime, planned-maintenance windows, backup practices, and the maximum period in which a manual process must remain available. Vendor-hosted systems can fail because of a cloud outage, a model-provider restriction, a change in API policy, or a commercial dispute. The agreement should therefore require an exit plan, exportable logs, documented configurations, data portability, and reasonable transition assistance. The buyer should specify whether the service may rely on a single large language model provider or a subprocessor outside the approved jurisdiction. Concentration risk should be documented, even if the buyer cannot eliminate it.
The contract should also regulate model updates and configuration changes. A vendor should not silently replace a model with one that has different behavior, training data, safety restrictions, or geographic coverage. “Material change” should be defined by reference to effects, not merely whether the supplier says the change is material. A reasonable trigger could include a new base model, a change affecting accuracy by more than 2 percentage points on a defined test, a new data source, a new subprocessor, or a material change in decision impact. The buyer should receive advance notice and a right to test, reject, or terminate without penalty for an unacceptable change.
Accountability, Redress, and Procurement Exclusions
The strongest clauses assign legal responsibility to the public agency and the vendor in a way that residents can understand. The contract should state that the public authority remains responsible for lawful decisions and cannot transfer its statutory duties to an algorithm. It should prohibit the vendor from making eligibility determinations, enforcement decisions, or final planning approvals unless a lawful governance process explicitly authorizes that use. The agreement should require explanations suitable for affected residents, including the general factors used, the data sources, the human review available, and the route to correction. “Explainable AI” should not mean that the provider supplies a generic statement such as “the model considered multiple inputs.”
A public buyer should also establish remedies for discrimination, privacy violations, and unequal service quality. The contract can require testing for disparate impact, a process for resident complaints, and reports showing whether the system performs differently across neighborhoods or protected groups. The absence of a statistically significant result should not automatically prove fairness, especially when datasets are incomplete or the sample is too small. A neutral acceptance threshold might be a 95% confidence interval, but agencies should define the test in advance and supplement it with error analysis and community feedback. Legal obligations differ across jurisdictions, so counsel must determine what language is enforceable rather than copying a private-sector certification as a substitute for due process.
Procurement exclusions should be narrowly drafted and tied to lawful objectives. As the Pentagon dispute illustrates, a government buyer can face litigation when a supplier challenges a contracting decision or alleged political coercion. The answer is not to ignore the risk or to pretend every exclusion is invalid. It is to use objective criteria, document the technical need, offer a consistent process, and provide notice of the decision. Contract clauses should prohibit retaliation against whistleblowers, employees, and contractors who report safety or legal concerns. They should preserve the buyer’s right to suspend a contract during an investigation without automatically accusing the vendor of wrongdoing.
Comparing Contract Strategies and Alternatives
There is no single clause package that suits every urban authority. A low-risk internal assistant used for drafting memos may justify lighter controls than a system connected to permit records, housing applications, or emergency dispatch. The buyer should also consider whether to purchase a finished service, configure an existing model, develop a system with an integrator, or acquire the underlying technology. Each option creates different intellectual-property, audit, and continuity problems. The table below compares four common approaches; it is a decision aid, not a substitute for procurement or privacy advice.
| Approach | Main advantage | Main weakness | Best fit |
|---|---|---|---|
| Off-the-shelf AI service | Fast deployment and low upfront cost | Weak customization and possible data-use concerns | Low-risk drafting or search with strict settings |
| Enterprise API with private retention | Greater control and easier scaling | Provider still controls model behavior and updates | Document analysis, coding, and service support |
| Dedicated or on-premises system | Stronger operational and data control | Higher cost, maintenance, and specialist staffing | Sensitive records or mission-critical internal use |
| Hybrid system | Balances control, capability, and cost | More complicated contracts and integration | Urban planning offices using sensitive and public data together |
Common Mistakes and Timing of Action
One common mistake is treating procurement as a model selection exercise. The buyer may compare parameter counts or benchmark rankings while failing to examine who supplied the training data, what the system cannot do, and how municipal records are retained. Another mistake is accepting the vendor’s standard terms because the contract appears cheaper or because staff believe AI rules will soon catch up. Standard terms may allocate update risk to the customer, require deletion only of readily identifiable files, or permit use of prompts for service improvement. Another error is drafting a broad autonomous-use mandate and adding a “human review” sentence later. Governance should precede procurement, not follow a public complaint.
Timing matters because contract negotiations are slower than software adoption. A public authority should begin before a vendor approaches it, ideally at least 90 days before a planned pilot and 6 to 12 months before a large enterprise deployment. The schedule will vary, but procurement, privacy review, security assessment, accessibility testing, labor consultation, and public accountability can each delay a launch. If an existing contract is being renewed within 12 months, the authority should send a reservation of rights and request usage logs, subprocessor information, incident history, model-change records, and deletion practices. A renewal should not be automatic merely because the service is already installed.
A second timing error is waiting for a catastrophic incident before creating an appeal mechanism. Residents need notice and a correction route before the system affects permits, inspections, or benefits. The authority should publish the system’s purpose, approved uses, limitations, accountable department, and complaint channel, while avoiding sensitive security details. It should also tell staff when AI recommendations are being used, how to record an override, and when to ignore an output. A system that performs well in a controlled test may still need revised thresholds after seasonal or demographic changes, so monitoring must continue after launch.
Cost, Pricing, and Making the Contract Work
Pricing varies widely because the buyer may pay per user, per API call, per document, per workload, by subscription, or through a custom license. Public contracts should state the total cost of ownership, including data preparation, integration, security review, evaluation, accessibility, model updates, human review, and exit. A low subscription fee can conceal internal labor costs and the expense of correcting errors. The contract should prohibit unapproved usage-based price increases during the term, define overage rates, and set a spending cap. It should also require free export of records and logs in a common format, or state the cost of migration before the buyer loses leverage.
Cost controls should be tied to service quality, not just consumption. A vendor may reduce cost by switching models, shortening retention, limiting review, or routing requests through a different subprocessor. If a change causes performance to fall below a defined threshold, the authority should receive credits or a termination right. Conversely, the authority should avoid unrealistic guarantees, such as requiring perfect predictions, because no responsible vendor can promise that. A useful contract distinguishes ordinary statistical variation from a material service failure and defines the calculation method.
The final review should be interdisciplinary. Procurement counsel should examine enforceability, privacy staff should assess data flows, security officers should test architecture, domain staff should design evaluations, accessibility specialists should examine interfaces, and frontline workers should review the operating procedure. A short negotiation schedule with assigned owners is more useful than an open-ended request for “AI governance documents.” The contract should include a governance meeting at least quarterly during the first year, followed by monthly review for high-impact systems. Success is not the absence of risk; it is the presence of documented ownership, evidence, review, and a workable exit.
What to Require Before Signature
The direct answer is to require a contract that treats AI as a regulated operational supplier rather than an ordinary software feature. It should prohibit unauthorized training and secondary use, define ownership and deletion, provide measurable service levels, preserve human review, regulate model changes, impose security and incident duties, and create a practical route for appeal and termination. The exact thresholds should reflect the system’s consequences. A tool that summarizes non-sensitive public documents can reasonably use lighter terms than one that ranks housing applications, predicts eviction risk, or recommends enforcement actions, but even low-risk deployments should have a data-use restriction and an accountable owner.
Before signature, request the proposed model and subprocessor list, data-flow diagram, retention schedule, training-data statement, security evidence, evaluation results, accessibility report, incident history, and exit plan. Require written confirmation that the city’s data will not be used for provider model training unless a separately negotiated, lawful permission applies. Set a negotiation deadline so procurement does not become indefinite, and preserve the right to obtain local legal advice. The relevant lesson from federal AI clause debates in 2025 and 2026 is not that one model of contracting has won. It is that undefined terms create disputes, and public buyers should define operational boundaries before they become contested after launch.