New York City's MyCity program was supposed to create a one-stop way for residents to obtain benefits and services. A December 2025 Comptroller audit concluded that after more than four years and $100 million, the city had “very little else to show” beyond a common platform and digitized childcare application. The audit also examined the program's AI chatbot.

According to the auditors, the chatbot produced inaccurate and inconsistent responses. In city reports covering July and August 2025, 50 of the 70 users who gave thumbs-up or thumbs-down feedback were negative. In an OTI review of 48 questions the chatbot should have answered, it failed to answer 23.

Asked about city services, the city chatbot sometimes apologized that its knowledge was limited to New York City government topics.

Capitalization as infrastructure

The audit team found that small wording changes could produce materially different results. A question phrased one way failed; a rephrased version worked. Even “NYC” versus “nyc” affected responses. Auditors warned that sensitivity to syntax could disproportionately affect people with learning disabilities or people who speak English as a second language.

City technology officials disputed parts of the audit and said weekly accuracy exceeded 95 percent with nearly zero hallucinations. The auditors challenged the calculation and said their independent testing continued to find inconsistency. This disagreement is useful: AI metrics are not self-explanatory. The denominator, excluded categories, test questions, and definition of success matter.

CRR remedy

Public-service chatbots need published test methods, accessibility testing, error logs, regular adversarial evaluation, and a visible route to a human. A beta label does not transform an incorrect benefits answer into harmless experimentation.

The chatbot asks that the next $81 million requested for the program include enough money to teach it where City Hall is.

← Return to all dispatches