Government Offices Great George Street, London, photographed in 2011. Source: Wikimedia Commons; Carlos Delgado, CC BY‑SA 3.0. Landscape display crop.
AI can make a complaint easier to write while making it harder to resolve. A study by Chris Schmitz, Lewis Hammond and Alan Chan documents 84 potential cases of “agentic flooding” across 11 jurisdictions, largely involving people submitting AI‑generated text themselves. Its clearest warning concerns complexity: 76 cases involved more demanding submissions, 50 involved greater volume, and 42 involved both. The August paper cannot establish how much of the change AI caused. The findings, covered by TechCrunch on September 10, suggest a sharper question for public services adopting AI: how much work does each additional submission create? The researchers’ findings separate those pressures.
Subtracting the study’s overlapping categories produces 34 cases involving complexity alone and eight involving volume alone. Half combined both pressures. These are calculations from the paper’s published totals, not estimates of how common AI overload is across government. The distinction matters because a submission counter cannot show whether a caseworker now needs twice as long to understand each case. An agency could receive the same number of applications and still find its staff falling progressively further behind.
Consider a hypothetical team receiving 100 submissions a day. If each requires 20 minutes of work, the incoming workload is about 33 hours. If more elaborate submissions raise the average to 30 minutes, the same inbox now creates 50 hours of work daily. If volume also rises to 120 submissions, the workload reaches 60 hours, an 80% increase from the starting point. These are illustrative numbers, but the arithmetic exposes a practical weakness in counting applications alone: workload depends on both the number of cases and the effort each case requires.
That changes what an effective writing assistant should optimise. A useful complaint makes the relevant dates, evidence, disputed decision and requested remedy easy to locate. Adding a longer argument has value only if it helps establish something the decision‑maker needs to know. For teams building such assistants, a worthwhile test would give a reviewer the underlying documents alongside the generated submission, then measure how quickly the reviewer can verify its claims. The output can sound persuasive and still perform poorly on that test. Evidence that is easy to check is a more useful product than paperwork that merely looks formidable. The paper's original twelve‑service charts also show why measures need careful interpretation: the dashed markers identify ChatGPT's release, without establishing that AI caused the changes around them.