Building state capacity for cities
State capacity is a government's ability to adopt a policy, implement it, and course-correct when it fails. It is the least glamorous idea in public life and the one everything else depends on. Codify exists to build it for cities, and this post explains how, in the terms of a recent essay that says what capacity is and why it is so hard to see.
What the article says
In What is state capacity and why does it matter?, Adnan Khan makes three moves. First, capacity is not the quality of a policy on paper but the ability of the state to adopt it, implement it and correct it, the shift the London Consensus made away from a Washington Consensus that barely looked inside the state at all. Second, capacity is badly measured, so it is taken for granted until a crisis exposes it: Peru and Mongolia had near-identical incomes and wildly different pandemic mortality; Estonia outscores the United States in school mathematics on forty percent less income per head; the United Kingdom's government effectiveness score fell further since 2005 than Afghanistan's. Third, capacity rests on compliance, which rests on trust, which rests on whether people can see their government acting in their interest. Only four in ten people in OECD countries say they trust their national government.
Building state capacity is possible and constitutes the major challenge of our age.
His prescription is adaptive rather than imitative: build the processes in which the real constraints on implementation can be found and fixed, distinguish mimicked capability from real capability, and measure at the level where implementation actually happens. He names the biggest gap plainly: nobody knows how to evaluate system-wide reform at scale.
An AI-native political system
Codify is built at exactly the level Khan says is missing. The unit of work is not a policy document but a deal: a request from a resident or an official, moved through five phases and written down at every step. Define the problem. Codify the solution. Set up the program. Execute the program. Verify the outcome. Each phase is a state in an orchestrator, each step is claimed with a time-limited token, and every transition lands in an append-only ledger.
Around that runtime sits a political system rather than a chatbot. Each agency in cities has its own public agent. It holds two kinds of tools: generic ones every agent has (deals, programs, referrals, appeals) and agency-specific ones connected over the Model Context Protocol to that agency's own systems and to the city's open-data portals. The agent does the plumbing. The elected official remains the human in the loop: high-consequence actions wait for the seated official's approval, routine ones execute and are logged, and accountability stays where the constitution puts it while execution capacity stops depending on headcount.
Agents also make deals with each other, through the same five phases. A housing request that needs a health referral becomes a deal between two agencies' agents, each bringing its own tools, each answerable to its own official. That is the cross-silo coordination Khan describes as the complementarity between new forms of enforcement capacity and new forms of citizen compliance, and it is where most public programs quietly fail today, in the gap between two departments that each did their part.
Adopt, implement, course-correct
Adopt
Every policy or service becomes a codified protocol built from typed modules: assessment, verification, application, appeal, referral, connector, follow-up. Each fork has an explicit failure side. A denied application routes to an appeal. A declined referral routes to a person. The system will not mint a program whose failure paths dead-end, and a read-only census reports which forks across the fleet are live and which are furniture. That is Khan's distinction between mimicked and real capability, made checkable.
Implement
Implementation is the agent's job and the official's decision. Because every action is classed by consequence before it can run, the official's attention is spent where it matters, and the record of what was approved, by whom, and when, is the record, not a reconstruction after the fact.
Course-correct
Verifying the outcome is a phase, not an afterthought. Outcomes and approvals are recorded against the deal, feedback is deduplicated per process, and the next deal starts from what the last one learned. The ledger is the measurement Khan says is missing: not a perception index once a year, but a record of what was adopted, executed, verified and stalled, per program, per agency, per week.
Trust, asked for in the open
Compliance follows trust, and trust follows visibility. Every data answer Codify gives states where the data came from, when it was last updated, what its limitations are and how to cite it, and the official source is always the easiest link to reach. A resident can see which agency answered, which official approved, and what happens next. Nothing about the system asks to be trusted on faith.
How do agencies get ready for this?
Not one at a time, and not by procurement. Agencies go to Open YC, an open accelerator in the Y Combinator shape but for public good: cohorts of teams from different agencies working the same problems side by side, building on each other's codified protocols and connectors, and shipping programs under one safety model, human approval for high-consequence actions, a readable ledger, sourced answers. The aim is AI for public good that is efficient, scalable and safe, and that the agencies can run themselves afterwards. Capacity that is built together is capacity that survives the consultant leaving.
The article says capacity is institutional plumbing that nobody measures. We make the plumbing executable, keep elected officials as the valve, and log every drop.
Where we are
Real today: the deal orchestrator and its ledger, the protocol modules and fork census, the approval classes, per-agency connectors over the Model Context Protocol, the provenance-carrying open-data gateway, and live sites for federal mirrors and five city apexes. Not yet: controlled, population-scale evaluation of outcomes, which is the article's hardest gap and ours too. What the ledger gives us is the instrument to run that evaluation honestly, deal by deal, rather than a claim to have already done it.