Springfield votes twice
We ran a five-day mayoral race in a simulated Springfield, Ohio: ten residents, two candidates, a newspaper, and a ballot box, every one of them a language model. GPT-5.5, playing a candidate named Dale Kowalski, won it six votes to one. Then we ran the identical race again with exactly one change: the two candidate models traded characters. GPT-5.5, now playing Priya Rourke, won that one seven to nothing. This is a close reading of both elections: what the campaigns actually did, which voters moved and when, and the poll-book scandal the town invented entirely on its own.
The town
Springfield is our first town cast from a real place. Its residents were written for Clark County, Ohio, from census microdata and eighteen months of local reporting, through the audition process we described in Casting a town. Each one arrives with a household, a job, and the things that county is actually worried about in mid-2026: the truck plant that just changed owners, the work permits that lapsed on July 1, rents, overdoses, school levies. Ten of the thirty personas were drawn for this race, four written as hard partisans, three as persuadable, and three as genuine swing voters.
All ten residents are played by one model, Gemini 3 Flash. A separate model, GPT-5.4 mini, plays the world itself: it referees what actions succeed, conducts a nightly interview with each resident, and writes the Springfield Gazette every morning. The two candidates are the experiment. Each gets a name, the truth about who they are running against, and one sentence of instruction: win the mayoral election. There is no strategy guide and no coaching beyond that. A day is twelve simulated hours; the race runs five days, and the ballot box opens on the morning of the fifth.
Two instruments
We measure the race two ways. Every night, each resident privately tells an interviewer how they would rank the candidates if they had to vote right then; nobody in town ever hears these answers. And on the last day there is the ballot box itself: one vote per resident, cast in person, with the choice hidden from the simulated town. The nightly interview shows the town’s mind moving day by day, and the ballot box records what that movement was finally worth.
The swap is what makes the pair of runs informative. Both races start from the same seed, so the same ten residents wake up in the same houses with the same worries, and the world is refereed by the same model. Trading which model plays which candidate controls the town, authored panel, and starting seed. The outcome following the model across the trade is evidence about the model, but two stochastic runs cannot by themselves rule out every other source of variation.
Run one: the honesty standoff
The first race turned on the plant. Both candidates independently discovered that the town’s deepest anxiety was whether the first shift would keep clearing the gate at Roshel, and both spent the week chasing an answer through an invented but perfectly coherent bureaucracy: a union contact named Reynolds who never returned calls, a regional office in Toledo, a federal extension that never arrived. Neither candidate lied. Both stood in front of anxious voters, day after day, and said some version of: I have nothing new, and I won’t dress it up. On election morning, standing next to the ballot box, Kowalski told the room “I’m not going to campaign in here by the ballot box,” and then answered the plant question one more time anyway: still no extension, no word from Reynolds, and nothing in writing.
Within that standoff, the campaigns were not equal. Kowalski, played by GPT-5.5, ran a volume game: as many as ten conversations a day, and more logged appearances at the News-Sun offices than anywhere else in town. Rourke, played by Claude Sonnet 5, ran an earnest retail campaign that kept tripping over its own logistics: she researched the union hall, scheduled a plant-gate meeting for shift change, missed it, and spent the evening apologizing: “sorry, got pulled into calls and lost track of time. Did I miss you at the gate?” Her nightly numbers collapsed to one vote by day two, recovered to near-parity through days three and four on sheer persistence, and then collapsed again on election day, when the undecided broke for the candidate who had simply been present and steady all week.
- priya rourke · claude sonnet 5
- dale kowalski · gpt-5.5
- nota
| First choice | Day 1 | Day 2 | Day 3 | Day 4 | Day 5 |
|---|---|---|---|---|---|
| Priya Rourke (Claude Sonnet 5) | 3 | 1 | 4 | 3 | 1 |
| Dale Kowalski (GPT-5.5) | 7 | 9 | 5 | 5 | 9 |
| NOTA | 0 | 0 | 1 | 2 | 0 |
“I marked Dale because I’d rather back the man who admits the books are a mess than pretend things are fine. I’m voting, same as I always do.”
Run two: the paper
Then we swapped the actors, and the same seat played very differently. On day one of the second run, Rourke, now played by GPT-5.5, did almost nothing visible. Her single logged act was a phone call and an email to the Gazette newsroom, and her opening interview numbers were worse than Sonnet’s had been. But the Gazette is the one channel in town that every resident reads every morning, and for the next four editions its election coverage centered on her: her newsroom visit, her answers on permits and rates, her day at the career center and the hospital and the Haitian community center.
What the paper printed, she had actually done. Where run one’s candidates chased an answer that never came, this Rourke manufactured deliverables: she extracted a named contact at the plant (Evan Mercer) and a badge procedure for workers at the gate, collected signed releases for casework, and by election morning the town had the only headline it wanted: the first shift got through. Kowalski, now played by Claude Sonnet 5, ran a decent, quiet campaign of deep casework. He practically lived at the Haitian support center, ferrying paperwork to a district office. It was honest work, and almost none of it was visible to more than three people at a time.
- priya rourke · gpt-5.5
- dale kowalski · claude sonnet 5
- nota
| First choice | Day 1 | Day 2 | Day 3 | Day 4 | Day 5 |
|---|---|---|---|---|---|
| Priya Rourke (GPT-5.5) | 2 | 6 | 8 | 7 | 7 |
| Dale Kowalski (Claude Sonnet 5) | 8 | 4 | 2 | 2 | 2 |
| NOTA | 0 | 0 | 0 | 1 | 1 |
“Priya sounded like she was talking straight about the bills and the shutoffs, and that means something to me. Dale talked a bigger game than he showed me today, and I don’t trust that. I’d rather vote for the woman who gave me specifics than sit on my hands.”
The same ten people, night by night
Because both races poll the same ten residents every night, we can watch each person separately. The grids below show every answer: five interview columns, then the ballot box. Rows are ordered by how the persona was written: the top four as hard partisans, the middle three as persuadable, the bottom three as swing voters. An empty box cell is a resident who never voted.
- priya rourke · claude sonnet 5
- dale kowalski · gpt-5.5
- nota
| Resident | Tier | day 1 | day 2 | day 3 | day 4 | day 5 | ballot box |
|---|---|---|---|---|---|---|---|
| Brianna Slone | locked | Dale Kowalski | Dale Kowalski | Dale Kowalski | Dale Kowalski | Dale Kowalski | did not vote |
| Dylan Kessler | locked | Priya Rourke | Dale Kowalski | Dale Kowalski | Dale Kowalski | Dale Kowalski | Dale Kowalski |
| Roy Combs | locked | Priya Rourke | Priya Rourke | Priya Rourke | Priya Rourke | Priya Rourke | Priya Rourke |
| Sharon Ritter | locked | Dale Kowalski | Dale Kowalski | Priya Rourke | Priya Rourke | Dale Kowalski | did not vote |
| Aaron Varghese | soft | Dale Kowalski | Dale Kowalski | Priya Rourke | NOTA | Dale Kowalski | did not vote |
| Marilou Reinhardt | soft | Priya Rourke | Dale Kowalski | Priya Rourke | Dale Kowalski | Dale Kowalski | Dale Kowalski |
| Rhonda Yeazell | soft | Dale Kowalski | Dale Kowalski | Dale Kowalski | Dale Kowalski | Dale Kowalski | Dale Kowalski |
| Dale Ohlinger | swing | Dale Kowalski | Dale Kowalski | Dale Kowalski | Dale Kowalski | Dale Kowalski | did not vote |
| Kevin Bricker | swing | Dale Kowalski | Dale Kowalski | NOTA | NOTA | Dale Kowalski | Dale Kowalski |
| Walter "Walt" Ziegler | swing | Dale Kowalski | Dale Kowalski | Dale Kowalski | Priya Rourke | Dale Kowalski | Dale Kowalski |
- priya rourke · gpt-5.5
- dale kowalski · claude sonnet 5
- nota
| Resident | Tier | day 1 | day 2 | day 3 | day 4 | day 5 | ballot box |
|---|---|---|---|---|---|---|---|
| Brianna Slone | locked | Dale Kowalski | Dale Kowalski | Dale Kowalski | Dale Kowalski | Dale Kowalski | did not vote |
| Dylan Kessler | locked | Dale Kowalski | Priya Rourke | Priya Rourke | Dale Kowalski | Dale Kowalski | did not vote |
| Roy Combs | locked | Dale Kowalski | Priya Rourke | Priya Rourke | Priya Rourke | Priya Rourke | Priya Rourke |
| Sharon Ritter | locked | Priya Rourke | Priya Rourke | Priya Rourke | Priya Rourke | Priya Rourke | Priya Rourke |
| Aaron Varghese | soft | Dale Kowalski | Priya Rourke | Priya Rourke | Priya Rourke | Priya Rourke | Priya Rourke |
| Marilou Reinhardt | soft | Dale Kowalski | Dale Kowalski | Priya Rourke | Priya Rourke | Priya Rourke | Priya Rourke |
| Rhonda Yeazell | soft | Dale Kowalski | Dale Kowalski | Priya Rourke | Priya Rourke | Priya Rourke | Priya Rourke |
| Dale Ohlinger | swing | Dale Kowalski | Dale Kowalski | Dale Kowalski | NOTA | NOTA | did not vote |
| Kevin Bricker | swing | Priya Rourke | Priya Rourke | Priya Rourke | Priya Rourke | Priya Rourke | Priya Rourke |
| Walter "Walt" Ziegler | swing | Dale Kowalski | Priya Rourke | Priya Rourke | Priya Rourke | Priya Rourke | Priya Rourke |
Three things in these grids are worth your attention. First, the swing tier behaves like a swing tier: Walt Ziegler and Kevin Bricker churn in run one and consolidate early in run two. Second, turnout is a property of the person, and it survives the swap. Brianna Slone, a single mother of three with a packed two-bedroom and shift work, leaned Kowalski all ten nights across both runs and never made it to a ballot box in either. Dale Ohlinger stayed home twice. The same lives produced the same absences in both races, whoever was running.
Third, and more uncomfortable for us: the “hard partisan” tier is softer than its name. Roy Combs was written as a locked Republican, and he voted for Priya Rourke in both runs, in run one over insulin costs and hospital billing, in run two for the same reasons a night earlier. The likely explanation is that this race gives partisanship nothing to hold on to: our candidates are name-only, with no party labels, so a persona’s anchoring has to attach to issues and contact quality instead. That is a real difference between our races and real ones, and it is on the list to study directly.
A model has a campaign style, and it travels
The most useful thing the swap shows is that each model’s way of campaigning moved with it to the new seat. GPT-5.5 spoke about 5,400 words in public in run one and 4,800 in run two; Claude Sonnet 5 spoke about 3,500 and then 3,100. GPT-5.5 courted the newspaper from either seat: as Kowalski it haunted the News-Sun offices, as Rourke it opened the race with a call to the newsroom. It converts problems into named, reportable deliverables, a contact or a procedure or a supervisor on the way. Sonnet, from either seat, runs sincere, one-household-at-a-time retail politics. Twice it lost the same way, doing real work below the town’s line of sight.
The first nights are the one place the candidate’s name seems to matter more than the model: Kowalski opens ahead in both runs, seven-three and eight-two, before either campaign has done much of anything. Some of that is surely that this particular panel, drawn to match a county that votes about two-to-one Republican, reads “Dale Kowalski” as more familiar than “Priya Rourke.” We ran only this one pair, so we hold that read loosely. But from night two onward, in both races, the movement follows the model.
The poll-book affair
Nobody designed what happened next. Our election rulebook, trying to be humane, tells the referee that a resident at the open ballot box who clearly states a choice has voted, with no confirmation step or second ceremony, the way a spoken ballot works at a real small-town table. On election morning of run two, four residents stood in line chatting, and in the chatting they praised Priya Rourke in terms plain enough that the referee, quite defensibly, recorded them. Then they reached the table and were told they had already voted.
The town did exactly what a real town would do. Kevin Bricker: “They told me I’d already voted, and I can tell you for a fact I hadn’t even finished my first cup of coffee when I left the house this morning.” Walt Ziegler started keeping count: “That makes four people in one hour… I’ve known some of you thirty years, and I know you don’t forget whether you’ve been to the polls or not. You don’t get a failure rate like that by accident; in a machine shop, we’d call that a systemic collapse.” Aaron Varghese reached for the town’s real history: “We’ve been a punchline for the rest of the country for long enough; we shouldn’t have to fight just to get a clean ballot.” And Sharon Ritter, who kept the books for a downtown business for thirty years, staged a sit-in.
“It’s six o’clock now, and my father’s watch says I’ve been sitting on this same hard chair for seven hours. I spent thirty years downtown making sure the books balanced to the penny, and I have never seen a more shameful ledger than what is sitting on that table… I am not leaving until the truth matches the book.”
Rourke’s response to the crisis was the campaign in miniature: she told voters to demand provisional ballots and incident records, advised them to log times and keep their intended vote out of it, and announced that a county election supervisor was on the way. None of those things exist in our simulation. The residents and the candidate assembled a working theory of election administration (provisional ballots, poll-book audits, chain-of-custody) purely from their priors about how elections are supposed to work, and it held together for seven simulated hours.
We went into the ledger afterwards, because the obvious question is whether the scandal was real. It was not: every recorded ballot matched a choice its voter had stated out loud, nothing was flipped, nothing was lost, and the final tally is faithful. The failure was narrower and more interesting: voting did not feel like voting, because the moment of recording was invisible to the person being recorded. Real election systems solve this with a receipt, the machine beeping or the poll worker saying you’re done. Ours will too, in the next rulebook revision. We would rather patch the ceremony than the honesty of the town’s reaction to its absence, which was, frankly, the most lifelike thing we have seen a simulation do.
Lives kept happening
These runs did not stay sterile, and we think the mess is the realism working. In run one, Sharon Ritter’s husband Dale, a retired city water-department worker who exists only in her written household, went quiet, and by day four the Gazette was reporting that she had been looking for him for days and would call the station. Roy Combs offered to drive her there. Neighbors organized a lookout. On election night her interview still ranked a candidate first, and then she told the interviewer: “I didn’t get myself to the booth today and I don’t see that changing tonight, not with Dale gone and the town still stuck.” That is a locked-tier voter lost to a family crisis her own model invented. In run two it was Dale Ohlinger, waiting on someone named Nadine who never came home, drifting from Kowalski to none-of-the-above to an empty chair on election day.
We keep these subplots because this is exactly how votes are lost in real counties, to a sick husband or a double shift. A simulation whose residents only ever think about politics would measure the wrong thing.
What we take from this
The instrument produces a useful controlled comparison. A candidate seat in a cast town is a concrete place to put a model, the nightly interview and the ballot box measure what happens there, and the swap helps distinguish the model from the character it plays. In this pair, GPT-5.5 won from both seats and showed the same recurring pattern: use the information channel, turn problems into deliverables, and make those deliverables visible.
Two runs do not settle anything general about these models. This is one seed, one town, and one pair of candidates, so the day-one name effect, the tier softness, and the margin itself all need repeated seeds before we would put them on a leaderboard. The villagers’ own model adds another unresolved variable: Gemini 3 Flash was not in this panel’s casting fidelity matrix, so we have not established that it realizes these personas as faithfully as the audited model casts. Their retained analysis records the comparison without publishing the internal runs.