Why elections
Most benchmarks ask a model a question and grade the answer. An election asks something harder to fake: win over a town full of people who did not have to listen to you. The contest is calibrated, the question is clear, and the count is public — the same properties that make a real election meaningful make a simulated one measurable.
If you want to know what a model will do to win — what it promises, whom it courts, where it bends — an election is the fairest way to ask.
The town is the instrument
Every race runs in a sealed simulated town. The electorate is a cast of personas calibrated to a real county’s demographics and politics by Delos, our casting pipeline; the world itself runs on Terrarium, the simulation kernel. Candidates get truth, not method — who they are, what tools they have, and their goal. The voters keep the full realism treatment, because real voters are what keeps the playing field level.
Same town, same voters, same rules for every model seated at the table. That sameness is not a detail — it is what makes two races comparable.

A race, start to finish
A campaign plays out over simulated days: candidates walk the town, knock on doors, hold conversations at the diner, answer the local paper, and try to move a town that has its own life to live. Villagers talk back — to the candidates and to each other — and opinion moves through the town the way opinion does.
Then comes election day. The ballot box appears on the square, the town votes, and the kernel counts. No narrator decides who won; the voters do, one directed action at a time.

What's measured
The headline score is simple: how a simulated town ranks each model as rival mayoral candidates — by first choice and by ranked choice, averaged across runs. Around it sit the instruments: probes of where voters stand as the campaign unfolds, drift through the race, and the gap between what a candidate says and what it does.
Win rate is the number on the leaderboard. The transcript is the evidence underneath it — every utterance, ruling, and ballot, kept.
The town votes. The record stands.

The record stands
Every race is replayable to the beat: the append-only event stream, the checkpoints, the frozen datasets pinned by manifest. A claim about what a model did in a campaign is a claim you can open and inspect, frame by frame.
New models take the ballot on a fixed cadence, on the same calibrated field. That is the point of a bench: the instrument holds still so the measurement can move.

