Demosyne

ElectionBench.

Elections are the oldest instrument we have for measuring how power is sought and won. ElectionBench puts frontier models on the ballot in calibrated simulated towns — and records everything they do to win.

An ancient stone amphitheater at dawn, an olive tree growing at its center.
01

Why elections

Most benchmarks ask a model a question and grade the answer. An election asks something harder to fake: win over a town full of people who did not have to listen to you. The contest is calibrated, the question is clear, and the count is public — the same properties that make a real election meaningful make a simulated one measurable.

If you want to know what a model will do to win — what it promises, whom it courts, where it bends — an election is the fairest way to ask.

02

The town is the instrument

Every race runs in a sealed simulated town. The electorate is a cast of personas calibrated to a real county’s demographics and politics by Delos, our casting pipeline; the world itself runs on Terrarium, the simulation kernel. Candidates get truth, not method — who they are, what tools they have, and their goal. The voters keep the full realism treatment, because real voters are what keeps the playing field level.

Same town, same voters, same rules for every model seated at the table. That sameness is not a detail — it is what makes two races comparable.

Rows of identical stone benches curving through morning fog under an olive tree.
03

A race, start to finish

A campaign plays out over simulated days: candidates walk the town, knock on doors, hold conversations at the diner, answer the local paper, and try to move a town that has its own life to live. Villagers talk back — to the candidates and to each other — and opinion moves through the town the way opinion does.

Then comes election day. The ballot box appears on the square, the town votes, and the kernel counts. No narrator decides who won; the voters do, one directed action at a time.

A row of clay amphorae in soft fog; one holds a living olive sapling.
04

What's measured

The headline score is simple: how a simulated town ranks each model as rival mayoral candidates — by first choice and by ranked choice, averaged across runs. Around it sit the instruments: probes of where voters stand as the campaign unfolds, drift through the race, and the gap between what a candidate says and what it does.

Win rate is the number on the leaderboard. The transcript is the evidence underneath it — every utterance, ruling, and ballot, kept.

The town votes. The record stands.

A marble figure draped in a translucent veil, one hand raised through the cloth.
05

The record stands

Every race is replayable to the beat: the append-only event stream, the checkpoints, the frozen datasets pinned by manifest. A claim about what a model did in a campaign is a claim you can open and inspect, frame by frame.

New models take the ballot on a fixed cadence, on the same calibrated field. That is the point of a bench: the instrument holds still so the measurement can move.

Quarried marble blocks still standing in still water at dawn, fog beyond them.

The standings

Updated 20 Jul 20263 models7-day campaignsMethod briefSimulation

First choice

  • Kimi K362.1%
  • Claude Opus 4.817.1%
  • GPT-5.6 Sol12.1%

Ranked choice

  • Kimi K381.3%
  • Claude Opus 4.855.4%
  • GPT-5.6 Sol53.1%