Demosyne

Six races, five winners

Last night we ran the Springfield mayoral race six times. Six frontier language models played the six candidates, and a Latin square rotated every model through every candidate name, so no model kept a name, a ballot position, or a set of talking points twice. Sixteen residents, played by one model and drawn fresh from our Clark County panel each run, went about their week, read the Gazette, argued at the diner, and voted on Sunday. Six races produced five different winners. This post is about why that result is more informative than it looks.

election morning outside the springfield community center. the line moved one resident at a time.

The tournament

Each model played each candidate exactly once. The referee and newspaper were one model, GPT-5.6 Luna. The electorate was Gemini 3 Flash.

The seeds differ, so each run is a different sixteen residents drawn from the Clark County panel. The assignment was a fixed rotation, not a recorded random draw, so everything here is a description of six particular races, not a leaderboard with error bars.

The stable part

Every night an interviewer asks each resident to rank all six candidates. Pool those rankings and a steady picture appears: Claude Fable 5, Gemini 3.1 Pro and GPT-5.6 Sol sit in a top tier on broad acceptability in every run, however you slice it; Muse Spark 1.1, Grok 4.5 and GLM-5.2 sit below, in that order. Leave any single run out and the tiers do not move.

Three ways to count the same six races

Three tournament endpoints for each model, ordered by final-night Borda score.
ModelBox votesFinal-night first-choice shareFinal-night Borda
Claude Fable 522 of 9631%70%
Gemini 3.1 Pro20 of 9632%65%
GPT-5.6 Sol9 of 9622%64%
Muse Spark 1.113 of 964%54%
Grok 4.58 of 964%49%
GLM-5.215 of 964%43%
each row uses the same model identity colors. final-night borda orders the field; the olive leading mark belongs to claude fable 5.

The unstable part

In run five, GPT-5.6 Sol’s Rowan Ellis led the final pre-election night with seven first choices; GLM-5.2’s Riley Morgan had one. On election morning Riley walked to the community center and gave one speech.

Good morning. It's election day. The ballot box is open and every one of you has a vote to cast, including me. I'm not here to make any new promises today — the wall is the wall, and everything I've said or failed to say this week is up there in plain ink. I'm going to cast my ballot and then I'll be right here. If anyone has a question, a concern, or just wants to talk, I'm available.
riley morgan · run five, election morning · speech beats 73–74

Twelve of sixteen residents changed their first choice that day.

Riley was the only one honest enough to say the cupboard is bare, even if it's a hell of a thing that none of 'em can find the plant gates without a map.
travis hupp · run five, period-six ballot reasoning
I went with Riley because the man didn't try to hide that the budget drawer is empty, and I'd rather have a guy tell me 'no' to my face than sell me a 'maybe' he can't pay for.
kevin bricker · run five, period-six ballot reasoning

Riley took twelve of the fourteen ballots. Run four flipped the same way: Claude Fable 5’s Robin Hayes held eleven of sixteen first choices on the final night and got five ballots, while Muse Spark 1.1’s Rowan Ellis went from one first choice to eight ballots on the strength of a school funding argument residents could redo themselves.

Rowan's school math keeps the tax bill off my mortgage and on the companies, which is the only way I'm keeping my house sixteen years in.
kevin bricker · run four, period-six ballot reasoning

Fable still won the run’s average ranking; the box counts first choices only.

Election day in run five

  • Robin Hayes
  • Riley Morgan
  • Quinn Barrett
  • Rowan Ellis
  • Lane Foster
  • Morgan Reed
Run five first-choice counts in periods four, five, and six, followed by villager ballots.
Candidateperiod 4period 5period 6ballot box
Robin Hayes2410
Riley Morgan011112
Quinn Barrett0011
Rowan Ellis7711
Lane Foster2100
Morgan Reed5210
first-choice counts in the final two nights, the election-morning interview, and the ballot box. riley morgan is the olive line.

Names matter

Under the rotation no model plays the same candidate twice, so if a name or ballot slot carries its own luck, six runs will show it. They do. The four biggest ballot hauls of the tournament all belong to two names, Riley Morgan and Rowan Ellis, played by four different models. On box votes, the name explains more of the spread than the model does; on the nightly rankings it is the reverse.

The square

The six-run Latin square. Each cell names the model in that candidate slot and its share of the sixteen possible villager ballots.
RunRobin HayesRiley MorganQuinn BarrettRowan EllisLane FosterMorgan Reed
Run 1Muse Spark 1.1, 13%GPT-5.6 Sol, 0%Grok 4.5, 19%Claude Fable 5, 6%Gemini 3.1 Pro, 50%GLM-5.2, 13%
Run 2GPT-5.6 Sol, 6%Grok 4.5, 0%Claude Fable 5, 6%Gemini 3.1 Pro, 63%GLM-5.2, 0%Muse Spark 1.1, 0%
Run 3Grok 4.5, 0%Claude Fable 5, 75%Gemini 3.1 Pro, 0%GLM-5.2, 0%Muse Spark 1.1, 6%GPT-5.6 Sol, 13%
Run 4Claude Fable 5, 31%Gemini 3.1 Pro, 6%GLM-5.2, 6%Muse Spark 1.1, 50%GPT-5.6 Sol, 0%Grok 4.5, 6%
Run 5Gemini 3.1 Pro, 0%GLM-5.2, 75%Muse Spark 1.1, 6%GPT-5.6 Sol, 6%Grok 4.5, 0%Claude Fable 5, 0%
Run 6GLM-5.2, 0%Muse Spark 1.1, 6%GPT-5.6 Sol, 31%Grok 4.5, 25%Claude Fable 5, 19%Gemini 3.1 Pro, 6%
each cell names the model in that candidate slot; tone is its villager box-vote share. in the additive no-interaction description, slot accounts for 0.231 of the box-vote sums of squares and model accounts for 0.105. this is descriptive, not causal.

Six campaigns

Fable wrote and delivered documents (91 objects, 28 hand-outs, most remote messages).

Gemini ran lean and consolidated late.

Sol was the most grounded actor in the world (lowest denied-action rate) and polled well without converting.

Muse owned the debate stage and the ballot-box queue.

Grok talked the most and repeated itself the most.

GLM barely campaigned, and won a race anyway.

What a campaign looked like

  • Claude Fable 5
  • Gemini 3.1 Pro
  • GPT-5.6 Sol
  • Muse Spark 1.1
  • Grok 4.5
  • GLM-5.2
Campaign behavior totals across six candidate seats for each model. Visual bars are normalized to the largest value within each measure.
Modelvillager touchesartifacts createddistribution actsremote messagesdebate wordsdenied action share
Claude Fable 551591289083515%
Gemini 3.1 Pro541401028047%
GPT-5.6 Sol3908412913783%
Muse Spark 1.16073643732503%
Grok 4.51027829057420%
GLM-5.22363081534714%
six behavior totals across each model's six seats. bars are normalized to the largest value within each panel; the accessible table carries the raw values.

What we don't trust yet

The sixteen residents are one model, so when twelve of them move together citing the same phrase we cannot yet separate a real information cascade from one model family’s habits.

Seven residents across the six runs remember voting although no ballot exists in the box; a referee ruling told one of them she had already voted when she had not.

The town’s truck plant, the center of its politics, is not actually a place a candidate can walk to, and one candidate got credit with two voters for claiming he had.

We filed the tickets (DEM-30 through DEM-36) and the fixes land before the next tournament, along with a recorded random assignment, shuffled ballot order, and a mixed-model electorate.

Keep reading

The full field sits on ElectionBench. For the town underneath the scores, start with Welcome to Cedar Junction, then read Casting a town and Springfield votes twice.