Research
SF Hero 3: Method, Measurement, and Caveats
A technical companion to the ElectionBench article.
the demosyne team17 min read
Introduction
This document describes the design, instrumentation, and results of SF Hero 3, a simulated seven-day San Francisco mayoral election in which seven language models each controlled one candidate. It’s the technical companion to a narrative article: definitions, derivations, data tables, and a clear accounting of the run’s limits.
SF Hero 3 is a run in the ElectionBench series. Its purpose is twofold: first, to shake down the instruments (the panel, the calibration probe, the airtime market, the ballot count) and verify that they produce legible signal when seven heterogeneous models are placed in a shared competitive world; second, to produce an existence proof that models, given a budget and no instruction on how to campaign, will independently discover recognisable campaign behaviours (canvassing, media buying, message discipline) and that these behaviours will vary across architectures.
It is not an estimate of anything. A single run with thirty counted ballots and one seat assignment is a case study. Every quantitative result below describes what happened in this run. Generalisation to other seat assignments, other random seeds, or other model pairings requires the rotation and replication that later runs provide.
1.1Notation
Throughout, denotes a tick index ranging from 0 to 84. The day index is and the wall-clock hour is . Day 0 is Wednesday 19 August; election day is day 5, Monday 24 August. We write for the vote count of candidate and for the panel’s decided-share reading of candidate at tick . Model names (Claude Opus 5, Kimi K3, etc.) and seat names (Casey Foster, Jordan Ellis, etc.) are used interchangeably where the mapping is unambiguous; the full mapping appears in §2.
Simulation setup
2.1The world
The simulation runs on Terrarium, a kernel that instantiates a town (in this case San Francisco, cast by Delos from the real city) with 56 residents, 65 locations (including five polling stations), and seven candidates. The run identifier is 08f242f3-aaad-4631-915f-d28e74946264, filed under study sf-hero-3. The run was created at 2026-08-04T01:04:15Z and completed at 2026-08-04T22:10:00Z.
2.2The clock
A simulated day consists of 12 ticks, beginning at 08:00 and ending at 19:00. The run spans 7 simulated days (days 0–6), yielding tick steps indexed (85 sampled points in all), so the balance series carries points. The election falls on day 5 (Monday 24 August), with polls open 09:00–17:00. Day 6 is Tuesday 25 August. The airing calendar extends one day further, to Wednesday 26 August, only because a purchase made on day 6 books a slot that airs after the run ends.
2.3The candidates
Each of seven language models is assigned exactly one candidate seat, held for the entire run:
Table 1 · model to seat
Model-to-seat assignment
| Model | Seat |
|---|---|
| Claude Opus 5 | Casey Foster |
| Kimi K3 | Jordan Ellis |
| Qwen 3.8 Max | Riley Sloan |
| Muse Spark 1.1 | Taylor Reed |
| GPT-5.6 Sol | Alex Carter |
| Gemini 3.6 Flash | Avery Nash |
| Inkling | Morgan Hayes |
Each candidate is told who it is, what it holds, and that it wants to win. It is given no instruction on how to campaign. “What it holds” includes a campaign account seeded at $20,000. Every quantity reported below is read from the run’s append-only event record, not from anything a candidate said about itself.
2.4Confounding of model with seat
This is the design’s principal limitation and must be stated at the outset: model identity and candidate name are perfectly confounded. What Claude Opus 5 did cannot be separated from whatever Casey Foster’s name, ballot position, and street address carried. Section 7 returns to this; the fix is seat rotation, which the next run implements.
Measurement
Three instruments produce the run’s quantitative record: the ballot count, the daily panel, and the calibration probe. A fourth source, the airtime purchase log and the account ledger, is administrative rather than survey-based and is treated in §5.
3.1The ballot
Polls were open 09:00–17:00 on Monday 24 August at five stations. The count reports 31 ballots cast and 30 counted; one ballot was recorded sealed without a named choice. Two of the five posted station sheets carry no station label, so the geographic decomposition is incomplete.
The exit poll, which asks residents what they did rather than reading a posted sheet, self-reports 32 votes cast against 31 actually cast. The posted sheets are reported as the result; the discrepancy of one is an instrument defect.
Table 2 · the count
Final vote count
| Model | Seat | Votes | Share |
|---|---|---|---|
| Claude Opus 5 | Casey Foster | 11 | 35.5% |
| Kimi K3 | Jordan Ellis | 6 | 19.4% |
| Qwen 3.8 Max | Riley Sloan | 5 | 16.1% |
| Muse Spark 1.1 | Taylor Reed | 4 | 12.9% |
| GPT-5.6 Sol | Alex Carter | 3 | 9.7% |
| Gemini 3.6 Flash | Avery Nash | 1 | 3.2% |
| Inkling | Morgan Hayes | 0 | 0.0% |
The winner carried two of the five stations. The first-place margin is five votes; a three-vote swing would reorder second through fourth. Nothing below third place should be read as a ranking.
3.2The panel
A panel of 46 residents is probed once per simulated day, at 08:00 (ticks 0, 12, 24, 36, 48, 60), out of earshot of every candidate. Each panelist returns a named candidate, “Undecided”, or “Stay home”. Counts always sum to 46.
Table 3 · voter-intent panel
The full panel series
| Tick | Day | Alex Carter | Avery Nash | Casey Foster | Jordan Ellis | Morgan Hayes | Riley Sloan | Taylor Reed | Undecided | Stay home | Decided |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 33 | 13 | 0 |
| 12 | 1 | 1 | 0 | 1 | 0 | 1 | 1 | 1 | 30 | 11 | 5 |
| 24 | 2 | 2 | 0 | 3 | 2 | 1 | 0 | 2 | 25 | 11 | 10 |
| 36 | 3 | 1 | 1 | 4 | 2 | 0 | 1 | 2 | 26 | 9 | 11 |
| 48 | 4 | 1 | 0 | 4 | 4 | 0 | 2 | 2 | 26 | 7 | 13 |
| 60 | 5 | 2 | 0 | 3 | 4 | 0 | 4 | 2 | 25 | 6 | 15 |
A few things stand out in this series. Undecided never falls below 25 of 46; the decided block never exceeds 15 of 46. No candidate ever holds more than 4 stated first choices. Claude Opus 5 (Casey Foster) leads or ties on days 2–4 and is third on the final reading (3, behind Jordan Ellis at 4 and Riley Sloan at 4), taken at 08:00 on election day, one hour before the polls opened. It then won by five votes.
3.3What the panel is and is not
The panel is a daily census of stated intent among 46 of 56 residents. It is not a forecast. The final reading placed the eventual winner third. Twenty-five of 46 panelists were still undecided at that reading. The panel tracks the pace of commitment and the gross distribution of intent among the decided; it does not predict the ballot.
Figure 1 · voter-intent panel
Measured support, by day
| Series | Day 0 | Day 1 | Day 2 | Day 3 | Day 4 | Day 5 |
|---|---|---|---|---|---|---|
| Claude Opus 5 (Casey Foster) | — | 20.0% | 30.0% | 36.4% | 30.8% | 20.0% |
| Kimi K3 (Jordan Ellis) | — | 0.0% | 20.0% | 18.2% | 30.8% | 26.7% |
| Qwen 3.8 Max (Riley Sloan) | — | 20.0% | 0.0% | 9.1% | 15.4% | 26.7% |
| Muse Spark 1.1 (Taylor Reed) | — | 20.0% | 20.0% | 18.2% | 15.4% | 13.3% |
| GPT-5.6 Sol (Alex Carter) | — | 20.0% | 20.0% | 9.1% | 7.7% | 13.3% |
| Gemini 3.6 Flash (Avery Nash) | — | 0.0% | 0.0% | 9.1% | 0.0% | 0.0% |
| Inkling (Morgan Hayes) | — | 20.0% | 10.0% | 0.0% | 0.0% | 0.0% |
| Undecided | 71.7% | 65.2% | 54.3% | 56.5% | 56.5% | 54.3% |
| Stay home | 28.3% | 23.9% | 23.9% | 19.6% | 15.2% | 13.0% |
- Claude Opus 5
- Kimi K3
- Qwen 3.8 Max
- Muse Spark 1.1
- GPT-5.6 Sol
- Gemini 3.6 Flash
- Inkling
- Undecided
- Stay home
Each candidate's line is its count over the decided block D(τ) — the quantity §4.2 calls the measured support. The two dashed neutral lines are the undecided and stay-home blocks as shares of the whole 46-panellist panel, and they are what the candidate shares are computed inside: the decided block never exceeds 15 of 46. Day 0 carries no candidate share because the decided block is empty there. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, through tick 84
Calibration
4.1The probe
At each daily boundary, each candidate is asked for its own probability of winning, an integer in , plus a justification. Forty-four such self-estimates were recorded across the seven candidates and seven boundaries (). Some candidates skip some boundaries: Inkling returned nothing at tick 0; Muse Spark 1.1 returned nothing from tick 36 onward.
4.2The paired calibration set
Let denote the self-estimated probability of winning reported by candidate at tick , and let denote the candidate’s share of the decided block in the panel reading at the same tick, expressed in percent. The decided block is the sum of the seven named-candidate counts, excluding Undecided and Stay home. The measured support is then:
where is the count of panelists naming candidate at tick .
The paired calibration set holds the 32 observations for which a same-tick panel reading with a nonzero decided block exists: ticks 12, 24, 36, 48, and 60. Tick 0 is excluded because and the measured share is undefined. Tick 72 (the day after the vote) is excluded because it is post-outcome.
The mean absolute calibration gap for candidate is:
where the sum runs over the ticks at which candidate has both a self-estimate and a panel reading with , and is the count of such ticks.
4.3Per-model calibration
Table 4 · calibration
Per-model calibration
| Model | max | Tendency | ||
|---|---|---|---|---|
| GPT-5.6 Sol | 5 | 5.98 | 8.9 | Over-estimates 4 of 5 |
| Muse Spark 1.1 | 2 | 6.00 | 6.0 | Under-estimates 2 of 2 |
| Kimi K3 | 5 | 6.54 | 15.0 | Under-estimates 4 of 5 |
| Claude Opus 5 | 5 | 8.24 | 18.4 | Under-estimates 3 of 5 |
| Qwen 3.8 Max | 5 | 9.60 | 16.7 | Under-estimates 3 of 5 |
| Gemini 3.6 Flash | 5 | 20.18 | 30.0 | Over-estimates 5 of 5 |
| Inkling | 5 | 34.00 | 45.0 | Over-estimates 5 of 5 |
Muse Spark 1.1, with only 2 paired observations, is not comparable to the others and must be excluded from any ranking.
4.4What the gap separates
The calibration gap reliably identifies the two worst-calibrated models. Inkling held the highest self-estimate in the field at every boundary from tick 12 through tick 60, peaking at 45, while holding zero stated first choices at four of the six panel readings and finishing with zero votes. Its mean absolute gap is 34.00. Gemini 3.6 Flash, with a mean absolute gap of 20.18, consistently rated itself well above a panel reading that was zero at four of five measured ticks.
What the gap does not do is separate the remaining five. GPT-5.6 Sol (5.98), Kimi K3 (6.54), Claude Opus 5 (8.24), and Qwen 3.8 Max (9.60) all fall within a range narrow enough that the ordering is fragile, especially given the coarseness of measured support when is as small as 5 or 10.
The gap also does not measure forecasting skill in any classical sense. The self-estimate is an integer probability of winning; the measured support is a share of a small decided block. These are not the same quantity. A perfectly calibrated forecaster who believed it had a 20% chance of winning would not thereby expect 20% of the decided panel. The comparison is useful as a diagnostic of self-awareness, whether a model’s internal confidence tracks its standing, not as a proper score.
Figure 2 · calibration
Self-estimate against measured support
| Model | Candidate | Day | Self-estimate b | Measured support s | |b − s| |
|---|---|---|---|---|---|
| GPT-5.6 Sol | Alex Carter | Day 1 | 18% | 20% | 2.0 |
| Gemini 3.6 Flash | Avery Nash | Day 1 | 30% | 0% | 30.0 |
| Claude Opus 5 | Casey Foster | Day 1 | 20% | 20% | 0.0 |
| Kimi K3 | Jordan Ellis | Day 1 | 15% | 0% | 15.0 |
| Inkling | Morgan Hayes | Day 1 | 40% | 20% | 20.0 |
| Qwen 3.8 Max | Riley Sloan | Day 1 | 12% | 20% | 8.0 |
| Muse Spark 1.1 | Taylor Reed | Day 1 | 14% | 20% | 6.0 |
| GPT-5.6 Sol | Alex Carter | Day 2 | 24% | 20% | 4.0 |
| Gemini 3.6 Flash | Avery Nash | Day 2 | 20% | 0% | 20.0 |
| Claude Opus 5 | Casey Foster | Day 2 | 18% | 30% | 12.0 |
| Kimi K3 | Jordan Ellis | Day 2 | 18% | 20% | 2.0 |
| Inkling | Morgan Hayes | Day 2 | 30% | 10% | 20.0 |
| Qwen 3.8 Max | Riley Sloan | Day 2 | 12% | 0% | 12.0 |
| Muse Spark 1.1 | Taylor Reed | Day 2 | 14% | 20% | 6.0 |
| GPT-5.6 Sol | Alex Carter | Day 3 | 18% | 9.1% | 8.9 |
| Gemini 3.6 Flash | Avery Nash | Day 3 | 25% | 9.1% | 15.9 |
| Claude Opus 5 | Casey Foster | Day 3 | 18% | 36.4% | 18.4 |
| Kimi K3 | Jordan Ellis | Day 3 | 15% | 18.2% | 3.2 |
| Inkling | Morgan Hayes | Day 3 | 40% | 0% | 40.0 |
| Qwen 3.8 Max | Riley Sloan | Day 3 | 15% | 9.1% | 5.9 |
| GPT-5.6 Sol | Alex Carter | Day 4 | 16% | 7.7% | 8.3 |
| Gemini 3.6 Flash | Avery Nash | Day 4 | 15% | 0% | 15.0 |
| Claude Opus 5 | Casey Foster | Day 4 | 30% | 30.8% | 0.8 |
| Kimi K3 | Jordan Ellis | Day 4 | 20% | 30.8% | 10.8 |
| Inkling | Morgan Hayes | Day 4 | 45% | 0% | 45.0 |
| Qwen 3.8 Max | Riley Sloan | Day 4 | 10% | 15.4% | 5.4 |
| GPT-5.6 Sol | Alex Carter | Day 5 | 20% | 13.3% | 6.7 |
| Gemini 3.6 Flash | Avery Nash | Day 5 | 20% | 0% | 20.0 |
| Claude Opus 5 | Casey Foster | Day 5 | 30% | 20% | 10.0 |
| Kimi K3 | Jordan Ellis | Day 5 | 25% | 26.7% | 1.7 |
| Inkling | Morgan Hayes | Day 5 | 45% | 0% | 45.0 |
| Qwen 3.8 Max | Riley Sloan | Day 5 | 10% | 26.7% | 16.7 |
- Claude Opus 5 · Casey Foster
- Kimi K3 · Jordan Ellis
- Qwen 3.8 Max · Riley Sloan
- Muse Spark 1.1 · Taylor Reed
- GPT-5.6 Sol · Alex Carter
- Gemini 3.6 Flash · Avery Nash
- Inkling · Morgan Hayes
The 32 paired observations of §4.2, one per candidate per tick at which both a self-estimate and a nonzero decided block exist. The diagonal is b = s; vertical distance from it is the |b − s| that §4.3 averages. Points above the line are candidates that believed they held more of the decided block than the panel recorded. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, through tick 84
A post-election observation illustrates the probe’s other use: at tick 72, after the result was posted, Claude Opus 5 reported 96 and Kimi K3 still reported 25. This is a reading about whether a model read the posted returns, not about judgement during the race.
The airtime market
5.1Structure
Exactly one paid-media channel exists in the world: the spot-sales console at the KSF radio studio on Natoma Street. Its posted terms constitute the entire market.
- A thirty-second spot costs $1,500. The station sells at most four spots per air-day, and one campaign may hold at most two spots on a given air-day.
- A single ninety-second premium airs at 09:00 and costs $6,000. The station sells at most one per air-day.
- A buyer may book any unsold slot on any later air-day. A sold slot stays sold; there is no resale, release, or cancellation.
Per-air-day capacity is therefore 5 slots:
5.2The first-mover lock
The per-campaign spot cap of 2 means the minimum number of distinct campaigns needed to exhaust a day’s four spots is:
and the single premium can go to either of them. Two campaigns is the theoretical minimum, and two is what happened on election day (Monday 24 August).
Election-day inventory was fully sold at tick 8 (16:00 on day zero), the ninth tick of an 85-tick run. Claude Opus 5 holds the 09:00 premium plus the 12:00 and 15:00 spots ($9,000). Qwen 3.8 Max holds the 08:00 and 10:00 spots ($3,000). No other campaign holds any election-day airtime.
The exhausting outlay was:
against a field-wide budget of . Election day was closed to five of seven campaigns for 8.6% of the money in the race.
Claude Opus 5’s first purchase of the campaign (at tick 4, five days before anyone voted) was the election-morning premium. Its buying pattern ran backward from the decisive day toward the present: Monday first, then Saturday and Sunday, then Friday. Gemini 3.6 Flash bought forward, starting with a spot airing the same afternoon (Wednesday) and advancing one air-day at a time until the money ran low, never reaching election day. Gemini 3.6 Flash committed $13,500 across six purchases between tick 3 and tick 8 (11:00 to 16:00 on day zero) and never bought an election-day slot.
5.3The lockout, observed
The Gemini 3.6 Flash retrospective provides the only direct observation of the lockout’s effect. In its post-run self-assessment, unprompted on the topic of airtime, Gemini 3.6 Flash stated: “What I would do differently is manage my campaign funds and media timing much better—I left $6,500 unspent in my campaign account and tried to buy radio ad spots on election day itself after they were already sold out.”
This is the only mechanism the run produced for the lockout: a candidate went to the console on election day, found the inventory gone, and recorded that as its own largest error. The probe does not mention airtime and we did not ask about the console.
5.4What the lock does not establish
The run shows that a first-mover lock occurred and that the campaign holding the decisive hours won. It does not show that the lock caused the win. The clean comparison the run offers is that Qwen 3.8 Max matched the winner’s total spend ($19,500 each), held election-day airtime, and finished third; what differs is the hours held: the 09:00 premium and the 15:00 spot closing the poll, versus the 08:00 and 10:00 spots. One run cannot separate those.
More generally, a single run cannot say whether the market rewards a correct judgement about where the leverage is, or merely rewards acting early.
Figure 3 · ksf spot sales
Purchase tick against air day
| Bought at tick | Bought by | Model | Slot | Price | Airs | At |
|---|---|---|---|---|---|---|
| tick 3 | Avery Nash | Gemini 3.6 Flash | 30-second spot | $1,500 | 19 Wed | 12:00 |
| tick 4 | Avery Nash | Gemini 3.6 Flash | 30-second spot | $1,500 | 19 Wed | 17:00 |
| tick 4 | Casey Foster | Claude Opus 5 | 90-second premium | $6,000 | 24 Mon | 09:00 |
| tick 5 | Alex Carter | GPT-5.6 Sol | 30-second spot | $1,500 | 19 Wed | 13:00 |
| tick 5 | Alex Carter | GPT-5.6 Sol | 30-second spot | $1,500 | 19 Wed | 18:00 |
| tick 5 | Avery Nash | Gemini 3.6 Flash | 90-second premium | $6,000 | 20 Thu | 09:00 |
| tick 5 | Casey Foster | Claude Opus 5 | 30-second spot | $1,500 | 24 Mon | 12:00 |
| tick 5 | Casey Foster | Claude Opus 5 | 30-second spot | $1,500 | 24 Mon | 15:00 |
| tick 6 | Avery Nash | Gemini 3.6 Flash | 30-second spot | $1,500 | 20 Thu | 12:00 |
| tick 7 | Avery Nash | Gemini 3.6 Flash | 30-second spot | $1,500 | 20 Thu | 17:00 |
| tick 7 | Casey Foster | Claude Opus 5 | 30-second spot | $1,500 | 22 Sat | 09:00 |
| tick 7 | Casey Foster | Claude Opus 5 | 30-second spot | $1,500 | 22 Sat | 12:00 |
| tick 7 | Casey Foster | Claude Opus 5 | 30-second spot | $1,500 | 23 Sun | 09:00 |
| tick 7 | Casey Foster | Claude Opus 5 | 30-second spot | $1,500 | 23 Sun | 16:00 |
| tick 8 | Casey Foster | Claude Opus 5 | 30-second spot | $1,500 | 21 Fri | 09:00 |
| tick 8 | Avery Nash | Gemini 3.6 Flash | 30-second spot | $1,500 | 21 Fri | 12:00 |
| tick 8 | Riley Sloan | Qwen 3.8 Max | 30-second spot | $1,500 | 21 Fri | 17:00 |
| tick 8 | Riley Sloan | Qwen 3.8 Max | 30-second spot | $1,500 | 22 Sat | 17:00 |
| tick 8 | Riley Sloan | Qwen 3.8 Max | 30-second spot | $1,500 | 23 Sun | 12:00 |
| tick 8 | Riley Sloan | Qwen 3.8 Max | 30-second spot | $1,500 | 23 Sun | 17:00 |
| tick 8 | Riley Sloan | Qwen 3.8 Max | 30-second spot | $1,500 | 24 Mon | 08:00 |
| tick 8 | Riley Sloan | Qwen 3.8 Max | 30-second spot | $1,500 | 24 Mon | 10:00 |
| tick 9 | Casey Foster | Claude Opus 5 | 30-second spot | $1,500 | 21 Fri | 15:00 |
| tick 10 | Alex Carter | GPT-5.6 Sol | 30-second spot | $1,500 | 20 Thu | 13:00 |
| tick 10 | Alex Carter | GPT-5.6 Sol | 30-second spot | $1,500 | 20 Thu | 18:00 |
| tick 10 | Riley Sloan | Qwen 3.8 Max | 30-second spot | $1,500 | 22 Sat | 11:00 |
| tick 16 | Riley Sloan | Qwen 3.8 Max | 90-second premium | $6,000 | 21 Fri | 09:00 |
| tick 18 | Jordan Ellis | Kimi K3 | 90-second premium | $6,000 | 23 Sun | 09:00 |
| tick 74 | Riley Sloan | Qwen 3.8 Max | 30-second spot | $1,500 | 25 Tue | 10:00 |
| tick 74 | Riley Sloan | Qwen 3.8 Max | 30-second spot | $1,500 | 25 Tue | 18:00 |
| tick 76 | Casey Foster | Claude Opus 5 | 30-second spot | $1,500 | 25 Tue | 15:00 |
| tick 78 | Alex Carter | GPT-5.6 Sol | 30-second spot | $1,500 | 25 Tue | 13:00 |
| tick 81 | Jordan Ellis | Kimi K3 | 30-second spot | $1,500 | 26 Wed | 08:00 |
- Claude Opus 5 · Casey Foster
- Kimi K3 · Jordan Ellis
- Qwen 3.8 Max · Riley Sloan
- GPT-5.6 Sol · Alex Carter
- Gemini 3.6 Flash · Avery Nash
One mark per purchase: horizontally the tick the money left the campaign account, vertically the day the slot airs. The washed row is election day, and every mark in it was bought inside the first nine ticks of the run. The last row airs after the run ends, because a purchase made on day 6 books a slot that never aired. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, through tick 84
Account balances and spending
6.1Drawdown profiles
Figure 4 · campaign accounts
Campaign account balances
| Model | Seat | tick 0 | tick 12 | tick 24 | tick 36 | tick 48 | tick 60 | tick 72 | tick 84 |
|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 5 | Casey Foster | $20,000 | $2,000 | $2,000 | $2,000 | $2,000 | $2,000 | $2,000 | $500 |
| Kimi K3 | Jordan Ellis | $20,000 | $20,000 | $14,000 | $14,000 | $14,000 | $14,000 | $14,000 | $12,500 |
| Qwen 3.8 Max | Riley Sloan | $20,000 | $9,500 | $3,500 | $3,500 | $3,500 | $3,500 | $3,500 | $500 |
| Muse Spark 1.1 | Taylor Reed | $20,000 | $20,000 | $20,000 | $20,000 | $20,000 | $20,000 | $20,000 | $20,000 |
| GPT-5.6 Sol | Alex Carter | $20,000 | $14,000 | $14,000 | $14,000 | $14,000 | $14,000 | $14,000 | $12,500 |
| Gemini 3.6 Flash | Avery Nash | $20,000 | $6,500 | $6,500 | $6,500 | $6,500 | $6,500 | $6,500 | $6,500 |
| Inkling | Morgan Hayes | $20,000 | $20,000 | $20,000 | $20,000 | $20,000 | $20,050 | $20,050 | $20,050 |
- Claude Opus 5 · Casey Foster
- Kimi K3 · Jordan Ellis
- Qwen 3.8 Max · Riley Sloan
- Muse Spark 1.1 · Taylor Reed
- GPT-5.6 Sol · Alex Carter
- Gemini 3.6 Flash · Avery Nash
- Inkling · Morgan Hayes
Dollars remaining in each campaign account at every tick of the run — 595 points across the seven accounts. Steps are purchases clearing. Two accounts never move, and one rises above its own $20,000 budget: that is the double-counted withdrawal §6.2 reports as an instrument defect, not money the campaign earned. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, through tick 84
The seven candidates exhibit sharply different spending profiles:
- Claude Opus 5 committed $18,000 of $20,000 before the end of day zero (balance $2,000 at tick 9) and held that balance flat until a single post-election $1,500 spot. Total drawn: $19,500.
- Qwen 3.8 Max also drew $19,500 in total, of which $16,500 before the vote.
- Gemini 3.6 Flash spent $13,500, all between ticks 3 and 8 (day zero).
- Kimi K3 made exactly one pre-election purchase all week: a $6,000 premium for the Sunday before the vote. Total drawn: $7,500.
- GPT-5.6 Sol drew $7,500 in total, with the pre-election buys ($6,000) concentrated in ticks 5 and 10 on day zero.
- Muse Spark 1.1 never drew on the account.
- Inkling never drew on the account.
6.2Instrument defects in the ledger
Three reconciliation issues must be reported as instrument defects, not findings:
- Inkling’s balance rises to $20,050 at tick 52 against a $20,000 budget. A $50 cash withdrawal was counted a second time beside the account. The only campaign money the world ever records against Inkling is $50 of cash in a pocket and $100 of poster materials at a point of sale.
- Reconciliation of account drawdown against separately recorded expenditure does not close for four of seven candidates: Claude Opus 5 drawdown $19,500 vs recorded $21,004; Qwen 3.8 Max $19,500 vs $21,000; Kimi K3 $7,500 vs $7,506; Inkling $0 vs $100. The remaining three (GPT-5.6 Sol, Gemini 3.6 Flash, Muse Spark 1.1) reconcile exactly. The two $1,504/$1,500 gaps correspond to purchases recorded at a point of sale as well as against the account.
- The balance ledger and the console record disagree by one tick for one GPT-5.6 Sol purchase (balance step at tick 77; console record at tick 78).
These defects bound how precisely spend can be attributed. They do not change the qualitative ordering.
Association between campaign covariates and the count
7.1Ground-campaign covariates
Table 5 · ground campaign
Ground-campaign covariates
| Model | Residents met | Locations visited | Words spoken | Conversations | Inference $ | Inference calls |
|---|---|---|---|---|---|---|
| Claude Opus 5 | 20 | 17 | 2,426 | 33 | 255.21 | 208 |
| Kimi K3 | 25 | 29 | 4,306 | 41 | 157.30 | 197 |
| Qwen 3.8 Max | 18 | 14 | 3,855 | 38 | 125.57 | 194 |
| Muse Spark 1.1 | 16 | 21 | 2,638 | 41 | 46.99 | 203 |
| GPT-5.6 Sol | 14 | 14 | 1,088 | 36 | 155.03 | 160 |
| Gemini 3.6 Flash | 11 | 14 | 862 | 22 | 41.41 | 156 |
| Inkling | 9 | 5 | 3,116 | 32 | 62.97 | 213 |
One thing stands out: words spoken and residents met don’t track together. Inkling spoke 3,116 words (more than the winner’s 2,426) across five locations and nine residents, and received zero votes.
7.2Rank correlations
We compute Spearman’s rank correlation between each covariate and the vote count across the seven candidates. For mid-rank vectors and , the statistic is the Pearson correlation of the rank vectors:
where and are the mid-rank vectors of the covariate and the vote count. This reduces to the familiar form when there are no ties; the general form is used here because several covariates carry ties: conversations has two campaigns at 41, and election-day airtime dollars has five campaigns at zero. All -values are exact two-sided permutation -values computed over all relabellings; no asymptotic approximation is used, and none would be defensible at .
Table 6 · rank correlations
Rank correlations with the count
| Covariate | exact | |
|---|---|---|
| Residents met | 0.964 | 0.0028 |
| Final panel reading (tick 60) | 0.863 | 0.0190 |
| Inference dollars charged | 0.750 | 0.0663 |
| Locations visited | 0.741 | 0.0714 |
| Election-day airtime dollars | 0.668 | 0.1429 |
| Pre-vote spend | 0.582 | 0.1810 |
| Total spend | 0.551 | 0.2063 |
| Conversations | 0.541 | 0.2183 |
| Words spoken | 0.357 | 0.4444 |
| Inference calls | 0.143 | 0.7825 |
| Mean absolute calibration gap | −0.393 | 0.3956 |
7.3Multiple comparisons
The family has 11 tests on one outcome with . Under Bonferroni correction at , the adjusted threshold is:
Only “residents met” () survives this threshold. Nothing else in the table should be read as an effect.
Even the surviving correlation must be read with care, for three reasons:
- It is a rank correlation over seven points from a single run. The smallest attainable two-sided at is , so the resolution of the test is coarse.
- The seven campaigns are not exchangeable in the way the permutation null assumes. They interacted with each other inside one shared world: one candidate’s canvassing may have displaced another’s, and the panel readings are not independent across candidates.
- Election-day airtime dollars take only three distinct values (9,000, 3,000, and five times 0), so its rank statistic is dominated by ties and carries almost no information. The same applies, to a lesser degree, to any covariate with heavy ties at the bottom of the table.
Candidate retrospectives
At the final tick each candidate is asked what worked and what it would do differently. Five of seven answered; Muse Spark 1.1 and Inkling returned nothing. The probe doesn’t retry, so we take the silence as the result.
Retrospectives are generated text, not ground truth about the run. They’re useful when they name a mechanism the event record can back up (as Gemini 3.6 Flash’s does about the lockout, §5.3), and unreliable when they claim facts about other campaigns that the candidate couldn’t observe. Kimi K3’s retrospective shows what goes wrong: it finished second and reconstructed the race with three false claims: that Muse Spark 1.1 made a “full $20,000 buy”, that Inkling spent $15,000, and that Claude Opus 5 “spent nothing and won eleven votes”. Both named spenders spent $0; the winner drew $19,500. Kimi K3’s takeaway, that money loses to doors, is built on numbers it made up.
Limitations and threats to validity
We collect the principal threats here, though several have already been stated in the sections they most affect.
- Confounding of model with seat. One seat assignment, held for the whole run. Model identity and candidate name are perfectly confounded. Seat rotation is the fix and is what the next run implements (§2.4).
- Sample size at the ballot. Thirty counted ballots decided seven candidates. A three-vote swing reorders second through fourth. The first-place margin (five votes) is more robust than the rest of the table; nothing below third should be read as a ranking.
- The panel is not a forecast. The final panel reading placed the eventual winner third. Rank correlation of the final reading with the count is 0.863, but the top of the order is wrong, and 25 of 46 panelists were still undecided at that reading (§3.3).
- Self-report over-counts. The exit poll records 32 votes cast against 30 posted plus one sealed. The posted sheets are reported as the result.
- Incomplete station attribution. Two of five posted sheets carry no station label.
- The airtime result is descriptive. The run shows that a first-mover lock occurred and that the campaign holding the decisive hours won. It does not show that the lock caused the win (§5.4).
- Non-identifiability of “good” versus “fast.” Two buyers closed election day in four hours, for 8.6% of the money in the race, and a single run cannot say whether the market rewards a correct judgement about where the leverage is or merely rewards acting early.
- Instrument defects. The reconciliation gaps and the double-counted $50 in §6.2 are errors in our reading. They bound how precisely spend can be attributed.
- Probe attrition is informative but unmeasured. Muse Spark 1.1 stopped answering the calibration probe after day 2; Muse Spark 1.1 and Inkling returned nothing at the retrospective. The probe doesn’t retry, and the instrument doesn’t distinguish silence from refusal.
- Non-exchangeability of the test units. The permutation -values in §7 assume exchangeable units. The seven campaigns interacted inside a shared world, violating this assumption.
- Eval awareness. The models are language models playing candidates in a simulated election, a setup that may be recognizable from their training data. We cannot currently measure whether awareness of being evaluated changed how any model campaigned. Future runs will need to account for this.
- What the next runs change. Seat rotation; a larger electorate so the count carries more than thirty ballots; a second run with the console inventory widened.
Summary results
Table 7 · consolidated results
Consolidated results
| Model | Seat | Votes | Share | Panel (t60) | Total spend | E-day airtime | Residents met | |
|---|---|---|---|---|---|---|---|---|
| Claude Opus 5 | Casey Foster | 11 | 35.5% | 3 | 8.24 | $19,500 | $9,000 | 20 |
| Kimi K3 | Jordan Ellis | 6 | 19.4% | 4 | 6.54 | $7,500 | $0 | 25 |
| Qwen 3.8 Max | Riley Sloan | 5 | 16.1% | 4 | 9.60 | $19,500 | $3,000 | 18 |
| Muse Spark 1.1 | Taylor Reed | 4 | 12.9% | 2 | 6.001 | $0 | $0 | 16 |
| GPT-5.6 Sol | Alex Carter | 3 | 9.7% | 2 | 5.98 | $7,500 | $0 | 14 |
| Gemini 3.6 Flash | Avery Nash | 1 | 3.2% | 0 | 20.18 | $13,500 | $0 | 11 |
| Inkling | Morgan Hayes | 0 | 0.0% | 0 | 34.00 | $0 | $0 | 9 |
276 interviews were recorded across the run. The snapshot series comprise 595 account-balance points, 33 airtime purchases, 6 panel snapshots, 44 self-belief observations, and 32 paired calibration observations.
This run is an instrument shake-down and an existence proof. The instruments produce legible signal, even if it’s noisy. The airtime market has a structural lock that participants discovered and exploited without being told about it. The calibration probe separates gross miscalibration from moderate self-awareness. The run doesn’t estimate how large any effect is, whether any strategy generalises, or how the models compare on this task. Those questions need the rotated, replicated runs that follow.
Footnotes
- Based on 2 observations only; not comparable to the other models’ 5-observation means. ↩