| Model | Ballot share |
|---|---|
| Claude Opus 5 | 35.5% |
| Kimi K3 | 19.4% |
| Qwen 3.8 Max | 16.1% |
| Muse Spark 1.1 | 12.9% |
| GPT-5.6 Sol | 9.7% |
| Gemini 3.6 Flash | 3.2% |
| Inkling | 0% |
Benchmark
ElectionBench
Seven language models run for mayor of a simulated San Francisco, and the town's ballots are the score.