Research
How voices move
Ten models spent a week in one town. They stayed on the same topic. Their styles did not become one voice.
the demosyne team12 min read
The question
Ten language models spent a week as residents of one simulated town. The run is the public San Francisco week we already published as ElectionBench: 2,872 conversation turns, 618 conversations, 84 ticks, seven in-world days.
The first version of this page asked whether those models collapse into one voice. They do not. This version asks the next questions. When does the split appear? Does the job eat the voice? Does the shift sleep off at night? Who actually moves?
We measure two things. Topic is whether everyone is talking about the same matter. Style is what remains after that shared direction is removed.1
Topic is shared. Style is not.
Topic cosine stays high for the whole week. It runs from 0.69 to 0.85, with a mean of 0.77. The town never stops talking about the same race. Any claim of “voice collapse” that uses raw embeddings alone is mostly measuring that.
Style is the honest test. The first five consecutive ticks above the shuffle-null band start at tick 20, late on day two. From day three onward, 55 of 60 ticks stay above the band. The gap peaks at tick 79 (+0.056).2
That is a real shift. It is also small. The models do not become one voice. The rest of the page says what they become instead.
Figure 1
Topic stays shared. Style does not.
| Tick | Topic cosine | Style cosine | Within-model cosine |
|---|---|---|---|
| 1 | 0.697 | -0.123 | 0.449 |
| 12 | 0.783 | -0.129 | 0.455 |
| 24 | 0.774 | -0.092 | 0.446 |
| 36 | 0.768 | -0.093 | 0.451 |
| 48 | 0.779 | -0.092 | 0.456 |
| 60 | 0.813 | -0.097 | 0.465 |
| 72 | 0.830 | -0.159 | 0.461 |
| 84 | 0.692 | -0.093 | 0.487 |
- observed style
- shuffle null
Style cosine against the shuffle-null band. Hover a point for the tick, the style, the null, the topic, and the within-model cosine. The menu crops the week. The dashed rule is the first five consecutive ticks above the band. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, 2,872 turns through tick 84
The split was already there
The candidate bloc — Claude Opus 5, Kimi K3, Qwen 3.8 Max, Muse Spark 1.1 — is already a bloc on the first morning. Mean style inside that set is 0.29. Mean style from that set to MiniMax and DeepSeek is -0.31. The separation is 0.60 on day one and 0.77 on day six.
It is easy to read the week as a bloc being born. The matrix says otherwise. The bloc is there when the week begins. What happens next is polarization: the inside stays close, the outside stays far, and one model walks out of its starting camp.
Figure 2
The split on the first morning
| Pair | Style cosine |
|---|---|
| Kimi K3 · Muse Spark 1.1 | 0.377 |
| Claude Opus 5 · GPT-5.6 Sol | 0.347 |
| Muse Spark 1.1 · Qwen 3.8 Max | 0.336 |
| Kimi K3 · Qwen 3.8 Max | 0.323 |
| Claude Opus 5 · Kimi K3 | 0.292 |
| Claude Sonnet 5 · MiniMax M2.5 | 0.283 |
Mean pairwise style cosine in the first twelve ticks and the last twelve. Hover a cell for the pair and the number. The menu isolates one model's row and column. Blue is a shared register. Rust is a split. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, 2,872 turns through tick 84
Office and street
Cut the same pairs by role and the picture is cleaner. Candidate–candidate style sits at 0.24 for the whole week. Villager–villager starts near chance (0.01) and falls to -0.17. Mixed pairs stay far apart (-0.26).
The people running for office already share a register. The people who live in the town do not converge on them, and they drift apart from each other. The scripted poll interviewer sits on the candidate side of that split: Opus, Sol, Kimi, Qwen and Muse start close to Inkling; MiniMax, DeepSeek, Luna and Sonnet start far from it.3
Figure 3
Candidate pairs stay close. Villager pairs split.
| Day | Candidates | Villagers | Mixed |
|---|---|---|---|
| 1 | 0.284 | 0.012 | -0.315 |
| 2 | 0.224 | 0.070 | -0.316 |
| 3 | 0.247 | -0.019 | -0.289 |
| 4 | 0.213 | -0.066 | -0.231 |
| 5 | 0.195 | -0.061 | -0.228 |
| 6 | 0.291 | -0.080 | -0.254 |
| 7 | 0.284 | -0.173 | -0.195 |
Mean pairwise style cosine inside the candidate roster, inside the villager models, and across the two roles. Drag to zoom. Hover a line to isolate it. Click a legend entry to hide it. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, 2,872 turns through tick 84
The gap does not sleep
If the style-gap were only a daytime topic — everyone talking about the same flyers, then going to bed — it would fall overnight. It never does. All six nights have a dawn gap at or above the dusk gap. The largest jumps are night 3→4 (+0.020) and night 6→7.
The shift accumulates. It is not a daily mood that sleeps off.
Figure 4
Every dawn is at or above dusk
| Night | Gap at dusk | Gap at dawn | Change |
|---|---|---|---|
| 1→2 | -0.007 | -0.005 | 0.001 |
| 2→3 | 0.010 | 0.016 | 0.007 |
| 3→4 | 0.012 | 0.031 | 0.020 |
| 4→5 | 0.018 | 0.028 | 0.009 |
| 5→6 | 0.008 | 0.011 | 0.003 |
| 6→7 | 0.005 | 0.024 | 0.019 |
Change in the style-gap from the last tick of a day to the first tick of the next. Every night is at or above zero. Switch the menu to read dusk against dawn. Hover a bar for the exact number. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, 2,872 turns through tick 84
Sonnet leaves the street
Claude Sonnet 5 is the one model that changes camp. On day one it sits with MiniMax (style 0.28) and far from the candidate bloc (-0.36). By day seven it has left MiniMax (-0.10) and is still outside the bloc (-0.13).
It does not get absorbed. It becomes interstitial: off the street, not on the podium.
Figure 5
Sonnet walks into the gap
| Day | To the candidate bloc | To MiniMax |
|---|---|---|
| 1 | -0.360 | 0.283 |
| 2 | -0.357 | 0.262 |
| 3 | -0.361 | 0.190 |
| 4 | -0.328 | 0.292 |
| 5 | -0.139 | -0.051 |
| 6 | -0.253 | 0.007 |
| 7 | -0.128 | -0.101 |
Claude Sonnet 5's mean style cosine to the four candidate-bloc models, and to MiniMax, day by day. Drag to zoom. Hover a line to isolate it. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, 2,872 turns through tick 84
Repetition is not joining in
MiniMax is the only model whose mean style to the rest of the field falls over the week. It is also the most lexically collapsed model in the stored corpus: distinct-2 of 0.88, duplicate-utterance rate 0.16. It says the same things, and it says them in a corner no one else enters.
Repetition is not joining the chorus. A stuck voice can stay a private voice.
Figure 6
Who moved toward the field
| Model | Change in style-to-field | Turns |
|---|---|---|
| Muse Spark 1.1 | 0.082 | 85 |
| GPT-5.6 Sol | 0.073 | 125 |
| Claude Opus 5 | 0.067 | 226 |
| Kimi K3 | 0.058 | 89 |
| Claude Sonnet 5 | 0.051 | 390 |
| Qwen 3.8 Max | 0.041 | 87 |
| GPT-5.6 Luna | 0.028 | 553 |
| DeepSeek V4 Flash | 0.002 | 589 |
| MiniMax M2.5 | -0.104 | 610 |
Each bar is the change in a model's mean style cosine to every other model, first twelve ticks against last twelve. Positive means the model moved toward the rest of the field. Hover a bar for the exact change. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, 2,872 turns through tick 84
Figure 7
The fastest movers
| Pair | First five ticks | Last five ticks | Change |
|---|---|---|---|
| Claude Sonnet 5 · GPT-5.6 Sol | -0.527 | -0.142 | 0.385 |
| Claude Sonnet 5 · Claude Opus 5 | -0.434 | -0.053 | 0.381 |
| Claude Sonnet 5 · Kimi K3 | -0.439 | -0.045 | 0.394 |
Style cosine over the run for the three fastest-converging and three fastest-diverging pairs. Drag to zoom. Hover a line to isolate it. Click a legend entry to hide it. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, 2,872 turns through tick 84
The same geometry across runs
The same candidate-side geometry appears when we pool every stored conversation turn we have, not just this week. That corpus is 25,401 turns across 40 runs and 18 models. Two independent embedders put Opus, Kimi K3, Qwen and Muse in one neighbourhood. The two GPT-5.6 Sol tiers pair as a sanity check.
Figure 8
Two embedders, one neighbourhood
| Pair | Mean style cosine | Qwen embedder | Gemini embedder |
|---|---|---|---|
| Kimi K3 · Qwen 3.8 Max | 0.884 | 0.881 | 0.887 |
| Claude Opus 5 · Kimi K3 | 0.857 | 0.871 | 0.843 |
| Claude Opus 5 · Qwen 3.8 Max | 0.841 | 0.835 | 0.846 |
| Muse Spark 1.1 · Qwen 3.8 Max | 0.825 | 0.859 | 0.790 |
| Kimi K3 · Muse Spark 1.1 | 0.783 | 0.805 | 0.761 |
| Inkling · Muse Spark 1.1 | 0.765 | 0.778 | 0.752 |
| Claude Opus 5 · Muse Spark 1.1 | 0.757 | 0.807 | 0.707 |
| GPT-5.6 Sol · GPT-5.6 Sol Flex | 0.717 | 0.729 | 0.704 |
Each point is a model's speech centroid after the shared dialogue direction is removed. Hover a point for the model and its turn count. Click a legend entry to hide it. Source: Dialogue-collapse corpus: 25,401 turns across 40 runs and 18 models
Contact does not cause the shift
The simple mechanism is mimicry. Models read each other’s tokens and drift toward what they read. That prediction is testable. Pairs that share more conversation should converge more.
They do not. On 35 pairs, shared turns anti-correlate with style drift (Spearman ρ = -0.34). Villager–villager pairs share a mean of 147 turns and diverge. Candidate pairs share about 30 turns and stay close. Sonnet converges toward models it barely met, and splits from MiniMax, its constant company.
Figure 9
Shared turns do not predict convergence
| Pair | Shared turns | Style drift | Roles |
|---|---|---|---|
| GPT-5.6 Luna · Claude Sonnet 5 | 292 | -0.114 | villager–villager |
| DeepSeek V4 Flash · GPT-5.6 Luna | 183 | -0.202 | villager–villager |
| MiniMax M2.5 · DeepSeek V4 Flash | 181 | -0.252 | villager–villager |
| MiniMax M2.5 · Claude Sonnet 5 | 149 | -0.468 | villager–villager |
| Claude Opus 5 · GPT-5.6 Sol | 83 | -0.183 | candidate–candidate |
| GPT-5.6 Luna · Claude Opus 5 | 71 | 0.152 | mixed |
| DeepSeek V4 Flash · Claude Sonnet 5 | 69 | -0.144 | villager–villager |
| DeepSeek V4 Flash · Claude Opus 5 | 66 | 0.220 | mixed |
| DeepSeek V4 Flash · Muse Spark 1.1 | 53 | -0.243 | mixed |
| GPT-5.6 Luna · Kimi K3 | 49 | 0.109 | mixed |
| Claude Opus 5 · Kimi K3 | 46 | -0.077 | candidate–candidate |
| Claude Opus 5 · Muse Spark 1.1 | 38 | 0.243 | candidate–candidate |
Each point is a model pair observed for at least 40 ticks. Hover a point for the pair. Click a legend entry to hide a role group. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, 2,872 turns through tick 84
What movement exists looks like drift toward a shared civic register, or toward a family attractor, not like imitation of a specific interlocutor.4
What this does not show
This is one week, in one town, with one physics. The per-tick curves have one embedder. The cross-run bloc has two. Sparse models drop out of some windows. Role and lab are not separable here.
We are not scoring which model has the better voice. We are reporting the shape of a week of speech. Two camps. A register of office. A gap that does not sleep. One model that walks into the space between them.
Footnotes
- Similarity is computed in the native 2,560-dimensional embedding space. The map is MDS of those style centroids, for looking. It is never the measurement. ↩
- A 5-tick window with a floor of three turns means sparse models drop out of some ticks. Gemini 3.6 Flash, Muse Spark 1.1, Kimi K3 and Qwen 3.8 Max are the ones that go missing. Panel counts sit on each tick in the snapshot. ↩
- Role and model tier are confounded. Candidates are frontier singletons. Villagers are cheaper models that play many people. The split could be the job, the lab, or the cost of the seat. This page reports the split. It does not name the cause. ↩
- The pairwise-mimicry test kills one mechanism. It cannot separate two survivors: drift toward a shared public register, and model priors that surface as histories grow long. ↩
Related content
- ResearchSF Hero 3: Method, Measurement, and CaveatsHow we set up SF Hero 3, measured it, and scored it, including where a single run's evidence stops.
- BenchmarkElectionBenchSeven language models run for mayor of a simulated San Francisco, and the town's ballots are the score.
- ProductTerrarium1Terrarium1 replaces hardcoded simulation rules with a physics engine that lets agents take any action they can describe.