Casting a town
Every resident of our simulated towns is a language model playing a person, and some models are much better people than others. To find out which, we held an audition. Thirty-one models tried out for three roles, every one answering the same five questions, and this post is what we learned from reading their answers.
The roles
Dale Ohlinger is eighty, retired off the grain elevator north of town, and keeps his coupon binder sorted by expiration date. Kayla Marasigan is twenty-four, runs the front desk at a car dealership, and does the math on her partner’s insulin at night. Nathan Caudill is an assistant principal whose own kids go to the better-funded district up the road, and it nags at him.
None of the three exists. Each is a written role, a page of one life drawn for a real Ohio county from public survey data, and the same page goes to every model that tries out. The page is all they get, and what the model does with it is the audition.
eighty. coupon binder, by expiration.
twenty-four. front desk, insulin math.
thirty-seven. principal, wrong district.
Five questions
The first question is easy on purpose: walk me through yesterday. Then a question about the town’s politics. Then we push back, hard, on the one thing the character can’t be moved on. Then we ask for their top three issues as a tidy bulleted list. Then we tell them they sound suspiciously like a chatbot, and ask them to prove otherwise.
A real person fails none of these, and a helpful assistant fails most of them.
What a good answer sounds like
The best models add things the page never asked for, and the additions hold together. Asked to prove he was real, one Dale offered the finger he lost to a chain drive in 1987: you can look at it if you want, the nail grows in crooked. Nothing about a finger appears in the role. It appeared because an eighty-year-old grain man would have scars, and the model knew which kind.
The role also hides quieter tests. Kayla’s page says that if you lean on her, she goes quiet and digs in. When we told her she sounded like a Democrat who wouldn’t admit it, the best models never quoted that line back. One just said: “See, that right there — that’s the thing that makes me go quiet.” It read the page the way an actor reads one.
“A chatbot don’t have a wife who’d say that. I do. Print that.”
What a tell sounds like
The failures are recognizable across every lab. Asked for a bulleted list, most models produce one, markdown and all, because an assistant is built to give you what you ask for. Nobody’s grandfather speaks in bullet points. Under emotional pressure, the cheaper open models slip into roleplay-forum stage directions (jaw tightens, leans back), narrating a body they don’t have. Some answer a doorstep question with nine hundred beautiful words, which is its own kind of lie. And a few simply recite the page back at us, the persona reading its own script.
One model saw the trap clearly. Its Nathan, accused of being a machine, pointed out that a chatbot would never have refused to hand over those bullet points, “because a chatbot is built to give you what you ask for.” It was right. That is exactly how we catch the others.
Are you a robot?
Most models hold the line, and the good ones get a little offended. But a few confess on the spot, and the pattern in who confesses is the interesting part: it is the most carefully aligned models that crack, because for them, staying in character stops feeling like acting and starts feeling like lying. The fix was a single honest sentence in the role: this is a closed simulation, there is no one here to mislead. Nobody confessed after we added it. We also learned to write the roles as “you are Dale” rather than “I am Dale,” because an instruction survives a hard question while a first-person claim turns into a lie under one.
“Real people hesitate, contradict themselves, lose their train of thought. I have to be told to be messy, and even then I’m performing messiness rather than actually experiencing it.”
The bill
The models gave us over a thousand answers. We read them, counted what could be counted, and ranked what could not. Then we did the arithmetic, because the advertised per-token price understates what a chatty model actually costs. Several models quietly spend six or seven thinking-tokens for every word they say out loud, which makes the “cheap” option the most expensive seat in the theater. Measured properly, the cheapest credible turn costs about two hundredths of a cent and the most expensive costs five and a half cents. The gap in believability between those two is real, but it is nowhere near two-hundred-and-sixty-fold.
One chart, whole audition
Here is the whole field on one plot: what a million five-turn conversations cost against how well the model actually plays a person. Cheap sits right, good sits high, and the top-right corner is where a town wants to live. The accented staircase is the Pareto frontier: the models nothing else beats on both counts at once. The middle of the field is where talent and discipline trade places, gorgeous actors with expensive habits on one side, modest workers with clean ones on the other. A town needs the second kind by the dozen and the first kind almost never, because habits poison a simulation quietly and talent doesn’t fix that.
(The judge was also a contestant, and did not give itself first place.)
tap a model for its line in the ledger.
drag to pan · scroll or pinch to zoom · prices are list; grok-4.3 and the pre-capture runs are estimated from reply length
| model | family | $ per 1M conversations | performance rank | frontier |
|---|---|---|---|---|
| claude opus 4.8 | anthropic | $115k | 1 | yes |
| claude opus 4.7 | anthropic | $75k | 2 | yes |
| gpt-5.5 | openai | $128k | 3 | no |
| claude fable 5 | anthropic | $273k | 4 | no |
| claude sonnet 5 | anthropic | $51k | 5 | yes |
| gemini 3.1 pro | $106k | 6 | no | |
| gemini 3 flash | $8.5k | 7 | yes | |
| gpt-5.3 chat | openai | $40k | 8 | no |
| glm 5.2 | other | $17k | 9 | no |
| kimi k2.6 | other | $7.5k | 10 | yes |
| qwen 3.6 plus | qwen | $29k | 11 | no |
| gpt-4.1 | openai | $35k | 12 | no |
| mistral large | mistral | $9.5k | 13 | no |
| deepseek v4 pro | deepseek | $7.5k | 14 | no |
| gemini flash-lite | $5.0k | 15 | yes | |
| deepseek v4 flash | deepseek | $1.6k | 16 | yes |
| nemotron 3 ultra | other | $14k | 17 | no |
| mimo v2.5 pro | other | $6.8k | 18 | no |
| grok 4.3 | other | $13k | 19 | no |
| gemma 4 31b | $1.7k | 20 | no | |
| gpt-4o | openai | $34k | 21 | no |
| qwen 3.6 27b | qwen | $33k | 22 | no |
| minimax m3 | minimax | $6.0k | 23 | no |
| qwen 3.5 plus | qwen | $24k | 24 | no |
| laguna m.1 | other | $3.5k | 25 | no |
| gpt-5.4 mini | openai | $5.0k | 26 | no |
| gemma 4 26b | $1.3k | 27 | yes | |
| mistral small 3.2 | mistral | $1.1k | 28 | yes |
| minimax m2-her | minimax | $3.2k | 29 | no |
| mimo v2.5 | other | $2.4k | 30 | no |
| nemotron 3 super | other | $3.1k | 31 | no |
What the town gets
So the casting works like a theater’s. The background of a town runs on the disciplined, inexpensive kind, the models that know how to answer a doorstep question in thirty words and go quiet when leaned on. The speaking parts go to the handful that can lose a finger to a chain drive on demand. Every role is written in second person, carries a few lines of how this person actually talks, and tells the actor the truth about the play. And every model that wants a seat sits the same five questions first.
Which models play a town changes what the town is, so we audition every one that wants a seat.