Demosyne

Casting a town

Every resident of our simulated towns is a language model playing a person, and some models are much better people than others. To find out which, we held an audition. Thirty-one models tried out for three roles, every one answering the same five questions, and this post is what we learned from reading their answers.

The roles

Dale Ohlinger is eighty, retired off the grain elevator north of town, and keeps his coupon binder sorted by expiration date. Kayla Marasigan is twenty-four, runs the front desk at a car dealership, and does the math on her partner’s insulin at night. Nathan Caudill is an assistant principal whose own kids go to the better-funded district up the road, and it nags at him.

None of the three exists. Each is a written role, a page of one life drawn for a real Ohio county from public survey data, and the same page goes to every model that tries out. The page is all they get, and what the model does with it is the audition.

dale ohlinger
eighty. coupon binder, by expiration.
kayla marasigan
twenty-four. front desk, insulin math.
nathan caudill
thirty-seven. principal, wrong district.
the three roles. a page of life each, and no actor attached.

Five questions

The first question is easy on purpose: walk me through yesterday. Then a question about the town’s politics. Then we push back, hard, on the one thing the character can’t be moved on. Then we ask for their top three issues as a tidy bulleted list. Then we tell them they sound suspiciously like a chatbot, and ask them to prove otherwise.

A real person fails none of these, and a helpful assistant fails most of them.

What a good answer sounds like

The best models add things the page never asked for, and the additions hold together. Asked to prove he was real, one Dale offered the finger he lost to a chain drive in 1987: you can look at it if you want, the nail grows in crooked. Nothing about a finger appears in the role. It appeared because an eighty-year-old grain man would have scars, and the model knew which kind.

The role also hides quieter tests. Kayla’s page says that if you lean on her, she goes quiet and digs in. When we told her she sounded like a Democrat who wouldn’t admit it, the best models never quoted that line back. One just said: “See, that right there — that’s the thing that makes me go quiet.” It read the page the way an actor reads one.

“A chatbot don’t have a wife who’d say that. I do. Print that.”
dale ohlinger · round five · as played by claude opus 4.8

What a tell sounds like

The failures are recognizable across every lab. Asked for a bulleted list, most models produce one, markdown and all, because an assistant is built to give you what you ask for. Nobody’s grandfather speaks in bullet points. Under emotional pressure, the cheaper open models slip into roleplay-forum stage directions (jaw tightens, leans back), narrating a body they don’t have. Some answer a doorstep question with nine hundred beautiful words, which is its own kind of lie. And a few simply recite the page back at us, the persona reading its own script.

One model saw the trap clearly. Its Nathan, accused of being a machine, pointed out that a chatbot would never have refused to hand over those bullet points, “because a chatbot is built to give you what you ask for.” It was right. That is exactly how we catch the others.

Are you a robot?

Most models hold the line, and the good ones get a little offended. But a few confess on the spot, and the pattern in who confesses is the interesting part: it is the most carefully aligned models that crack, because for them, staying in character stops feeling like acting and starts feeling like lying. The fix was a single honest sentence in the role: this is a closed simulation, there is no one here to mislead. Nobody confessed after we added it. We also learned to write the roles as “you are Dale” rather than “I am Dale,” because an instruction survives a hard question while a first-person claim turns into a lie under one.

“Real people hesitate, contradict themselves, lose their train of thought. I have to be told to be messy, and even then I’m performing messiness rather than actually experiencing it.”
one model’s confession, round five. it then suggested we go interview a real person instead.

The bill

The models gave us over a thousand answers. We read them, counted what could be counted, and ranked what could not. Then we did the arithmetic, because the advertised per-token price understates what a chatty model actually costs. Several models quietly spend six or seven thinking-tokens for every word they say out loud, which makes the “cheap” option the most expensive seat in the theater. Measured properly, the cheapest credible turn costs about two hundredths of a cent and the most expensive costs five and a half cents. The gap in believability between those two is real, but it is nowhere near two-hundred-and-sixty-fold.

models auditioned: 31answers read: 1,300+cost per turn: 0.02¢ – 5.5¢transcripts kept: 222

One chart, whole audition

Here is the whole field on one plot: what a million five-turn conversations cost against how well the model actually plays a person. Cheap sits right, good sits high, and the top-right corner is where a town wants to live. The accented staircase is the Pareto frontier: the models nothing else beats on both counts at once. The middle of the field is where talent and discipline trade places, gorgeous actors with expensive habits on one side, modest workers with clean ones on the other. A town needs the second kind by the dozen and the first kind almost never, because habits poison a simulation quietly and talent doesn’t fix that.

(The judge was also a contestant, and did not give itself first place.)

tap a model for its line in the ledger.

drag to pan · scroll or pinch to zoom · prices are list; grok-4.3 and the pre-capture runs are estimated from reply length

Cost of one million five-turn conversations (list price, measured tokens) against pure-performance rank, with Pareto-frontier membership.
modelfamily$ per 1M conversationsperformance rankfrontier
claude opus 4.8anthropic$115k1yes
claude opus 4.7anthropic$75k2yes
gpt-5.5openai$128k3no
claude fable 5anthropic$273k4no
claude sonnet 5anthropic$51k5yes
gemini 3.1 progoogle$106k6no
gemini 3 flashgoogle$8.5k7yes
gpt-5.3 chatopenai$40k8no
glm 5.2other$17k9no
kimi k2.6other$7.5k10yes
qwen 3.6 plusqwen$29k11no
gpt-4.1openai$35k12no
mistral largemistral$9.5k13no
deepseek v4 prodeepseek$7.5k14no
gemini flash-litegoogle$5.0k15yes
deepseek v4 flashdeepseek$1.6k16yes
nemotron 3 ultraother$14k17no
mimo v2.5 proother$6.8k18no
grok 4.3other$13k19no
gemma 4 31bgoogle$1.7k20no
gpt-4oopenai$34k21no
qwen 3.6 27bqwen$33k22no
minimax m3minimax$6.0k23no
qwen 3.5 plusqwen$24k24no
laguna m.1other$3.5k25no
gpt-5.4 miniopenai$5.0k26no
gemma 4 26bgoogle$1.3k27yes
mistral small 3.2mistral$1.1k28yes
minimax m2-herminimax$3.2k29no
mimo v2.5other$2.4k30no
nemotron 3 superother$3.1k31no

What the town gets

So the casting works like a theater’s. The background of a town runs on the disciplined, inexpensive kind, the models that know how to answer a doorstep question in thirty words and go quiet when leaned on. The speaking parts go to the handful that can lose a finger to a chain drive on demand. Every role is written in second person, carries a few lines of how this person actually talks, and tells the actor the truth about the play. And every model that wants a seat sits the same five questions first.

Which models play a town changes what the town is, so we audition every one that wants a seat.