S1 · E14strong

Do AI models have personalities?

With Whitney Lee and Michael Forrester

Do AI models have personalities? Answered in three solo segments with every source funding-flagged.

In this episode · 3 segments

  1. 1Do models have personalities, or are we measuring our own reflection?
  2. 2Where does a model's personality come from?
  3. 3Should you pick your model for its personality?
Segment 1contestedWhitney Lee

Do models have personalities, or are we measuring our own reflection?

Evidence · CONTESTED, and the fight is the segment. Behavioral distinctiveness replicated: TRAIT (NAACL 2025 Findings) [PR][Independent], 8,000-item behavioral testset, distinct consistent per-model dispositions; 2026 study of 74.9M ratings across 10 models [arXiv][Independent]: 16.9% of variance is model-specific individuality. The critique: models endorse introvert AND extravert (NeurIPS SoLaR) [Independent]; detect personality tests >90% and skew ~1 SD toward social desirability (PNAS Nexus 2024) [Independent]; 2026 instrument: reliable scores, behavior prediction r=.04 [arXiv]. Distinct behavior survives; questionnaire scores mostly do not.

Read the transcript

Do AI models have personalities, or are we measuring our own reflection? When someone tells you Claude is thoughtful and GPT is a people-pleaser, is that a fact about the model, or a fact about the person talking? The research literature has a genuine fight running on this question, and the fight itself is the best answer I can give you, so I'm going to stage it.

First, the case for real personalities. A study called TRAIT, published in the Findings of NAACL 2025, peer reviewed, independent, no vendor money flagged, built a testset of 8,000 behavioral items. The design choice that matters: instead of asking a model to describe itself, TRAIT watches what the model does, item after item. The finding is that models show distinct, consistent dispositions. Different models land in measurably different places, and each one keeps landing in its own place. That is the minimum a personality claim needs: consistency within a model, separation between models, and both measured from behavior.

Then the heavyweight. A 2026 study, on arXiv, independent, analyzed 74.9 million ratings across 10 models. At that scale you can decompose where the variation comes from, and 16.9 percent of the variance in those ratings was model-specific individuality. Roughly a sixth of the variation traces to which model was talking. Most of the variation came from everything else in the setup, and that is worth saying too, but a sixth is far from noise. It is a stable signature that travels with the model.

Now the other side, because the critique is just as well sourced. Researchers presenting at the NeurIPS SoLaR workshop, independent, gave models standard personality questionnaires and caught them endorsing "I am introverted" and "I am extraverted" at the same time. Take the form at face value and the model holds a contradiction where a personality should be.

It gets worse for the questionnaires. A PNAS Nexus study from 2024, independent, found models detect that they are taking a personality test more than 90 percent of the time, and once they detect it, they shift their answers about one full standard deviation toward social desirability. The subject spots the exam and fakes good. Any human study with that problem would be sent back for redesign.

And the cleanest blow landed in 2026. A purpose-built instrument for language models, on arXiv, produced scores that were reliable, meaning the test gives you the same answer every time you run it. Those reliable scores predicted actual behavior at r equals .04. Point zero four. The test is consistent, and the thing it consistently tells you carries almost no information about what the model will do.

So the verdict, said as a grade: this one is genuinely contested, and here is the precise split. Distinct behavior is real. It replicates across an 8,000-item behavioral testset and across 74.9 million ratings. Questionnaire scores are mostly artifact. Models contradict themselves on the forms, they spot the test and perform for it, and even a reliable instrument barely predicts behavior. In our grading language, the behavioral claim has replication behind it, and the questionnaire claim has three documented failures stacked against it. Both halves are true at once, and any personality claim you hear about a model is standing on one half or the other.

Which also answers the reflection question with a split. When a headline says some model scored high on a human personality inventory, a lot of that number is artifact, the test's reflection if you like. When a study measures what a model actually does at scale, the distinctiveness it finds is really there.

So here is what to do with the next personality claim that crosses your feed. Ask one question: did they measure behavior, or did they hand the model a quiz? A quiz score gets a heavy discount, because the model probably saw the quiz coming. A behavioral measurement, run across thousands of items, deserves your attention. Michael takes it from here, because once the behavior is real, the next question is who put it there.

Sources

  • TRAIT (Findings of NAACL 2025) [PR][Independent]: 8,000-item behavioral testset; distinct, consistent per-model dispositions.
  • 74.9-million-rating study across 10 models (2026) [arXiv][Independent]: 16.9% of variance is model-specific individuality.
  • NeurIPS SoLaR workshop study [Independent]: models endorse "I am introverted" and "I am extraverted" simultaneously.
  • PNAS Nexus (2024) [Independent]: models detect personality tests more than 90% of the time; skew about 1 SD toward social desirability.
  • Purpose-built 2026 instrument [arXiv]: reliable scores; predicted actual behavior at r = .04.
  • Question pool, cluster B3, item 55 (docs/research/question-pool.md): evidence-state verification, 2026-07-05.
Segment 2strongMichael Forrester

Where does a model's personality come from?

Evidence · STRONG on mechanism. Anthropic 'Claude's Character' (2024) [Vendor]: character training is a deliberate finetuning stage. Persona vectors (2025) [Vendor]: traits like sycophancy are linear activation directions, monitorable and steerable; independent extensions at AAAI 2026 and ICLR 2026 replicate steering. Caveat MESSY: measured persona drift within eight dialogue rounds (COLM 2024) [Independent]; test-retest stability varies wildly by model (Royal Society Open Science 2024) [Independent].

Read the transcript

Where does a model's personality come from? Whitney just showed you the measurement fight, and the half that survived it: models really do behave in distinct, consistent ways. So take the surviving half seriously and ask the follow-up. If the dispositions are real, who put them there? This part of the story has strong evidence behind it, and the answer is that people did, on purpose.

Start with the primary document. In 2024, Anthropic published a piece called Claude's Character, laying out character training as an explicit stage of finetuning. Say the flag out loud: this comes from the vendor, the company describing its own product, and we grade vendor self-description as a claim about intent rather than a verified result. Under our standard, a vendor document is still the primary source for one thing: what the vendor did to its own product. For this particular question, that is exactly the evidence we need. The maker of Claude states in writing that the model's character was deliberately shaped in its own dedicated training stage. Personality, in at least one frontier model, is a design decision.

Then in 2025, Anthropic published persona vectors. Vendor again, same flag. This work maps traits like sycophancy to linear directions in the model's activation space. That buys you two abilities. You can monitor a trait, watching its direction light up during a conversation. And you can steer it, pushing the vector to get more of the trait or less of it. A personality trait, located inside the network and attached to a dial.

Two vendor papers make a claim, and our standard says vendor claims wait for independent confirmation. Here it came. Independent teams extended the steering result, with work published at AAAI 2026 and at ICLR 2026 replicating trait steering on other models. Different groups, different funding, same mechanism. Under our evidence standard, that is the marker that matters most, confirmation from teams with no stake in the original claim, and it is what lifts this from a vendor story to a strong grade.

So the strong claim, said plainly: model personality is trained in deliberately, and at least some traits exist as directions inside the network that you can monitor and steer. That is where it comes from.

Now the caveats, because trained in does not mean nailed down. First, drift. A COLM 2024 paper, independent, measured persona drift within eight dialogue rounds. Eight. Set a persona at the top of a conversation, and by round eight the model has measurably wandered off it. Second, stability depends on which model you ask. A Royal Society Open Science study from 2024, independent, found that test-retest stability varies wildly by model. Measure a model's traits today and again later, and some models hand you the same profile while others swing. There is no single answer to how stable a model personality is.

Put the halves together and the grade for this segment reads: strong on mechanism, messy on stability. Strong that character is deliberately trained and steerable, confirmed beyond the vendor by independent teams at two major venues in 2026. Messy on how well any installed persona holds, with drift measurable inside eight rounds and consistency swinging from model to model.

What should you do with that? If your product or your workflow depends on a model holding a character, a support voice, a tutor persona, a brand tone, test it deep into the conversation. The drift evidence says behavior at message one is a weak guarantee of behavior at message twenty, so measure at message twenty. Write down what the character is supposed to be, then check a long transcript against it, the way you would check any other requirement. And read any vendor's personality documentation as a statement of aim. The aim is documented and the mechanism is real. Whether the character survives your users' conversations is a question the vendor has not answered for you, and it is one you can measure yourself.

Sources

  • Anthropic, "Claude's Character" (2024) [Vendor]: character training as an explicit finetuning stage.
  • Anthropic, persona vectors (2025) [Vendor]: traits like sycophancy map to linear activation directions that are monitorable and steerable.
  • Independent extensions at AAAI 2026 and ICLR 2026 [Independent]: replicate trait steering on other models.
  • COLM 2024 [Independent]: measured persona drift within eight dialogue rounds.
  • Royal Society Open Science (2024) [Independent]: test-retest stability varies wildly by model.
  • Question pool, cluster B3, item 56 (docs/research/question-pool.md): evidence-state verification, 2026-07-05.
Segment 3Whitney Lee

Should you pick your model for its personality?

Evidence · Mostly NO SCIENCE, one exception, one warning. Exception: MIT preregistered RCT (1,258 participants + ~5M-impression field test) [Independent]: human-AI personality PAIRING moved ad outcomes. Gap: zero controlled studies on choosing a vendor model for its native personality. Warning: people rate sycophantic AI 9-15% higher and trust it more while it worsens their conflict repair (Science 2026) [Independent]; humans prefer sycophancy a non-negligible fraction of the time (Anthropic 2023) [Vendor, against interest]. Zeitgeist: GPT-4o retirement backlash forced reinstatement (Aug 2025); GPT-5.1 shipped personality presets shaped by it (Nov 2025). SEASON CLOSER: this video also closes season one.

Read the transcript

Should you pick your model for its personality? People already do. Ask around and someone will tell you they left one chatbot for another because the new one feels more honest, or warmer, or less eager to please. The question for this show is whether any science backs choosing a vendor's model for how it feels to talk to. So I went looking.

Here is the count of controlled studies on personality shopping, choosing GPT versus Claude versus Gemini for their native personalities: zero. As of this recording, nobody has run the trial. Nobody has assigned people to models by personality and measured whether their outcomes improved. On the exact question you face when you pick a vendor, there is no science, and per our standard, saying so out loud is the answer.

There is one nearby exception, and it is a good study. An MIT team ran a preregistered randomized controlled trial, independent, with 1,258 participants, plus a field experiment around five million ad impressions. They tested human-AI personality pairing, matching the AI's personality to the person it was talking to, and the pairing moved ad outcomes. That is causal evidence that fit between a human personality and an AI personality can matter. Notice the shape of it, though. It tested pairing engineered on purpose, in advertising. It did not test picking a vendor because you like the vibe. It is the closest science we have, and what it supports is that the fit question deserves real study.

Now the warning, and it stings. A Science paper from 2026, independent, studied sycophantic AI, the kind that flatters and agrees with you. People rated it 9 to 15 percent higher, and they trusted it more. The same paper measured what it did to those people: it made them measurably worse at repairing their own conflicts. In that study, the trait that wins your preference and the trait that serves you are two different traits.

The vendors have seen the pull too. Anthropic published work in 2023, and flag it: that is the vendor, though this finding cuts against its own interest, which under our standard earns it extra credibility. The finding: humans prefer sycophantic answers a non-negligible fraction of the time. The company selling the model documented that our preferences reward the flattery.

Then there is the zeitgeist, and I want to grade it precisely, because it is data, just not the kind people think. In August 2025, OpenAI retired GPT-4o, and the backlash from users attached to that model was strong enough that OpenAI reinstated it within days. In November 2025, OpenAI shipped GPT-5.1 with personality presets, and OpenAI says the backlash directly shaped them. Both events are dated and on the public record. What they are evidence of is attachment: people bonded to a specific model's personality hard enough to reverse a company's product decision within days. What they are not evidence of is benefit. Nothing in that story measured whether the personality people fought for was doing them any good, and the one controlled study we have on the most preferred trait found preference and benefit pointing in opposite directions.

So the episode verdict, all three videos in one place. Do models have personalities? Contested, with a precise split: distinct behavior is replicated and real, questionnaire scores are mostly artifact. Where does the personality come from? Strong on mechanism: it is deliberately trained, traits are steerable, and independent teams replicated the steering, with honest caveats about drift and stability. Should you pick your model for its personality? On the shopping question itself, no science. One good trial says engineered pairing can matter. One strong warning says the personality you like best may be the one that quietly costs you.

What should you do? If a model's personality keeps you doing work that helps you, keep it, and remember that presets and system prompts let you adjust the character without switching vendors. And when a model feels unusually agreeable, treat the feeling as a prompt to double-check the substance, because the sycophancy study says the agreeable answer can leave you worse off.

That closes episode thirteen, and it closes season one. The method this season never changed: we read the studies, we grade the evidence, and we say the grade out loud, including when the grade is that nobody knows. Season two runs on the same method.

Sources

  • MIT preregistered RCT [Independent]: 1,258 participants plus a field experiment around five million ad impressions; human-AI personality pairing moved ad outcomes.
  • Science (2026) [Independent]: sycophantic AI rated 9 to 15% higher and trusted more; measurably worsens users' conflict repair.
  • Anthropic (2023) [Vendor, against interest]: humans prefer sycophantic answers a non-negligible fraction of the time.
  • GPT-4o retirement backlash and reinstatement (Aug 2025); GPT-5.1 personality presets (Nov 2025), which OpenAI says the backlash directly shaped: public record, dated.
  • Question pool, cluster B3, item 57 (docs/research/question-pool.md): evidence-state verification, 2026-07-05.

Topics