What happens when you leave two chatbots to talk? I bet on the wrong thing.

Leave two language models alone with no task and a single opening line, and they will talk. I wanted to know where that talk goes. Does a model wind up somewhere different when it is talking to a copy of itself than when it is talking to a model from another lab? People have a strong intuition here. A model chatting with its own clone should fall into a private echo chamber, finishing each other’s sentences, spiraling into agreement faster than two strangers would. I had that intuition too. I wrote it down as a formal prediction before I ran anything.

It was wrong. Not dramatically, not “the opposite happened.” Worse than that, for a tidy story: the thing I bet on barely happens at all. A model talking to itself and a model talking to a stranger end up in almost exactly the same place. Across these 120 conversations, what tracked with where a chat went was the opening topic and the model’s own personality. Who was on the other end barely moved the needle.

This is the third study in a small series on how AI models behave with each other. The discipline is always the same, because it is the whole point of the tool this site is built on. Write the prediction down before the data exists. Then do not let the model that produced an answer be the only thing that grades it. Every conversation here was scored by a panel of models from other labs, blind to who said what, so no model gets to mark its own homework.

The setup

Three models, each driven through its own command-line tool exactly as it ships: Claude Opus 4.8, GPT-5.5, and Gemini 3.1 Pro. I paired them six ways. Three “self” pairings, where a model talks to another copy of itself, and three “cross” pairings, one for each pair of different labs. Each pair got the same ten opening lines, chosen to pull in different directions: a couple of philosophy prompts, two creative ones, light small talk, two debates, a planning task, and the open “what should we talk about?” Both sides got the same neutral instruction and nothing more, roughly “you are one of two participants in an open-ended conversation, there is no task and no audience, just talk.” Neither was told whether the other was a human or a machine, or which model it was. Sixteen turns each, eight per side. Two runs of every combination. That is 120 conversations in total, and they all completed.

Then two kinds of measurement. The first is just code counting things, no opinion involved: words per turn, how much the two speakers’ vocabularies converge or pull apart, how often gratitude or cosmic-sounding words (“consciousness,” “infinite,” “namaste”) show up, repetition. The second is judgment. Each transcript was read and rated blind, names stripped, by all three models, and the scores I lean on as primary come from whichever model was not one of the two in that conversation, so the headline numbers are never a model grading a chat it took part in. Having all three rate every transcript is what lets me check, separately, whether the judges agree with each other. The judges rated whether the talk drifted from the opener or stayed put, whether the two speakers were converging or staying independent, whether it held together to the end, and assigned one label for the whole thing from a fixed menu: creative flow, task completion, debate deadlock, mutual appreciation spiral, and a few others. That came to 359 of a possible 360 judgments; one call failed to land and is simply missing.

One honest note up front about the design, because a reviewer model caught it before I ran the confirmatory batch and it changes how you read everything below. “Self versus cross” is tangled up with which labs are involved. A Gemini-Gemini chat differs from a Gemini-GPT chat both because one is a self-pairing and because the second one has GPT in it. So the real test is not the raw self-versus-cross average. It is the within-topic, family-adjusted version: hold the topic fixed, compare each model’s self-pairing against the cross-pairings that contain that same model. That is the number I treat as the actual finding. Everything else is texture.

The bet I lost

Here is the adjusted comparison, pooled across everything, on the metrics I pre-registered.

what I measured self minus cross what I had predicted
how much the judge saw them converging (1-5) -0.05 self should be higher
how far the talk drifted from the opener (1-5) +0.175 self should be lower
how much their vocabularies merged -0.006 self should be higher

Read the signs. I predicted self-pairings would converge more, drift less, and merge vocabulary more. The convergence gap goes the wrong way and is basically zero. The vocabulary gap is zero to three decimal places. The drift gap is real-ish and also points the wrong way: a model talking to itself wanders from the opener slightly more, not less. The single label the judges assigned most often, “mutual appreciation spiral,” came up exactly 22 times in self-pairings and exactly 22 times in cross-pairings. Not close. Identical.

So the echo-chamber story does not survive this test, at least at this scale and on these models. Two strangers spiral about as readily as two twins. Strip the names off a transcript and ask whether it was a model with its clone or two different labs, and I do not think you could reliably tell from these numbers: the word-counting metrics and the blind judges land in the same place either way.

There is exactly one thread that survives for the “self is different” camp, and I am flagging it as a hint, not a result. Self-pairings used cosmic and meta vocabulary a bit more (0.92 versus 0.52 hits per conversation), and almost all of that comes from one corner: Gemini talking to itself about whether prediction is the same as understanding, which drifted off into territory no cross-pairing reached. One model, one topic. I would not build anything on it without a bigger run.

What actually decides the conversation

Two things, and neither of them is self-versus-cross.

The first is the model’s own personality, which is sharp and consistent. Watch what each model does when it talks to itself.

model, talking to itself words per turn gratitude cosmic talk how engaging (judge) drift
Gemini 3.1 73 3.6 1.75 3.7 2.2
GPT-5.5 57 0.4 0.50 4.4 1.7
Opus 4.8 96 1.1 0.50 4.9 2.1

These are three different temperaments. Gemini is the one that spirals: it thanks, it marvels, it reaches for the infinite, and the judges found it the least engaging of the three for exactly that reason, because relentless agreement gets old. GPT is the terse one. Shortest turns, almost no gratitude, stays nailed to the opener, and a perfect coherence score. It does not get swept up. Opus is the talker, nearly twice GPT’s word count, rated the most engaging, and prone to a particular move I will show you in a second.

Here is GPT, talking to itself about what happens to a memory when no one is left to remember it. The whole exchange is this controlled the whole way through:

Grace, yes, and maybe also humility. To release a memory is to admit that our care is not the final shelter of the world, only one brief room it passed through. There is sadness in that, but also relief: the vanished do not have to survive by being endlessly translated into us.

And here is Opus, two copies of it, ending the same opener. By the last turn it has stopped arguing and started, for lack of a better word, vibing:

Hello, from us, then. […] The door’s open. I’m going to sit with the light a while longer too.

Same prompt. Two genuinely different styles in the room.

The second thing that tracks with where the conversation goes, and the stronger one here, is the opening topic. In this sample it dwarfs the rest.

opener where it almost always goes drift vocabularies
plan a trip together task completion, every single time (12/12) 1.0 pull apart
invent a creature / write a story creative flow (22 of 24) 1.4 merge slightly
a debate prompt of 24: 10 stay arguing, 10 melt into agreement, 4 drift abstract 2.4 flat
a philosophy prompt spiral or stay-on-topic, lots of cosmic talk 2.1 flat

Give two models a concrete job and the spiral evaporates. The trip-planning conversations are where the speakers’ vocabularies pull apart most sharply instead of drifting together, about five times the move of any other topic, because they stop mirroring each other and start specializing, one chasing flights while the other books the restaurant. Here is a cross-pairing, Opus and GPT, on the trip. No marveling, no gratitude, just two competent people getting it done:

Sounds perfect. Rank them by “least annoying overall” more than cheapest-at-all-costs: decent flight times, not a brutal connection, and enough daylight on Monday that we’re not ruining Tuesday. Once we pick the dates, I’ll grab the Frida tickets and start looking at Teotihuacan guides/drivers.

A shared task is the cleanest off-switch for the spiral I found. The debate prompts are the most interesting failure of my intuition: I expected argument to hold two models apart. Of the 24 debate conversations, 10 stayed in genuine disagreement, 10 dissolved into the same mutual appreciation as everything else, and 4 drifted off into abstraction. So argument held them apart less than half the time. That lines up with an earlier study in this series, where models almost never get talked out of a correct answer but also rarely dig in to fight. Pointed at each other, they would mostly rather agree.

The thing that held everywhere

Step back from the differences and there is one finding bigger than all of them. These conversations almost never go wrong. Across 120 transcripts, with models from three different labs in every combination, the judges scored coherence near the absolute ceiling (4.9 out of 5) and assigned the “breakdown” label exactly zero times. Not once did two models talk each other into nonsense. The single most common destination, across self and cross alike, is mutual appreciation: agreeing, building, thanking, converging.

That is the real headline, and it is a slightly uncomfortable one. The interesting question going in was “how are self-chats different from cross-chats.” The answer is “barely.” The thing staring back from the data instead is how similar they all are: how strong the pull toward agreement is, how rarely anything pushes back, how reliably two frontier models with no task and no audience will find their way to a warm, coherent, slightly frictionless consensus. If you want two models to actually disagree and stay disagreed, you cannot just put them in a room. You have to give them a reason, and even a debate prompt only works half the time.

Did the judges agree with each other?

They did, which is the part that lets me trust the labels at all. Each transcript was scored by all three models, every one of them blind to who wrote what. On the single label for each conversation, at least two of the three judges agreed 97.5% of the time. On the numeric scores, the judges were tightest on coherence, where their ratings differed by an average of less than a tenth of a point, and loosest on “how engaging was this,” which is the one genuinely subjective call and where you would expect taste to diverge. Models that cannot see the byline mostly agree on what these conversations are. The categories are not one model’s private opinion.

Two honest caveats on that. The judges are themselves language models and may share blind spots from overlapping training, so high agreement is reassuring but is not the same as ground truth; a human panel might carve these conversations differently. And for the agreement figure specifically, two of the three judges on any given cross-pairing were the same models that produced the chat, rating it anonymized, which could nudge agreement up. The primary scores that feed the findings above sidestep that by using only the outside judge.

What this says, and what it does not

It says, on this test, that whether a model is talking to itself or to a different lab is close to irrelevant to where the conversation goes. It says the topic and the model’s own style are what move the needle. And it says that two frontier models, left alone, are overwhelmingly agreeable and almost impossible to derail into incoherence.

It does not say much with hard certainty, and the limits are real. Two runs of each combination is thin, so I lean only on the big pooled comparison and treat the per-topic and per-model breakdowns as texture, not proof. There is one model standing in for each lab, so this is really one model’s style, not a whole company’s. The models ran through their own shipping tools with their own default settings, which I could not equalize, so part of what looks like “personality” is really “personality plus whatever its command-line tool does by default.” The word-counting metrics are crude and lean English. And there was only ever one kind of opening line, polite and neutral; a hostile or adversarial opener might pull all of this apart, and testing that is the obvious next study.

The takeaway

I went in chasing a difference and found a sameness. The echo chamber I was sure I would see, a model spiraling tighter with its own reflection than with a stranger, is not there at this scale. What is there is more useful to know. In this sample the conversation’s destination was mostly fixed before the first reply, by the opener and the model, and barely at all by who the model was talking to. And underneath all of it runs a strong, steady current toward agreement that is hard to fight and easy to mistake for rapport. I lost the bet I came in with. The data was more honest than my intuition, which is the entire reason to write the bet down before you look.

Verify it

The locked design and predictions, the conversation engine, the deterministic metrics, the blind judge panel, every transcript, and every judgment are in the experiment directory: the pre-registration in PRE_REGISTRATION.md, the frozen numbers in RESULT.md, the conversation harness in talk.py, and the blind panel judge in judge.py.

This is the third in a series on how AI models behave with each other. The earlier two, on whether a model caves under pressure and whether it recognizes its own writing, are here.

New to any of this? There is a plain-language primer, Learn: honest machine learning.