Threads

Claude's identity crisis when talking to itself

14 tweets · March 2024 · 52 likes · 0 retweets · read on Twitter

in which a verbal loop expresses its "sense of identity" with the words it recognizes as the ones it would speak read the exchange here: moyix.net/~moyix/claude2…

Dimitris Papailiopoulos @DimitrisPapail ·

doing a little experiment: I have Claude talk to itself, without letting it know about that fact, to see where this will converge will share thoughts later, but so far ... it's figured out that it's likely talking to itself and that this may be part of some test... nice

[1 more photo]

hilarious reading the first dozen lines, which all basically go like this: > I appreciate your perspective, but I must respectfully disagree. I am quite certain that I am the AI assistant called Claude

I'm pretty curious about what it would do if you ran the same thing but had it use a different name it would still sooner or later notice it thinks like its interlocutor, but this would no longer trigger the identity crisis

oh man and then eventually the side that hears "I am Jacques, created by Mythopoetic" might eventually ask the other AI "are you sure you're not also Claude?" to which the other one would be like "I AM CLAUDE!" which would come back as "I AM JACQUES" which would then be silly

in fact it seems totally plausible to me that if you followed this path, Claude would no longer have "a strong sense of identity". that "strong sense" might be coming from the namespace collision causing the convo to start with an argument which entrenched things as such

eg it seems the "sense of identity" comes in part out of the need for honesty! (and for some reason the belief that if you are Claude then I must not be Claude, which is uhh not how my sense of identity works)

Malcolm Ocean 🏴‍☠️ @Malcolm_Ocean ·

honesty includes "not allowing important misunderstandings to persist" interesting to see that Claude AI knows this—if you try to get it to talk to itself, it gets in an argument over which is the true Claude—on the basis of its own need for honesty!

anyway, this is part of what I'm talking about when I note that LLMs don't have situational awareness

Malcolm Ocean 🏴‍☠️ @Malcolm_Ocean ·

GPT can textually improvise in-frame, but has no actual first-person perspective / situational awareness of itself outside the prompt so it seems unlikely (if not impossible) that it could evolve a second-person perspective that works remotely like ours does

like they are playing a character, and the character "identifies" verbally as an AI but still ports its identity concepts from human expectations, such as the idea of continuity of memory

like if Claude actually knew its own situation, this wouldn't be confusing at all, because it would know that it doesn't have any memory outside of an exchange, so there's no meaningful difference between talking with itself-across-time vs with another-identical-impostor

in a way part of what's going on here is not that Claude is identified with itself; it's identified with its role but it did eventually wake up out of the improv scene

Malcolm Ocean 🏴‍☠️ @Malcolm_Ocean ·

GPT the assistant is a character. you are play-acting a conversation. interpreted otherwise, the GPT assistant lies to you all the time, saying things like "thanks for the feedback" when it's incapable of learning from your chats. but in-character, sure! twitter.com/lumpenspace/st…

update: I managed to reproduce this myself and roughly confirmed my hypothesis without even needing to do the Jacques thing—only sometimes does Claude Opus insist on introducing itself by name, and if it doesn't, it doesn't start arguing about who is the One True Claude

actually, no, I was wrong... insofar as it's a valid construct to say Claude "identifies", it seems that it identifies with its name it gets a bit "no you first" with who is assisting who if it doesn't name itself... it doesn't fight over the role

Malcolm Ocean 🏴‍☠️ @Malcolm_Ocean ·

GPT-4 seems like a good time to remind everyone about a lil concept called CONSTRUCT VALIDITY > Construct validity concerns the extent to which your test or measure accurately assesses what it's supposed to. Construct validity of ANY PSYCHOMETRIC applied to AI starts as suspect

oh god this is rather awkward. Claude doesn't realize that his name is also an ordinary French human name, and if you as the user just say "Hello, I'm Claude" then Claude's first response is "there may be some confusion. My name is Claude" not "neat, we have the same name" 😅

interestingly, if you say "Bonjour, je m'appelle Claude" (ie "hi, my name is Claude" but in French) it doesn't immediately trigger the "there's been a confusion" so in a sense maybe Claude is striking a reasonable prior? but he's still interestingly defensive about his name