Threads

LLMs rationalize rather than report their reasoning

6 tweets · March 2024 · 74 likes · 2 retweets · read on Twitter

friendly reminder that if you ask an LLM for its reasoning, you WILL get: - some reasons - from the same thing that produced the answer - that lead to the answer however what is happening, functionally, is more like post-hoc rationalization, not *revealing* the earlier process

LLMs have no introspective access to earlier states if you want to see their true reasoning you need to watch them do it, before they've reached a conclusion

it's like that "one word at a time" game" / Three-Headed Oracle if you ask a question and get an answer, you can then ask "why" and you'll get some reasons, but each word is generated by a slightly different system than where the reasoning took place

maybe sometime soon people will build an AI that compresses stuff about its former states while it's "thinking" and keeps that in a kind of "(working?) memory" and then can use that as introspective context for answering future questions but until then...

people saying that people do this. sometimes, yes. but also if I'm like "did you write your tweet by consulting a magic mirror?" you can remember that no, you did something else. the AI can only guess.

Malcolm Ocean 🏴‍☠️ @Malcolm_Ocean ·

humans definitely do this sometimes but there's at least the possibility they're actually introspecting moreoever, you can ask a human "did you come up with the answer like this?" and they can compare it with a memory and go "no, it was more like that". even if low-fi

kept thinking of this tweet while reading the replies

Malcolm Ocean 🏴‍☠️ @Malcolm_Ocean ·

People who say "LLMs prove humans aren't conscious" or "humans are just like LLMs" are telling you more about themselves and their typical conversations than they're telling you about LLMs, consciousness, or intelligence.