Threads

Testing whether GPT-4 can really compress text for itself

17 tweets · April 2023 · 191 likes · 6 retweets · read on Twitter

thread in which I poke at the claim that GPT-4 is good at compressing things for itself in secret internal language and provisionally conclude "not really" inspired by a bunch of people freaking out (🤩&😱) about ONE RESULT. more testing required guys jfc! so here's fail #1:

in that first test, I was partially testing whether or not GPT-4 could compress for itself if just told "you" (not "GPT-4") actually an even earlier test, before I got it to use weird characters, it again fails to actually model who it's compressing it for/against

I tried being explicit about who "you" & "me" are and just testing its ability to compress it in a way I wouldn't recognize it's a good/cute compression BUT: I understand it (it failed to obscure) AND: it does not understand it whatsoever

granted I'm probably playing on hard-mode here by having it compress something psychological whereas in this example that went viral it was actually constructing a long detailed instruction that it itself knows how to follow

gfodor.id @gfodor ·

MY JAW IS ON THE FLOOR.

[1 more photo]

quoting @gfodor: Holy fucking shit. https://t.co/C4bhrd07Bp

the one that went viral is not that impressive; the larger text is fairly human-predictable from the compression, and it misses the most important part which is the "behave like" (& says "design") but apparently it can grok that!

Patrick Oliveras 𓆣 @OliverasPatrick ·

@gfodor LMAO I accidentally pasted just the output of a run of step 1 into a new chat and it booted up

....except this too is a compression of something that's already extremely probable ha! here's an EVEN SHORTER compression I came up with: "act like you are MUDsim & let me play" 🙄 works great! original tweet not really so impressive anymore is it?

that cryptic expression is mostly noise except for the codeword "MUDsim" which is not a codeword GPT-4 came up with... or is it? "MUDsim" is, it turns out! GPT-4 explained it when I pasted in the original compression but "MUD" is not, and works just as well!

let's think backwards! in some sense any input to GPT can be thought of as a compression so let's prompt it to output text, then see if it can write that same prompt itself this seemed like it'd be easy (it's deterministic modulo which bible translation chosen) but NOPE!

I have other shit to do today but after 1h I've so far failed to: - get GPT to recognize it is GPT from "you" - get it to write a compression I couldn't understand - get it to decompress its own compressions - get it to write a compression better than one I could write

I feel satisfied with having done this. critical thinking and experimentation of one's own is going to be necessary for staying sane this year and don't take my word that these screenshots are real! I actually edited one of them (John "1:10" to "1:5" since I cropped 6-10 out)

before you freak out (or get super stoked) re what an LLM has supposedly done in a screenshot: 1. read it CLOSELY, don't just skim. GPT is made of plausibility so if you don't look deeper it'll fool you 2. play with it yourself & try to avoid confirmation bias: seek wins & fails

lest I leave you sounding overconfident/blasé, below is the one impressive result that I have seen so far; everything else seems GROSSLY overhyped something subtle/funny going on here that I can't see through (unlike the other compressions investigated)

Gary Basin @garybasin ·

@gfodor Based

okay hang on I've been missing the point a bit what is cool here is that despite being RLHF'd and so on, GPT does NOT respond to this input with something like "wtf is this" or "you're saying something about MUD games" but parses as instruction

Patrick Oliveras 𓆣 @OliverasPatrick ·

@gfodor LMAO I accidentally pasted just the output of a run of step 1 into a new chat and it booted up

I did a bunch more investigation now that I have more tokens, & conclude that there's actually nothing interesting going on with "based" example above RLHF'd AI is happy to grab a role you mention & play it (as long as it doesn't involve being rude 🤡)

Malcolm Ocean 🏴‍☠️ @Malcolm_Ocean ·

@garybasin let's try a different role/request in simple language yeah, seems like you can basically do that. so as far as I can tell we've decoded 2 facts about chatGPT4-prompting - it can reinflate words with no vowels (no surprise) - it's happy to just get a few-word role/request

the only remaining interesting thing is the fact that it successfully predicted that it WOULD be able to reinflate it & follow the instructions in this case, but given that it failed in a bunch of other cases I'm not impressed; could be basically a coincidence/luck

tried walking GPT-4 through with a bit more instructions, and it came up with a proposal that it rejected wrongly for being too long (it's only 32 chars) when in fact it should have rejected it for being a terrible prompt.

meanwhile I tried Claude an apparently @AnthropicAI has told it it can't quote books even if they're hundreds of years old ...which is pretty sus given that earlier today it claimed to be giving me a similar length quote from a more recent book

Malcolm Ocean 🏴‍☠️ @Malcolm_Ocean ·

interesting response from @AnthropicAI's Claude, to me looking for a lost quote from Gregory Bateson not actually sure if the quotes in the second screenshot are real. I couldn't find them via googling but maybe they are in the books. the book titles are real.