Threads

Construct validity and what AI benchmarks actually measure

25 tweets · March 2023 · 176 likes · 7 retweets · read on Twitter

GPT-4 seems like a good time to remind everyone about a lil concept called CONSTRUCT VALIDITY > Construct validity concerns the extent to which your test or measure accurately assesses what it's supposed to. Construct validity of ANY PSYCHOMETRIC applied to AI starts as suspect

To be clear, construct validity of any psychometric starts out as suspect for humans as well. Oh, you've written a survey where people fill in how happy they are. Does that correlate with the thing we care about? You can't take that for granted!

🌷🐰 sonya serendipitously 🐇🎀 @sonyasupposedly ·

"there is no ancient human practice or context in which people anonymously tell the pure, innocent truth with language, in response to questioning, with no thought for the motives of the questioner or the effect of their answers." carcinisation.com/2020/12/11/sur… by @literalbanana

And it's not straightforward to answer! At best, you can compare it to some other test, that is maybe more costly to run (a 1h interview, transcribed and scored by multiple trained people, vs a 10 Q scale survey). Or correlate with something else—sweat response, for stress?

Do IQ tests measure intelligence? Is that true by definition? Well, we use the term "intelligence" in many different ways, but even aside from that, you can't separate how motivated someone is to even TRY to get a good score on the test!

Malcolm Ocean 🏴‍☠️ @Malcolm_Ocean ·

@SpencrGreenberg the obvious thing imo is investigating the extent to which IQ tests mostly measure "The Desire To Pass Tests" or some property downstream thereof. not sure how you'd test it though, since that seems to induce paradox hotelconcierge.tumblr.com/post/113360634…

So let's consider Theory of Mind. It has some fairly well-established tests for humans, although even these are sus—recent clever studies on infants have shown signs of ToM earlier than they could even FAIL a verbal test

Malcolm Ocean 🏴‍☠️ @Malcolm_Ocean ·

ICYMI: this "change in location" test is used to measure Theory of Mind in kids. iirc most kids fail it at age 4, and pass by age 7. the range might be tighter, but knowing the exact value isn't important since there are signs the science has holes. still interesting to consider

And so even well-established tests for humans like so: GPT-4 can do well on "theory of mind" tests made for humans WHO KNOWS HOW IT SOLVES THEM? we know it's awfully good at putting narrative fragments together to produce a plausible next paragraph... 🥸

Henry Shevlin @dioscuri ·

I continue to be blown away by the leap in performance on difficult tasks from GPT-3.5 to GPT-4. Here are some advanced theory of mind questions (specifically 'Strange Stories' tasks - from linked paper). GPT-4 does shockingly well. srcd.onlinelibrary.wiley.com/doi/full/10.11…

[2 more photos]

the idea of tests like these is that you create a situation where someone is motivated to achieve X, and they can only achieve X by means of skill Y but if you give the same test to an entity who can achieve X by some other skill Z, you have no idea how much skill Y they have!

"GPT-4 can solve Theory of Mind tests" does not imply "GPT-4 has theory of mind" it MAY have theory of mind, but there may be other ways to solve these tests, including ones we can't fathom

you can test skills in rats (memory, etc) by getting them to navigate a maze. if it can get the food, it has the skill give the same maze to a bird who can fly over the maze, and it doesn't matter how fast it gets the food—that doesn't tell you if it has the relevant skill

kids below a certain age can't answer questions about why people might do certain things, because they haven't developed theory of mind; above those ages, they consistently can if GPT-4 can answer those questions, what does that say about its theory of mind? on its own, nothing!

what does theory of mind even MEAN? like what intuitions do we have for what it means for someone to have it? as far as I can tell it's not just a propositional knowing but a perspectival knowing. knowing to anticipate others knowing things you don't

perhaps GPT doesn't even have theory of people! it may not KNOW there are people "out there". it just knows there's text and it knows what text comes next. it doesn't know it's "in there" but it'll recite the "I'm an AI in a box, sure wish I could smell roses" trope if asked!

GPT can pretty generate beautiful metaphors for what it's like to be an LLM but if you could somehow ask "what's it like to be you?" WITHOUT it knowing that the "you" it is asking is "an LLM" and without knowing what an LLM is, it would have no answer because it has no "you"

GPT can textually improvise in-frame, but has no actual first-person perspective / situational awareness of itself outside the prompt so it seems unlikely (if not impossible) that it could evolve a second-person perspective that works remotely like ours does

(I played around with Codex, which wasn't beaten into being obsessed with a helpful assistant role like ChatGPT, and when I started with just a prompt asking "what's the situation?" I got "I think we're in deep trouble" & some improv later that character said it was Arthur Dent)

To be clear, this is not just about memorization. It's not even just about "pattern-matching" vs other methods of solving. It's about the relevance of any tests.

Henry Shevlin @dioscuri ·

Obviously data contamination is always a worry for these things, but here's GPT-4 acing a similarly difficult social cognition question I just concocted myself:

The point of my construct validity rant is that EVERY test made for humans should be suspect when applied to AIs. is the performance on the tests impressive? yes! absolutely! it's mindblowing! but the question "what does that MEAN?" remains!

it probably still even makes sense to think of it as a developmental milestone! but that doesn't mean it's the same developmental milestone that 7-year-olds experience when they have the insights that lead them to pass those tests

I believe this rant is not simply the "your princess is in another castle" retreat that is common, where anything AI can do becomes "not actually TRUE intelligence" it's really smart! we just need to caution against assuming it's like us in any particular way. we need to check.

understanding evolution and the Interface Theory of Perception should help here in general, brains evolve (biologically) and minds evolve (learning) to solve problems as cheaply as possible

Malcolm Ocean 🏴‍☠️ @Malcolm_Ocean ·

finally read the Interface Theory of Reality paper and it's fairly short & quite good! quite easy to read. good metaphors, a lil math. resonates a bunch with some of what I've been orienting to regarding collective perception sites.socsci.uci.edu/~ddhoff/interf…

sometimes also this means humans use LLM-like auto-complete-style reasoning! it's cheap and sometimes gets the job done but that's obviously not ALL we do, and if someone tells you otherwise I encourage you to make an observation about your own live sensory perceptions

you can also play around with exercising situated *perspectival* theory of mind. look at the person talking to you and imagine what the world is like from their perspective first: what images are entering their eyes second: all of what they see means something different to them

huh—that 2nd piece is like this thread! separate from "is it like something to be GPT?" info-theoretically, I can step thru my code's if-statements in first person & likewise I can imagine what LLM processing is like, & I'm pretty sure it CAN'T be what I call ToM. type error!

if you've read all the way here and you think it would be good if the main message (about construct validity) got to more people, you may want to tweet your own remix on the idea, or maybe RT the top tweet in the thread:

Malcolm Ocean 🏴‍☠️ @Malcolm_Ocean ·

GPT-4 seems like a good time to remind everyone about a lil concept called CONSTRUCT VALIDITY > Construct validity concerns the extent to which your test or measure accurately assesses what it's supposed to. Construct validity of ANY PSYCHOMETRIC applied to AI starts as suspect

here's an even better thread on the same topic, with some LLM-specific suggestions for how to design better tests

Michael C. Frank @mcxfrank ·

People are testing large language models (LLMs) on their "cognitive" abilities - theory of mind, causality, syllogistic reasoning, etc. Many (most?) of these evaluations are deeply flawed. To evaluate LLMs effectively, we need some principles from experimental psychology.🧵