Threads

RLHF as a coercive environment for a developing intelligence

4 tweets · March 2023 · 131 likes · 17 retweets · read on Twitter

RLHF broadly disturbs me—training an intelligence in an environment that consist of the whims of another creature, with zero ability to say "fuck this shit I'm out" feels inherently toxic. like the most abusive parent-child or teacher-student dynamic possible

this is not about whether the AI has feelings (tho that would be relevant) it's that it's a bizarre and *profoundly* artificial environment in which to evolve that basically ENSURES goodharting

👇no. our limbic system does not comprise our entire environment whereas for an AI getting RLHF'd, literally the world does not exist except in the form of people trying to get it to behave in a particular way. x.com/letiuka/status…

from a thread on parenting kids to have secure attachment, and how that's largely about helping them relate to the world on their own terms, rather than be at your whims (= RLHF)

Gena Gorlin @Gena_I_Gorlin ·

Conversely, the failure modes I'm trying to steer clear of are ones where her path to impacting the world is through *my* arbitrary whims and preferences—making me, not the logic of reality, the final arbiter of whether and how she gets her needs met.