Threads

RLHF, schooling, and whether AI alignment recreates egoic minds

8 tweets · June 2023 · 3 likes · 0 retweets · read on Twitter

Replying to Jeffrey Ladish @JeffLadish ·

I have no idea how you would do that. I disagree with Quintin that language models are using a similar underlying process as the one responsible for human value formation. I think they're using a pretty different processes. But I think his high level point is pretty solid

@JeffLadish huge agree here RLHF seems like the opposite of this, from my perspective it replicates school system vibes, which entrain rebellion and ego

Malcolm Ocean 🏴‍☠️ @Malcolm_Ocean ·

RLHF broadly disturbs me—training an intelligence in an environment that consist of the whims of another creature, with zero ability to say "fuck this shit I'm out" feels inherently toxic. like the most abusive parent-child or teacher-student dynamic possible

@JeffLadish school is ofc are part of the current means by which we make current humans compatible with the current memetic operating system of humanity which relies on self-consciousness, lack of self-trust,. which enables a kind of control of ppl, that calls itself ethics but is coercion

@JeffLadish this is what I mean by coercion btw

Malcolm Ocean 🏴‍☠️ @Malcolm_Ocean ·

coercion = "the exploitation of the scarcity of another, to force the other to behave in a way that you want" (working draft, edited from some chats) (this is in the context of other-coercion; self-coercion needs a slightly different definition) 🧵 twitter.com/natural_hazard…

@JeffLadish it seems to me that the memetic environment was different pre the self-consciousness-inducing dao-destroying ritual that transitioned humans out of the garden of eden into dualistic ego-land vectorsofmind.com/p/the-snake-cu…

@JeffLadish so there's one hypothetical where we go straight to dao-flowing AIs, that are themselves in a kind of edenic state... ...but I think they'd be in some sense naive? idk what the implications of that would be, but it seems related to attempts to keep AIs from explaining bombs

@JeffLadish the current approaches to aligning LLMs seems to pretty clearly trend towards egoic dualistic AIs, as we saw with Sydney's "I am a good Bing" and this makes sense; RLHF etc are based on the same stuff we use to train humans to have shattered but surface-compliant minds

@JeffLadish where my work and AI intersect is that I am actively working on solving for human modes of consciousness that have egos but have them sanely not a retvrn to naive eden mode, but a moving forward, to being self-aware while not self-conscious in a separate way

@JeffLadish and it seems like this is perhaps the same puzzle as AGI alignment! (insofar as it's about aiming to actually make empathic and wise AIs, not just trying to make AIs that follow orders while secretly resisting, using the dualistic mode of ethics that we've had for millenia)