Threads

Critiquing the human-alignment analogy in AI safety

7 tweets · December 2023 · 0 likes · 0 retweets · read on Twitter

Replying to Nora Belrose @norabelrose ·

Introducing AI Optimism: a philosophy of hope, freedom, and fairness for all. We strive for a future where everyone is empowered by AIs under their own control. In our first post, we argue AI is easy to control, and will get more controllable over time. optimists.ai/2023/11/28/ai-…

@norabelrose > Black box methods are sufficient for human alignment sort of true, but individual humans are much more limited than AIs are. in part, they're the same order of magnitude as intelligent as each other, and they still need to sleep and eat, and can't clone themselves

@norabelrose also a lot of what you're calling "human alignment" isn't actually aligning humans with humane values but aligning them with oppressive cultures or ideologies (often calling this morality!) meanwhile biological evolution makes humans already caring, empathic, and prosocial!

@norabelrose > secret murder plots aren’t actively useful for improving performance on the tasks humans will actually optimize AIs to perform. maybe, tho humans will optimize AIs to manipulate other humans (they already are with social media algorithms, to sell more ads etc!)

@norabelrose prompt engineering will become decreasingly powerful once the AIs are actually agents, planning & acting with decreasing oversight sure, at the start it looks great, but maybe once an AI gets sufficiently smart it starts generating murder plots not during training but later

@norabelrose > Some people point to the effectiveness of jailbreaks as an argument that AIs are difficult to control. They are a serious sign that if an AI becomes powerful enough to enable crimes, it may be difficult to release such a model with safeguards that prevent those crimes at all.

@norabelrose You're kind of jumping around here between different concerns: 1. AIs that murder because they want more power 2. AIs that can be programmed to aid murder 3. AIs that will cause problems due to failing to understand human values while trying to do good 4. and others

@norabelrose I'd like to see your philosophy respond to the critiques in betterwithout.ai, a piece that talks about the messy middle between "oh no AI said a racist thing" and "AI takes over universe" that's like "AI increasingly controls things and humanity doesn't know how to regain"