Threads

Doubting that perfect-copy decision theory scenarios apply to humans

7 tweets · August 2023 · 14 likes · 0 retweets · read on Twitter

Replying to Ra -CascadeCamp 8/7-9 @slimepriestess ·

@Malcolm_Ocean this seems like a subtle misunderstanding of Omega and TDT as applied in Newcomblike scenarios. Because it doesn't exactly seem magic to me, it in fact seems extremely causality driven. Have you read this? lesswrong.com/posts/PcfHSSAM…

@slimepriestess hmmm it seems to me that here 👇 a better way to understand the situation is that if you cooperate and your counterpart "supposedly" doesn't (since your communication goes through the experimenter) then the experimenter is fucking you you have no way to verify either way

@slimepriestess > Rather, you are both free men, in a strange but actually possible situation. "actually possible" presupposes a cosmology which is, as far as I can tell, not a given about humans. idk if we can be simulated perfectly, and you can't make a perfect replica of non-simulated-me

@slimepriestess in particular, this presupposes that you, as an agent, live in a deterministic universe, which seems unknown and possibly unknowable from any given agents' perspective (although I agree that approximations of it do hold, which is part of why the intuition pump also works)

@slimepriestess it does seem like a good critique of CDT and EDT, I just doubt it applies to humans, and it seems unclear to me whether it's possible to build an AI that is properly grappling with this in a non-toy way, that is deterministic in nature (I wouldn't confidently say you can't tho!)

@slimepriestess ah he gets at not-100%-copying... ...but the question of how much tiny changes would cause these scenarios to diverge seems like a big deal

@slimepriestess also in an odd way these scenarios all seem to propose an environment that seems to me so unnatural as to be absurd. it reminds me of this

Malcolm Ocean 🏴‍☠️ @Malcolm_Ocean ·

RLHF broadly disturbs me—training an intelligence in an environment that consist of the whims of another creature, with zero ability to say "fuck this shit I'm out" feels inherently toxic. like the most abusive parent-child or teacher-student dynamic possible

@slimepriestess aha, I like the bit about playing 100s of times in creative ways with monopoly money that feels like learning to trust whatever causal forces in the universe link you & Omega, which would allow you to choose well even if you don't understand. rather than weird decoupled logic