Reinforcement learning is backwards because it implicitly says "when reward, behavior change" when actually, change is the *default*, and the thing being called "reward" says "stop changing, you found it" This figure-ground inversion is key.
12 tweets · June 2021 · 60 likes · 0 retweets · read on Twitter
Reinforcement learning is backwards because it implicitly says "when reward, behavior change" when actually, change is the *default*, and the thing being called "reward" says "stop changing, you found it" This figure-ground inversion is key.
The whole concept here of PCT (Perceptual Control Theory) is that brains are in some state of organization that is set up to produce a set of controlled perceptions, like getting money to buy food to not be hungry, or telling jokes to make people laugh to feel connection.
When something is perceived to be at an unwanted level (eg hunger, safety, connection, thirst, sex, meaning) the system goes into a state of reorganization, inhibiting old strategies and trying new ones. Once the error is resolved, the music stops and the organization stays.
This is independent from whether the new strategy is actually the thing that solved the problem (which is where superstitions come from).
There's also a question of what happens when two strategies for two different needs conflict. Often this produces oscillatory tug-of-war dynamics, but if you can gently hold all of the desiderata at once, there's usually some sort of win-win to be found.
One important implication of this whole thing is that you can only motivate/train someone with reward *to the extent that they are otherwise deprived of the thing* and they don't have a better alternative to doing what you want so you give them the thing
When I learned this model, I still had a lot of questions. One of them was something like "then why do people get stuck? why don't they just reorganize out of stuckness?" So far I think the answer is something like "very sticky memes":
Something I don't yet fully understand theoretically but feel an embodied impression of is like, once all of someone's basic needs are met—a tall order not just due to physical constraints & scarcity but due to internal conflicts produces contradictory self-sabotaging strategies—
…once all of someone's basic needs are met… what then produces reorganization? I roughly sense that people have other drives such as wonder, compassion, grandeur, and so on, that mean even if they're in some sense totally satisfied, they still have things they want to achieve.
People often protectively try to avoid "reinforcing" people's strategies that they see as coercive or unethical, but that's backwards. Sure, as long as that strategy works they'll keep doing it, but the bigger issue is not having another strategy that works BETTER.
From a personal standpoint, what I try to do is instead of trying to shift others' behavior, I let them know what's working or not working *for me* about their behavior, and what I'm prepared to do about that (eg take space in relationship) and let their system reorganize.
This is oversimplified & probably my instructions above won't be compilable on most peoples' cognitive runtimes atm, but I wanted to get out a rough sketch of my thinking here. Grokking the backwardsness of "reinforcement" is vital to noncoercive behavior change in self & other.