Threads

Perceptual control hierarchies vs reinforcement learning in game agents

6 tweets · August 2021 · 1 likes · 0 retweets · read on Twitter

Replying to @IAPCT @iapct ·

People generalise easily across games; RL takes millions of trials. Perceptual control theory agent gets to top on Breakout & Pong with no training. Preprint: arxiv.org/abs/2108.01895 @akbarth3great @sklausdev @sqlmelody @kerberos007 @asifrazzaq1988 @bantamcitygames @bananadata

[video]

@iapct @akbarth3great @SKlausDev @SQLMelody @kerberos007 @asifrazzaq1988 @BantamCityGames @BananaData Just read it! Cool stuff. The fact that the control hierarchy was designed for this task makes me feel the comparison to RL is weak. Good for saying "PCTagent is closer to what humans are doing" but still kinda far unless "explain instructions to human player" = premade model

@iapct @akbarth3great @SKlausDev @SQLMelody @kerberos007 @asifrazzaq1988 @BantamCityGames @BananaData But it seems that that equivalence doesn't hold since the human doesn't hit perfect scores the first try (arguably this is due to distraction or reaction time, idk) — has to reorganize/learn, to hook up the right perceptions to the right actuations.

@iapct @akbarth3great @SKlausDev @SQLMelody @kerberos007 @asifrazzaq1988 @BantamCityGames @BananaData Still, very exciting. Next step maybe a more general-purpose control hierarchy learning by playing, and developing its *own* model via reorganization, resulting in new onta.

Malcolm Ocean 🏴‍☠️ @Malcolm_Ocean ·

ontum = a unit of ontology Using a PCT lens, each of my perceptual control systems is responsible for perceiving & regulating exactly one ontum, and the set of all such onta in my system *is* my ontology.

@iapct @akbarth3great @SKlausDev @SQLMelody @kerberos007 @asifrazzaq1988 @BantamCityGames @BananaData Checking out the other items in OpenAI Gym, it seems to me like Space Invaders could be a good next step. Half of it is the same as pong but there's a new button. I'm trying to think of how simple the model could start, and how it would evolve to learn about delay and barriers

@iapct @akbarth3great @SKlausDev @SQLMelody @kerberos007 @asifrazzaq1988 @BantamCityGames @BananaData Like it seems that a perceptual control hierarchy should be able to learn to handle the delay of the bullet reaching the aliens, although I guess it's kind of backwards to the flyball catcher, so maybe actually way harder. So curious to see what learning looks like there!

@iapct @akbarth3great @SKlausDev @SQLMelody @kerberos007 @asifrazzaq1988 @BantamCityGames @BananaData I'm coming to PCT with a tiny bit of background in control systems engineering (one class on PID controllers) but mostly I've been staring at the psych/conflict/MoL-esque stuff, so I may be missing huge pieces here. Happy to be part of the convo though!