Session details
Date: Jan 27, 2023
Series: ModelStream #008.1
Guests: Tom Ringstrom
Paper: Reward is Not Necessary: A Compositional Theory of Self-Preserving Agents with Empowerment Gain Maximization
此页面由机器翻译成中文。 查看英文原版。
ModelStream #008.1
Jan 27, 2023 · with Tom Ringstrom
▶ Watch on YouTube ↗Date: Jan 27, 2023
Series: ModelStream #008.1
Guests: Tom Ringstrom
Paper: Reward is Not Necessary: A Compositional Theory of Self-Preserving Agents with Empowerment Gain Maximization
Transcript
The full transcript is available on GitHub. This excerpt is generated by automated speech recognition and may contain errors.
Hello and welcome. It's January 26, 2023. We're here in Active Inference Model Stream number 8.1. Today, we're appreciative to have Thomas Ringstrom, who will be presenting on Reward is Not Necessary, a compositional theory of self-preserving agents with empowerment gain maximization. There will be a presentation followed by a discussion. So, Thomas, thank you for joining. Really looking forward to this. Off to you. Yeah, thank you very much. It's really nice to be here and talk to this group. I'm a computer science PhD student at the University of Minnesota. My interests are in sort of what are the computational properties that we would need to have agents which flexibly plan in sort of high-dimensional product spaces of variables. And also, how do we get agents to perform complex tasks in an intrinsically motivated way, especially in high-dimensional product spaces? So, this presentation is going to argue that reward, and I know I'm talking to a sort of active inference crowd, but some of the same points apply to active inference perhaps too. But this presentation is mostly about how there's going to be some major problems, I think, using reward maximization objectives. And by moving to sort of reward-free objective functions, we can get really nice factorizations that help us plan. So, let's just start off with a sort of simple picture, a simplified picture of an organism. So, we have a honey badger here. And so, this honey badger has internal states. It gets hungry and it gets thirsty. And there's also an external world that the agent lives in. And, you know, in order to sort of modify internal state spaces, the agent might have to perform some complex tasks. It might have to get several items in order to, you know, eat an apple or things like that. So, you could see that maybe some symbolic state space sort of mediates the connection between things you do in the world and transformations that you make on some internal state space. And, of course, real, you know, real world organisms and humans live in a high dimensional world. So, you could imagine that there's an encoder and a decoder. And really, you know, there's a sort of high dimensional physiological or interoceptive domain that is these sort of Y tilde with the dot, or Y tilde, I should say, and Z tilde. And then there's, you know, a high dimensional world. You could just say that the that's X tilde. And really, you just need to sort of map these down to some discrete state space in which these domains can interact. So, yeah, here's like a simple encoder and decoder. And, and so if, if an agent has a kind of simple ontology like this, you can imagine that it's sort of generatively entrained to the world that it lives in, which is a notion that's probably pretty familiar to sort of active inference people. So, you know, you have some encoder of these high dimensional states, and then you have some operator PS, which is in the middle of the head here, and that just sort of advances the latent states, and then you decode, and you'll get sort of expectations of, of the high dimensional world that the agent lives in. John P. And so the problem is, is that you can't really represent explicitly these latent space transition operators that would dictate, you know, the dynamics of, of all these variables. You can't represent it as an explicit object because the more state spaces that you keep track of in the world, the larger this object PS becomes. So in order to handle this, you'd have to represent it in a sort of factorized form. So you just would represent its factors, and you wouldn't explicitly sort of enumerate all of the state vector transitions under, under this transition operator. And so what we really need to think about is, you know, what are the sort of model based, or the sort of Bellman principles for decomposing hierarchical state spaces, so that we can create the right representations that help us plan in a flexible way, in a sort of time, time varying…