تم ترجمة هذه الصفحة آليًا من الإنجليزية. اطلع على النسخة الأصلية باللغة الإنجليزية.

Parr, Pezzulo, Friston 2022 Textbook Cohort 5, Chapter 8 part 1

Textbook Group meeting for Parr, Pezzulo, Friston 2022 .

Mar 4, 2024

▶ Watch on YouTube ↗

Session details

Date: Mar 4, 2024

Series: Parr, Pezzulo, Friston 2022 Textbook Cohort 5, Chapter 8 part 1

Paper: Active Inference: The Free Energy Principle in Mind, Brain, and Behavior

Transcript

AI-generated transcript excerpt

The full transcript is available on GitHub. This excerpt is generated by automated speech recognition and may contain errors.

all right welcome back cohort 5 we're in our second discussion on chapter so is there anything anybody wants to begin with can look at the text we can look at any question we can add a new question we could look at pmdp anything um anyone can write in the chat or raise their hand or just go for it yeah so I have a sort of more more open-ended question I've been sort of um thinking about you know doing a little little modeling project and um what I would like to se what what I'm thinking about is that I'm just wondering how I can use this active inference Machinery to model some kind of um kind of like a recognition task like how would you approach doing something like an amness classification with um um active inference but what I am interested in is is bringing this you know one key difference um from the active inference like um approach to um the problem that's you know basically completely solved in the using artificial neural networks but in a in a way that um you know when you if you feed amnest to um like a Transformer based classifier you patchy it so it turn it into small little chunks of image and you feed that into a Transformer and the Transformer kind of like sees the whole sequence and then finally learns how to classify an image I'm just sort of wondering um thinking about like how to find some sort of connection between between doing something like that except getting the neural network or some other kind of you know um system to choose where to look to to to sort of model the behavior to be almost like like a sakad like I'm seeing a part of an image and then I have to make a decision where to look next to um you know um resolve as much uncertainty as I can about what what am I seeing and it doesn't have to be something like an amist it could be like a really really simple images to start with you know like 8 by8 pixels some generated I don't know kns and Crosses or whatever but I I don't really even know where to start with that um it's a it's a very open-ended big big thing but if you have any any suggestions worth look I would love uh some pointers yeah great great questions well certainly what I'll show now is not the end of the story but I think it'll provide some motifs that you can pick up on here's a 2020 one paper with axle constant so in this setting it was an icade Visual and one of their Novelties was that they made a bounding box that was kind of like the zone that was visible in the Cade so the Cross Focus is the center of vision and then the affordance of the iate is to move to another cross um and then that would move the bounding box um and it's a cultural pattern recognition task like based upon identifying these higher order patterns of Pottery so that's one visual way I think this is a super fascinating thing because in almost all visual recognition tasks um either the whole image is passed through if it's sufficiently small like mnist or it's ch tokenized or chunked or however um but here we bring the action into Vision with actual guided cognitive Vision moving to uncertainty so that might be um way more efficient in certain settings um for example recognizing a person like first there's like the movement just empirically like there's like movement to the body and then there's a movement to the face and then identifiable parts of the face that's the pyade model and then on the mnist the recent uh structure learning paper actually specifically they they do amnest um and they [Music] um so the labeled data in amnest are the true identity of the digit and then they use their structure learning technique to um develop the the Styles handwriting styles of the digit so they're not doing structure learning on how many digits there are as as far as my reading but they do look at different styles like these are different styles of the digits so then it is learning a it's doing it's saying how many styles of one are there and then it's finding it through um structure learning which is the proposal…