ModelStream #016.1

Divide-and-Conquer Predictive Coding: a structured Bayesian inference algorithm

Dec 6, 2024 · with Eli Sennesh, Tommaso Salvatori

▶ Watch on YouTube ↗

Session details

Date: Dec 6, 2024

Series: ModelStream #016.1

Guests: Eli Sennesh, Tommaso Salvatori

Transcript

AI-generated transcript excerpt

The full transcript is available on GitHub. This excerpt is generated by automated speech recognition and may contain errors.

Hello and welcome. This is Active Inference Model Stream 16.1. It's December 6th, 2024. We're here with Eli Sinesh and Tomaso Salvatore discussing divide and conquer predictive coding. So thank you both for joining. Eli, to you for the presentation. All right. Hi, everyone. So I recall being here once before, you know, presenting that old paper, Interoceptionist Modeling Allostasis as Control. And I think at the time I even mentioned that there wasn't really an inference algorithm that, like, I could just go and apply to all of these kinds of, like, active inference and predictive coding problems. And over roughly the last year, let's say, you know, we've started work, this whole team of myself, Hao Wu, who is unfortunately not on the stream due to having a nine to five. And Tomaso Salvatore, who's here with us. You know, we've actually been working on this, you know, like, how would we take these theories and, you know, this predictive coding idea and just turn it into a Bayesian inference algorithm you can use for stuff. So, you know, as background, like, to really refresh ourselves. This is the super short, you know, NeurIPS version of our talk. You know, there's, like, basically theoretical neuroscience, neuro AI, or now, I don't know, Tomaso, like, what keyword do you guys use it versus? I should have asked you maybe. Or, we do use neuro AI, actually, bioplausible deep learning. That's a sentence I often use. Yeah, so... I think that that dates back from even, like, before my versus times, I guess. Yeah, so the broad theory is that, you know, predictive coding sort of solves two tasks for the brain, and that's credit assignment and Bayesian inference. Now, like, predictive coding is also a thing in machine learning now. They do a lot of it at versus. Tomaso's done a bunch with it. You know, but it doesn't always... For various reasons we can go into, like, it sometimes doesn't always scale as well as backprop, basically. And so, we basically started over again, I guess you could say, with the concept of predictive coding to build our divide-and-conquer PC algorithm. And we basically find that you can use this at roughly the same scale you would with VAEs. So, of course, in order to, you know, test it out, we had this deep latent Gaussian model. It's actually the same one from, I think, the Monte Carlo predictive coding paper out this year. We really did this a lot, like, leaning on, you know, comparing to other predictive coding papers rather than to some arbitrary baseline. Because, you know, we figured, let's... If we pitch this as being about bioplausibility, then we can sort of maybe actually get results that are good enough to publish. Any... Any... Yeah. So, you know, we racked up our first win there. You know, we went and compared to other algorithms that were related or inspired us. So, you know, basically we would say, all right, you have... Yeah. Oh. You know, after, you know, demonstrating that this quote-unquote wins, like, okay, it works well enough. The... Which... Oh. Oh. This entire slide, yes. Okay. Yeah. So, our DCPC, in addition to being a predictive coding algorithm, is a particle algorithm. So, it's training a particle cloud, so to speak, to approximate the posterior distribution in a Bayesian inference problem. So, as such, we compared it to this algorithm particle gradient descent, PGD, from which we were drawing pretty direct inspiration. And then we also compared it to, you know, Langevin predictive coding out at ICML this year. And so, we compared on, you know, the pressure inception distance that people use in deep learning, rather than on, you know, the variational bound to the surprise. You know, we didn't compare on log evidence, because that's, you know, you can often get very bad visual samples from a model that has a very good log evidence, actually. So, in order to, you know, sort of say, like, okay, that, you know, log evidence is not the only thing that counts, we, you know, compared…