此页面由机器翻译成中文。 查看英文原版。

Livestream #057.2

Active Data Selection and Information Seeking

Jun 4, 2024 · with Thomas Parr

▶ Watch on YouTube ↗

Session details

Date: Jun 4, 2024

Series: Livestream #057.2

Guests: Thomas Parr

Transcript

AI-generated transcript excerpt

The full transcript is available on GitHub. This excerpt is generated by automated speech recognition and may contain errors.

All right, hello and welcome. It's May 24th, 2024. We're in ACTIMF livestream number 57.0, doing background and context video for the Active Data Selection and Information Seeking paper and series. So welcome to the ACTIMF Institute. We're a participatory online institute that is communicating, learning, and practicing applied active inference. This is a recorded and archived live stream. Please provide feedback so we can improve our work. All backgrounds perspectives are welcome and will follow video etiquette for live streams. Head over to activeinference.org to learn more about any of the projects including the live streams. So today we're going to do together a background first pass on a very interesting paper from Thomas Parr, Carl Friston, and Peter Zeidman, Active Data Selection and Information Seeking from 2024. In this video we're going to introduce ourselves, talk about big questions, go through the keywords of the paper, then most of the sections section by section, and as always with the dot zero it's just like a first pass and we'll look forward to speaking with hopefully some of the authors in the coming weeks and also looking what people ask about. So Christopher, let's introduce herself and go from there. Thanks a lot also for helping in the dot zero preparation. Happily. Yeah, so I'm Christopher Bennett. I'm a bioinformatics scientist. I do a lot of data mangling, data analysis, and that sort of thing. This paper was of great interest to me as we kind of go into this more data driven era in making sure that with such large data sets that we have, making sure that we can actually select relevant data for any of our applications going forward, be it machine learning more, whatever we're trying to do. And I'm Daniel. I'm a researcher in California and also was drawn to this on one hand on the applied side, the idea of more efficient and effective data sampling, and then on the more theory side, the connection with epistemic value information game. So here are some of the big questions. Why don't you add some detail to this? Absolutely. So there's five major big questions that I had after reading this. For the most part, it boils down to doing our sampling. You can do sampling over time and sampling of different data sets in different ways. Is there a way that we can intelligently select the data that we're going for the time that we're trying to select? Is there a way that we can understand how the time aspect of sample through time instead of just doing a like a dynamic or more dynamic instead of doing a static like we're going to do time zero, time seven, time 14, time 21? Can we say, hey, the differences between time one and time two are very, time point one, time point two are very interesting. It's a lot of data in there alone. So we'll sample one and two or and then maybe sample 10. Is there a way that we can intelligently select the time points that we are sampling from when we get into the time series aspect? There's a number, the paper mentioned a number of different time dimension models that you can add to the core model that they're utilizing, one of which was a hidden Markov model. Another was, they mentioned a differential equation in the actual model itself in the equation itself. Is there one, is there other situations that one performs better over than the other? Or is what they have selected to use in the paper the optimal solution in most cases, if not all cases? You know that it, when it comes to clinical trials, that was a section in this, that they discussed. There's a lot of FDA regulations of the clinical trials and it's very heavy red tape right now. Is there a way that, there's minimum ends that you need in many clinical trials to actually be considered passing? Is there a way that you can bound this model that they're, they've developed in to something that you can guarantee a minimum number, a minimum sampling that the FDA requires, or any regulatory body? Another point is the next…