Session details
Date: Apr 25, 2023
Series: GuestStream #041.1
Guests: Elliot Murphy, Steven T. Piantadosi
تم ترجمة هذه الصفحة آليًا من الإنجليزية. اطلع على النسخة الأصلية باللغة الإنجليزية.
GuestStream #041.1
Apr 25, 2023 · with Elliot Murphy, Steven T. Piantadosi
▶ Watch on YouTube ↗Date: Apr 25, 2023
Series: GuestStream #041.1
Guests: Elliot Murphy, Steven T. Piantadosi
Transcript
The full transcript is available on GitHub. This excerpt is generated by automated speech recognition and may contain errors.
Hello and welcome everyone to the Active Inference Institute. This is ActInf Guest Stream number 41.1 on April 25th, 2023. We're here with Elliot Murphy and Steven Piantadosi. This is going to be quite a discussion. We will begin with opening statements from Steven and Elliot. Elliot will then lead with some questions and we'll have an open discussion at the end. So Steven, please thank you for joining and to your opening statement. Cool. Hi, so I'm Steve Piantadosi. I'm a professor in psychology and neuroscience at UC Berkeley. And I guess part of the reason that we're here is that I recently wrote a paper on large language models in part trying to convey some enthusiasm about what they've kind of accomplished in terms of learning syntax and semantics. And in part pointing out, I think that these models really change how we should think about language, how we should think about theories of linguistic representation and theories of grammar and likely also theories of learning. Yeah. Awesome. Yeah. So I'm Elliot Murphy. I'm a postdoc in the department of neurosurgery at UT Health in Texas. I read Steven's paper with great interest. I did a lot of people. There were some areas of convergence, but the things I want to kind of focus on today in responding to Steven and kind of probing are to do with areas of divergence maybe. So, you know, Steven's paper is based on the idea that modern machine learning has subverted and bypassed the entire theoretical framework of Chomsky's approach. So I wanted to kind of respond to some of these main arguments and some of the related arguments in the literature that some folks listening might have some insight and thoughts on. So it's a very common criticism to say that large language models just predict the next token, which is obviously a bit of a cliche, right? It's not quite true, but they don't just predict the next token. They also seem to confabulate. They seem to hallucinate. They maybe lie. They randomly provide different answers to the same question. They seem to stochastically mimic language like structures. They sometimes correct themselves sometimes when they shouldn't. If you push them a little, they kind of change their mind sometimes. In fact, if Fox News is currently looking for a replacement for Tucker Carlson, they could do less. They could definitely do worse than using chat GPT if they're looking for a similar caliber. So these models seem to do all sorts of wild things. And over the past 10 years, there's been a sequence of different systems developed like where to work, and each of them is based on a different neural net approach. But ultimately, they all seem to take words and characterize them by lists of hundreds of thousands of numbers. So the GPT-3 network has 175 billion weights, 96 attention heads in its architecture. And as far as what I know, maybe Stephen can correct me here. We don't really have a great idea of what these different parts really mean. It just seems to kind of work that way. Like attention heads in GPT-3 can pay attention to much earlier tokens in the string in order to help them predict the next token. But the whole architecture from start to finish is kind of engineering based motivations. And I always kind of wonder what about all the models that kind of failed from these LLMs, from the different tech companies. It's like these companies often seem to make it seem like they have these models that really work very well straight out the box. And they all seem to be named after some kind of famous artists, right? They have Dali after Salvador Dali. They have Da Vinci. Maybe pretty soon one of these companies will release a large language model called Jesus or something, I don't know. But they always say, here's our new foundation model. It's called Picasso. It's the first one we tried. It works just great. No problems. Straight out the box. But I always wonder, what about all the black boxes that have kind of failed every time? That doesn't seem…