The mechanism is irrelevant if the results are similar. No, it doesn't have to be implemented the same way as a human to qualify as a human-level intelligence. That would be silly.
How is text-based memorization going to let you learn non-linguistic skills, whether animal/human level or beyond (based on new data senses)?
How is text-based memorization going to substitute for learning? The two are not the same. Perhaps this is more applicable to robotics than a text generator, but I also doubt an LLM could learn text-based skills like programming or math if it had not been pre-trained on them via SGD & RL, and had to instead rely on some poor-man's-learning context-based recall instead. What else can't it learn? How is the LLM intern, the "drop-in replacement remote worker" going to do on day #2?
Instead of pretending that an LLM can be human-level, or super-human, or become generalist, why not just admit that this is not the final form of AI. An LLM is not an animal/human-like intelligence, it is something different - a language model, with it's own strengths and weaknesses.
Despite all the AGI hype, an LLM seems to have more in common with a pre-trained single-purpose system like AlphaGo than a brain, but with the rubric/reward-based policy function baked into the weights.
In another 10-20 years some new idea, hopefully more brain-like, will have superseded LLMs and they will indeed be labelled as "LLMs" as the AI/AGI label becomes attached to the new more brain-like creative intelligence. Perhaps it'll be sooner than 10-20 years, but I doubt it given the current 10-year fixation with LLMs which doesn't appear to be slowing down anytime soon. Perhaps Sutskever is working on something a bit different?
>How is text-based memorization going to let you learn non-linguistic skills
I don't know. How is Astra a step change in computer use and spatial reasoning to the extent it can play games, paint good looking stuff with e.g canva and a whole number of other things ?
>How is text-based memorization going to substitute for learning? The two are not the same.
Of course if you call it something else then you can say it's not the same.
>How is the LLM intern, the "drop-in replacement remote worker" going to do on day #2?
Just fine I imagine ? ICL and the memory tools around a harness are pretty good. I'm not sure what sort of magic you're expecting from the human, but they're not getting any improvement in that time frame a frozen transformer can't match.
>Instead of pretending that an LLM can be human-level, or super-human, or become generalist, why not just admit that this is not the final form of AI.
It doesn't seem like I'm the one pretending here.
>Despite all the AGI hype, an LLM seems to have more in common with a pre-trained single-purpose system like AlphaGo than a brain, but with the rubric/reward-based policy function baked into the weights.
If you say so.
>Perhaps it'll be sooner than 10-20 years, but I doubt it given the current 10-year fixation with LLMs which doesn't appear to be slowing down anytime soon.
The architecture that keeps delivering results isn't slowing down ? You don't say.
There's no shortage of people, even researchers, who for one reason or the other are convinced we are in need of some paradigm shift.
But guess what? Talk is cheap. You beat the current paradigm or you don't.
> How is Astra a step change in computer use and spatial reasoning to the extent it can play games and paint good looking stuff with e.g canva ?
Presumably because of pre-release training, because some alien outside of the model, armed with the reinforcement learning algorithm, came in and programmed its weights.
> The architecture that keeps delivering results isn't slowing down ? You don't say
Sure, nothing wrong with that, as long as you don't misrepresent the limitations of the approach.
> There's no shortage of people, even researchers, who for one reason or the other are convinced we are in need of some paradigm shift.
> But guess what? Talk is cheap. You beat the current paradigm or you don't. Do you seriously think that Meta never scaled JEPA ?
The idea has not been taken very far, so what is there to scale? It's not a complete cognitive architecture. So far it's also been using a pre-trained transformer as the learning component, which makes it of limited interest.
The animal intelligence approach, even in it's most fledgling form (that you would apparently dismiss), requires a complete agentic architecture, including new learning algorithms and generative behavior, before it can be compared to LLMs. We know that, done right, our brain architecture is more capable than an LLM, so even if any hypothetical attempts to reproduce it were not highly performant, we know that the idea itself is sound. You might compare with Uszkoreit's initial poor-performing implementation of his new language model architecture - should he have given up?
>Presumably because of pre-release training, because some alien outside of the model, armed with the reinforcement learning algorithm, came in and programmed its weights.
They didn't program anything. They gave it data at best.
>The idea has not been taken very far, so what is there to scale?
LeCun was the head of Meta AI for over a decade and his baby that he keeps harping on about wasn't taken very far ? Come on. You're smarter than that. It went the way all the alternate architectures have gone since the transformer, a sidegrade at best, probably not even that.
>You might compare with Uszkoreit's initial poor-performing implementation of his new language model architecture - should he have given up?
What are you talking about? There was no poor performing implementation of transformers that was published that he needed to push through. Are you talking about self attention experiments before the finished transformer? That's literally just research. And if he languished on that for a decade then yeah I'd tell him to probably look at something else, but of course he didn't.
You regard the LLM post-training process as "giving it data" ?!
Have you looked at all the published JEPA research both while LeCun was at Meta, and since (up to and including the latest AdaJEPA from June)? Please enlighten us as to exactly which line(s) of research you think were "scaled" at Meta, and then tell us which of these constituted anything even remotely resembling a complete testable intelligence?
FYI, it's been a long time since FaceBook/Meta even had a single head of AI. Since 2018 it has been split into two groups, FAIR and Generative AI, with LeCun being in the FAIR group, not as head, but as Chief AI scientist. LeCun only invented JEPA in 2022 (shortly after FaceBook became Meta), first writing about it in his "A Path Towards Autonomous Machine Intelligence" paper, perhaps unhappy with the work of the GenAI group, which he had no control over, that presumably was getting all the compute.
I was referring to Uszkoreit's personal telling (on YouTube) of the origin story of the Transformer, his motivations with the design, his initial personal failure to implement his idea in a performant enough manner to beat the current LSTM SOTA, and Noam Shazeer then throwing the kitchen sink at it and eventually coming up with the Transformer design.
>You regard the LLM post-training process as "giving it data" ?!
Yeah. Presumably, lots of synthetic data is being generated, experiments being run, but post-training is still a largely automated process.
>Have you looked at all the published JEPA research both while LeCun was at Meta, and since (up to and including the latest AdaJEPA from June)? Please enlighten us as to exactly which line(s) of research you think were "scaled" at Meta, and then tell us which of these constituted anything even remotely resembling a complete testable intelligence?
I have. My point isn't that his ideas are trash or that he should stop working on them. My point is it's not "gone very far" because he's taking it as far as he can, which isn't very far. He's not had a lack of influence, resources or will, either from his time at Meta or now with his billion dollar startup. He's had far more of it than most, if anything. That there's not much to show for it so far is not for a lack of trying.
>I was referring to Uszkoreit's personal telling (on YouTube) of the origin story of the Transformer, his motivations with the design, his initial personal failure to implement his idea in a performant enough manner to beat the current LSTM SOTA, and Noam Shazeer then throwing the kitchen sink at it and eventually coming up with the Transformer design.
So it's what I thought. This is just regular research unless an inordinate amount of time was spent on it and that's not the case.
> My point is it's not "gone very far" because he's taking it as far as he can, which isn't very far.
I wouldn't really agree - I'm no fan of LeCun, but the problem with JEPA isn't that it's a bad idea, or can't go very far, but just that it's not much of an idea in the first place!
It's no secret that our brain basically works by prediction, and what we're predicting is necessarily the external world as we perceive it though our own senses, aka latent representations, aka JEPA.
So, you COULD take this unoriginal smidgen of an idea and built it out to a full model of a human/animal brain, whether or not it's LeCun's intention to do so (he seems more interested in just the representational / world model aspect to it), but he certainly hasn't done so yet, nor created any research manifesto indicating that as his intent.
The fact that JEPA implementations to date are using pre-trained Transformers doesn't seem inherent to the approach - one could, with more effort, still predict latent representations (i.e. sensory feedback) but do so using a new real-time learning algorithm based on prediction failure.
LeCun seems more of an academic / research director than a builder, and I would never have put much stock in him being the one to build an animal brain.
Well I also don't think JEPA is a bad idea or anything. I don't think most of the alternative architectures or tweaks i've seen are bad ideas. On the contrary, some of them seem very cool and i'm all for interesting new ideas. I guess my opinion is more on the supposed necessity of it all.
I think the transformer is a powerful general learner. I think it's enough, and on the matter of intelligence, sometimes i think we make the mistake of zoning in too much on potentially spurious details. We still don't know much about Intelligence at the end of the day, what is and isn't really important, what is and isn't simply a detail of the environment and what kind of seemingly bizarre but surprisingly equivalent mechanisms can arise in the face of vastly different environments.
It's not a very high bar - even a rat has the basic ability to learn.