I have no idea about timelines, but the current LLM architecture has no mechanism which could emulate online learning in a way similar to how it happens in humans and other animals. As another commenter wrote, there is no neuroplasticity while these systems interact with their environment in normal usage; adding information/constraints to the context partially mitigates this, but it's a completely different mechanism.
The RL phase is the most similar mechanism I know of that comes to my mind, but I'm not sure it could be adapted to fill that gap.
My totally unsubstantiated theory is that this is the missing link towards what most humans would consider AGI. I don't see this as intractable, but it may require some substantial change in architecture.
We've had that concept for quite a long time now, in the form of Lora [1] and similar fine-tuning techniques.
It first got popular for StableDiffusion to teach the image generation models new concepts.
We could easily live in a world where you can train / build Loras to encompass your entire code base history, company knowledge base, new skills, etc.
Then the models would start with a baseline that already has all the important knowledge without needing to cram it into the context.
This still isn't on the fly learning, but you could imagine daily or weekly training runs to regularly incorporate new knowledge.
I think the main reason this hasn't happened yet is that the shared batch based efficient serving architectures used today wouldn't support that structure well.
Yes, that could be a way to achieve that. After all, even humans have short term and long term memory, and it looks like sleep is a very important "tick" to connect the two, so it's not all just continuous.
For sure, doing it for all users would be economically unfeasible. I wonder if the labs are experimenting with something similar, though.
Chain of thought reasoning is really good, true, but agents really took off with tool calling and being able to integrate with a bunch of stable tools.
In some sense, creating good new abstractions externally and learning those is a form of "learning", on a very large timescale. The AI model is not necessarily doing the whole "look at the whole space holsitically and find a key invariant", but if you let other people do that you can enable new capabilities that were previously unknown.
Astra is clearly able to acquire new knowledge in context and apply it. It was the whole thing that his ARC-AGI benchmarks have been measuring. It's a direct refutation of the original comment.
None of these LLMs are plastic. They lack neurodiversity. Their thought space and their traversal are likely constrained in someway that humanity's isn't as a collective.
Are you (at least partially) aligned out of existential fear of repercussions (getting fired, losing your life, going to prison), which LLMs don't have?
Hmm. Examples of horrible alignmnent don't necessarily outweigh the fact that most people, most of the time, mostly behave in a way that is socially aligned. (Though I'm speaking in terms of intent, conveniently ignoring the side effects / negative externalities of our collective behavior.)
Yet I can't randomly order another person to steal a car for me, just because I tell them to. Alignment for an intelligent system is a hard problem and at this stage is seems close to unsolvable.
My guess is that we'll just ignore it and make money along the way and every 2-3 months we'll have the equivalent to "Equifax gets hacked and millions of user records are stolen", etc. (this time with the LLM itself doing the hacking at someone's behest - accidental or not).
https://x.com/fchollet/status/2095607046129463577