It has not been even 4 years since ChatGPT hit and LLMs + Transformers + Whatever they do has gotten us to solving millennium problems.
4 years ago, a program that could create photorealistic pictures, talk to you in any language of the world and solve the hardest math problems that we know, we would have called it AGI.
Now I don't know if what we have is AGI or not but I do not understand how you can see what has happened in the last 3 years and say "it will not get us there" no matter what "there" is.
> 4 years ago, a program that could create photorealistic pictures, talk to you in any language of the world and solve the hardest math problems that we know, we would have called it AGI.
I keep seeing this idea and I don't understand the reasoning behind it.
I think it could be a bit like saying if you showed someone 500 years ago a smartphone they would likely conclude at first it was magic. But once you had some time to let them use it and tell them how it all worked on a high level they would eventually obviously realise, no, it's not magic.
I guess just in the same way if you presented current LLM tech out of nowhere a few years ago to someone who'd never seen it, I concede they may be likely to imagine it was AGI in that first conversation, depending on their background.
But after using it for a bit and learning what an LLM is etc they'd land exactly where everyone is today - a great technology useful for some things, not AGI, not magic.
Wait, are you literally saying that modern LLMs are not "magic" (or AGI) because we know how it works?
If that is the case, then it trivially follows that humanity will never be able to build AGI, because every time we build something new, we will understand how we built it, which disqualifies it from being an AGI.
But that doesn't seem like a useful definition to me.
(Edit: I fully agree with your comment, just wanted to pick on a small point)
Thing is, we don’t understand how LLMs work. Not really. And everyone at the forefront of this space has said so. We know how to grow them but that is not anything like the same thing as understanding why or how they actually function. Our level of understanding of LLM intelligence is only slightly ahead of our understanding of the neuroscience of human intelligence, and no one would say we understand that at all.
So I think describing LLMs as magic at this point is absolutely appropriate.
The fact that a monkey can pick up an artifact a size of a small mirror and do a video chat, see and talk to another monkey on tge other side of the planet is absolutely bizarre. There's no good reason why this should be allowed.
I'm finding it more bizarre, if compared to creating a virtual talking monkey from scratch, with a little bit of Python and a large calculator.
What ? If that experiment was done on monkeys I don't think they had to use Android / iOS / iPadOS interface themselves. All they had to do is push a button
I always here things like "oh it's useful but dumb on some things", but it's just vague.
What is the test? What is a question that it fails at compared to humans? And no, you can't just say "find me the cure for cancer", but I believe there is probably enough intelligence in the weights that there is likely a cure in there with enough compute and the right questions.
The fact that you need to ask "the right questions" is why it's not AGI. A general intelligence should be able to ask of its own volition the interesting questions required to advance its goals.
I can’t predict how a novel intelligence could prove to me that it is intelligent. A novel intelligence would have to work out how to do that for itself
When it can replace a white collar remote worker without humans team realizing they work with the machine. And no, Agents are not like this whatever initial prompt and set of skills you give them. They still won't progress, learn and apply that knowledge like a human.
> What is the test? What is a question that it fails at compared to humans?
Human intelligence does not consist of answering questions.
And every day I have to say "No, you stupid clanker ...", so if this is "general intelligence", it's the intelligence of an idiot savant who is good at calculations but sucks at most other forms of intelligence.
> 4 years ago, a program that could [...] we would have called it AGI
If you had told someone in the 1800s that a machine could instantly multiply 100 digit numbers, that would have been considered dazzlingly intelligent. And yet we are not that dazzled by our calculators today (despite how useful they might be!).
Are you trying to explain how things once considered dazzling get normalized over time? Because otherwise this is a non-sequitur and has no bearing on the trivially verifiable, exponential explosion of capabilities we have seen in the last 4 years.
I keep saying this, until ChatGPT came out 4 years ago it was basically unimaginable that a single model could do any of, let alone all, the things they are doing today. Like, seriously, go take a look at the state of the art in NLP and NLU, the very first challenge in getting computers to even “understand” natural language, let alone other things like reasoning. Everything it does automatically was once a heavily experimental deep research field with long glorious careers for the researchers.
And now it’s all gone because the Bitter Lesson won again. If that’s not general enough to qualify for the G in AGI I don’t know what it is. And we’re sitting here going, “But it sometimes writes bad code though.”
It's not non sequitur at all ... it points out that the argument that, because there has been a rapid and surprising advance, therefore it's AGI, is intellectually bankrupt.
> If that’s not general enough to qualify for the G in AGI
The way i understand the way you define AGI is, that the lower bound of the Spiky profile is at least on Par with the best of humanity.
This way no human can say "But it cant do X as well as Mr./Mrs. Y"
The way that many other here see it is that the fact that it has lots of spikes in many different directions and some lower bounds above the level of a 10 year old, is indeed enough generality.
Probably the G in AGI is a spectrum, and it it is now a maturing process to lift the lower bounds of the spiky profile of to the best of humanity.
If this will ever work with an alien intelligence is debatable.
Even in the "Culture Series" the ASI had lower bounds of some skill sets that were below the best of Humanity, and thus the "Contact" department was formed, to use these humans to cover the shortcomings the ASI could not.
It might indeed be our chance for a symbiosis to make up for some of these lower spike bounds.
In the end who is to say that our intelligence is not spikey, and me are just normalizing it to a circle..
> It's not non sequitur at all ... it points out that the argument that, because there has been a rapid and surprising advance, therefore it's AGI, is intellectually bankrupt.
But that was not the original argument, which was: this incredible and rapid advance and everything we've seen since that was broadly considered decades away before ChatGPT; so it is unwise to make flat assertions that this will not lead to AGI because we have no idea where this is leading, not to mention what AGI even is.
The response to that was "In 1800 calculators would have been considered intelligent." ¯\_(ツ)_/¯
> Sorry to hear it, but some people do.
Funny, I've never heard a cogent definition anywhere so far, least of all in this thread! To reiterate my point:
The state of AI/ML pre-ChatGPT: research and industry churning away experimenting with all kinds of tweaks and tunings with all kinds of models to make narrow gains in very narrow fields. NLP, NLU, computer vision, speech recognition, text-to-speech, planning, symbolic reasoning, analytics, translation, sentiment detection, etc. were all areas of heavy research . "Code generation" that we see today was a pipe dream until 2015 or so. Ordinary users experienced AI/ML almost exclusively as a narrow feature embedded in specific products, such as recommendation systems or spam filters or ads.
Post-ChatGPT: Any researcher or any user anywhere can prompt a model in natural language, or audio and video. with any arbitrary task, and get largely useful results.
The very concept of a single model being able to even attempt an unlimited variety of tasks, doing surprisingly well on many of them and continuously getting better at everything, coming from the previous deeply fragmented state of the art, I think reasonably qualifies as "general."
If there's another take on what AGI means given the background of the field, I'd love to read it!
In any case, I think this misses OP's point that LLM capabilities have rapidly made progress towards being more generally intelligent and capable, which is not true of most tech advances.
This is a motte & bailey moment. Parent comment stated something much sharper, that I responded to:
> I do not understand how you can see what has happened in the last 3 years and say "it will not get us there" no matter what "there" is.
--
Your statement is something much weaker, and I would still question what exactly "general" means when AI capabilities are commonly accepted to be so "jagged".
P.S. It's hard to deal with so much inconsistency. This person says they fully agree with
> I do not understand how you can see what has happened in the last 3 years and say "it will not get us there" no matter what "there" is.
and at the same time says
> When people have a very narrow 'confidence interval' about their AI predictions, in either direction, it's difficult to trust them.
and says
> The last 3 years of progress have been so explosive and, yes, general that it seems crazy to fully rule out dramatic future progress
and then moves those goalposts to "AGI, RSI, the apocalypse, whatever".
Fine ... let's not trust folks who are so confident that the progress in last 3 years entails that "it will get us there" where there includes AGI. Again, it's an argument against a strawman to say "it seems crazy to fully rule out dramatic future progress" -- everyone agrees that there will be dramatic future progress. But it's obnoxious naysaying to insist that "the last 3 years of progress" entails future AGI, when numerous people have explained why they think it doesn't. [Generally speaking, it's not a good idea to characterize as "crazy" people you're trying to persuade.]
And qualifiers matter ... so when someone says [emphasis added]
> yet we are not that dazzled by our calculators today (despite how useful they might be!)
it is yet another strawman argument to respond
> Speak for yourself, I am dazzled by calculators!
And on the "general" front, there's a big difference between AI that can address a wide range of problems -- what one might call gAI for varying values of "general" -- and AGI, an artificial replication of the sort of generalized intelligence found in human beings, who can learn to fish by being shown how to fish, or learn to repair an engine by being shown how to repair an engine, or learn how to play a game by being told the rules of the game, etc.
I read all this an am baffled about how you're seeing inconsistency.
All of these statements say the same thing: "People are sure about a low upper confidence bound for future AI. My upper confidence bound about AI is high; the future is massively uncertain"
>I think this misses OP's point that LLM capabilities have rapidly made progress towards being more generally intelligent and capable, which is not true of most tech advances.
A PC of today can accomplish many more "general" tasks than one of 40 years ago.
Much of the "why" is because of the huge infrastructure built up around them in the meantime. The abilities of LLMs to accomplish those same tasks through the PC is heavily piggybacking on that (both in the specific, with the existence of all the APIs and tools; and in the generic, using search engines to find specific sources and using that for instruction or troubleshooting).
In the world of "agents" much of the improvement appears to have been on a specific set of skills: impersonation of an 'I' that wants to accomplish a goal, and synthesizing existing information from documents with trial-and-error execution loops to move rapidly toward a solution much faster and with less boredom than a human would. The quality of the output when there is not a rapid-evaluation-and-validation harness lags considerably.
It's incredibly powerful automation but doesn't appear to be trending towards Matrix-style conscious AIs. The quality of an individual method written by the agent also is not particularly advanced compared to GPT-4 in early 2023, as far as I can tell—I was dabbling with trying to make such harnesses back then, where a major challenge was that the model itself was bad at staying on-track in a conversation, so instead much of that logic was moved to deterministic code, which was much more limited as it was super-tedious to enumerate all the necessary tool calls/etc to find its way out of corners. Staying on task is much better now, as is "read compiler error, fix try next thing" harness loop-handling. But the output remains—across Fable, Astra, whatever else I've tried—"iffy" in terms of the actual code structure on the first pass output. You can set it then on a different task to review and clean up the code, and it can do that well too, but it is a curious gap of generality where the "create" focus is much more limited than the "review" one (and conversely the "review" focus can make suggestions, but if it goes deep down the well of implementing them, loses that big-picture again).
If it kills us all, it will because someone decided to give the trial-and-error-loop-machine access to nukes or similar. The blame for that is on the "someone" not on some sort of "rogue" AI.
(I wonder if re-watching Terminator/Terminator 2 would support this sort of interpretation of it. Unlike in the Matrix, I don't think we get much sentient-AI POV/infodumping. Is it a plausible universe for "someone made ChatGPT control a fleet of soldier robots and gave it a bad harness with an insufficient sandbox"?)
In the past 4 years the definition of what AGI would even be has been changed so many times that new marketing claims fit the description that the term itself no longer has any meaning.
However, it is reasonable to assume that the current architecture will just not be it. It works well for certain problems, and falls apart for several others.
Unrelated to your comment, more a general comment on the online discourse:
What I don't understand is why people get so defensive about and offended by the thought that more progress in the basics of the approaches could be needed. That doesn't change the performance of LLMs now, and it also doesn't change what could be possible in the future.
Why are people so deeply and personally invested in the current state of the art?
> It works well for certain problems, and falls apart for several others.
Which one does it fall apart?
I am not personally invested in the LLM design at all, I just don't see how you can say that an approach will not solve X when we still see an increasing rate of improvement, not even a slowing of improvement.
This an absurdly circular argument. You are simply asserting that, if we had had today's systems 4 years ago, we would have would have called them AGI, even those who don't call them that today.
> Now I don't know if what we have is AGI or not
But you just asserted that we do.
> I do not understand how you can see what has happened in the last 3 years and say "it will not get us there"
They explained how. Your argument, OTOH, is nonsensical, like suggesting that, because a bike path was built rapidly, it will soon become a highway.
You misunderstand. Our definition of AGI is constantly shifting the more capable AI becomes to the point where it is always clearly not there yet and clearly we will not get there with current tech.
> But you just asserted that we do.
No, I say we would have called the current state of things AGI in a pre chatGPT time.
> They explained how. Your argument, OTOH, is nonsensical, like suggesting that, because a bike path was built rapidly, it will soon become a highway.
I'd say it would be something like:
"If a bike can drive 100 km/h, it's clearly a superbike"
[3 years later]
"This bike drives 100 km/h, so it is a superbike"
"No, only a bike that is going 500 km/h is a superbike and this bike is still very far off"
Superhuman performance at chess probably would have blown people’s minds in the 1950’s. We’ve since learned that sometimes intelligence can be narrow and sometimes it can be spiky, even if you can have a decent conversation.
It’s hard to point to anything and say it’s impossible. AGI doesn’t break any laws of physics. But some things like driverless cars can still be a long slog to get to widespread deployment.
Waymo is the only self driving car that works well enough to even try widespread deployment with pricing that isn't going to be an operating cash flow disaster. But they're going to have to at least double their footprint to put a meaningful dent in the billion dollars a year Waymo spends on R&D.
Waymo isn't LLM powered. Everyone who thought LLMs would make Waymo Driver obsolete were wrong. In part because Waymo can't be spiky. In part because practical applications of AI require a lot of time and effort.
I am amazed at how usable and useful coding agents have become since they were a hot mess about a year ago. I would not be at all amazed if there are only one or two other use cases that have the same favorable evolutionary trajectory.
Those are just the same capabilities than before, but with a much bigger compute power and training data behind it.
AGI can't be reached by "training harder" as, the way I see it at least, it requires a qualitative leap, not just quantitative.
We are getting a machine that better navigates across the information in its training data, we are not getting a machine that can think out of that training process, even if it can fool a few people at that.
> 4 years ago, a program that could create photorealistic pictures, talk to you in any language of the world and solve the hardest math problems that we know, we would have called it AGI.
And yet, hand it simple task out of domain, it fails badly. If a system 4 years ago did all you describe without being trained to do those exact things with unfathomable volumes of data, i'd agree, but we don't have that. We have something that can replicate seen patterns effectively.
There is no reason to view the transformer vased LLM family as anythi g other than a translation machine. From the language of english spec sheet to python language script, for from image description to image, etc. The only real read we can take from current behaviour of llm's is that the relationship between domains we think of as seperate are likely joined more tightly than previously assumed.
It has not been even 4 years since ChatGPT hit and LLMs + Transformers + Whatever they do has gotten us to solving millennium problems.
4 years ago, a program that could create photorealistic pictures, talk to you in any language of the world and solve the hardest math problems that we know, we would have called it AGI.
Now I don't know if what we have is AGI or not but I do not understand how you can see what has happened in the last 3 years and say "it will not get us there" no matter what "there" is.