Wouldn't "general intelligence" require so much more than scoring well (or even amazingly) on benchmarks?
Like what about having some "AGI model" embodied in something (maybe humanoid), and test it by having it step in an assortment of cars and park them. Does bodily-kinesthetic intelligence account for nothing? Humans are intelligent creatures and can dynamically adapt to the physical shape of a variety of vehicles and their movement characteristics. And there's so many things like this that are extremely basic, which some people dismiss since practically every human has the capability to do it, but actually requires a high degree of intelligence.
What you've described is just a new benchmark, though. It'll be called CarParkBench, various embodied LLMs will then be run against that benchmark, and some will score better than others.
I do see where you're going, but that's already what's happening: we have so many different benchmarks because there's no real single way to test for general intelligence.
Also, it takes a human probably at least a decade of world experience, growth, learning, etc, to pass your benchmark. I'm quite confident that it will be very soon that an embodied LLM will pass your new benchmark, much sooner than a human would take if born today.
I think the issue is less about creating a new benchmark and more that the existing benchmarks shouldn't be called anything related to AGI unless they measure AGI.
If a model couldn't go to work as e.g. a first year apprentice plumber on their first day and perform anywhere remotely close to the median, but can pass a benchmark that claims to measure AGI, the benchmark is wrong and the model is not exhibiting general intelligence yet. ApprenticePlumberBench sounds like it's genuinely better than ARC-AGI at measuring AGI and that's a bit silly.
(Edit: I wrote ARC-GIS the first time around, for some silly reason)
It's not measuring AGI at all, it starts from human "core knowledge" so it is parochial. It is made of tests that still fail so by definition next version will also start low. Moving goalpost.
This is a big reason why I feel like even though LLMs are _effectively_ AGI in some regard, they also are a hack around what most people figured AGI would look like before the advent of LLMs. Humans can do metacognition, output multimodally at the same time (verbal _and_ physical intelligence go together to produce an expressive face while one talks), have a good sense for what they do and don't know, continuously take in and respond to the world around them in a (mostly) uninterrupted fashion without "turns", learn knew knowledge and retain it for their whole lives, etc. When you reduce a human to a text generator, yes obviously SOTA LLMs perform way better, but rather than invent something that can operate as an always-running "being", we've grafted a harness around an intelligence that is bound purely to speak only when spoken to. Maybe organic intelligence is already that, playing out at a super high refresh rate, but I don't know.
Current AI is arguably much more capable of multimodal output than humans. It can produce an incredibly vast variety of audio, images, and video. Humans are limited to producing the sounds we can make with meatflaps in our throats, and contorting various parts of our bodies to produce crude symbols and shapes.
(Very capable!) Embodiment, persistent operation and continuous learning are indeed things that still set us apart from AI. None of those are fundamentally difficult to solve, though.
More importantly, none of those are particularly relevant for being "intelligent": If a criminal threatened to kill your family unless you solve some difficult problem that requires only intelligence and you could choose any single person, animal, or AI to help you with it, which would you choose? Be honest.
The thing I trust the most to solve tricky problems reliably is a specific very skilled programmer I've known for twenty years.
I wouldn't say he never makes mistakes, but his success rate is a damn sight better than any LLM I've ever interacted with (and I drive Opus daily, due to corporate demands to use LLMs).
Same partial answer as I gave your sibling commenter:
"OK, so the task ends up being to write a 40-page analysis of the result of a specific experiment in quantum chromodynamics, to be finished within 2 minutes. Does your choice for a human work, or would you have rather chosen any of the frontier AI assistants? Be very, very honest. Remember that the lives of your family are on the line and nobody will judge you for a lack of allegiance to humans."
This task is intentionally designed to ensure a human cannot do it.
The initial scenario is utterly, insanely absurd to begin with, but I tried to go along in good faith and gave you the true answer.
The result was a bad-faith rhetorical trap, so I'm done with this thread.
In another attempt at good faith, as part of bowing out I will add some actual response to your anti-useful cheap rhetorical trap:
I do not trust LLMs to get things right in high-stakes scenarios. I have seen the current models spit out falsehoods and errors regularly in the handful of fields I have expertise in, and have no reason to think they would do otherwise outside my expertise.
The scenario you describe is an absurd fiction, and no human making the absurd threat could evaluate the paper in less than hours (realistically even an expert would need days, and a nonexpert could not do it at all [short of becoming an expert]).
So, there's no point trusting a bullshit machine to save my family - it might very well get them killed, and whether it was right or not, what would actually matter would not be its correctness, but what the presumable bullshit machine evaluating my offered input spits out.
So, the best move I could realistically make would be to put a stab at prompt injection into the input.
For that job, I probably would actually prefer aforementioned programmer over any other option, come to think of it - I suspect he'd have better success than even another model (especially considering the safeguards the models no doubt have to try to keep users from using the models to inject other models).
Again - I'm disappointed in your worthless rhetorical cheap shot.
I suspect you'll have much better success convincing people LLMs are intelligent if you engage in good faith, listen to their perspective, and address their actual thoughts, instead of devising the sort of inanity that comes out of high school debate clubs, where people literally want to score points instead of find truth.
> This task is intentionally designed to ensure a human cannot do it.
It is one of a vast array of things the hypothetical kidnapper could come up with, some of which humans I agree will do better at (currently) and some of which AI will do better at. We clearly agree that in that array there is at least one task that a frontier AI would be better at than any human you could pick.
It is a thought experiment, so there is nothing fundamentally wrong with it being extreme or unrealistic (thought experiments very often are), but for the sake of goodwill let's 'weaken' it a bit: the AI or human always has an hour to come up with the answer, the question and answer are in their preferred language, and the answer fits on four pages. You can't help them, though; The criminal 'prompts' them. They can use the internet as an informational resource, but they can't communicate/ask for help/post anything (with the spirit of this being: no loophole in letting somebody else do the task or parts of it for them). And of course all subject matter of all complexity is fair game (including but not limited to quantum chromodynamics experiments).
Given that situation, do you think the programmer you mentioned would be more successful than a frontier AI in more than 50% of the possible intelligence tasks?
Edit, addendum: Please, if you can, also let said programmer read this thread and give his opinion on it. It sounds like he would have interesting things to say on this.
Before "AI," humans have created a vast array of "multimodal output" (computer art, instruments, dance, architecture, etc.). Why are you giving the AI a harness and a plethora of tools and not the human in this comparison? Without these, the LLM too would be utterly useless.
Also, that situation is extremely contrived. If a criminal threatened to kill me if I misspelled a word, I would choose a dictionary. If a criminal challenged me to predict a next token, I'd choose the LLM. For all real precarious dangerous situations, I would obviously choose a human. Like immagine the hilarity (or tragedy) that would pursuit if ChatGPT tried to handle a hostage situation or a plane hijacking.
and those meatflaps are normally called vocal folds/cords btw
> Why are you giving the AI a harness and a plethora of tools and not the human in this comparison?
I am not. Multimodal models generate that output directly, without tools. Which 'tools' does AI use to generate all those images, songs, and videos do you think?
> Also, that situation is extremely contrived. If a criminal threatened to kill me if I misspelled a word, I would choose a dictionary.
Of course it is contrived, it is a thought experiment. Does not make it less valid. It is essential that you don't know what task it is going to be, just that it is a task requiring a lot of intelligence. This way question dodging loopholes like "I'd choose a dictionary" are impossible (and people will always try to find some cheesy exit rather than facing reality). You have to commit to something or somebody that has broad and general intelligence; you do not have the luxury of choosing the perfect tool for a very narrow task.
> For all real precarious dangerous situations, I would obviously choose a human.
OK, so the task ends up being to write a 40-page analysis of the result of a specific experiment in quantum chromodynamics, to be finished within 2 minutes. Does your choice for a human work, or would you have rather chosen any of the frontier AI assistants? Be very, very honest. Remember that the lives of your family are on the line and nobody will judge you for a lack of allegiance to humans.
Again, don't go for shitty loopholes. Engage with the thought experiment in good faith and thus as it is stated, not some conveniently distorted version of it.
> and those meatflaps are normally called vocal folds/cords btw
What? Next you're going to tell me that meatspinner is also not the name for the human male reproductive organ.. Maybe I need to get a refund on my Temu Gray's Anatomy.
That said, a look at the state of self driving and the recent robot olympics shows that advancement on that has accelerated enormously, though whether it's reflected in any of the LLMs is something else entirely.
>AGI has a pretty precise definition, covering only cognitive tasks.
OpenAI's own charter defines AGI as "Highly autonomous systems that outperform humans at most economically valuable work". This is actually fairly sensible and involves obviously a ton of non-cognitive, emotional, social and physical activity. In other words, if you can replace most or all human beings with a machine, you have something that's generally intelligent.
That's obviously not even remotely where we're at, AI chatbots do well on narrow usually text based or programmatic problems, but can't even replace a barista or a plumber.
There is rapid progress in generality in humanoid robotics though. I think within the next year or less we will get the ChatGPT moment for humanoid robots. If you look at progression of capabilities such as the recent Skild AI demos.
There are simply no technologies today that can replicate the density, precision, and versatility of human touch sensors. Until then, there is simply no way to create generally capable robots that can operate at the level of a human.
And unlike LLMs, advancement is held back by physical limitations like materials science, so progress has been and will continue to be much slower.
What task do you think that humanoid robots can't do? Also, we don't need fully equivalent touch to get useful performance.
If you look you will see a really broad range of tasks accomplished already, including thing like manipulating screws, picking up pills, inserting wire harnesses, folding clothes, putting away dishes. And there are several companies with built in or component advanced touch sensors like Figure or leading edge touch sensor companies like SynTouch and GelSight.
Peel an orange? Crack an egg? Thread a needle? (Heh, drive a car...) There's a huge range of tasks that a non-specialized, general purpose robot simply cannot do yet. I'd be easier to enumerate the things they can do than the things they can't given the current state of the art.
Sure, build an orange peeling machine and it'll do great. But that's not what we're talking about here.
As for those demos videos we often see, those are very highly choreographed demonstrations. Show me a real life humanoid robot operating free form on a factor floor and doing those things and I'll be impressed.
And to be clear, this is not meant to understate what's been accomplished. I'm just saying the path for advancement is a lot harder and based in physical rather than computational limitations, which are much harder to overcome and go much slower. We simply cannot look at the growth curve of LLMs and expect robotics to advance at the same rate.
peeling an orange and cracking an egg already demonstrated. Figure 02 worked on BMW's actual Spartanburg production line for about 1,250 hours running 10-hour shifts.
keep paying attention, you will see how wrong you are about it being physical limitations as the physical AI continues to rapidly improve.
Surely part of the problem is that intelligence seems to be implicitly conditioned on embodiment, to the point that "covering only cognitive tasks" seems inherently ill-defined or arbitrary. Everything we do is a cognitive task. At some point, the criticism will be "sure, it can solve research-grade math problems, but it can't fold my laundry".
Even our large language models have an implicit embodiment in the domain of text (and more recently, multimodal inputs). That seems sufficient for certain things, and insufficient for others. I suspect that AGI that does everything a human can do eventually turns out to be fairly analogous to humans in terms of sensory input and domain output, even if the scale is radically different (e.g. thousands of robots uploading (touch, sight, audio, smell, etc.) sensory data to a single model, and each being actuated individually).
Like what about having some "AGI model" embodied in something (maybe humanoid), and test it by having it step in an assortment of cars and park them. Does bodily-kinesthetic intelligence account for nothing? Humans are intelligent creatures and can dynamically adapt to the physical shape of a variety of vehicles and their movement characteristics. And there's so many things like this that are extremely basic, which some people dismiss since practically every human has the capability to do it, but actually requires a high degree of intelligence.