If a foundation model company burns billions of tokens to brute force an LLM into finding a new training algorithm that e.g. allows recurrent networks without catastrophic forgetting...
I won't really care that it didn't have a "real measure of understanding". I'll care that it has made an even more dangerous technology, which needs work to make it aligned.
We've then improved that through systems similar to prolog intentionally searching a tree.
Then systems added heuristics for which paths in that tree are likely to be taken.
The LLMs are just using slightly more accurate heuristics for this task.
But the real measure of understanding are tasks that are not so strictly constrained.