AFAIK, the “large” qualifier came when transformers allowed to scale the size of language models compared to the recurrent models that where in fashion before. And although BERT isn't large by today's standard, it was large enough for the time.
idk the definition is fuzzy. thats why people use the "modern" qualifier to talk about decoder-only style and this is also not clean since you now have reasoning models which are separate
When I made this, the point was to show that you dont need pretraining (which is what makes an LLM) to perform well on complex tasks
And yes it is not a language model either. I did not train it on any language data. Only ARC puzzles