> Hutter Prize being where you are paid if you can compress wikipedia small enough. LLMs do very well at that, if, big if, you ignore the cost of initial weights.
then.
If all you care about is compressing Wikipedia, but ignore the size of the actual data, what is it that you are actually trying to do?
Well, then Wikipedia itself is a very good compression that only needs the title to perfectly predict the full article.