For example, LLM pretraining datasets are on the order of tens or hundreds of terabytes. If an LLM-based code for that data is ~twice as efficient as gzip, you could afford to transmit the weights of even a very large LLM and still come out ahead.
In other words: LLMs actually are excellent compressors of their training sets in the formal information-theoretic sense.