I had a long ranting comment I deleted. I just don't like this trend of people presenting work in a way that makes you think some combo of 1) they discovered from scratch themselves 2) it's new 3) they didn't try to cite or acknowledge where they learned it/point to good sources 4) they don't really care about trying to teach something deeply, they want shiny stuff that makes them seem deep.
This post references specific parts/calculations, but you'd never know it was not news if you didn't know better.
The author of the post uses standard terminology like entropy coding and arithmetic coding, and cited a paper "in 2023, Google DeepMind released a paper arguing that language modeling and compression are two views of the same thing" which discusses it further.
This blog post is great. Well explained, and clearly took a lot of effort.
I don't interpret it as them claiming to have to discovered it independently.
You expect every blog post to find the earliest relevant paper to cite, just so one could look at the year (without reading said paper - which would have made clear that the connection isn’t recent) to assess novelty? I don’t think that’s reasonable.
It’s a blog post. If it was, say, a peer reviewed paper by Hinton or LeCunn that fails to cite Schmidhuber, that would be reasonable criticism in my opinion. (Spoiler: they fail to cite him)
Why would blog posts not be subject to such criticism?
Either the author knew of prior work that argues the same thing and they ignored it, or they didn't know. And if one writes a 1000+ word article premised on this idea, wouldn't one be presumed to know at least in which century the idea originated from?
Arguably these kind of blog posts should be more subject to such criticisms, because the blog posts purport to "teach" the general public about a concept in an authoritative tone (or at least the author seems to pose as knowledgeable in the subject), while for academic papers, everyone who actually reads the paper knows where the ideas came from anyway and it's mainly an issue of attribution (and maybe about fairly distributing the citation count...)
> Why would blog posts not be subject to such criticism?
You're asking why casual comments from amateurs made for fun on the internet shouldn't be held to the same standard as those made by funded career academic experts writing for other experts over months and meant as part of the permanent record of the field?
Personally, I think that's a bit like asking why a friend having you over for dinner isn't always an elegant 7-course meal with wine pairings. I guess you can expect that if you want, but to me it sounds like a child expecting to go to Disney every day: ignoring the economic realities of the situation is a recipe for eternal disappointment.
This is why I hate online arguments. I didn't say earliest, but to give a pointer to the general era, and make it clear how standard the concepts are. They are part of undergrad, it's basic things. I'm not asking for a deep lit review. But modern exposition to anything related to Ai / ML / stats have extreme recency bias and young'uns are led to believe there was nothing before the transformer paper.
I guess it depends on the actual goal of the post. Education? Point people to background. Surely if you have a deep understanding of the subject, you have favored references.
If it's not education, where people are supposed to know everything, who is the audience? If the audience is supposed to know all of the background, then the post is not saying something original to the audience.
I agree the quality of the text itself is good. But I knew the background and I left (well, actually began) already agreeing.
I am happy and want people to write and share stuff. I think it's a really good exercise if only for the author. To be clear, I think the post itself is good except for the gap I mentioned in my original post, which in the grand scheme of things may seem like a nitpick. But if the same type of content makes it to the front page often without me understanding who the audience actually is, I am also allowed to wonder why out loud.
If the blog were about calculus, and stated that an elegant proof of the Fundamental Theorem of Calculus could be found in such and such undergraduate textbook, would you be upset that the citation wasn't to either Newton's or Leibniz' work?
I would want them to point to point to what they think is the best resource. (Assuming the point is educating or explaining something to the reader.) Which, If everyone followed, would automatically point back to newton/leibnez.
"Understanding" is clearly linked to compression. Taking observations and coming up with a more compact representation that explains them, analogous to coming up with a compact set of axioms that generate facts, or a small Turing machine or short program that generates a list of strings.
Intelligence is a broader concept but definitely involves understanding how a system/envoronment works and making predictions about its unfolding, especially actionable ones that allow you to steer that state towards some goal states.
The two contributioms that come to mind are the Hutter cash prize for the best compressor of the English Wikipedia and the work on PPM compression-based
text classification by the late Prof. Ian Witten's group at Waikato (NZ) [1,2].
The model that best compresses the input string was likely generated by the distribution from which the compression model was 'trained'
I didn't say they claimed that. They left it rather vague, intentionally or not. I didnt mean to pick on just this post, I just happened to see a few such posts on the frontpage recently. I do think it was a good and obviously effortful blog post, overall.
The post says this is all part of gzip and LLMs, what are you saying? I’ve been using gzip my entire life. I read between the lines “this is common knowledge” throughout the piece. Throwing in some names and dates only makes this super clear story harder to read (and more like studying then the playful exploration this post was intended as).
A short paragraph at the end on the origin of these ideas can be an easy way to dispel the misconception of potential beginner readers that the insights are novel.
I'm of two minds here. The pro is that the "you could have invented this" walkthrough from first principles is more engaging than "and then so and so introduced this term in 1972 and the definition is such and such". This style is a reaction to that boring and dry teaching style and tries to push towards what eg Feynman pointed at in the Brazil critique.
The con is that you don't get to understand and see any of the history of the ideas or even the ballpark when it was discovered, you attribute it to the blog mentally and you don't know what is how new or old and can't reference it properly when talking to others.
Your theory is that anybody who writes anything is obligated to make sure you can find any related information with one Google search? Again, to me that looks like wanting to be spoon fed.
I have no idea why you think the world owes you endless 101-level discourse, but I hope you recognize you're setting yourself up for equally endless disappointment. If you take a little responsibility for your own education, you'll be happier.
My theory is that anyone wishing to elucidate a complex technical concept is better served with contextual details.
I think one method is superior to another.
Anybody, anything, obligated, and any related information are hyperbolic terms I would have avoided. I understand with a weak position people often resort to hyperbole.
I believe adding more terms to be googled is superior to fewer search terms. When working with ideas from first principles sharing the roots and development is superior to making no reference.
I have no idea why you think the world owes me anything. I was describing what I prefer.
I'm not sure why you think I'm not already endlessly disappointed by the state of the world and it's discourse.
not having expectations always helps.
I think that if you stop projecting ideals on strangers might help your mood.
If you were trying to describe a preference, you didn't do a very good job. You stated it as a universal.
But taking you at your word, if there are things that don't meet your personal tastes, maybe move on to the next thing to read? Not everything has to be for everybody. There are plenty of audiences where people are capable of digging deeper when they want it. Personally, I prefer writers who don't spell everything out. The kind of writing you favor I generally find tedious. I'd much rather read something that shows a little faith in the audience.
I was replying to a comment that was undecided and outlined two choices. I described the reasoning of my preferences and even provided a caveat that my reasoning applied to a subset of the audience.
The overall message and impression of a text has to give the right picture regarding what's new, where the info comes from, etc. It's not about any kind of Google search. It's not about extra info. It's about intellectual honesty.
Often the most straightforward way to walk through an idea while teaching it is not the same order that the ideas were developed, and might not even use the same set of ideas in building up to it, so it can be tricky to get both the best explanation of the idea and the historical context in at the same time without making things more confusing.
I like to see ideas presented as the evolved. Each solution is developed as a perceived reaction to the shortcomings of the previous. This becomes a contrast and comparison as to why one idea is appropriate for a particular context.
I agree. Simple albeit imperfect solution would be pointing even to a wiki page. That alone would signal acknowledgement/awareness that it's building on other stuff.
I'm speaking in general terms and I have not studied much on this subject so I don't have specific suggestions for this post. Please refer to earlier posts in this comment chain for a general idea of what citations would be useful.
First off: What is your goal here? My comment was one sentence and wasn't specific to this article. Why are you acting like I said things I never said? But okay I'll try to answer.
> What are your thoughts on the other comments in this chain that claim that the references and citations as they are in the linked post are acceptable?
> Why do you disagree with them?
Which comments in particular are you talking about?
And I mean, the post doesn't have "references and citations" plural that I can see. It has one clear reference. I see two comments defending that reference, and a whole bunch of comments saying that a post like this doesn't need references.
I don't need to disagree with any of those comments. My point isn't about whether a post that lacks citations is "acceptable". I was talking about whether a post could add historical citations without losing momentum, so it could do a better job of fitting into context. A dozen notes sprinkled around about when/where different ideas came from, not just the mention at the very end of what decade arithmetic coding is from and the mention of the 2023 paper.
I will say it's bad that the 2023 paper is described as "arguing that language modeling and compression are two views of the same thing", emphasis in the original. That very much makes it sound like a new idea when it isn't.
I don't mean to pick on you, it's just then just saw a bunch of negative comments on an article I enjoyed reading and found informative, so I picked a few to reply to for clarification of their position.
I'm confused by the comments that seem to imply that the author is taking credit for ideas that they're writing about, as I didn't get that impression from writing in the article at all.
I'm also confused by the idea that they didn't provide enough citations for concepts in a non-academic source. I feel like in the day and age of Wikipedia it should be sufficient to colourize / bold terms like "Entropy coders" or "Huffman Encoding" to indicate to the reader that these are important terms that they can research for more details if interested.
Overall I guess my position on this is that ~80/285 comments on this post are related to whether not this is a good post which I think is far too many. There are slight issues with the post but they aren't worth 80 comments.
Where I replied to wasn't that the author is taking credit but making it hard to get the context for where the ideas came from. (With the one citation implying a very wrong thing about how old that idea is.)
Specifically the comment I replied to suggested it was a style tradeoff but I don't see why you can't do citations in either style. That's all I came here for.
> I'm also confused by the idea that they didn't provide enough citations for concepts in a non-academic source.
It's useful information about how we got here. That's good in any long story about an idea.
> I feel like in the day and age of Wikipedia it should be sufficient to colourize / bold terms like "Entropy coders" or "Huffman Encoding" to indicate to the reader that these are important terms that they can research for more details if interested.
It works to some extent but some of the concepts don't have fancy words for them and then they don't get that searchability.
> ~80/285 comments
Sometimes tangents spark a discussion. I don't think those numbers are an issue. And they're mostly not calling it a bad post, they just have an objection.
They did not come up with the ideas themselves, so they got them somewhere. Follow the source and all the citations show up. It must be a modern thing where online blogging randos pretend they are all geniuses.
Better too assume they are just not aware. Technology is multi-layered cake of development. I have no doubt the only reason I know a lot of details is that I lived their development.
When standing on the shoulders of giants it's hard to tell what is below them.
If they are not aware they are geniuses who rediscovered a bunch of shit in one blog post of effort. Better not to assume anything and let the writer tell you what's going on, if just an endnote
You're reading this the wrong way I think, citations aren't given because its obviously a pedagogical article about well established stuff. Much like you wouldn't give citations in a blog post explaining calculus.
One could give citations regarding calculus it's pretty interesting. Since it was done twice by both Newton and Liebniz. There must have been cultural developments in the 1660's that demanded calculus be invented.
The first sentence says she came across it when reading about compression. I didn't read that as her claiming to have discovered the idea or that it was a new idea, I read it as "today I learned". I think somebody who was unfamiliar with how compression algorithms or language models work would find this an approachable and interesting introduction. Not everybody studied information theory.
If you followed the data compression scene in the 80s and early 90s, there were plenty of reinventions of LZ-ish and Huffman-ish algorithms (I also coded my own variant...), and people even tried to patent some of them, so at least for the basics I think it is something that many can discover independently; of course in these times, it's more likely they didn't.
I'm glad to see someone feels similarly. There is nothing wrong with ignorance, but there's no excuse mistaking learning for invention. Especially from someone bearing the title "Developer Educator"
I don't think it's the case here, but worth noting too that LLM-written blog posts adopt this tone seemingly by default.
Never the least bit of surprise, wonder, doubt, or frustration to get in the way of the steady staccato beat of metaphors, conclusions... and three-item lists.
> I'm glad to see someone feels similarly. There is nothing wrong with ignorance, but there's no excuse mistaking learning for invention. Especially from someone bearing the title "Developer Educator"
>> a Developer Educator at ngrok with a passion for nerd-sniping developers.
Perhaps tangential to your point, but I often write blog posts (although finish and publish far fewer than I start) where I write about something as it has occurred to me, informed by things I've absorbed no doubt, but without specific research. In such cases I explicitly avoid searching out prior work as a) seeing that something is well discussed and explored can take away the motivation to explore (in the same way reading puzzle solutions before starting might), and b) to avoid having green shoots of ideas shaped by the current of existing consensus. Now that doesn't mean I don't come back after doing my own thinking to see what the more well developed literature of people cleverer than me, who've thought far longer than me think; I just don't want to snuff out my own exploration at the start.
As I say most of these I never publish as I'm mainly using writing as a vehicle for thought, but when I do I'm never sure how to flag them. I don't want (imaginary, lets be honest) readers thinking I'm deluded into thinking I've found something new. I want to come up with a tag I can put on them which adds a pithy disclaimer card at the top or something so I feel more comfortable publishing them.
I find it bothersome that language works this way. You can spend your whole life discovering things that are well known by the rest of the world. But the minute that you mention to a large group that you “discovered” it, suddenly you’re taking credit for discovering it for all of mankind.
Yes, and it's made pretty obvious with all the background section and references.. which btw I'm not advocating for such baggage in a blog post. Just something.
And yet to this day, in AI threads, so many people act shocked and surprised if you dare follow the obvious implication and claim that understanding is a form of lossy compression.
Of course it is but again "X is just Y" is often used to mislead. A brain is just neurons! A computer is just transistors! An LLM just predicts the next token! It's just like a parrot! It's just like a blurry jpeg of the internet! Kinda yes, but what do you use this for? It's a bad intuition pump is it leads people to conclude demonstrably false things about capabilities.
Shorter description isn't understanding, let alone of it is lossy.
When you shorten a description in a lossy way, you are deciding a priori that some differences in the object don't matter, and it's not because you understand the object, but because it serves your goal of shortening the description.
You can compress syntax, losslessly even, with zero understanding of its semantics. Zero understanding not only imbued into the compressor/decompressor, but even the designer of the compressor doesn't require understanding the semantics. Actually, even of the syntax.
A compression program can compress a book written in a language that the author of the program doesn't understand, on a topic he knows little about.
Finding common characters and building a list of words is a low level type of understanding. Doing it better does actually start directly representing syntax patterns and that's a less-low level of understanding.
I think "losslessly even" is the wrong way to think about it. Lossless compression often requires less understanding than high quality lossy compression. If you can do a lossy compression that correctly decides what details are unimportant, that's a good sign of understanding.
> I think "losslessly even" is the wrong way to think about it. Lossless compression often requires less understanding than high quality lossy compression. If you can do a lossy compression that correctly decides what details are unimportant, that's a good sign of understanding.
This is the crux and reminds me of things like mp3 that exploit the nature of human hearing being limited to a frequency range.
We model the data. The model, hopefully, captures something real in the data. If it does, then it's fair to say that we understand the data better.
But it's frankly a philosophical question what's real or not. No model is going to capture absolutely everything about the thing it models - at that point, it would be the thing. The best we can hope for is that it captures everything we care about.
And no experiment or metric can tell you if you care about the right things. At best it can tell us if we care about a thing given other things we care about. "No cares in, no cares out".
To make it a little more concrete: you could compress a string from back to front. You could build an LLM to help you do that. If you care about file size, that's almost certainly a bad idea, the forward LLM will be better for that purpose. But are there purposes for which the backward LLM might be better? I think that's not so hard to imagine. Often we wonder about "what came before".
I'm not aware of a better definition of "understanding" that would allow me to tell whether some system "understands" some other system. Do you happen to know one?
See: A. M. Turing (1950) Computing Machinery and Intelligence. Mind 49: 433-460.
I mean, my interpretation is that the question Turing tried to answer is equivalent to "How can we determine whether machines understand humans/human thought?"
This only works when both systems can talk about pretty much arbitrary things, but if you want a more general method for less complex systems, perhaps having one system simulate another system is sufficient. (Which is also another Turing invention)
There are many people who would claim that passing the Turing test is insufficient to show "understanding" (compare for example the Chinese Room thought experiment).
Yes but it's (kind of?) a definition as you asked for.
At this point, I am unaware of a better definition. I know the Chinese Room argument (and I disagree with it), but I'm not aware whether the proponents of that argument have a better definition of understanding other than "well, the Turing Test isn't enough"...
---
PS: Interestingly the issue of compression is highly relevant regarding the Chinese Room argument -- the essential element in the Chinese Room argument is that the information is not compressed...
I was speaking in the context of humans. When someone teaches you, the content coming from the teacher is very compressed. One decompress it when they can generalize and apply it. So understanding is compressed, but is not the act of compressing. I mean it is not compressed from a larger data or made by compressing a larger data. The larger data it represents never existed. It is like the definition of a fractal...
I think of it as more "providing you with the model you can use to decompress". You only really learn something when you do it yourself, not just listen to someone talk about it.
I'm glad you're pointing this out because not only are these old insights, but I'm also pretty sure I've seen variations of this blog post years ago on even HN already.
The author acting as if they discovered this independently had me feel the exact same way. Kinda irritating and almost ... disrespectful? Not sure of the right words to describe it tbh
> The ts_zip utility can compress (and hopefully decompress) text files using a Large Language Model. The compression ratio is much higher than with other compression tools.
It's not only an old idea it's been totally done already.
This post references specific parts/calculations, but you'd never know it was not news if you didn't know better.