Having a clean DAG still breaks down with recursive Make: if binary X depends on library A, then X's Makefile either has a rule to recurse to A (and then parallel builds of A might happen, even when building from the top level, unless you use -j1) or it doesn't (and then you can only build from the top level Makefile and you might give up parallelism due to granularity of the top-level DAG).
At the time that essay was written, Make was the best runner available, and the only portable one. There are better options now.
What about those do you think are "shady"? Price discrimination in favor of small customers at the expense of large customers is somewhat common; businesses want customers to buy more of their products, but customers are not obliged to buy more if they like their current deals. Having limits on using finite resources seems even easier to justify.
Humans see differences in prices as unfair. The greatest example of this is price gouging during an emergency, but also look at the level of hate that scalpers get.
Using AI to do this, if anything, given the common negative sentiment, is seen as even worse. People think of this as AI using information asymetry to squeeze more out of users, not to cut people a deal.
Taking advantage of information asymmetry is generally looked down upon as well. Look at all the laws we have protecting kids from this. Businesses often don't seem similar protections because those are businesses with big legal teams (and when it is a big legal team vs a small mom and pop store without a single lawyer on payroll, people do start taking issues with it). The power difference between the average company using AI pricing and the average consumer falls pretty solidly in the 'we don't accept this' side of taking advantage of information asymmetry.
I could keep going, but I think these are already plenty enough reasons to why people look at AI price discrimination as not just a bad business practice they don't like, but an immoral/unethical one.
Sure, one can make economical counter arguments, but that's arguing on an orthogonal dimension that simply isn't relevant to where these feelings/thoughts come from.
Do you think AI subscriptions are analogous to either price gouging during an emergency or ticket scalping? I don't see the similarities.
It is fair to complain that subscription AI usage limits are opaque, but that is at best tangentially related to them being subsidized relative to API prices.
Shady as in not clearly communicated and change with random tweets being the only indicator, or a blog post if you're lucky.
x20? 25x? 5x? All in comparison to some other plan that also doesn't have clear limits. "Oh BTW we mean session limits, your weekly is 2x of 5x". "This model eats your limits twice as fast and you can use half of your weekly limits on it". "No this time we mean weekly only".
That doesn't resolve them from providing a stable service for something they market. I'm not buying something under the counter, and it's not advertised to be an alternative for people not willing to pay the full price. From my perspective, the subscription is the full price and I expect stable, reliable service.
You say that when the customer makes irrational requests.
Is it irrational to want consent went the subscription terms change? Netflix does it a month in advance, and cancels the plan when no consent given. What do you think is special with LLM companies so they get special treatment?
I'm from Europe so maybe this is why it feels so weird to me. Then again, they could stop offering their services to a region if they don't want to adjust their ways.
What gets me confused is how they aren't prosecuted for this. My theory is that people are too busy with the security threats and (currently) ignore the shady practices.
It depends. My personal experience is with RINEX, a text file format for GNSS observation data. gzip compresses large sets of files by about 4x. Someone named Yuki HATANAKA came up with a compact text format (CRX) that shrinks it by about a factor of 4.1x. You can gzip CRX for a total ratio of almost 12x, or bzip3 it for a total ratio of 17x. But if you combine Hatanaka's technique (essentially higher-order delta coding) with some others, you get a binary format with a compression ratio of 19x that reads faster than the CRX text.
When one's uncompressed RINEX files are 200 GB/day, it's worth some effort to shrink the files, especially if it means they are faster to read.
The 90% and 95% are against some blend of tasks meant to be broadly representative. A pricey model seldom fails a problem that cheap models do well, so there's stratification of tasks by difficulty. Someone doing novel research may be in the "hard" 15% of the blend, where P(solution) goes from one third to two thirds.
On the other hand, if it's cheap to tell whether you got a good solution, and you think the 90 and 95% apply to your task blend, then it's almost always worth trying the cheap model first.
> Shouldn't they be 64 bits on most modern systems then?
Arguably so, but then one would lose the ability to natively name 16-bit integer types because "short" would be 32 bits.
An earlier comment addresses x86-64. AArch64 (pedantically, the A64 instruction set used for AArch64's 64-bit execution mode) is similar, in that addresses are 64 bits wide but ALU instructions typically encode a width bit, called "sf", that selects either 32- or 64-bit data registers and arithmetic. See, for example, https://arm.jonpalmisc.com/latest_aarch64/add_addsub_ext .
long can't be smaller than int, so a 64-bit int leaves short as the only type between char and long long, and you need to pick whether it's 16 bits or 32.
Then again, we'd still have stdint around. And ultimately this doesn't matter because most code isn't portable anyway.
I don't think of the stdint types as being less native than the keyword integer types. C# for example makes keyword types aliases to System.* types, not the other way around.
I think you would need to have a compiler intrinsic type which is 16 bits. That is what the stdint.h file would define (u)int_16 to.
It would be odd for there to be a compiler intrinsic type to be unavailable until a header was included.
The compiler intrinsic type could be a mess like __int16_exactly_t but without it stdint.h would have some magic line which makes a compiler intrinsic available which wasn't before, or generates a new 16-bit type ex nihilo.
So you could have a "#pragma expose_extra_types" in stdint.h but that would not be the conventional approach.
I mostly wanted to make the point that a mere "at least 16 bits" type is not sufficient for some use cases, which is not in response to you but the comment you responded to.
I thought it was a really good, clear write-up of a very nice engineering discovery and optimization. If anything, I think an LLM would generate native superscript for the `2^(-r)` expression, so I didn't have an impression of LLM authorship.
Maybe an LLM did the search for what expressions to put in the blocks vs tail of the vectorized collision detection, but that seems like a good place to use an LLM as long as all of the expected bit relations are verified as checked properly.
Hi, author here. FWIW the blog is fully hand-written. I used Claude to help me generate the visuals.
What goes into head vs. tail is actually fully deterministically decided by the solver based on a P(tail) gradient, so expressions that are more effective in filtering out blocks before they can hit the (slower) tail are preferred in the prefix.
Allowing more than one lead byte is a complication of the existing standard. As many others have pointed out, we have plenty of coding space without that (e.g. by allowing 5- and 6-byte UTF-8 again), so the case for the extra complexity is currently not compelling.
> BOMs (and especially the hilarious UTF-8 BOM) are strictly a legacy Microsoft/Windows thing
How should a reader infer the bye order for a UCS-2 or UTF-16 file without a BOM? It seems like one would have to read until finding a code point that would be illegal under one ordering (but files might not include such a code point).
Similarly, a UTF-8 BOM is a useful flag to distinguish UTF-8 from other text encodings. You are right that the ambiguity goes away if those other encodings do, but people don't want to rewrite their legacy files. Some people don't want to use two bytes for common non-ASCII characters, so they are really attached to ISO-8859 or Windows-1252 or koi8r or whatever. CJK languages have their own encodings that are more efficient for their languages. UTF-8 is great for English speakers, but it's a compromise for everyone else, so they might reasonably want incompatible systems for their own use. UTF-8 BOM is a good "magic" sequence to detect encoding as long as people have non-UTF-8 files.
> How should a reader infer the bye order for a UCS-2 or UTF-16 file without a BOM?
Simple: switch to UTF-8 as the only encoding standard for sharing text data, keep UTF-32 as 'internal' runtime format for random access to codepoints, and get rid of all other legacy encodings (UCS-2, UTF-16, Extended ASCII with code pages, and all the other region specific encodings that popped up in the 70s and 80s because UTF-8 wasn't invented yet.
This general switch to UTF-8 should have happend in the mid-to-late 1990s (e.g. together with the web becoming popular), and Microsoft alone is to blame for dragging this shit along for the next three decades. If all Microsoft tools would only save text data as UTF-8 starting by the end of the last century, but still support reading all sorts of encodings for a decade or so, the transition would have been finished by 2010. Alas, that never happened.
And tbh, the file size argument for alphabets that don't fit into 7-bit ASCII doesn't really make sense anymore today where images and videos make up the vast majority of data volume.
At the time that essay was written, Make was the best runner available, and the only portable one. There are better options now.
reply