It may be so but the article does not claim that the strong foundation is the code, rather it seems to be product design, and design of other products, at that (the two predecessors, webapp and desktop app). No reason why you couldn't study existing products now and tell your agent to build something based on that.
This feels like a sleight of hand to me. The hard part of evolving Scribe and Web Scrapbook was discovering that a browser extension manipulating a local SQLite database was _the only_ architecture that could reconcile local offline persistence with live DOM scraping across arbitrary catalogs of academic data.
An agent can synthesize existing solutions but (because I see this failure mode at work constantly) it can't synthesize an architecture to resolve the sorts of tensions that the person prompting it doesn't yet understand (not that that is stopping anyone). You can't prompt it to build something if the operational primitives required to solve the problem haven't been mapped.
"Build a tool based on Scribe and Web Scrapbook" in 2003 would've made a fragile PHP wrapper because that's what the existing landscape looked like.
Good stuff. I share this view fully. I've been working on a project, which I am happy I had no agentic help on for the first year. When you lay the architectural foundation you need to start with a vision, derive constraints and then map a solution to all of them. This takes a level of intentionality that agents don't have. Probabilistic models will work against anything that is unconventional, so they tend to favor low-value solutions that have already been seen.
Wrote up the longer version of my journey on that project here:
Yes, exactly. And this is why I don't think that the LLMs can make significant process beyond what humans have done and published.
"But the math proofs," people will say. A lot of those seem to be spam-solving things with a huge swath of existing lemmas, and a some of these are being debunked and retracted.
Just today I was quizzing ChatGPT about a basic grammar question for a language that has huge training data but for which the grammar was not well documented. It kept giving me confidently wrong answers until I drilled and drilled it and then finally it found/gave back an explanation that perfectly fit a pattern given in one particular grammar, citing that as a source. It doesn't appear to have been able to figure out the inner structure on it's own. It appears only able to pattern match and put things together from what humans have already discovered and written.
Right, there's a difference between statistical interpolation and semantic induction. The whole point is that LLMs can't reason from first principles to drive missing rules. It keeps confidently feeding you approximations until it collides with some source that already mapped it.
Interesting take, given that the whole reason LLMs are interesting is that they're the first system we have that can work in semantic space. Statistical interpolation, that we've solved long ago.
I think that's equivocating on the word semantic. Word embeddings map concepts like "king - man + woman = queen" or cluster synonyms together in high-dimensional vector space and the ML literature very loosely calls this "semantic space." But a high-dimensional topology of token co-occurrences isn't semantics in the sense of computation or formal semantics. It's still "just" measuring distributional similarity. Some vector that represents "thread deadlock" lives near tokens like "mutex" and "race condition" and "starvation," but the model itself has no concept of concurrency and contention.
Claiming we "solved" statistical interpolation long ago just means curve-fitting and basic regressions on structured data. Transformers are a truly impressive achievement, scaling all of this to unstructured high-dimensional text topologies, but it's fundamentally the same math operations on statistical proximity.
Like how do we explain hallucinations here? Tokens that are hallucinated are semantically "close" in that vector space but they're completely false in reality. If LLMs operated in a true semantic space they wouldn't hallucinate CLI flags that don't exist.
> Some vector that represents "thread deadlock" lives near tokens like "mutex" and "race condition" and "starvation," but the model itself has no concept of concurrency and contention.
I propose that concepts of "concurrency" and "contention" are themselves vector in latent space. All concepts are. Recall that we're talking about a 10^4 - 10^5 dimensional space. You can fit in pretty much any conceivable association as some direction in there.
And try to zoom in on any concept you know. If you do, it should quickly become apparent that there's never any concept you can give a closed definition for. We can only define concepts, and we can only learn them, through generalizing from examples. Which is conceptually (pun not intended) regression - finding a vector along which examples live.
On the other hand one of the main advantages of console gaming was that you could resell the games after you are done. Now with the move to digital games, what is the advantage of buying a console when Steam has regular discounts and a massive catalog?
Oh boy, that time is long gone. Today you wait an hour for the game to be installed from disk and then for the day 1 patch to be downloaded and installed.
> Lower up-front cost because of the subsidized hardware
I do not believe any current-gen console is subsidised? Though old stock thereof is probably cheaper because they locked the component prices in before RAM prices started going crazy
It's written like one of those long form magazine articles. Which is fine for that kind of thing but maybe not great for software release notes. It goes buildup buildup buildup ... underwhelming payoff.
How is that relevant for the average buyer (as much as these things have an average buyer anyway)? Precision, accuracy, and battery life is helpful info but the inner components are an irrelevant detail. It's not like having a high quality part prevents them from using it wrong anyway.
The datasheet gives you much better fidelity data than an N=1 test. You'll never know what the tolerance is if you're just looking at one unit. Presumably the average buyer cares if they expect the temperature sensor to be 5% off on the unit they buy. The accuracy of all of these sensors have ratings, but they are often just copying the value on one component and not even considering a stackup. Devils lie in details.
Regarding accuracy, it's been about a decade, but I did some embedded projects related to accurate temperature measurements back in the day. From what I recall, anything sensor that is < $100 or so is going to be 5% margin of error at the very best, and regardless of price, you're going to have to calibrate it and figure out the curves in your ADC, it's not going to be linear. Additionally, just the act of measuring will heat your sensor slightly so that will need to be taken into account as well.
In short, if you need less than a couple degrees of accuracy, it's going to take a bit of tinkering. I wouldn't trust a $1 sensor for a tropical fish tank, for example, but probably good enough for a home thermostat.
That tracks somewhat with what I remember. I was working on process control for an industrial washing machine in a hot and humid environment. I think the extreme temp and humidity is what made it difficult, but those error margins do look great for the cost.
The datasheet will gives you absolute best case scenario. Like current consumption in sleep mode, or what register to set for best accuracy. It does not mean the Firmware will do exactly that.
The datasheet gives you the 3 sigma worst case for accuracy.
For power consumption a system power tree needs to be constructed and that will set the bounds that firmware can achieve. That would be better than trusting the manufacturer but a real test is best.
Avg buyer wont read a token of tests before buying whats marketed best.the whole point in test like these is to provide non-average buyer with info average buyer doesnt care for.
Can sort of make sense of the transfer and performance bit as: the location is encoded as part of the data type instead of some sort of automatic spillover to main memory that will behave differently based on data size.
Give it up. People have been predicting this for 15 years.
You can’t have a development machine locked down like iOS. As long as the Mac is the development machine at apple it will need to be able to run arbitrary code.
What support? Asahi had to fight to exist, especially early on when they essentially had to reimplement much of the graphics pipeline. And now its in mainline linux - short of a complete hostile lockdown of the bootloader, it’s not like they support other OS’s very much as is.
Asahi appears to be dead. It wont even boot on anything apple released in the past 2 years.
Sure, their M1-M3 support is okay (as long as you're fine with no hibernate or sleep support, so have to plug in the laptop or risk losing all your open stuff), and it was a lot of work by lots of volunteers... But at this point it's starting to look more like a retrocomputing project. As time ticks on, it'll become less and less viable using an ancient laptop, especially if local AI becomes commonplace.
Wtf do you mean? They just enabled M3 in the official installer and they're working on M4 bringup too. Keep in mind it took them like 16 months to even support M1 at first. Then they supported M2 within a few months, then M3 threw in a wrench that they're still working on, but there is no evidence of death at all, just a lot of work left to do.
M4 has no support and was released in May 2024, over 2 years ago, and looks at least a year from usability - doesn't yet boot. M5 also has no support. M6 also has no support.
I did say they're working on M4 bringup, which means it doesn't boot yet, yes. I don't see how that means they're "dead" though. Or that they'll never catch up ever. In fact, the very recent progress in this area suggests the exact opposite.
reply