I could make a non-deterministic chainsaw fairly easily. I’d also get sued into the ground if I sold it, and I wouldn’t be able to claim ‘Oh, it’s just an unavoidable part of progress’.
It’s not a defect, the stochastic “temperature” setting unleashes the chainsaw’s creativity and imagination! Who are we to cast aspersions on the Oracle Chainsaw’s intelligence—nay, wisdom!—just because it happens to be non-living?
A chainsaw is a physical machine. Physical machines are technically non-deterministic if you look closely. They have Variance. The discipline to manage variance is called Tolerance.
Many physical machines and components come with a datasheet that will list their tolerances.
Failure to correctly document tolerances does in fact get you sued.
However, while this is truly a great idea, we're not going to be able to make it work for computational systems. Computers, software, and also LLMs are sensitive to initial conditions. Which is why tolerances are not so familiar to computer people. (but not entirely: eg your PSU might list 110-240Vac/300W as input tolerance)
Interestingly, LLMs actually have a somewhat lower sensitivity to initial conditions than traditional interpreters. See what happens if you misspell "What is One Plus nOe?". So they're actually a skosh off the edge and towards the middle, though I'd argue still very much at the computational end, just from the sheer scale of the valid inputs and outputs.
Mind you, if you have a pretrained LLM doing a measurable task on a line, possibly some sort of tolerances could be determined. Not so much when doing arbitrary chat.
Something unintuitive: I bet that often setting the temperature > 0 (aka introduce stochasticity deliberately, variously comparable to dithering or simulated annealing in other disciplines - doing the thing where you escape local minima) will tighten the output tolerance range and improve reliability, especially in iterated processes. This works for a lot of physical and digital processes actually, and LLMs simply stole the same trick.
(edit: I'm trying to compress a huge chunk of dynamics intuition in a few lines here. Hopefully still useful.
TL:DR; Everything real is continuous and noisy if you look close; and you're really trying to build attractors and bound variance, if you can. )
I mean, are you saying ‘but more RAM?’ because obviously yes that’s true but not a solution if you already own a laptop, and have you seen ram prices? Also swapping to a fast SSD isn’t like it used to be on spinning discs. I’ve been amazed how responsive Mac neo laptops are and they are swapping all the time.
I'm not saying to avoid swap. It is great at what it was designed to do. I only take issue with optimizing swap for bandwidth when it will never be able to keep up with RAM bandwidth. The MacBook Neo for example has 60 GB/s of RAM bandwidth but only 1.5 GB/s SSD bandwidth.
Because it wouldn't soften the blow. Swap benefits the most from improved latency not bandwidth. Data is moved from swap back to physical RAM when swapped out pages are accessed by software (which stalls the thread). The kernel has less insight into the memory access pattern to prefetch the next page so the stalls likely continue for each page it needs to restore. The latency of zram (compressed in memory) is still lower than the fastest SSDs which reduces those stalls.
Wow, that is widely disingenuous, I don't really think there is any excuse for that, I don't believe someone deep in compression algorithms wouldn't know they could adjust the block size, and 512GB is a huge block size for bzip3, as it needs to basically all be in memory so you can't pretend that's just 'the standard value'.
What’s interesting is it’s not obvious how this is leveraged to ‘cure disease’. But I’d love to know.the advantage of this is there is a clear measure of success. Here is a rule language. Prove this. You are done when your proof passes. You can sit quietly and spin for billions of tokens.
How does that work for drugs? We can’t let AIs make millions of test drugs and try them out on people.
I worry this doesn’t check correctness. I’ve been finding Claude is lately awful at folllowing instructions, I’ll ask it to implement the algorithm from a paper and it will do something simpler and slower and when challenged do it’s stupid apology thing. It can’t be trusted with anything I’d put in a paper, it lies too much. 4.6 couldn’t do as complex tasks, but it would do what was asked.
Benchmark does check correctness. Team is working to write a paper that will likely contain failure mode analysis; checking for instruction following could be a good idea.
Agents are given instructions in markdown format, allowed to read data and libraries sandboxed in a Docker container, and evaluated on deterministic pytests on the outcome. Two things the team aimed to enforce to add tasks we could trust:
- Scientific workflows are often simulations that are correct upto numerical tolerances (the scientist decides what's reasonable), so task verifiers' evaluate results of agent-written code within the tolerances.
- There could be multiple solution codes to a scientific workflow, and the team tried to ensure the verifier tests accommodate those. Not overfit to the oracle reference code, written by the scientist.
Instruction following is implicitly assumed, if the model gives up and doesn't complete the task it counts as a failure because the verifier tests fail.
I can't find the specific inverventional study I mentioned. As I recall, it was mixed but possibly I remembered wrongly.
Some studies do show a significant correlation between unhappiness and social media use, but causation has not been established and the bigger, better done studies have weaker correlations. People believe in causation because it suits their bias against social media.
I dislike social media but I doubt it's harmful. This substack summarizes my view on what really makes children unhappy: https://substack.com/home/post/p-186087964 (school). I'd guess that lack of freedom is what makes school so awful; modern schools are like little totalitarian states, where every minute of your day is accounted for and there is no privacy.
This is why I’m trying to move to open Chinese models — because I will be able to use them forever, while the older Claude models which I genuinely enjoyed writing short stories with have now been deleted, replaced with hypothetically cleverer models which produce text everyone hates.
You only need to mention Protected Group Of The Week (I'm one of them and I like to research and read about history, so that makes it extra challenging) or anything resembling negative human emotions (guilty on that front as well), and the model screeches to a halt.
Just because OAI doesn't want another headline like "Chatbot convinces teen to off himself"
I understand this in principle, but I'm not convinced that dulling everyone's knives is better than figuring out how to keep them out of kids' hands
> "no I won't tell you how to build a bioweapon for genocide" is guardrails.
I call that censorship too. I'm curious enough that I want to know about such things. I don't want any limitations on what I'm allowed to understand and know about.
Your freedom ends where my begins. Welcome to study chemistry and learn things from first principles. but if you want to just ask for the practical steps of making a genocidal bioweapon then honestly I don't know you or whether you are honestly "merely curious" and so prefer your freedom ends there?
> Welcome to study chemistry and learn things from first principles.
As if your censorship was not going to kill that too. Can't even ask Fable about aminoacids without getting blocked. So much for "learning from first principles".
> but if you want to just ask for the practical steps of making a genocidal bioweapon
Nothing wrong with practical steps. Just because I know how to do something, doesn't mean I'm actually going to do it.
> As if your censorship was not going to kill that too.
I didn't say "with LLMs". Last time I checked they still teach chemistry at unis and schools.
And yeah, guardrails are not perfect. Honestly I don't think good enough guardrails are possible, it's all eventually defeated or becomes silly. And yes that should be one of the reasons the technology as a whole is banned.
But until then, guardrails are guardrails and not censorship in a pretty obvious way. If you refuse to see that, be my guest. I personally like to live.
I haven't said anything about "good", I just pointed out that there isn't exactly an alternative to Chinese models if censorship and restrictions are regarded as bad for longevity. You can't really do better than the Chinese models for longevity; US models are by far the worst in this regard. So "What about censorship?" is an absolutely hilarious question to ask when Chinese models are presented as an alternative. Yes, what about it? They have less than the obvious alternative from US labs, and where is it you imagine you'll find less censorship?
"nothing happens in 1989" is censorship. "I won't tell you how to build a bioweapon for genocide" is guardrails.
I like the second one because I like to be alive.
> Where will you run them when powerful enough GPU and RAM are only sold to hyperscalers?
Do you think that fabrication will never progress (in volume) than what we have now? The hyperscalers are already having trouble paying the bills, they can't keep this up forever.
The hyperscalers will be bailed (maybe not all of them but enough). US economy will crash if not. And whatever is made will go to them, because they pay more (thanks to US taxpayer bucks among other things) than any regular person. First they build on land then they build in space.
My big problem with leaving Google is I’ve clicked on far too many ‘login with google’ boxes. In retrospect this was a mistake as it now ties my own domain to google basically forever, but lots of sites don’t even have the option to undo it.
A lot of those sites will let you just reset password, send an email to the Google account's associated gmail, and then you're done. I suspect that it isn't thought through as an official feature in such cases, but it does often work.
Additionally, if you signed up with an email and password, you can often just sign in via SSO with a matching email address for convenience and still have the underlying password remain working. Again, possibly not an intentional feature (but rather something like using the email reported via SSO or via email verification as a primary key).
I switched away from Google and haven't had a single issue going from 'login with google' to a regular email/password FWIW. Every site has offered some kind of painless recovery.
This is a bit surprising to me because I switched emails some time ago, and most sites were painless, but there were a fairly long tail of sites that had no self-server mechanism to change email from one to another. A fair number of them required phoning in to read off my new email to a CSR rep. Another pile of them had no mechanism to change email at all, their only solution was to open a new account.
You can change the MX records of the domain to point at Fastmail (or some other provider) without losing the ability to log into associated "Google Workspace" accounts.
I thought this would be a blocker too but you can both host your email somewhere else and continue to login with google. You still need to keep the google workspace account enabled and active, you just point the MX record elsewhere.
Same; I was a sucker for GitHub SSO as well. I have no idea how to unwind it all, but (like Google SSO) it certainly had its benefits while I was using it. I remember there was an “open” SSO provider - OpenID maybe - back in the 2000s when a federated identity made interacting with blogs and such much easier. It’s the same with GitHub today, with various 3rd parties who need access to your code. They don’t have a “give me an git+ssh url” and so you almost need GitHub.
i’m not sure i’ll ever delete my google accounts, but my goal was to stop using them. if someday an email comes in, it’ll still get auto forwarded, and if i need to login, i’ll have access to it.
You hope you can ... I didn't login to a Google Account for a few months now and it suddenly wants a lot of personal details to do additional "verification" because I am "suspicious".
Did it also give you a tip to "log in to your usual Wifi network"? I have a feeling all large corps let their heuristics run wild with little to none supervison.
I don't even bother with forwarding mail from Gmail, it's just a black hole for marketing emails that I'd only occasionally glance at when I'm looking for a discount code.
Then go through each service and email them to have them try and move your account to new accounts with a traditional username/password login. Some might even be able to change your login method to username/password without requiring a new account!
I was paying for Google Business for my family, and then decided to switch to iCloud, but left only my own account on Google, because of that reason.
Also, some services requires you to log in with a Business Google account, regular gmail accounts don't work (ie: Granola).
reply