AC at 68F is obnoxious and wasteful. At least data centres enable me to communicate with distant colleagues, watch enriching movies, and use very strong AI tools to accelerate my research.
It's actually significantly more environmentally friendly than heating to above 68F in the winter in most cases, since most AC is done with heat pumps but most heating is oil/gas (100% fossil fuel consumption) or resistive electric heaters (dramatically less efficient than heat pumps). On top of that AC use aligns strongly with solar generation output while heating load does the opposite, which means that even when heating is done with electricity it comes disproportionately from gas turbines at natural gas peaker plants.
Data centers also enable the NSA to store the metadata and contents of every phone call you make and police to (without a warrant) to track anyone who's walked in front of your doorbell or many street corners in your city. They also store all the location data your phone collects for Google or Apple, nearly your entire internet history collected surreptitiously because many sites automatically use your GMail or Facebook logins, and your entire purchase history collected via Visa or MasterCard.
And even better: all of that data is aggregated together and stored in its entirety several times over across multiple data centers because everyone shares your data with everyone else, such that if any single of more than a dozen data centers is ever compromised, it all gets leaked to whomever compromised it.
My building is leaky. The association won't fix it. How about this compromise, you set yours to 74 to compensate for my 68. Fair enough, or will you say I should move?
It's interesting that Claude also over-uses en-dashes. It's very willing to create compound-noun-phrases, especially in that compressed-summary-paragraph it often writes. The 0-days-vibes-vulns that started this thread looks a lot like that, but it could be Claude directly, or just Claude's style influencing people who spend too much time with it.
Those are just hyphens, actually. En dashes are for ranges (e.g., 1–4, which is admittedly hard to distinguish from 1-4), not compound words. Point stands though—LLMs do love compound words and dashes in general.
People are missing that Willison is among the very best people we have in the role of (for lack of a good name): early access to frontier models, evaluate them in real scenarios, no wishful thinking, hype, or doom, communicate the possibilities. Yes he could have fixed this himself but then he would have learned nothing about the AI, and we wouldn't have read a fascinating and important article.
there is absolutely zero value in spending time to learn about new models as in few months new model will be out and whatever you learned about the current one will be useless.
Also with models getting better and better you have to know less and less to achieve same results.
As the models get better you need to know more about their capabilities, because otherwise you risk prompting Claude Fable 5 like it's GPT-4o and complaining loudly about how it's all hype and nothing about these models is improving at all (yes, I do see people say that.)
Getting the best results out of these models requires skill, experience, intuition, and domain expertise. There's always room for improving every one of those.
Isn't the whole point of a better model that it should be better at understanding you than the previous one? So the same prompt should return a better answer.
Prompting differently to the new model seems entirely backwards when trying to determine if the model has improved.
It doesn't matter how good the models get, they still won't be able to act on unclear directions.
Learning to provide unambiguous, clear directions is a skill. A lot of people who report bad experiences with models aren't yet good at that skill.
More importantly though, the key to successful communication is having a good understanding of what the other side of the conversation already knows and understands.
Saying "use uv and inline script dependencies" won't mean anything to a model with a knowledge cutoff date prior to the launch of uv!
I think this is true when models were going from bad to pretty good like happened last year. But when they start to get good, and can work deeper and with more nuance, how you prompt also can change the results quite a bit. Note this is also true of asking smart humans to do things; personality and approaches vary, they don’t exist on a single axis continuum of quality
>> Getting the best results out of these models requires skill, experience, intuition, and domain expertise.
domain expertise has nothing to do with llms. On the contrary, to have it you need to avoid llms.
>>you risk prompting Claude Fable 5 like it's GPT-4o
Thats fine because when GPT came out you had to treat it like a baby, GPT2 and around that time "Prompt engineering" was a thing.
Now its all dead.
After opus 4.8 all you have to do is say "fix it" or add /plan. All that time spend on learning previous models is time wasted.
And in a year or two with developed harness you will be out of the loop, errors are incoming - llm fixes them or adds new features based on some transcripts etc.
Even if model development stops now - there is nothing to learn really. Sure you may need to adjust prompt style a bit. You will do it naturally just like when you communicate with a new person. There is no "knowledge" to it, it is very smart.
> domain expertise has nothing to do with llms. On the contrary, to have it you need to avoid llms.
It has everything to do with LLMs.
Go ask Claude Fable to write you a two page position paper on how the European economy recovered after World War II, suitable for submission to a conference for economists.
It will do exactly that (well, probably, Fable can find all sorts of reasons to refuse) - and the value of what it wrote to you will be virtually zero, unless you yourself have deep expertise in economics and history.
Way back before instruct models it was pretty difficult, but for the last couple of years I haven't needed anything more complex than the type of text that I might send in a detailed email to a colleague.
I agree but this particular example showed nothing about leveraging skill, experience, or intuition. If anything, this is another straightforward example of a one shot ask.
edit: that said, I understand this particular post is about model capability
There’s zero value? Surely you don’t believe zero, it’s potentially the most powerful predictive AI in the world ever made? Maybe only incremental steps sure. But also their IPO is coming, you don’t want people evaluating them beforehand?
you know, women make a big deal about you meeting their father/parents, and honestly, I'm too autistic to really fucking have put any importance until now as to why that was remotely important, but if N+1 is coming for your job, it seems it might be worth your while to know the capabilities of N, no?
"interpolate" has a technical meaning - in this meaning, LLMs almost never interpolate. It also has a very vague everyday meaning - in this meaning, LLMs do interpolate, but so do humans.
That’s the thing, I have never seen detailed costs of what people are spending their money on. I know that for Claude there’s a $200 monthly subscription through which assigned credits one burns pretty fast, at which point (and I may be wrong on this, because I’ve never used the thing) one can run extra code on a “pay as you use it” basis? Again, I might be wrong on this.
I’ve also seen it mentioned a lot of people having 2, 3 or even more subscriptions, which I’m pretty sure that can easily go South when it comes to costs.
But, again, and the most important point, I’ve never seen a detailed post on what people spend on this AI thing on a monthly basis (let’s say).
> Since these companies can’t improve their AI models without fresh data created by human beings
Totally wrong. Self-play dates back to Arthur Samuel in the 1950s and RL with verifiable rewards is a key part of training the most advanced models today.
Not totally wrong. Self play works well with if your problem can be easily simulated in an RL environment where the model can easily explore different states. RLHF or similar techniques is not that since we don't have exactly have a simulation environment for language modelling
Right now there are companies which hire software devs or data scientists to just solve a bunch of random problems so that they can generate training data for an LLM model. Why would they be in business if self play can work out so well?
> Right now there are companies which hire software devs or data scientists to just solve a bunch of random problems so that they can generate training data for an LLM model.
Current models don't yet use RLVR with self-play though, at least as far as we know. They use RLVR with large numbers of manually created RL environments.
Yes but Anglo-Catholic doesn’t mean an Anglo who is Catholic. It means an Anglican who is pretending to be Catholic but without acknowledging the pope’s authority.
> It also recouped more than the trial's net cost of 72 million euros ($86 million) through increases in arts-related expenditure, productivity gains and reduced reliance on other social welfare payments, according to a government-commissioned cost-benefit analysis.
I'd love to see the breakdown on this, because in my experience with Government comms, if it was a straightforward economic win like an FDI or industrial announcement, they'd headline the figure. Unquantified phrases like "reduced reliance on other social welfare payments" are usually spin at best.
I understand your point, but in response to GP (they should spend this money on houses for other poor people instead), the reduced reliance on other social welfare is totally legitimate to count.
I agree, but another commenter linked the cost-benefit analysis and it really is creative accounting to get to a positive net social return.
The net fiscal cost after accounting for increase tax revenue and social protection savings was €72 MM.
This was then offset to get to a positive net social gain by €80 MM in "wellbeing gains", as measured by a single survey question called the WELLBY test:
> “Overall, how satisfied are you with your life nowadays, where 0 is "not at all satisfied" and 10 is "completely satisfied"?
The €80 MM in "wellbeing gains", which is the sole decider of whether this pilot was a net positive or a huge net negative to society, is because on average, the 2,000 pilot scheme participants had a very approximate 0.7–1.1 increase in score when asked the above question during the pilot as compared to before the pilot. Each 1 point was deemed to be worth €15,340.
They just totally made up[1] a number, tripled it, doubled that, and finally applied a multiplier before using it as the basis to support their preconceived… let’s not mince words: agenda.
1. That’s not very kind. I’m sure they didn’t “just make it up”. There’s bound to be an entire bureaucracy if not dedicated to, at least tasked with, conjuring up this sort of codswallop.