Is anyone else just exhausted by the pace of all this. The models change constantly and relentlessly and so does the pricing, basically weekly at this point between all the labs.
It feels nearly impossible to have any rigorous approach when choosing a particular model and price point for a task and more like blindly picking one. The time period needed to actually get familiar with various models to a degree you can intuitively choose appropriate ones for a task is moot when it will likely be superseded faster than the needed time.
I guess if companies are footing the bills most employees just opt for whatever the most expensive model they can get away with. Even then choosing between the various leading models is the same kind of frustrating task. Every release every company has the same random collection of graphs and charts claiming the best performance on X, Y, and Z.
You really don’t need to watch it that closely. If the model you’re using today is working well, just stick with it.
If one day you open up Claude Code and it’s Opus 5.1 now instead of Opus 5, no big deal. It probably will work about the same as it did before. Maybe a little better.
Or if you’re on Codex and some new cool Claude model comes out, no worries. There will probably be a similar new model for Codex within a few weeks. Maybe even within a few days.
One suggestion is to make a list or make a skill to have your agent keep a list of things you do not feel work well with today's models. And then, when new models come out, periodically, revisit items on that list to see if you get better results.
You can. The point of having your agent keep track of it is that it will likely notice things you won't, and it can automate cataloging it with relevant metadata (prompts, environment, examples, etc.) that make it trivial to automate rerunning those tests when new models launch.
> The models change constantly and relentlessly and so does the pricing, basically weekly at this point between all the labs.
A dev in my team saw a new model and changed one application to use said model (essentially changing the contents of a url). One week later I received an escalation from the CTO of the company that our pace of weekly usage was in the millions of dollars (rather than low hundred thousands). Turns out that the new model was 5x more expensive but no one noticed.
Okay, well, that seems like a natural problem. I could understand if he went from one of the Gemini Flashes to the next (when they rebranded Flash to Flash Lite and came up with a new much more expensive Flash). Now that would be a mess.
I could see how this might feel frustrating to someone who doesn't enjoy experimenting with new things all the time.
In practice, you can get away without keeping up with everything all the time. For personal use, pick a provider and get on their ~$20/month plan. Learn their high/medium/low model hierarchy. Start with their highest or second-highest model (GPT-5.6, Opus, etc) and observe your quota usage. If you're doing a lot of manual code review and analysis, the $20/month plan goes very far even on the highest models. If you're trying to vibecode everything as fast as possible it's a different story.
If you keep running into quota limits, experiment with the next model down for easier tasks or adjusting the effort level. If the results are good enough, you've found your fit. If they're not, you might need the next plan up.
For API/business use, you have to be checking your token spend as you go to calibrate to how much each task costs and where you fall in your budget. There are a lot of different tools that make this easy to visualize.
For data tasks, you should have an eval with a golden dataset that you can run against new models for a nominal amount of token expenditure. It should be as simple as pointing the eval script at a new API or model and checking the score versus price.
Another suggestion to get the most bang for your buck: use the best model you have access to with max reasoning for planning, implement with a smaller model/lower reasoning, then review with the big model. Repeat as needed.
Input tokens are much cheaper than output tokens. Not only because of baseline price—caching makes a huge difference too. There are many ways to take advantage of this asymmetry to get similar quality for a fraction of the cost!
"Within a year" is a bit of an exaggeration but it's true that the pace of PC tech during the 90s was much, much faster than it is now. CPU power was doubling every two years, and today we're at roughly eight years. Add onto that the rise of video cards in the late 90s.
It was both. 90% of people never needed nor purchased a bleeding-edge computer. The mid-tier was "good enough" and far closer to affordable for most people; though, that bar also moved upward every year.
If you bought a mid-tier computer that was good enough for what you needed, then you probably didn't shop/compare for the next few years and didn't notice. But if you shelled out $7-10k for a top-of-the-line system and paid attention to progress, you'd easily see that become the mid-tier $1000 option within two years or less. This is how it was in the 90's PC boom, at least. Likely the same for the decades before, not sure how it went in the 2000's.
if you shelled out $7-10k for a top-of-the-line system and paid attention to progress, you'd easily see that become the mid-tier $1000 option within two years or less
This is not how I remember that period at all. Do you have any examples?
My first PC 386 was in todays money easily $5000+ (basic 2d GPU + screen)... A lot of hardware in our family was handed down to my folks, because you lost so much on selling, that it was better to keep using them as they had less demands.
386 to 486 to the first Pentium (with the bug!)... You did not upgrade in place, it was often a new system. Sure, you maybe kept your screen, keyboard etc but ... The only upgrade we had on the same MB, was a coprocessor upgrade. Remember those? Each new generation of CPU was a new motherboard. Upgrading CPUs in the same MB really became a thing only later on.
GPUs had a shelf life of barely a year. Its been 35 year but i remember TNT to TNT2 having like 9 month in between. Moving from 2D to 3D involved a constant cost as GPUs evolved fast and the latest games required latest hardware.
We have not talked about the ISA, AGP, and PCI fun ... The “bus wars”.
DOS to Windows 3.1 (and OS/2 somewhere in between) to 95 ... with software being pushing hardware, just like games did.
This is why people are spoiled with cheap PC hardware where its cheap, and easily lasts 4+ years. Even with the bad memory price and more expensive GPUs, your can stil buy a $1500 system that will last you years (with maybe some lower game settings later on ... or the catalog of 10.000s games that will easily run on a mid tier GPU).
PC hardware has become boring but extreme stable. You can run GPUs for year, switch MBs without issues while keeping large amounts of old hardware. That was NOT the 80s and 90s that i remember.
Really depended when you bought them .. 386 is like the point where people really started to buy PCs. My 386 was more to the end period, so the jump to 486 was very short. Do not forget there was a ton of 386's and 486s released also. I remember the SX, DX, ... the coprocessor that you plugged into the socket and then the CPU into it.
The problem is that your too focused on the CPU only. There was GPU improvements (2d) then 3D, the constant improvements in sound cards. Printers ...
There was very strong depreciation in that time. I think that the poster before with his $10.000 > 1000 is somewhat exaggerating but like i said, my $5000 range system was not worth reselling. So did several other systems we acquired at the time, it simply became hand-me downs (for family who needed a PC for Lotus or WP but not a "advanced" PC).
Maybe my memory is off, but hardware in that time was way more expensive, then it is today. While todays hardware lasts WAY longer. The wife is still on a 8250u/8GB laptop, from like 8 years or so, and i can not pry it out of her hands. Needed a new battery lol ...
There really is rarely pressure to upgrade these days. That is not how i remember the 90's... Where upgrades tended to have much bigger impacts. Especially as gamers. But even on side equipment like CD players, CDRs and later DVDs.
Todays PCs is like ... raytracing, no interest. 4k? 1440p is perfectly fine. CPUs? Over powerful for most folks. The only new thing is AI and that is a totally different issue.
tbh this is how i remembered that time as well. If you look at recommended system requirements for something like Max Payne (in 2001) vs Unreal Tournament 2003, everything had basically doubled
I had a 600MHz/64MB/9GB laptop that came with Windows mistake edition. I managed to survive first year of uni on it by switching to Vector Linux, which was really fast compared to Windows. (Of course, it had issues playing sound from more than one source, this was oss days).
Then one day the hard drive appeared to die. I eventually realised the issue was located around the 1.5gb mark, so I recreated my Linux partitions after 2gb and it worked fine for the rest of the year.
They overclocked well though, I think you could run the 300Mhz chips at >400Mhz.
I also believe you could get motherboards that supported 2 Celeron chips. I have no idea how effective/useful it was, but it was certainly a cheap/interesting way to get multiple CPU's.
The new releases and breakthroughs do the opposite for me - I feel energised by them. I felt like nothing truly that interesting had happened in tech for quite some time, now it's like the space race (except there is no one moon to reach).
I appreciate boring tech as much as the next well worn engineer and I'm not saying this is all positive but it's so sure as hell thrilling and you don't have to be an astronaut to immediately benefit (or suffer I guess) from it.
Hey, I'm on the team at LiteLLM that's building the auto-router and our goal right now is to abstract that decision making away from the end user. The biggest thing we're trying to figure out right now is how do we do that without frustrating the end user - as a developer myself I would hate for my agent to be dumbed down below the threshold needed to complete a task.
In theory though, there is a minimum viable model for any given task, and we think that is a problem that the big labs will avoid because they profit from charging more per task. We're trying heuristic and LLM-based approaches but it's still a work in progress, so if this is something you'd be interested in trying would highly recommend trying ours out -- any and all feedback at this point is extremely valuable to us.
It is exhausting to keep up with model releases yes, much like it was for a while during the Cambrian explosion of FE frameworks, eventually tech seems to work out to consolidation.
But more so it seems there is Fear of missing out (FOMO) in our behaviours. The reality is, if whatever model you are using are good for your purpose, well, keep on it.
Yeah I'm a bit exhausted at this point. I just finished benchmarking GPT 5.6 Sol and Fable 5.0 like two days ago. My data became obsolete literally one day after.
It feels nearly impossible to have any rigorous approach when choosing a particular model and price point for a task and more like blindly picking one. The time period needed to actually get familiar with various models to a degree you can intuitively choose appropriate ones for a task is moot when it will likely be superseded faster than the needed time.
I guess if companies are footing the bills most employees just opt for whatever the most expensive model they can get away with. Even then choosing between the various leading models is the same kind of frustrating task. Every release every company has the same random collection of graphs and charts claiming the best performance on X, Y, and Z.