Yeah agreed, there were some minor gains, but new releases are mostly benchmark ...

Yeah agreed, there were some minor gains, but new releases are mostly benchmark overfit sycopanthic bullshit that are only better on paper and horrible to use. The more synthetic data they add the less world knowledge the model has and the more useless it becomes. But at least they can almost mimic a basic calculator now /s

For api models, OpenAI's releases have regularly not been an improvement for a long while now. Is sonnet 4.5 better than 3.5 outside pretentius agentic workflows it's been trained for? Basically impossible to tell, they make the same braindead mistakes sometimes.