Two benchmarks (artificial analysis and vals) show a increase in cost (a insane increase for vals compared to Opus 4.8).
Already posted this before, so here is the link.
https://news.ycombinator.com/item?id=49041158
You see the issue, if you try to scale effort down, you also need to compare how other competing models compare.
Two benchmarks (artificial analysis and vals) show a increase in cost (a insane increase for vals compared to Opus 4.8).
Already posted this before, so here is the link.
https://news.ycombinator.com/item?id=49041158