The harness is making a big difference, lackluster performance with pi but somehow very good performance with opencode. There’s some rl there for sure, for a smaller model it’s likely going to perform much better in a harness it understands the best.
I've been using a harness that I made myself, and making improvements to the harness has vastly improved its performance. It's actually been a useful model to identify flaws in the harness.
I think something also went wrong with the Ox provider last night (at least on OpenRouter), for a few hours it wouldn't accept tools. Zero change to the harness while I slept and it was back working again the next morning.
Yet to find a model that cross-model review doesn’t find a bunch of things wrong with. I’m running simultaneous review with whichever of Grok4.6/GLM5.3/Fable/Sol didn’t write it, and each model tends to find items the others didn’t.