[meta] I wonder why people have such wildly different bar for what is "good" agentic coding?
In a way, it's absolutely amazing that we've went from "Playing 'Set a Timer' on Apple Music" intelligence to something that may pass the Turing Test, but in practical terms the small models are still far from what I'd call "good" for more than a tech demo.
To me, 7B models are just a fuzzy echo of Wikipedia. Gemma models at 4 bit are too clumsy to even reliably generate JSON for tool calls or copy a line of code to apply a patch.
Qwen needs so much detail and babysitting to stop it from doom looping or losing the plot, that the instructions that I need to give are usually longer than the code I end up keeping.
Is there some magic prompt that I don't know? Do other people just have a lot more patience, or way lower expectations?
I had similar doubts. I think expectations differ because the workload differs. For small scripts, glue code, or simple CRUD changes, smaller models such as Qwen3.6-27B can work wonders than they do on a larger, messier code base.
Those who have never known anything better are okay with much less. For example, anyone who used Fable when it came out are saying that it is very difficult to go back to lesser models now. Even our strongest aren't good enough in comparison.
I used fable, and directly compared it against sonnet 4.8 and gpt 5.5 on various tasks. It was generally better, but still not perfect. Going back to sonnet/gpt has been perfectly fine.
There is a lower bar (that gets lower over time), but ime, the config you are describing is too low still.
qwen/gemma in the 27/35B range @fp8 are better than gemini-2.5, but less than gemini-3.1, you can run DS4-flash @fp8 on two DGX spark, and things keep becoming better. DiffusionGemma came out recently with 4x token gen speeds.
tl;dr - the models you appear to be trying with are too small or too quant'd
We aren’t wealthy enough to have the hardware that would make this good.
The people who have the money to buy a spare maxed out Mac mini just don’t get it. I see lots of folks with RTX 6000’s in threads like these. Or any RTX card that ends in “90”.
Cloud AI is what allows the proles to participate in the broader AI conversation, but not these AI conversations.
But cloud is what will enslave them to the corporation's will.
Google (of all companies!) demonstrated you can get useful stuff with reasonable performance with model running local on their smartphones.
Depending on your expecations you can get the local models running on a recent enough laptop, you just need 16GB of ram to be comfortable. It certainly exceeded my expectations (but i don't use the LLM to write code, only to do the real boring stuff: docs.)
As one of those folks with a 6000 Pro who tries out basically every local model I can get my hands on; local models still aren't ready to be relied on for larger-scale software engineering. They're quite capable when it comes to writing and editing code though.
Because as with every AI, when it passed, the goalposts have been moved. Naysayers say it doesn't count, because winning once isn't enough, it needs to win 50%+ of the time, or winning against an uneducated human who can't ask right questions doesn't count etc.
In a way, it's absolutely amazing that we've went from "Playing 'Set a Timer' on Apple Music" intelligence to something that may pass the Turing Test, but in practical terms the small models are still far from what I'd call "good" for more than a tech demo.
To me, 7B models are just a fuzzy echo of Wikipedia. Gemma models at 4 bit are too clumsy to even reliably generate JSON for tool calls or copy a line of code to apply a patch.
Qwen needs so much detail and babysitting to stop it from doom looping or losing the plot, that the instructions that I need to give are usually longer than the code I end up keeping.
Is there some magic prompt that I don't know? Do other people just have a lot more patience, or way lower expectations?