I find the use of the word intelligence to be a bit of a misnomer. Is something intelligent if all its doing is pattern matching? Is the evolution that led to owl butterflies appearing like an owl intelligent? Im not sure.
As an aside, its amusing we simulatenously have this article on the front page as well as [Generative AI coding tools and agents do not work for me](https://news.ycombinator.com/item?id=44294633) also on the front page. LLMs are really dividing the community at the moment and its exhausting to keep up with what I (as a dev) should be doing to stay sharp
Keep in mind that as usual, mostly the extreme views are getting posted. The urge to both write and click on "sometimes I find LLMs useful for partial solutions in the right context" is low compared to "AI will replace all developers in 2 years". It may not be as dividing as we read here. It certainly isn't, when looking at what my co-workers do. You can chill and learn it like any other new tech. (Without following every detail day to day)
I will die on the hill that for the foreseeable future, LLMs inside an IDE are just fancy autocomplete.
In a more general interface they're also nice for getting a birds-eye view on a topic you're unfamiliar with.
However, just as a counterexample of how dumb they really are: I asked both Gemini 2.5 Pro and Opus 4 if there were any extra settings for VSCode's UI density and without hesitation both of them made up a bunch of 'window.density' settings.
If they can't even get something so extremely basic and well-documented right, how are you going to trust them with giving you flawless C or Typescript?
Well the article briefly addresses this: it's about the iteration: given a problem and sufficient processing power we can attain an intelligent,correct answer by quickly iterating from prompt to results.
There's also a measurement vector for zero-shot LLM responses. But excelling at zero-shot is not a requirement for making LLMs useful.
The market is pointing the way, agents increase iteration capabilities, increasing usefulness. Reasoning models/architectures are another example where iterations make advances - the LLM iterates "in-band" and self-evals so that there's a better chance of a correct outcome.
All that in a mere 3.5 years since launch. To call it an autocomplete is very short sighted. Even if we reached LLMs ceiling, the choice of AI-oriented workflows (TTS, TDD, YOLO...), tooling, protocols and additional architecture adjustments (gigantic context windows, instant adaptors, speed, etc) will make up for any lack of precision the same way we work around human flaws to help us succeed in most tasks.
> will make up for any lack of precision the same way we work around human flaws to help us succeed in most tasks.
A human won't flip-flop from "You're right! That doesn't exist" and then straight back to "You're right, that does exist!" based on how a question is asked.
People really hold LLMs their capabilities in way, waaaay too high esteem. You have to walk a tightrope with them, and you always will unless they can fix the hallucination problem, which is quite unlikely due to how LLMs work.
Both are true. What the parent poster is saying is that the models don't have to be perfect and can "hallucinate" because software is a unique domain in that:
* Validation: You can validate against objective signals either you or your tooling define (e.g. unit tests, compile errors, etc).
* Cost of Failure is Low: You can undo bad work, and feed errors back into the model as a signal to reduce future errors. Its not like physical domains (e.g. building a house, bridge, etc) where "undoing" is expensive and wasteful.
The models just need to be "good enough" that with enough tries the error accumulated over long jobs doesn't grow -> by adding data back into the model giving it feedback at each step you can curb this. How these agent tools sometimes achieve that is that they integrate with your build tooling stack, your IDE, your unit tests, etc etc -> they have a lot of "guard rails" to effectively curve the risk of hallucinations and/or rather when there is one to bring things back in line because they are long running processes.
TL;DR if you can't reduce the risk of bad model outputs you can mitigate the impact of the risk through retries and guard rails that feed back into the model. That's what these tools do to reduce the error rate. People are complaining about the risk of bad model output without looking at the other side - is there a way to make the effective consequence of that minimal and correct course? That's what these agent tools try to do - they want to work "like a human".
Don't get me wrong; there's a lot of signals that are hard to capture related to taste (e.g. I add a feature, it subtly changes the design of another X features already done) - and I personally find it easier to fine tune/go manual after a certain point but YMMV.
You can feed errors back in, but on anything of meaningful scale, context windows are going to be the largest bottleneck.
Context window size limits aside, Claude Code seems to atrophy or misinterpret existing context even prior to compactions very frequently. The more tokens within the context windows the shittier and shittier it performs—not that it's great to begin with.
That there is the issue. Right now they may not be good enough and create issues - but they are still a lot more useful than they were. Any improvements will feed straight into the tooling without much effort.
Don't get me wrong; I would love to be wrong. But I do think models will get better. There's just too much money thrown at the problem and SWE seems to be the target especially I think for Anthrophic where thats their main market/use case. They don't seem to have the same diversified user base.
If I wasn't a SWE I would think there's no need to become one tbh. Just need to wait a little bit more.
Not only are they just fancy autocomplete, they are so intrusive that it increases friction for me instead of lowering it
Having to pause a second to decide if I want to tab complete a line is one thing
Having to pause a minute to evaluate an entire suggested function is jarring and completely wrecks my momentum, especially if I wind up rejecting the suggestion
And especially because if I reject the suggestion, the damn thing keeps re-suggesting stuff while I'm typing the rest out
"Fancy autocomplete" doesn't seem very scathing when the autocomplete you're talking about is "give me feedback and find bugs in this 3000 line file: <paste>".
Ah yes, it's just using pattern recognition from its training data to generalize abstract concepts about software so that it can apply them to my specific, complex file to find a bug that eluded multiple software engineers for a month.
(IMO)The 2 sigma's are getting receptive audiences as pre-announces of layoffs at Amazon, MS and others yield the analog of a 50 VIX in the options markets, where enumerating mid/good/bad scenarios goes from difficult to impossible.
| its exhausting to keep up with what I (as a dev) should be doing to stay sharp
That for me is the biggest thing I am feeling about LLMs at the moment, things are moving so quickly but to what end? I know this industry is constantly evolving and in some ways that is very exciting but I also feel like it is this exponential runaway that requires very deliberate attention focused on the bleeding edge to stay relevant, when a lot of my time in my day job doesn't facilitate this (which I have identified and have made the effort and will be changing company in a month).
My own two cents on LLMs (as a junior / low mid level early career software engineer) is that they work best as a better version of Google for any well explored issue, and being able to talk through problems in a conversational manner has been a game changer. But I do fear sometimes that I am not gaining the same amount of knowledge as I would before LLMs became mainstream, it's a shortcut that in the long run I fear is going to reduce the average problem solving ability and original / novel thinking ability of software engineers (whether that is even a requirement in most SWE jobs is up for debate).
> its exhausting to keep up with what I (as a dev) should be doing to stay sharp
I have observed the JavaScript ecosystem producing one new framework after another. I decided to wait for the dust to settle. Turns out vanilla.js is still fine for the things I need to do.
I think this staying sharp is FOMO instilled by influencers and people selling guides/courses. Most of the stuff will be implemented by Anthropic, OpenAI etc.
You can run local models but it is like playing matchbox cars in your backyard and imagining you will be F1 driver some day.
Big guys have APIs you pay for to do serious work that’s all you need to know.
A bit unfair to call local models Matchbox cars compared to F1. There are plenty of uses for LLMs locally that don't require the largest models, it's not like it has to be all-or-nothing. For example, as a general browser assistant to help summarize articles, explain context, etc. the gemma-3-4B model does very well and is lightning fast on my old 3060 Ti.
You just wrote exact confirmation. Running 3060 Ti gemma-3-4B to play with as your local assistant is toying around.
Make the same as a startup or a company and you most likely will be out of business in 3 to 6 months because big guys will have everything faster and better in no time. GPT-o3 price drop of 80% most likely made running 3060Ti more expensive if you check your energy bill.
Not looking to do that though! You can call it toying around if you want, but I think you're really limiting your perspective by dismissing smaller models.
> I find the use of the word intelligence to be a bit of a misnomer. Is something intelligent if all its doing is pattern matching? Is the evolution that led to owl butterflies appearing like an owl intelligent? Im not sure.
Is a random number generator intelligent? I don't think people perceive or understand intelligence equally. I don't think we have an answer to what exactly is intelligence or how to create it.
> LLMs are really dividing the community at the moment and its exhausting to keep up with what I (as a dev) should be doing to stay sharp
You could try at your comfortable pace. I only started using agents very recently. The dangerous thing is to go to extremes (all in on AI or completely refusing the tech)
Would it surprise you seeing eg. seeing on the front page articles about both Nobel prize winners and Darwin award winners?
What is intelligence, after all? We expect AI to be as smart as Einstein or Terens Tao, but so far, we see that LLMs are pretty good at behaving just like humans, that is, most times stupid.
Personally, I dont think so. I can understand a mathmatical axiom and reason with it. In a sequence of numbers I will be able to tell you N + 1, regardless of where N appears in the sequence. An LLM does not "know" this in the way a human does. It just applies whatever is the most likely thing that the training data suggests.
But technically you can do that only because you recognize the pattern, because the pattern (sequence) is there and you were taught that it’s a pattern and how to recognize it. Publicly available LLMs of now are taught different patterns, and are also constrained by how they are made.
Maybe there’s something for LLMs in reflection and self-reference that has to be “taught” to them (or has to be not blocked from them if it’s already achieved somehow), and once it becomes a thing they will be “cognizant” in the way humans feel about their own cognition. Or maybe the technology, the way we wire LLMs now simply doesn’t allow that. Who knows.
Of course humans are wired differently, but the point I’m trying to make is that it’s pattern recognition all the way down both for humans and LLMs and whatnot.
When talking about AI, intelligence is meaningless if you don't defined it beforehand. The common sense meaning of intelligence fails on this kind of discussion.
I keep seeing the same “middlebrow dismissals” of LLMs in HN comments, it’s getting pretty repetitive to have to cover all of this over and over, but here goes (I recognize GP is only saying one of these, I’m just trying to preempt the others).
- “LLMs don’t have real intelligence” - We as a society don’t have a rigorous+falsifiable consensus on what “intelligence” is to begin with. Also many things that we all agree are not intelligent (cars, CPUs, egg timers, etc.) are still useful.
- “But people are claiming they’re intelligent and that they’re AGI” - OK, well what if those people are wrong but LLMs are still useful for many things? Not all LLM users are AGI believers, many aren’t.
- “But people are forcing me to use them.” - They shouldn’t do that, that’s bad. It doesn’t mean LLMs are bad.
- “They’re just pattern-matchers, stochastic parrots, they can’t generalize outside their training data.” - All the academic arguments I’ve seen about this become irrelevant when I ask an LLM to write me code in a really esoteric programming language and it succeeds. I personally don’t think this is true, but if in fact they are categorically no more than pattern-matchers, then Pattern Matching Is All You Need to do many many jobs.
- “I have an argument why they are categorically useless for all tasks” - the existence of smart people using these things of their own accord, observing the results and continuing to use them should put a serious dent in this theory.
- “They can’t do my whole job” - OK, what if they can help you with part of your job?
- “I’m a programmer. If I use an AI Assistant, but still have to review its code, I haven’t saved any time.” - This can’t be categorically disproven, but also isn’t totally true, and in the gaps in this argument lie amazing things if you’re willing to keep an open mind.
- “They can’t do arithmetic, how can they be expected to do everyday tasks.” - I’ll admit that it’s weird that LLMs are useful despite failing at arithmetic, but they are. Rain Man had trouble with everyday tasks, how could he be expected to do arithmetic? The world is counterintuitive sometimes.
- “They can’t help me with any of my job, I do surgery all day” - Thank you and my condolences. Please be aware though that many jobs out there aren’t surgery.
- “The people who promote them are annoying. I call them ‘influencers’ to signal that they are not hackers like us.” - Many good things have annoying fans, if you follow this logic to its conclusion you will miss out on many good things.
- “I’ve tried them, I’ve tried them in a variety of ways, they’re just really not for me.” - That’s fine. I’d still recommend checking in on the field later on, but I can totally admit that these things can take some finagling to get right, and not everyone has time. They will get easier to use in the future.
- “No they won’t, we’ve hit a plateau! Attention isn’t all you need!” - If all LLM development were to stop today, all AI cloud services shut down and only the open weights LLMs were left, I predict we’d still be finding novel usage patterns for them for the next 3-5 years.
As an aside, its amusing we simulatenously have this article on the front page as well as [Generative AI coding tools and agents do not work for me](https://news.ycombinator.com/item?id=44294633) also on the front page. LLMs are really dividing the community at the moment and its exhausting to keep up with what I (as a dev) should be doing to stay sharp