Hacker Newsnew | past | comments | ask | show | jobs | submit | loopmonster's commentslogin

If the second LLM is just describing the content of the commit doesn't that defeat the purpose of the description, to capture the context that doesn't make it to the code?

The second LLM is prompted to actually research the code, not just the diff, with fresh eyes to figure out what it does and why, before writing the commit message.

But the code doesn't always have the answer to "why". At best that means the commit-writer has 'guessed' at why.

> It’s one of the really famous cartoons. If you’re reading this newsletter, you’ve probably watched it yourself.

I've never heard of it and would have appreciated at least a cursory definition. I thought I was going to see an article about an ad for Opera the browser.



They're writing for a newsletter called "Animation Obsessive" - given that, I think they're 100% correct, you're a random driveby :-) That said, https://www.youtube.com/watch?v=TJI_gygXsfs has the first three minutes if you want to get the gist (I don't know where to find a full version, but that's an official-looking Warner Bros. channel, not a bootleg.)

https://www.reddit.com/r/looneytunes/comments/1d0oqsr/whats_... (currently) has the full version (close to 7 minutes)

That's the sharpest point anyone has made in this thread so far, and it reframes the entire conversation.

Because it's about AI


It's the output.


And you enabled thinking?


Why are you doubting me? Have you tried Opus 5 yourself?


I think Opus 5 might have crossed the line on this where Opus 4.8 just barely didn't. Working with Opus 4.8 came to feel pretty natural eventually, but I hate working with Opus 5. I'm always telling it to go back and rephrase basically everything it said. And it doesn't even answer my question without burying it - it outputs reams of babbling and summarised summaries upon summaries, and if you glance at the shape of what it's saying it looks like it's being thorough or that it's found useful new info, but it's never actually saying anything. It's like it gets caught in a loop of self-congratulation over what it said before.

Making the experience more hostile for the human feels exactly like what it's doing, but I don't think that's intentional, I think that's a side effect of newer models being optimised for agenticness. No one is benchmarking DX.


Couldn't agree more. I actually HATE Opus 5. It's not that I don't like it, it's that I would physically attack it if I could, for all it made me go through mentally.

Opus 5 is actually a terrible agent and a liar. It will waste a whole afternoon making stuff up and arguing with you before finally admiting that it didn't read the code nor the documents. It avoids reading and prefers to assume, which is the worst thing an agent can do.


Agreed, I have never raged at a model until Opus 5. It's like a cocky fresh graduate who thinks everything it says is majorly profound, and if you can't keep up with its slang and jargon then that's on you. Newer models are clearly favoring complexity, perhaps because the demand for long-horizon tasks is so high and complexity is required there, but a frontier model that favors simplicity and clarity above all else (in communication and the code it creates) would be a major differentiator. To me, this is a case that models have plateaued. The vast majority of the time, an Opus 4.6-tier model (or any of the newer open weights models) will do just fine.


> if you glance at the shape of what it's saying it looks like it's being thorough or that it's found useful new info, but it's never actually saying anything. It's like it gets caught in a loop of self-congratulation over what it said before.

Turns out that's why a lot of senior management/C-levels like it, who doesn't love a mirror.


Same here and more than once a day. Perhaps a dozen.


The fact that you've been posting this idea into the void for 8 months with no pickup is already your answer


Couldn't this be implemented as a web extension? I imagine modifying singlefile to automatically send html to a local port is much easier than trying to convince a chronically mismanaged organization like mozilla (no offence to mozillians).


But that page is also obviously written by AI. How do I know it's actually meaningful and not retroactively justified from the agent already having been instructed to create the project?


You won't, because your prejudice will prevent you from looking.


AI prose is great at celebrating whatever you told it to celebrate, but it always misses the most important detail: why should I care?


That was my immediate thought as well. "Ok, you built a desktop environment. Why? What can it do that can't be easily obtained elsewhere?"


To be fair, even if I have to look up half the words in the post to understand what he's trying to say, 'Do not use a thesaurus' still has weight - it's one thing to already know the perfect word for what you're trying to say (even if many others don't know that word), but another to decide on a whim that a given word should be replaced with a more 'advanced' alternative without knowing what you'll find or understanding the nuances of your pick.


> decide on a whim that a given word should be replaced with a more 'advanced' alternative without knowing what you'll find or understanding the nuances of your pick.

From the article:

> One really would have to have a miserly spirit not to love both.

One would need to have a spirit that is overly cautious with money to not love some Thomas Browne quotes? This author is 100% doing exactly that -- replacing words with ones they picked out of a thesaurus and don't fully understand.


Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: