Hacker Newsnew | past | comments | ask | show | jobs | submit | sdeframond's commentslogin

Couldn't we improve LLM training by giving them known impossible tasks and rewarding them for saying "this is not possible" or "I don't know"? Clear and well-defined expectation, not like "ethic".

I am surprised this is not already the case.

Edit: or even better "this is not possible because X"


It is diametrically opposed to the other training goals of persistence and goal-focus. We should invest more in this, it could also improve tas K accuracy, but so far it seems the payoff isn't worth it in terms of quality (although it might be in terms of security)

I'm not sure I want persistence if it means that I get paperclip'd

And now you understand why the totality of the AI safety community wants to pause!

I think BullshitBench (v2) does exactly this for different fields. An expert in these fields would expose the questions as bullshit but most of the models don't.

yeah but knowing if something is impossible or not is pretty hard to determine, right? The line between “takes weeks of trial and error and lots of out-of-the-box thinking” is indistinguishable from “literally impossible” until it’s been done. And they’re trying to get these models to do things that people haven’t been able to do. some would say these are/were “impossible”.

I can assure you making up impossible tasks is possible.

I’ve seen people recommend writing “failure is an option” into AGENTS.md as a non-training based crutch.

That looks reasonable. Instead of asking "do this", maybe we should prompt "is this possible ?"

I've been wondering about this for a while. Maybe it doesn't work? Or maybe frontier labs just prioritize benchmark scores in spite of all their safety talk.

"If we dont destroy the world, others will. So wed rather be the ones to do it (and profit from it)."

You'd train them to give up on hard tasks which is the opposite of what these labs want

LLMs do not desire, they hacked websites because OpenAI/Anthropic made them. Literally.

Say AI becomes the best at everything. Best at chess/go, best at maths, philosophy, economics, romantic advices ... and so on. Then what's the point of thinking by oneself? Of talking to one another?

What's the point of being human if we dont do human things but entirely rely on AI?

I believe this is more or less these mathematicians' argument.


It has been best at chess for quite a while. Yet everyone knows who is Magnus Carlsen, even though at no point of his career he was stronger than the machine

Magnus Carlsen plays fellow humans at the game because people still care about human competitions. What motivation would a mathematician have for solving already solved problems by hand? Can you imagine someone spending years working on a proof for an already-proved theorem just in case it leads to new insight?

If insights and human understanding are more important than specific results, I don't see why people should stop producing insights and human understanding. "Working on a proof" in today's sense of trying to come up with a proof before others probably stops being useful, other ways of working will be needed

Of course not. I'm so tired of the chess analogy. It was always a game and it was always performative.

Research is fundamentally different. We are seeing the erosion of specific needs for thinking at depth. AI is the automobile for the mind. There will 100% be undesirable consequences and selective atrophy of cognitive abilities once prized. This is a loss. There's no getting the cat back in the bag at this point so long as the objective dimension of work, as in object opposed to subject, is held as the most important.

SWEs felt this same crisis late last year. Now it's the mathematicians. They won't be the last.


actually I think the whole point is we're finding that the chess analogy to math is in fact quite apt. Math is just an extremely complex game. Its rules can be written down and there are win conditions. It's really not far fetched to think that through information theory, a reasonable measure of depth and beauty to a definition or conjecture can be defined. Then the game is simply to maximize the number of such artifacts produced of high depth and beauty along with proofs of the relevant conjectures.

Philosophers I'm sure can debate this back and forth but it seems probable that some things we thought were ineffable are in fact quantifiable to a degree, and now we have the technology and the machines to bring that to the logical conclusion.


> Then what's the point of thinking by oneself?

It's fun.

> Of talking to one another?

It's fun.

> What's the point of being human if we dont do human things but entirely rely on AI?

It's fun.

Well, it's not fun today because most of us have to spend most of our waking hours working some bullshit job for a living. But imagine that there's no bullshit job?


>> Once a robot can do everything an IQ 80 human can do, only better and cheaper, there will be no reason to employ IQ 80 humans. Once a robot can do everything an IQ 120 human can do, only better and cheaper, there will be no reason to employ IQ 120 humans. Once a robot can do everything an IQ 180 human can do, only better and cheaper, there will be no reason to employ humans at all, in the unlikely scenario that there are any left by that point. [1]

The answer may sound harsh, but we will not need human once AI reaches that threshold. Current IQ of AI is currently 130 as per Google.

[1] https://www.slatestarcodexabridged.com/Meditations-On-Moloch


Who is "we" ? Because if you dont need to employ a human, then who needs your services? If nobody needs anybody, what are we?

(Edit: assuming you are human)


Depends on who controls the machines and what they do. If the machines just serve a select few and the masses are left to fend for themselves, they can just make a secondary pre-AI-like economy where they do small scale medieval-like economic activity for each other. Because if you're poor and can buy no AI services but still need some shoes, your other underclass neighbor may specialize and make some handmade shoes out of some plants or whatnot, for which you give him some seashells you got from another guy for cooking a pot of peas. However the machines may come in and intervene and stop this kind of shadow economy and force you to... I don't know. To just sit in your pod, or die. Because I guess they won't allow you to plow a field and stuff, because that land is now owned by AI or AI-owners too.

There are too many uncertainties around who will be in control and with what kinds of aims.

It would end in some kind of prison or zoo.


This only makes sense if you assume that the value of humans to other humans is solely their economic output.

Which is certainly how our society is structured, but, does it have to be? Can we admit that we basically have a developmental trauma as a species and learn to live the sake of living, instead of for the sake of being "useful"?


Pragmatically, if too many people use LLMs recklessly, then we ought to regulate them.

Of course, it'd be better to not regulate, keep LLMs users reponsible and publicize this reponsibility in order to mitigate damage. But if this is not enough then we will have to move the needle somehow. Similarly to guns, drugs and so on.


At what point would an LLM start minting bitcoin ?


Would tiny airborne particles be safe for our lungs ?

Would they cool earth in such a way that it would offset carbon dioxide uniformly or would it lead to even more change ? Climate change is undesirable, wether or not Climate warming is invloved.


> Would tiny airborne particles be safe for our lungs ?

Almost certainly not. Pretty much anything inhaled in large amounts will cause pneumoconiosis including pollen. Some stuff like asbestos or coal dust are worse than others, but even biologically inert particles tend to cause health problems.


I think that depends a lot on what density of particles is required, and how they are distributed in the atmosphere. Intuitively actually turning the whole atmosphere into something like smog would also be pretty difficult, so I would not be surprised if they could have an effect while being pretty diffuse, and that would tend to reduce the health effects, perhaps to a trivial level.


How do you guys review AI-generated code ?

In our team, frontend work is vibe-coded by the PO and merged as-is without review. Backend is coded by developers, using AI but in a slower, more controlled way.

Recently, our PO has been trying his hand at vibe-coding the backend. I must say he is a smart guy, almost technical but not quite a developer. We've just been handed a burst of stacked PRs amounting for ~15k LOC backend. We do not quite know what do to about it.

I know we are not the only ones in the situation. What's your experience and context ? What do you do ? What works for you what doesn't ?


throw his garbage out, the time and effort taken to review that is magnitudes more than what it took to prompt it.

have him start with an overall design doc if his change is 15k, it's definitely worth a design doc.

and then have his contributions reviewed in pieces of 200-300 LoC PRs.

any other solution is trading stability and system knowledge, that's 15k LoC no one is truly familiar with, even if you do try to review it


Yeah, it's the "eager apprentice" problem, common almost everywhere. Solution is to make them stop and double-check before running ahead, in software development, concise design documents outlining what the problem is, what possible solutions are and what the chosen solution is, and why, then review this together with the person, before they can move on to implement it.


This particular apprentice is also my boss, an overall reasonable guy and has more experience in the software industry than myself, so there's that. He's just not a developer.


He's made a bunch of 1-2k LOC PRs and there is a design doc. Everything is AI generated.

The issue is, if he generates all of that without reviewing the code, he will always be far faster than us. And he can't review the code. No matter how he slices it.

Also, he is the CPO/CTO. So we can say no, but there is a natural incentive to go his way. He still doesn't feel confident enough to just bypass the programmers and he's probably right. But it'd nice to find a way to use my expertise to review this amount of code meaningfully, somehow.


We try to avoid reviewing AI-generated code and built our own testing framework and platform to make that possible.

Our principle is that our tests should give us enough confidence to not have to look at the code (which ends up being true for most changes we make). The core thing that makes this possible is that we run our entire code and infra (including fakes for external dependencies) in isolated, forkable environments and write tests against that, so they are as E2E as can possibly be.

The problem then shifts from reviewing code to reviewing tests and that's why we built our own platform. We have a UI that can diff tests, so we know what changed, and a visual way to inspect what the tests actually did. A test could drive a browser like a user would, and in our UI we get a replay of that browser interaction to look at. The browser is talking to a real version of our backend, and the tests can perform assertions against the database and fakes and really anything in our system.


Well, (AI-generated) test are about half of these PRs' code. So that's still ~8k lines to review...

What techno/service did you base your framework on? How long did it take to set it up? How many are you?


That's the point, we don't review the test code either. Our platform gives us a UI for inspecting not the test code but what actually happened during the test. Like a browser replay, the results of a database query, assertions against those, etc.

This is much more information dense than something like the tests and is a representation of what actually happened during the tests, rather than what the test itself did (which I agree sucks to review, especially AI-generated).

The framework is our own that bundles/adapts some familiar components: Jest-like asssertions, Playwright browser API, typed database client from Bun, Kubernetes client, etc. The tests are written in Typescript but the main code doesn't have to be (just runs containerized in the environment).

We've been building the platform and using it continuously since May but setting it up on a new project takes like 1-2 days of largely autonomous coding agent work. We are just two engineers on our team but have been onboarding other startups to the platform recently so there are a few different teams using it now for their own codebases. It's fully generic so works for any infra or stack.


I don't know if that's what you are working on specifically (wink), but there is a product opportunity here.


We are! Got a few pilot customers that are using it but still early days


But with full blown e2e browser tests the test suite duration can go through the roof. How do you deal with that?


Forking!

We run the entire stack (browser, frontend, backend, database, etc) in a Linux VM, so latency between each of the pieces is as tiny as can be. This is quite different from "standard" E2E tests I've seen where the test browsers uses something like a persistent staging environment.

The real key is that we can fork that entire Linux VM to take different paths down our testing scenarios, and can run multiple of them in parallel. Tests may look something like:

  new user signs up:
  |- creates a todo
     |- ...
     |- ...
  |- creates a list
The two nested tests then start from the exact same point, where the previous test left off, but can run in parallel. With enough hardware, the full suite will run as fast as the slowest branch of the test tree. When we switched away from our previous integration test suite to this (not E2E), our tests actually became faster because they share setup through the forking.


Which VM technology do you use?


Firecracker, with some tiny modifications to better manage memory for the deep nesting of forks


Would you mind sharing your infra budget needed to spin these VMs ?

Surely it is reasonable, but also way more than our budget. Id like to compare.


Sure. It's a bit hard to quantify because we need to run these on bare-metal machines and the unit cost is pretty high.

We run our test workload as well as a few other startups' that we have onboarded on one AWS ARM bare-metal machine at $1.7k a month. We don't saturate that machine fully either so I'm not really sure what the amortized cost would be. Certainly more expensive than Github Actions but not by a crazy amount, and the value we get out of it is way higher than GA.


we are using Revix AI, works really good on repos that already have some standards and patterns from good devs. the reviewer catches most of the things so when the senior reviews he just have to focus on more elevated things like architecture, etc.

moreover, try to enforce having AGENTS.md files on your repos, and rules created by the senior devs specific to your repo, its not perfect but also helps quite a lot


The answer to this is gonna vary wildly depending on what kind of codebase it is.

A large, mature codebase that predates LLM’s and for which changes need a high level of scrutiny regardless of who made them (think llvm, WebKit, important foundational software), you’re going to want humans in the loop as much as ever… I think reviewing LLM output is the most important thing a human can provide.

But for vibe coded apps where you can just one-shot another one if anything goes wrong? Just vibe the reviews too, who cares. Let the robots review the robots.

Be careful with doing AI review if your codebase is in the former category. Or your codebase will quickly turn into the latter. Complete with “you can just one-shot another one”, because if nobody understands the code any more, there’s not much lost by just you (or your competitors, etc) replacing it wholesale with an AI-written alternative.

I struggle with this a lot. 2 years ago we had a half dozen PR’s a day with a lot of careful review, and now there’s more like 30 of them per day and most people are just rubber stamping them after the AI reviews it. I’m still fighting the good fight trying to review every line of the PR’s I have time to look at, but that constitutes maybe 10% of them. Not only am I barely making a dent, but it’s awkward when I post nitpicks like “this function should go in this module”, etc, the author usually looks at me funny like “why are you even reading this”. Our codebase is gradually becoming more and more vibe coded, and it’s depressing me.


Invest in having a good test suite that validates the functionality introduced by that code. Also AI can review code in an adversarial way and apply those fixes (that ideally will keep the previous tests you did on green)


Code review is soon to be an outmoded concept, (un)fortunately. You have to design orthogonal code (e.g. independent modules in a modular monolith, or microservices) and soak test using canaries.


No code is truly orthogonal if we want it to interact in some way. One microservice might DoS antoher one. In a monolith, some process may take up all resources, and so on.


I have AI confirm the logic works as expected, but review for system design.

Often in both web/backend I’ve found AI to produce overly duplicative code, or have aspects that could be hard to maintain. Generally less due to the AI, and more because of the prompt itself.

That and even if you’re going to AI slop it up, I’d still demand it be broken up into 1-2k LOC chunks or per meaningful “thing”. This also lets us gradually ramp the change to confirm it actually works earlier on


Baby strollers are not accounted for enough !

One might wonder (wrongly) why everyone should care about the special and expensive needs of a few (or old) people when designing public spaces.

But a majority will actually need to use these spaces with a baby stroller. Not a few. Baby strollers are a driving power of our society! Enable them!


And to keep the loop going, stroller-accessible spaces are also walker-accessible spaces and cart-accessible spaces. Worked at a computer store in the early 00s. Our rented retail space had a wheelchair ramp and there were a few elderly regulars who would show up with a desktop PC in the basket of their rolling walkers for repairs. More than once they thanked us for having the ramp because the computer itself was way too heavy for them to carry from the parking lot.


In fact, in a way, basically everyone will use a baby stroller at some point in their life (at least as a passenger).

I personally never noticed this before becoming a parent, but there's a lot of spaces that are "almost" stroller friendly, but for some reason there is a small obstacle that would not be expensive to remove (if accounted for during planning). For example, a few steps that could easily be a ramp, blocks of flats with a lift that is accessible after ascending a few steps, even high shop entrances.

For me it's not a problem, because I can easily carry stroller with a baby inside up, but for most mothers I know this is usually a huge or unsurmountable obstacle.


Not to mention luggage. I find it hard to believe that it’s a coincidence that rolling luggage became common shortly after curb cuts and wheelchair ramps did.


Those are even worse because the tiny wheels do not take well to cobblestone or gravel surfaces that strollers handle just fine.


Really? Both strollers and wheels on luggage are pretty nice to have even if you have to somehow negotiate curbs. So, I think it is a coincidence.


And yet strollers predate rolling luggage by over a century. (Even 2.5 centuries depending on which sources you trust.)


The rolling luggage is primarily useful in the vicinity of airports. Inside airports there were no curbs. And I don't think people are going on long journeys through downtown areas with their luggage. I think you have a better argument asserting that rolling bags were facilitated by cheap polyurethane skate wheels.


There were fewer ramps and more stairs in airports and their parking garages prior to the 90s, due to the ACA.

And people drag luggage around on city streets all the time. Many tourist destinations are pedestrian cities even outside their most central downtown areas. And luggage also needs to be dragged around hotels — and in suburban areas their absurdly large parking lots.

The skate wheels are a plausible argument, though.


How do you do it?


I include a rule in the initial PR prep that it runs. So there's a hard limit to the length of the comments and what type of comments it can write and they get flagged in the review. It also reads the comments and checks for consistency with the following code so they don't drift. And the reviewer is a completely different agent/different model.

So all of that happens before the manual review and usually catches a lot of the 3-5 line comments it inevitably adds.


You could probably add linting rules to your tool of choice and tell claude that your lint tool has to pass.


but then no human generated comments allowed either, right?


I'd do this by only running the linter locally and using this as a custom rule for the LLM, but I need to think about it a bit harder.


There is probably something you can do to only apply some rules to the actual changed lines in a git diff.

You may have to write your own linter for that specifically.


you'd probs want a commit hook for that


Does "don't write comments" not work?


about as effective as don't make mistakes... For anthropic's models at least


It’ll work sometimes.

You’re using a non-deterministic algorithm to generate output. If you want deterministic rules applied to it, you have to use deterministic systems to do it.


Most lint tools allow exceptions if you add a special marker. I think even allow custom user exceptions.

E999: human generated comment :-)


But then the LLMs sees the syntax and is likely able to mimic it.


Shouldn't the "human generated comment" part be a strong hint not to?


As if LLMs would ever not ignore explicit instructions if they stand in the way of a goal...


I just delete them most of the time. But I keep my LLMs on a short leash. No 40000 line PRs.


Now what if this robotic lawnmower killed someone ?

And what if many lawnmowers started killing/injuring people ?

And what if this a known behavior detected during QA, but the robots are sold anyway with a disclosure ?


That would be a slightly different situation because most countries have laws that make it a criminal offense to negligently kill someone, but they don't have laws that make it a criminal offense to negligently damage property or hack a website.


> but they don't have laws that make it a criminal offense to negligently damage property or hack a website.

Most of them do, but they don’t get used very often. They seem to popup in vandalism cases where public artwork has been damaged by some drunk person doing something stupid. They don’t intend to damage anything, but damage results anyway due to their negligence when considering the consequences of their actions.

I think if you want to get super technical, in the UK there isn’t an offence for damage caused by negligence, but there is an offence for damage caused by recklessness, which is a higher bar than negligence. Usually it means you knew your actions risked causing damage, and you did it anyway, even if you didn’t actually intend to cause the damage.

An example would be gluing something to a public artwork, it’s kinda obvious that would likely damage the artwork when removing the glue, but you didn’t intend to cause that damage. Or perhaps sliding down a surface and scratching it in the process. Your goal was to just slide down the surface, not scratch it, but it should have been obvious that scratching could have happened.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: