Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I find it odd that discussions of alignment don't mention 'legality'

Certainly humans have complicated alignments and are guided by emotional morality - we might say many of these principles are hard to define and humans don't agree.

All true, and at the same time the principles get codified into laws, I would guess particularly in areas where harms may result.

Would it potentially be easier to train strong alignment-with-legality vs grappling with fuzzier questions of values?

 help



> I find it odd that discussions of alignment don't mention 'legality'

Because other wishy-washy stuff doesn't involve prison. Not that our new oligarchs with the politicians in their pockets have any real risk of it, but why take chances. Much safer to doodle about alignment and such abstractions in safer and softer contexts.


But that was kind of my point

Maybe it'd be easier to train the models on key parts of the legal code and give it a hard aversion to breaking the law - rather than training on vague value judgements and then hope the model doesn't break the law


I got that. I'm saying that it is deliberate. Wiggle room et all.

How does not explicitly trying to train the model specifically not to break the law give them any wiggle room if the model then goes and breaks the law?

Deniability on intention. Models can hallucinate, all bets are off. Sure, they can make it stricter but why do that and subject themselves to harder scrutiny based on the training criteria?



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: