Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Crucially, unsafe is also about the social contract. Rust's compiler can't tell whether you wrote a safety rationale adequately explaining why this use of unsafe was appropriate, only other members of the community can decide that. Rust's compiler doesn't prefer an equally fast safe way to do a thing over the unsafe way, but the community does. You could imagine a language with exactly the same technical features but a different community, where unsafe use is rife and fewer of the benefits accrue.

One use of "unsafe" that was not mentioned by Nora but is important for the embedded community in particular is the use of "unsafe" to flag things which from the point of view of Rust itself are fine, but are dangerous enough to be worth having a human programmer directed away from them unless they know what they're doing. From Rust's point of view, "HellPortal::unchecked_open()" is safe, it's thread-safe, it's memory safe... but it will summon demons that could destroy mankind if the warding field isn't up, so, that might need an "unsafe" marker and we can write a checked version which verifies that the warding is up and the stand-by wizard is available to close the portal before we actually open it, the checked one will be safe.



The social contract is manageable precisely because the unsafe subset of typical Rust codebases is a tiny fraction of the code. This is also why I'm wary about expanding the use of `unsafe` beyond things that are actually UB. Something that's "merely" security sensitive or tricky should use Rust's existing features (modules, custom types etc.) to ensure that any use of the raw feature is flagged by the compiler, while sticking to actual Safe Rust and avoiding the `unsafe` keyword altogether.


I agree with this mindset a lot. One example I like about how to deal with exposing a "memory and type safe but still potentially dangerous" API is how rustls supports allowing custom verification for certificates. By default, no such API exists, and server certificates will be verified by the client when connecting, with an error being returned if the validation fails. However, they expose an optional feature for the crate called "dangerous_configuration" which allows writing custom code that inspects a certificate and determines for itself whether or not the certificate is valid. This is useful because often you might want to test something locally or in a trusted environment with a self-signed certificate bit not want to actually deploy code that would potentially allow an untrusted certificate to be accepted.


When I read posts like yours, I wish unsafe had a different name like "human invariant" or whatever.

Something that would make it harder to water down the meaning of the clearly defined unsafe keyword to suddenly mean something else.

Using the unsafe keyword to mark a function as "potentially dangerous" is just wrong.

Just prefix your functions with something like "dangerous_call", but don't misuse unsafe!


This is a complete misuse of the unsafe as language concept in high integrity computing.


This person isn't wrong. A lot of serious Rust users don't agree with what the GP is suggesting. `unsafe` has an explicit meaning: the user must uphold some invariant or check something about the environment, otherwise it is memory unsafe.

I have several times in code review prevented people from marking safe interfaces as "unsafe" because they are "special and concerning", overloading the usage of unsafe is itself dangerous.


True, but I think allowing potentially UB-invoking code to not use "unsafe" (e.g. because the use is in the context of FFI, so the unsafety is thought to be "obvious" and not worth marking as such) might be even less advisable. This makes it harder to ensure the "social rule" mentioned by GP, that every potential UB should be endowed with a "Safety" annotation describing the conditions for it to be safe.


Your comment gave me an idea for a lint that might help prevent those mistakes. Right now rustc flags `unsafe {}` with an "unused_unsafe" warning. However it doesn't warn for `unsafe fn foo() {}`. Maybe it should.


I think as described you would get false positives, because `unsafe fn foo() { body that performs no unsafe operations }` can be unsafe to call if it interacts with private fields on datastructures used by safe (to call) functions that perform unsafe operations... I expect you would end up with a reasonably high number of false positives.

For an example, consider Vec::set_len in the standard library. Which only contains safe code, but lets you access uninitialized memory and beyond the length of your allocation by modifying the length field of vector: https://doc.rust-lang.org/src/alloc/vec/mod.rs.html#1264

You might be able to fix this with a lint that looked at a bit more context though, `unsafe fn foo()` in a module (or even crate) with no actually unsafe operations is very likely wrong. Likewise `unsafe fn foo()` which performs no unsafe operations and only accesses fields, statics, functions, and methods that are public.


The Rust Vec type has an unsafe function called `set_len` that changes the length of the vector without checking whether the new length is in bounds, or whether the memory containing any new values is initialized. The body of the function is the following:

    self.len = new_len;
No unsafe operations in sight. Should the compiler emit a warning here?


This made me think about what you could potentially do to get a similar sort of thing using the type system instead of the built-in effect of unsafe.

Seems like you could create a sort of userland-unsafe by using a closed trait[1] and requiring its use on a 'dangerous' method or struct or whatever:

   mod my_unsafe {
      pub trait MyUnsafe {}
   }
   pub struct MyUnsafe;
   impl my_unsafe::MyUnsafe for MyUnsafe;

   pub fn dangerous_function<U: my_unsafe::MyUnsafe>() {}

   fn main() {
     dangerous_function::<MyUnsafe>()
   }
Obviously this doesn't let you do a block of unsafe without having to repeat it like `unsafe {}` does, but it doesn't leave you much room to do the dangerous things without the shrinkwrap agreement either (and turbofish are so ugly at least for me they'd be a deterrent).

That said I find the named use-case kind of weird. The whole point of the library is to do these unsafe things, so it's kind of silly to be like "don't forget it's dangerous!"


Oh that's just the suggestion I mentioned here https://news.ycombinator.com/item?id=31009572 (I saw it on some internals.rust-lang.org thread), but with no new syntax. It also mirrors the usage of existing beyond-UB safety markers like UnwindSafe https://doc.rust-lang.org/std/panic/trait.UnwindSafe.html but with inverted polarity (instead of asserting something is safe, it asserts that the programmer acknowledges the new kind of unsafety)

I think an interesting part of unsafe Rust is the interplay between positive and negative polarities (unsafe/safe fn vs unsafe blocks basically) and I think that adding new syntax is the way to leverage this kind of idiom in the new kind of unsafe


Real formal verification is clearly a step up from rust's meaning of "safe", but I don't think it's wrong to try to add another rung to the verification ladder at a different height. Verification technologies have a lot of space to improve in the UX department, here Rust trades off a some verification guarantees for a practical system that is still meaningful.


Real formal verification in Rust is basically waiting for a proper formal model of Rust semantics (including semantics of "unsafe"). This is one of the most unappreciated things about Rust's future potential, AIUI; the only language in common use that supports both systems-level programming and formal proof is ATS.


Eh. You can write, and people have written, formal proofs about C code too. In fact, the difference between Rust and C in terms of safety isn't very large if you have an infinite budget to put toward proving all the code that you write correct.

The safety benefits of Rust appear when you aren't willing to formally prove all the code you write correct. That's because you only have to prove the unsafe code, plus the memory model, correct in order to guarantee memory safety for the entire program. This is less burdensome than in C, where to do the same you have to prove correctness properties about the specific code that makes up the whole program. Rust makes it easier to prove certain properties about the program (far easier in the case of memory safety), but it was always possible.


I’m doing my PhD on the formal verification of Rust, and while you’re right that safe code provides a lot of informal advantages it also dramatically simplified form reasoning.

In particular, the dirty secret of C verifiers is that they don’t handle pointers all that well. Either you find yourself doing a lot of manual proof work or you have to dramatically simplify the memory model.

In contrast, when verifying safe Rust, the rules of the borrow checker allow us to dramatically simplify the verification work. All of a sudden verifying a manual memory program with pointers (borrows) becomes as simple as verifying a basic imperative language. I’ve been working on a tool: https://github.com/xldenis/creusot to put this into practice

On the other hand, the moment you dive into unsafe, all bets are off and you find yourself wading through the marshes of (weak) memory models with your favorite CSL as your only friend.


> I’ve been working on a tool: https://github.com/xldenis/creusot to put this into practice

Note that there are other tools trying to deal with formal statements about Rust programs. AIUI, Rust developers are working on forming a proper team or working group for pursuing these issues. We might get a RFC-standardized way of expressing formal/logical conditions about Rust code, which would be a meaningful first step towards supporting proof-carrying code directly within Rust.


> Eh. You can write, and people have written, formal proofs about C code too.

Not without a formal model of C, and the C standard is only an informal, natural-language text. Rust having memory safety as its express goal (which is basically table stakes for any sort of workable language semantics) means that it's at least realistic to think about a formal semantics for Safe Rust. Then you "just" need to deal with the uncomfortable reality that lots of Safe Rust facilities actually bottom out into Unsafe Rust, which is why it turns out you must care about its semantics too. But the hope is ultimately that the small portion of real-world Rust codebases that's Unsafe Rust might not "infect" the Safe Rust to the point of making verification as practically unworkable as in C.


There are formalization of C, e.g. compcert, which also compiles to machine code in a way that should guarantee proofs all the way though execution (assuming the machine modeling is correct).

It seems like it is a long way off from being possible to do that for rust.


Nope, Ada/SPARK and Frama-C, both with plenty of experience in high integrity systems deployments into production than Rust, currently.


That isn't really true. "unsafe" specifically means "memory-unsafe". The checking is left to the humans, and so is the justification, but the meaning is not.

See https://doc.rust-lang.org/book/ch19-01-unsafe-rust.html

Check out also the "Rust considers it safe to" section of https://doc.rust-lang.org/nomicon/what-unsafe-does.html


I feel like there's plenty of instances where unsafe is used in rust to indicate library level UB, without actually being memory-unsafe. (Like, where using a method in a certain way may violate a runtime assumption of a library, causing application level UB).

For example, String::from_utf8_unchecked[1]. There's a comment in the documentation trying to justify why invalid strings could cause memory unsafety, but its pretty weak:

> This function is unsafe because it does not check that the bytes passed to it are valid UTF-8. If this constraint is violated, it may cause memory unsafety issues with future users of the String, as the rest of the standard library assumes that Strings are valid UTF-8.

Like, sure, we could make that argument for any runtime invariant. "Look, this is totally memory safe but if you call this method we assume this constraint is valid. We're marking this unsafe anyway because we want people to take care in using it. So uh, its memory-unsafe because violating this constraint could cause memory unsafety issues in the future or something? Yeah that'll do."

It seems like lots of folks quietly want unsafe to mean "this method skips runtime checks", with no specific grounding in memory-unsafety. And the line between those two ideas is super blurry in practice, even in the standard library.

[1] https://doc.rust-lang.org/std/string/struct.String.html#meth...


It is perfectly reasonable that `from_utf8_unchecked` is unsafe because the sequence of operations `from_utf8_unchecked(array).chars().next()` will call [1] and can trigger memory unsafety with the `unwrap_unchecked` call if the array is not valid utf-8. This means that at least one of the three functions must be marked unsafe.

[1] https://github.com/rust-lang/rust/blob/027a232755fa9728e9699...


Ensuring that values actually are a specific type isn't directly memory safety, but if you don't do it then you can't even have an exhaustive match without the danger of your program going wild.

I see enforcing utf8 as about on par with enforcing that a boolean be 0 or 1, and not 5.

How do you feel about this list of what's unsafe and undefined? https://doc.rust-lang.org/nomicon/what-unsafe-does.html


> I see enforcing utf8 as about on par with enforcing that a boolean be 0 or 1, and not 5.

UTF8 correctness is an invariant of the code. Its not something the compiler understands or cares about it.

But really, I get it! My point is that rust is rife with people marking functions "unsafe" when they mean a method will cause bugs / "runtime UB" if used improperly. Even the standard library does this, so it feels a bit rich to cry foul when 3rd party libraries (ab)use "unsafe" in the same way.

I certainly use unsafe like that, in order to mark methods which, if misused, will violate internal invariants of my code.

How does that list of unsafe features justify the unsafe marker on String::from_utf8_unchecked?


Well how do you feel about the comparison to booleans?

If I had a way to set some booleans "unchecked", in a way that could make the numerical value be 5, would you label it unsafe?

I really think it should be unsafe, because it's crazy to make ifs or switches on booleans be non-exhaustive, and it's also crazy to have an "invalid boolean" path anywhere one is used.

When you have an encapsulated value, with special accessors to make sure it's always in range, then you don't have to have an unchecked setter function at all. But if you do add one, I think it makes sense to make it unsafe. That function is basically allowing a blind memcpy over a piece of data. It ties into safety pretty strongly. That kind of access is a lot like dereferencing a raw pointer, and not making it unsafe means you allow a lot of those "invalid value" cases in the link to be possible in safe code.


I get that this was just an example, but a boolean 5 would actually be language-level UB that could have very real correctness implications!

The Rust compiler is allowed to (and sometimes required to) use the invalid values of a child value to represent other cases of an enum. For example, this is how it optimizes `Option<&T>` into "pointer where null is None and any other value is Some".

This would technically also allow it (I don't know whether it currently does this..) to optimize the following enum:

    enum Foo {
        A(bool),
        B(bool),
        C(bool),
    }
into the following single-byte representation:

    A(false) => 0,
    A(true) => 1,
    B(false) => 2,
    B(true) => 3,
    C(false) => 4,
    C(true) => 5,
With this representation, `A(5 as bool)` would be impossible to distinguish from `C(true)`, which would be complete and utter nonsense.


> I get that this was just an example, but a boolean 5 would actually be language-level UB that could have very real correctness implications!

Yes, that's why I used it as the example.

And you don't need anything that complicated, either. A simple if or switch statement might be optimized by the compiler for 0 and 1 and jump into random code for 5.


I agree. I think "unchecked" booleans which could contain 5 should be unsafe too.

I'd be happy to extend the definition of "unsafe" to mean "If you misuse this, you may violate some internal invariants of the library. And that may cause unexpected bugs."

Thats a broader, but in my opinion much more clear definition of unsafe, which covers how the keyword is actually used in actual code. (In std::String and elsewhere).

I've implemented high performance b-tree and skiplist implementations in rust. There's plenty of functions in both libraries which I've marked as unsafe because if you use them carelessly, you'll violate some internal invariants. Will the library break as a result? Yes. Will the resulting bugs include memory corruption? I really have no idea, and I don't really care enough to go in and test that. So I've marked the methods unsafe, and provided safe APIs for consumers to use instead.

Do you think this is an appropriate use for unsafe? If not, I'm curious to hear why.


I think it depends on the kind of invariant. If it leads to undefined behavior that wouldn't already be possible, then mark it unsafe or change the unsafe code to eliminate it. If it just makes it act badly, like dropping a node, that's likely not a good place for unsafe.

Or more simply: Worry about all kinds of undefined behavior. That makes many issues easier to find, and you don't need to chain on additional logic to figure out how it might corrupt memory.


I would argue it’s language-ub to convert an invalid slice into into a String, it’s bubbling up a requirement of str which relies on the contained utf-8 to be well-formed for memory safety.


the key thing here is that if you know that no safe code contains invalid utf8, it is safe to do unsafe indexing into tables that are only sized to deal with valid utf8. as such, this works in practice lead to ub in otherwise correct code, so it should be marked unsafe.


> Like, sure, we could make that argument for any runtime invariant.

One difference is that Rust has a separate type for String-like objects that do not respect the invariant, namely Vec<u8>. So those who wish to write memory-safe code acting on arbitrary bytes can simply use that. Also, it's quite normal for an invariant defined entirely in Safe Rust to impact memory safety in a very real sense; consider Send and Sync. These are seemingly arbitrary labels, but their semantics is nonetheless quite well defined.


> One use of "unsafe" that was not mentioned by Nora but is important for the embedded community in particular is the use of "unsafe" to flag things which from the point of view of Rust itself are fine, but are dangerous enough to be worth having a human programmer directed away from them unless they know what they're doing.

Rust gave up this notion when they decided that mem::forget could, in fact, be safe, since even though it was initially marked unsafe it could be trivially implemented in safe code using only safe constructs from the stdlib.

We need another keyword or concept to refer to unsafeness that isn't related to UB. Some people suggested tagging unsafe with something like

    // declare SummonsDemons as some kind of safety market that goes beyond preventing UB

    unsafe(SummonsDemons) fn f() {
        // summons demons
    }

    fn safecode() {
        // SAFETY: the demons are cool this time
        unsafe(SummonsDemons) { f() }
    }


> One use of "unsafe" that was not mentioned by Nora but is important for the embedded community in particular is the use of "unsafe" to flag things which from the point of view of Rust itself are fine, but are dangerous enough to be worth having a human programmer directed away from them

I thought there were subtle language differences such that if you took ordinary, perfectly valid as safe rust code and marked it as unsafe the result could be incorrect? Am I mistaken-- maybe that was some pre-1.0 property of the language and it's actually okay to go peppering around unsafe for typechecking like usage?

Separately, if people are creating unsafe interfaces to protect them against uses that fail to uphold their required invariants it would probably be really useful if they could be tagged with specific required unsafty-capabilities such that if a function requires foo-stability and bar-stability you don't accidentally call it with code that only guarantees foo-stability under a mistaken impression that foo-stability was all the unsafty of the function was related to.


Would you be able to provide a real example of that HellPortal thing? I'm not really following


I'm assuming "real example" means of such unsafe-means-actually-unsafe behaviour in embedded Rust, as opposed to a real example of summoning demons?

For example volatile_register is a crate for representing some sort of MMIO hardware registers. It will do the actual MMIO for you, just tell it where your registers are in "memory" and say whether they're read-write, read-only, or write-only just once, and it provides the nice Rust interface to the registers.

https://docs.rs/volatile-register/0.2.1/volatile_register/st...

The low-level stuff it's doing is inherently unsafe, but it is wrapping that. So when you call register.read() that's safe, and it will... read the register. However even though it's a wrapper it chooses to label the register.write() call as unsafe, reminding you that this is a hardware register and that's on you.

In many cases you'd add a further wrapper, e.g. maybe there's a register for controlling clock frequency of another part, you know the part malfunctions below 5kHz and is not warrantied above 60kHz, so, your wrapper can take a value, check it's between 5 and 60 inclusive and then do the arithmetic and set the frequency register using that unsafe register.write() function. You would probably decide that your wrapper is now actually safe.


So, I peeked into the documentation of volatile_register. The unsafe is there for a clear reason: the compiler can't verify that you aren't creating a mapping to some memory location that is used by the program itself. If you are allowed to safely create a mapping to, for example, a local variable and safely modify it using `register.write()`, you have UB, and that's all in safe code!

So, this isn't a case that the `unsafe` is there just as a warning lint, it has an actual meaning, protecting the memory safety invariants.


> I'm assuming "real example" means of such unsafe-means-actually-unsafe behaviour in embedded Rust, as opposed to a real example of summoning demons?

That was what I meant, thanks for the answer! Though if you have an example of the other thing I'd be open to that too


Similar to the sibling, stuff where you're dealing with parallel state in hardware, like talking to a device over i2c or something, where you know certain things are supposed to happen but you don't, like, know know.


I'm not convinced that it's a great use of Rust's `unsafe`, but since you want an example... dealing with voltage regulators maybe? Where an invalid value put into some register could fry your hardware? There's a ton of such cases in embedded.


You could probably make the case that any function that might physically destroy your memory is memory unsafe :)


Imagine the code screaming

"Rip and tear, until it's done!"


"From Rust's point of view, "HellPortal::unchecked_open()" is safe, it's thread-safe, it's memory safe... but it will summon demons that could destroy mankind if the warding field isn't up, so, that might need an "unsafe" marker and we can write a checked version which verifies that the warding is up and the stand-by wizard is available to close the portal before we actually open it, the checked one will be safe."

"The only thing they fear is you"

playing in the background.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: