> I see enforcing utf8 as about on par with enforcing that a boolean be 0 or 1, and not 5.
UTF8 correctness is an invariant of the code. Its not something the compiler understands or cares about it.
But really, I get it! My point is that rust is rife with people marking functions "unsafe" when they mean a method will cause bugs / "runtime UB" if used improperly. Even the standard library does this, so it feels a bit rich to cry foul when 3rd party libraries (ab)use "unsafe" in the same way.
I certainly use unsafe like that, in order to mark methods which, if misused, will violate internal invariants of my code.
How does that list of unsafe features justify the unsafe marker on String::from_utf8_unchecked?
Well how do you feel about the comparison to booleans?
If I had a way to set some booleans "unchecked", in a way that could make the numerical value be 5, would you label it unsafe?
I really think it should be unsafe, because it's crazy to make ifs or switches on booleans be non-exhaustive, and it's also crazy to have an "invalid boolean" path anywhere one is used.
When you have an encapsulated value, with special accessors to make sure it's always in range, then you don't have to have an unchecked setter function at all. But if you do add one, I think it makes sense to make it unsafe. That function is basically allowing a blind memcpy over a piece of data. It ties into safety pretty strongly. That kind of access is a lot like dereferencing a raw pointer, and not making it unsafe means you allow a lot of those "invalid value" cases in the link to be possible in safe code.
I get that this was just an example, but a boolean 5 would actually be language-level UB that could have very real correctness implications!
The Rust compiler is allowed to (and sometimes required to) use the invalid values of a child value to represent other cases of an enum. For example, this is how it optimizes `Option<&T>` into "pointer where null is None and any other value is Some".
This would technically also allow it (I don't know whether it currently does this..) to optimize the following enum:
> I get that this was just an example, but a boolean 5 would actually be language-level UB that could have very real correctness implications!
Yes, that's why I used it as the example.
And you don't need anything that complicated, either. A simple if or switch statement might be optimized by the compiler for 0 and 1 and jump into random code for 5.
I agree. I think "unchecked" booleans which could contain 5 should be unsafe too.
I'd be happy to extend the definition of "unsafe" to mean "If you misuse this, you may violate some internal invariants of the library. And that may cause unexpected bugs."
Thats a broader, but in my opinion much more clear definition of unsafe, which covers how the keyword is actually used in actual code. (In std::String and elsewhere).
I've implemented high performance b-tree and skiplist implementations in rust. There's plenty of functions in both libraries which I've marked as unsafe because if you use them carelessly, you'll violate some internal invariants. Will the library break as a result? Yes. Will the resulting bugs include memory corruption? I really have no idea, and I don't really care enough to go in and test that. So I've marked the methods unsafe, and provided safe APIs for consumers to use instead.
Do you think this is an appropriate use for unsafe? If not, I'm curious to hear why.
I think it depends on the kind of invariant. If it leads to undefined behavior that wouldn't already be possible, then mark it unsafe or change the unsafe code to eliminate it. If it just makes it act badly, like dropping a node, that's likely not a good place for unsafe.
Or more simply: Worry about all kinds of undefined behavior. That makes many issues easier to find, and you don't need to chain on additional logic to figure out how it might corrupt memory.
I would argue it’s language-ub to convert an invalid slice into into a String, it’s bubbling up a requirement of str which relies on the contained utf-8 to be well-formed for memory safety.
the key thing here is that if you know that no safe code contains invalid utf8, it is safe to do unsafe indexing into tables that are only sized to deal with valid utf8. as such, this works in practice lead to ub in otherwise correct code, so it should be marked unsafe.
UTF8 correctness is an invariant of the code. Its not something the compiler understands or cares about it.
But really, I get it! My point is that rust is rife with people marking functions "unsafe" when they mean a method will cause bugs / "runtime UB" if used improperly. Even the standard library does this, so it feels a bit rich to cry foul when 3rd party libraries (ab)use "unsafe" in the same way.
I certainly use unsafe like that, in order to mark methods which, if misused, will violate internal invariants of my code.
How does that list of unsafe features justify the unsafe marker on String::from_utf8_unchecked?