3.4 · OsString, CStr, CString: Non-UTF-8 String Types
Domain 3 — Strings and Text Processing Duration: ~15 minutes Library components:
std::ffi::OsStr,std::ffi::OsString,std::ffi::CStr,std::ffi::CString,std::ffi::NulError,std::ffi::FromBytesWithNulError
Introduction
String and &str guarantee valid UTF-8. Two important environments do not give this guarantee:
- The operating system: On Unix, a filename is an arbitrary byte sequence (not necessarily UTF-8). On Windows, a filename is WTF-16.
OsStrandOsStringrepresent these platform-native strings. - C libraries: A C string is a null-terminated
char *buffer with no encoding guarantee.CStrandCStringmodel this buffer for safe FFI.
The two pairs have the same ownership pattern as &str and String. The borrowed type (OsStr, CStr) is a slice that you use through a reference. The owned type (OsString, CString) holds the allocation.
This tutorial describes:
OsStrandOsStringfor platform-native strings.- How to convert between
OsStr,Path, andstr. CString::newand its check for NUL bytes.CStr::from_bytes_with_nulandfrom_bytes_until_nul.as_ptr, which gives the pointer that you pass to a foreign function.
OsStr and OsString
OsStr is a borrowed platform-native string. OsString is the owned version. Their internal representation is opaque and platform-specific. Do not write code that depends on it.
use std::ffi::{OsStr, OsString};
// OsStr::new borrows the &str. It does not copy the text.
let os_str: &OsStr = OsStr::new("hello.txt");
// Convert to &str when the content is valid UTF-8 (None if it is not).
let s: Option<&str> = os_str.to_str(); // Some("hello.txt")
// The lossy conversion always succeeds.
let lossy = os_str.to_string_lossy(); // Cow<str>: "hello.txt"
to_string_lossy returns a Cow<str>. If the bytes are valid UTF-8, it returns Cow::Borrowed and does not allocate. If the bytes are not valid UTF-8, it returns Cow::Owned. In that owned string, U+FFFD replaces each invalid sequence.
An OsString can grow:
// OsString::from copies the text into a new allocation.
let mut owned: OsString = OsString::from("my_file");
owned.push(".rs"); // appends to the end: owned is now "my_file.rs"
// into_string consumes the OsString and returns Result<String, OsString>.
owned.into_string() // Ok("my_file.rs"): the content is valid UTF-8
OsStr and Path
Path and PathBuf are thin wrappers around OsStr and OsString. You can convert between them without restrictions:
use std::path::{Path, PathBuf};
// The parts of a path are &OsStr values, not &str values.
Path::new("/usr/local/bin").file_name() // Some(OsStr::new("bin"))
Path::new("archive.tar.gz").extension() // Some(OsStr::new("gz")): only the last extension
// PathBuf is the owned type, as OsString is the owned type for OsStr.
let mut pb = PathBuf::from("/home/user");
pb.push("documents"); // adds one path component
pb.push("report.pdf"); // pb is now "/home/user/documents/report.pdf"
pb.as_os_str() // &OsStr: the OS string of the full path
In an API that receives filenames, accept &OsStr (or AsRef<OsStr>) and not &str. Then the API does not lose data on platforms that have non-UTF-8 filenames.
03_14_osstring_osstr.rs prints the lines below. The last two lines are from its describe_file function, which reads a filename with to_str.
OsStr: "hello.txt"
to_str: Some("hello.txt")
to_string_lossy: hello.txt
OsString after push: "my_file.rs"
OsString len: 10
into_string: Ok("my_file.rs")
Path::file_name: Some("bin")
extension: Some("gz")
PathBuf: /home/user/documents/report.pdf
as_os_str: "/home/user/documents/report.pdf"
main.rs is a Rust source file
data.csv is a regular file
All assertions passed.
CString and CStr
C functions typically accept const char *, which is a pointer to a null-terminated sequence of bytes. CString allocates such a buffer, and CStr borrows one.
CString::new
use std::ffi::CString;
// CString::new copies the text and appends the NUL terminator.
let cstr = CString::new("hello from Rust").unwrap();
// Internally: b"hello from Rust\0"
// ptr: *const c_char (c_char is i8 or u8, as the platform defines).
// You can pass ptr to a C function. It is valid only while `cstr` is alive.
let ptr = cstr.as_ptr();
CString::new returns the error NulError if the input contains an interior NUL byte. Such a byte would terminate the C string too early:
// The NUL byte is at index 3, after the 3 bytes of "bad".
let bad = CString::new("bad\0string"); // Err(NulError): its nul_position() returns 3
Recovering bytes
// Each method consumes the CString and returns its bytes as a Vec<u8>.
CString::new("rust").unwrap().into_bytes() // [114, 117, 115, 116]: no NUL
CString::new("rust").unwrap().into_bytes_with_nul() // [114, 117, 115, 116, 0]
Borrowing C Strings
CStr::from_bytes_with_nul
CStr::from_bytes_with_nul borrows a &[u8] as a &CStr. The slice must end with exactly one NUL. If the slice has an interior NUL or has no NUL at its end, the function returns FromBytesWithNulError:
// CStr is in std::ffi, as CString is.
let bytes: &[u8] = b"example\0"; // 8 bytes: the last byte is the NUL
// The &CStr borrows `bytes`. The function does not copy them.
let cstr: &CStr = CStr::from_bytes_with_nul(bytes).unwrap();
cstr.to_str().unwrap() // "example" (to_str returns Err if the bytes are not UTF-8)
CStr::from_bytes_until_nul (stabilized 1.69)
CStr::from_bytes_until_nul finds the first NUL byte and uses it as the terminator. Use this function for a C buffer of constant size that may contain unwanted bytes after the NUL:
// The buffer has two NUL bytes. Only the first NUL is the terminator.
let buf = b"hello\0ignored\0";
let found = CStr::from_bytes_until_nul(buf).unwrap(); // a &CStr that contains "hello"
found.to_str().unwrap() // "hello": the bytes after the first NUL are not included
CStr comparison
CStr implements PartialOrd and Ord, so you can compare C strings by byte value:
let a = CString::new("alpha").unwrap();
let b = CString::new("beta").unwrap();
// as_c_str borrows a CString as a &CStr.
a.as_c_str() < b.as_c_str() // true: the first bytes differ, and b'a' < b'b'
03_15_cstr_cstring.rs prints:
CString: "hello from Rust"
bytes_with_nul last: 0x00
as_ptr: 0x10542cd20 # (varies)
CString with interior NUL: true
NUL at position: 3
into_bytes: [114, 117, 115, 116]
CStr::from_bytes_with_nul: "example"
to_str: "example"
from_bytes_until_nul: "hello"
"alpha" < "beta": true
All assertions passed.
When to Use Each Type
| Type | Use when |
|---|---|
&str / String | The text is always valid UTF-8 (most Rust code) |
&OsStr / OsString | The text is a filesystem path, an environment variable, or other platform-native text |
&CStr / CString | You pass strings to C functions, or receive strings from C functions, through FFI |
Figure: Choosing the Right String Type
Do not use String for filenames if portability is important to you. Always use Path or PathBuf (their storage is OsStr and OsString). Convert to str only at the point where you need a str. For that conversion, use to_str() (which returns an Option) or to_string_lossy() (which always succeeds).
Summary
| Type | Borrowed/Owned | Encoding | Interior NUL allowed |
|---|---|---|---|
&str | Borrowed | UTF-8 guaranteed | Yes |
String | Owned | UTF-8 guaranteed | Yes |
&OsStr | Borrowed | Platform-native | Yes |
OsString | Owned | Platform-native | Yes |
&CStr | Borrowed | Null-terminated, unspecified | No (ends at first NUL) |
CString | Owned | Null-terminated, unspecified | No (CString::new rejects it) |
Code Examples
| File | Description |
|---|---|
03_14_osstring_osstr.rs | OsStr/OsString, to_str, to_string_lossy, and their relation to Path/PathBuf |
03_15_cstr_cstring.rs | CString::new, NulError, CStr::from_bytes_with_nul, from_bytes_until_nul, as_ptr |