Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

3.4 · OsString, CStr, CString: Non-UTF-8 String Types

Domain 3 — Strings and Text Processing Duration: ~15 minutes Library components: std::ffi::OsStr, std::ffi::OsString, std::ffi::CStr, std::ffi::CString, std::ffi::NulError, std::ffi::FromBytesWithNulError

Introduction

String and &str guarantee valid UTF-8. Two important environments do not give this guarantee:

  • The operating system: On Unix, a filename is an arbitrary byte sequence (not necessarily UTF-8). On Windows, a filename is WTF-16. OsStr and OsString represent these platform-native strings.
  • C libraries: A C string is a null-terminated char * buffer with no encoding guarantee. CStr and CString model this buffer for safe FFI.

The two pairs have the same ownership pattern as &str and String. The borrowed type (OsStr, CStr) is a slice that you use through a reference. The owned type (OsString, CString) holds the allocation.

This tutorial describes:

  • OsStr and OsString for platform-native strings.
  • How to convert between OsStr, Path, and str.
  • CString::new and its check for NUL bytes.
  • CStr::from_bytes_with_nul and from_bytes_until_nul.
  • as_ptr, which gives the pointer that you pass to a foreign function.

OsStr and OsString

OsStr is a borrowed platform-native string. OsString is the owned version. Their internal representation is opaque and platform-specific. Do not write code that depends on it.

use std::ffi::{OsStr, OsString};

// OsStr::new borrows the &str. It does not copy the text.
let os_str: &OsStr = OsStr::new("hello.txt");

// Convert to &str when the content is valid UTF-8 (None if it is not).
let s: Option<&str> = os_str.to_str();  // Some("hello.txt")

// The lossy conversion always succeeds.
let lossy = os_str.to_string_lossy();   // Cow<str>: "hello.txt"

to_string_lossy returns a Cow<str>. If the bytes are valid UTF-8, it returns Cow::Borrowed and does not allocate. If the bytes are not valid UTF-8, it returns Cow::Owned. In that owned string, U+FFFD replaces each invalid sequence.

An OsString can grow:

// OsString::from copies the text into a new allocation.
let mut owned: OsString = OsString::from("my_file");
owned.push(".rs");   // appends to the end: owned is now "my_file.rs"

// into_string consumes the OsString and returns Result<String, OsString>.
owned.into_string()  // Ok("my_file.rs"): the content is valid UTF-8

OsStr and Path

Path and PathBuf are thin wrappers around OsStr and OsString. You can convert between them without restrictions:

use std::path::{Path, PathBuf};

// The parts of a path are &OsStr values, not &str values.
Path::new("/usr/local/bin").file_name()  // Some(OsStr::new("bin"))
Path::new("archive.tar.gz").extension()  // Some(OsStr::new("gz")): only the last extension

// PathBuf is the owned type, as OsString is the owned type for OsStr.
let mut pb = PathBuf::from("/home/user");
pb.push("documents");    // adds one path component
pb.push("report.pdf");   // pb is now "/home/user/documents/report.pdf"
pb.as_os_str()           // &OsStr: the OS string of the full path

In an API that receives filenames, accept &OsStr (or AsRef<OsStr>) and not &str. Then the API does not lose data on platforms that have non-UTF-8 filenames.

03_14_osstring_osstr.rs prints the lines below. The last two lines are from its describe_file function, which reads a filename with to_str.

OsStr: "hello.txt"
to_str: Some("hello.txt")
to_string_lossy: hello.txt

OsString after push: "my_file.rs"
OsString len: 10
into_string: Ok("my_file.rs")

Path::file_name: Some("bin")
extension: Some("gz")

PathBuf: /home/user/documents/report.pdf
as_os_str: "/home/user/documents/report.pdf"

main.rs is a Rust source file
data.csv is a regular file

All assertions passed.

CString and CStr

C functions typically accept const char *, which is a pointer to a null-terminated sequence of bytes. CString allocates such a buffer, and CStr borrows one.

CString::new

use std::ffi::CString;

// CString::new copies the text and appends the NUL terminator.
let cstr = CString::new("hello from Rust").unwrap();
// Internally: b"hello from Rust\0"

// ptr: *const c_char (c_char is i8 or u8, as the platform defines).
// You can pass ptr to a C function. It is valid only while `cstr` is alive.
let ptr = cstr.as_ptr();

CString::new returns the error NulError if the input contains an interior NUL byte. Such a byte would terminate the C string too early:

// The NUL byte is at index 3, after the 3 bytes of "bad".
let bad = CString::new("bad\0string"); // Err(NulError): its nul_position() returns 3

Recovering bytes

// Each method consumes the CString and returns its bytes as a Vec<u8>.
CString::new("rust").unwrap().into_bytes()           // [114, 117, 115, 116]: no NUL
CString::new("rust").unwrap().into_bytes_with_nul()  // [114, 117, 115, 116, 0]

Borrowing C Strings

CStr::from_bytes_with_nul

CStr::from_bytes_with_nul borrows a &[u8] as a &CStr. The slice must end with exactly one NUL. If the slice has an interior NUL or has no NUL at its end, the function returns FromBytesWithNulError:

// CStr is in std::ffi, as CString is.
let bytes: &[u8] = b"example\0";   // 8 bytes: the last byte is the NUL
// The &CStr borrows `bytes`. The function does not copy them.
let cstr: &CStr = CStr::from_bytes_with_nul(bytes).unwrap();
cstr.to_str().unwrap() // "example" (to_str returns Err if the bytes are not UTF-8)

CStr::from_bytes_until_nul (stabilized 1.69)

CStr::from_bytes_until_nul finds the first NUL byte and uses it as the terminator. Use this function for a C buffer of constant size that may contain unwanted bytes after the NUL:

// The buffer has two NUL bytes. Only the first NUL is the terminator.
let buf = b"hello\0ignored\0";
let found = CStr::from_bytes_until_nul(buf).unwrap();   // a &CStr that contains "hello"
found.to_str().unwrap() // "hello": the bytes after the first NUL are not included

CStr comparison

CStr implements PartialOrd and Ord, so you can compare C strings by byte value:

let a = CString::new("alpha").unwrap();
let b = CString::new("beta").unwrap();
// as_c_str borrows a CString as a &CStr.
a.as_c_str() < b.as_c_str() // true: the first bytes differ, and b'a' < b'b'

03_15_cstr_cstring.rs prints:

CString: "hello from Rust"
bytes_with_nul last: 0x00
as_ptr: 0x10542cd20   # (varies)

CString with interior NUL: true
NUL at position: 3

into_bytes: [114, 117, 115, 116]

CStr::from_bytes_with_nul: "example"
to_str: "example"

from_bytes_until_nul: "hello"

"alpha" < "beta": true

All assertions passed.

When to Use Each Type

TypeUse when
&str / StringThe text is always valid UTF-8 (most Rust code)
&OsStr / OsStringThe text is a filesystem path, an environment variable, or other platform-native text
&CStr / CStringYou pass strings to C functions, or receive strings from C functions, through FFI

Figure: Choosing the Right String Type

Do not use String for filenames if portability is important to you. Always use Path or PathBuf (their storage is OsStr and OsString). Convert to str only at the point where you need a str. For that conversion, use to_str() (which returns an Option) or to_string_lossy() (which always succeeds).

Summary

TypeBorrowed/OwnedEncodingInterior NUL allowed
&strBorrowedUTF-8 guaranteedYes
StringOwnedUTF-8 guaranteedYes
&OsStrBorrowedPlatform-nativeYes
OsStringOwnedPlatform-nativeYes
&CStrBorrowedNull-terminated, unspecifiedNo (ends at first NUL)
CStringOwnedNull-terminated, unspecifiedNo (CString::new rejects it)

Code Examples

FileDescription
03_14_osstring_osstr.rsOsStr/OsString, to_str, to_string_lossy, and their relation to Path/PathBuf
03_15_cstr_cstring.rsCString::new, NulError, CStr::from_bytes_with_nul, from_bytes_until_nul, as_ptr