3.1 · String, &str, and the UTF-8 Contract
Domain 3 — Strings and Text Processing Duration: ~15 minutes Library components:
std::string::String, primitivestr,std::str::from_utf8,std::str::Utf8Error,FromUtf8Error
Introduction
String and &str are not aliases of each other. They are two different things. String is a heap-allocated, growable buffer of UTF-8 bytes. &str is a borrowed slice of UTF-8 bytes. Those bytes can be anywhere: on the heap, on the stack, or in static memory.
Each design decision in their API comes from two guarantees. The contents are always valid UTF-8. The length that len() returns is always a byte count.
This tutorial shows:
- The internal layout of
Stringand&str. - Why an index of type
usizeis a compile error. - How to iterate over chars, bytes, and (byte offset, char) pairs.
- How to slice safely, and which slice operations panic.
substr_range: how to get the byte offsets of a substring without a search.- How to convert between byte slices and
String/&str, with UTF-16LE and UTF-16BE included.
Internal Representation
String is a thin wrapper around Vec<u8> that keeps the UTF-8 invariant. It has three fields: a heap pointer, a byte length, and a byte capacity.
&str is a fat pointer: a heap (or static) address together with a byte length. For a string literal such as "hello", the compiler stores the UTF-8 bytes in the binary. The literal is a &'static str that points into that read-only segment.
Figure: String vs &str Memory Layout
let owned: String = String::from("héllo");
println!("{}", owned.len()); // 6 bytes: 'h'=1, 'é'=2, 'l'=1, 'l'=1, 'o'=1
println!("{}", owned.chars().count()); // 5 chars
let literal: &str = "world"; // a &'static str: the bytes are in the binary
println!("{}", literal.len()); // 5 bytes and 5 chars (all ASCII)
len() always counts bytes, not characters. This is the most common cause of confusion for new Rust programmers.
Why indexing by usize is not permitted
The expression owned[0] is a compile error. If the compiler accepted it, the expression would return one u8. But one u8 from the middle of a multi-byte sequence is not a valid Unicode scalar value. There is no safe way to give it to the caller. Thus str does not implement Index<usize>, and the type system rejects the expression.
03_01_string_str_representation.rs prints:
owned.len() = 6 (bytes)
owned.chars().count() = 5 (chars)
with_capacity(16): len=0 cap=16
after push_str: len=5 cap=16
slice &owned[0..1] = "h"
literal &str: "world", len=5
Indexing by usize is intentionally a compile error.
Use .chars(), .bytes(), .char_indices(), or explicit byte slices.
Raw pointer to first byte: 0x100e48d20 # (varies)
All assertions passed.
chars(), bytes(), and char_indices()
Three iterators supply all the access patterns:
| Iterator | Yields | Use when |
|---|---|---|
chars() | char (Unicode scalar value) | You count characters or process code points. |
bytes() | u8 | You process raw bytes or ASCII-only data. |
char_indices() | (usize, char): byte offset and char | You slice the original string afterward. |
let text = "café"; // 'é' is U+00E9: one char, two bytes in UTF-8
// chars: 4 Unicode scalar values
for ch in text.chars() {
println!(" {ch:?}"); // 'c', 'a', 'f', 'é'
}
// bytes: 5 raw bytes (195 and 169 are the two bytes of 'é')
let byte_vec: Vec<u8> = text.bytes().collect(); // [99, 97, 102, 195, 169]
// char_indices: the byte offset where each char starts, and the char
for (byte_pos, ch) in text.char_indices() {
println!(" byte {byte_pos}: {ch:?}");
}
// (0, 'c'), (1, 'a'), (2, 'f'), (3, 'é')
char_indices is the iterator to use when you want to slice the string afterward. The iterator guarantees that each byte offset is a valid char boundary.
Figure: UTF-8 Multi-Byte Encoding — "café"
03_02_chars_bytes_char_indices.rs prints:
chars():
'c'
'a'
'f'
'é'
count = 4
bytes():
[99, 97, 102, 195, 169]
char_indices():
byte 0: 'c'
byte 1: 'a'
byte 2: 'f'
byte 3: 'é'
suffix from 3rd char: "fé"
"rust" reversed: "tsur"
All assertions passed.
String Slicing and Byte Boundaries
A slice of a str with &s[start..end] is fast: it is only a pointer offset and a length adjustment. But start and end must each be on a char boundary, or the operation panics at run time.
let s = "héllo wörld"; // 'h' is byte 0, 'é' is bytes 1 and 2, the first 'l' is byte 3
let safe = &s[3..5]; // "ll": the two bounds are char boundaries
// let bad = &s[1..2]; // PANIC: byte 2 is inside 'é' (a 2-byte char)
is_char_boundary
// `s` is "héllo wörld" from the previous snippet.
// is_char_boundary(i) returns true when a char starts at byte offset i.
s.is_char_boundary(0); // true: 'h' starts here
s.is_char_boundary(1); // true: 'é' starts here
s.is_char_boundary(2); // false: the second byte of 'é'
s.is_char_boundary(3); // true: 'l' starts here
Non-panicking slicing: get()
str::get(range) returns Option<&str>. When the range is not valid, it returns None and does not panic:
// `s` is "héllo wörld" from the previous snippets.
let maybe = s.get(3..5); // Some("ll")
let bad = s.get(1..2); // None: byte 2 is not a char boundary
substr_range (1.98) is the inverse of slicing. You give it a &str that already points into source. It returns the byte offsets as a std::range::Range. The method uses pointer arithmetic, not a search. A separate "hello" with the same text gives None. Code that splits text uses it to get the spans for diagnostics:
let source = "fn main() { hello }";
let hello = &source[12..17]; // "hello": a slice that points into `source`
// substr_range compares addresses. It does not search `source` for the text.
assert_eq!(source.substr_range(hello), Some(std::range::Range { start: 12, end: 17 }));
Safe N-char prefix using char_indices
// Returns the first `n` chars of `s`. The slice never ends inside a char.
fn first_n_chars(s: &str, n: usize) -> &str {
// nth(n) gives the byte offset where the char with index n starts.
match s.char_indices().nth(n) {
Some((byte_pos, _)) => &s[..byte_pos], // the slice stops before that char
None => s, // `s` has n chars or fewer: return all of it
}
}
// first_n_chars("héllo wörld", 3) returns "hél" (4 bytes)
03_03_string_slicing.rs prints:
safe slice [3..5]: "ll"
is_char_boundary checks:
byte 0: is_char_boundary = true
byte 1: is_char_boundary = true
byte 2: is_char_boundary = false
byte 3: is_char_boundary = true
byte 4: is_char_boundary = true
first 3 chars: "hél"
get(3..5): Some("ll")
get(1..2): None
All assertions passed.
UTF-8 Conversion API
The standard library has a full set of functions that convert between byte sequences (&[u8], Vec<u8>) and string types:
| Function | Input | Output | Allocates? |
|---|---|---|---|
str::as_bytes() | &str | &[u8] | No |
String::into_bytes() | String | Vec<u8> | No (moves) |
str::from_utf8(&[u8]) | &[u8] | Result<&str, Utf8Error> | No |
String::from_utf8(Vec<u8>) | Vec<u8> | Result<String, FromUtf8Error> | No (moves) |
String::from_utf8_lossy(&[u8]) | &[u8] | Cow<str> | Only for invalid input |
String::from_utf8_lossy_owned(Vec<u8>) (1.99) | Vec<u8> | String | Only for invalid input |
FromUtf8Error::into_utf8_lossy() (1.99) | FromUtf8Error | String | Yes |
String::from_utf16le / from_utf16be (1.98) | &[u8] | Result<String, FromUtf16Error> | Yes |
from_utf16le_lossy / from_utf16be_lossy (1.98) | &[u8] | String | Always |
// as_bytes: borrows the UTF-8 bytes of a &str (no copy, no allocation)
let bytes: &[u8] = "hello".as_bytes();
// from_utf8: validates the bytes and borrows them as a &str (no allocation)
match std::str::from_utf8(bytes) {
Ok(s) => { /* s: &str, the bytes are valid UTF-8 */ }
Err(e) => { /* e: Utf8Error, e.valid_up_to() is the length of the valid prefix */ }
}
// from_utf8_lossy: always succeeds. It replaces each invalid sequence with U+FFFD.
let lossy = String::from_utf8_lossy(b"hel\xFF\xFElo"); // 0xFF and 0xFE are not valid UTF-8
// lossy == "hel\u{FFFD}\u{FFFD}lo"
The two strict conversions have one important difference. str::from_utf8 returns a &str that borrows from the input slice (no allocation). String::from_utf8 consumes a Vec<u8> and wraps it (no copy). When String::from_utf8 fails, it returns a FromUtf8Error that contains the original bytes. Call .into_bytes() on the error to recover them.
03_04_utf8_conversion.rs prints the lines below. The terminal prints each U+FFFD as the character �:
as_bytes: [104, 101, 108, 108, 111]
into_bytes: [114, 117, 115, 116]
from_utf8 valid: "héllo"
from_utf8 invalid: error at byte 0
String::from_utf8 valid: Ok("hello")
String::from_utf8 invalid: true
recovered bytes from FromUtf8Error: [255]
from_utf8_lossy: "hel��lo"
from_utf8_lossy (all valid): "clean"
All assertions passed.
Rust 1.99 adds the owned lossy conversions. String::from_utf8_lossy borrows a &[u8] and returns Cow<str>. When you own the bytes as a Vec<u8>, String::from_utf8_lossy_owned returns a String directly. If the bytes are valid UTF-8, the String reuses the buffer of the Vec, and the function copies no bytes. FromUtf8Error::into_utf8_lossy gives the same repair after a strict conversion fails:
// Valid input: the buffer of the Vec becomes the buffer of the String (no copy).
let text = String::from_utf8_lossy_owned(b"status: ok".to_vec());
assert_eq!(text, "status: ok");
// Invalid input: each bad sequence becomes U+FFFD, as with from_utf8_lossy.
let repaired = String::from_utf8_lossy_owned(b"status: \xFFok".to_vec());
assert_eq!(repaired, "status: \u{FFFD}ok");
// Strict conversion first, lossy conversion as the alternative. The error keeps the bytes.
let wire: Vec<u8> = b"id=\xF0\x90\x80;end".to_vec(); // a 4-byte char without its last byte
let recovered = match String::from_utf8(wire) {
Ok(text) => text,
Err(error) => error.into_utf8_lossy(), // consumes the error, returns the repaired text
};
assert_eq!(recovered, "id=\u{FFFD};end");
UTF-16 data in a network protocol is a byte stream with an endianness. Since 1.98, String::from_utf16le(&[u8]) and from_utf16be decode it directly. An odd length or a lone surrogate gives Err. The _lossy variants replace invalid sequences with U+FFFD and always allocate a String. Unlike from_utf8_lossy, they have no borrowed fast path.
Summary
| Concept | Key point |
|---|---|
String | Heap-allocated Vec<u8> with the UTF-8 invariant |
&str | Fat pointer (address and byte length) into valid UTF-8 |
len() | Always the byte count, never the char count |
Indexing by usize | Compile error. Use iterators or slices. |
chars() | Unicode scalar values |
bytes() | Raw UTF-8 bytes |
char_indices() | (byte_offset, char) pairs, safe for subsequent slicing |
get(range) | Slice that does not panic. It returns Option<&str>. |
substr_range (1.98) | Byte Range of a substring, from pointer arithmetic, not a search |
from_utf8 | Validates bytes. It returns Result<&str, Utf8Error> (no allocation). |
from_utf8_lossy | Always succeeds. It returns Cow<str>. |
from_utf8_lossy_owned (1.99) | Converts an owned Vec<u8> to a String. Valid input reuses the buffer. |
from_utf16le / from_utf16be (1.98) | Decodes UTF-16 bytes (&[u8]) of a known endianness into a String |
Code Examples
| File | Description |
|---|---|
03_01_string_str_representation.rs | String layout, len() compared with chars().count(), with_capacity, raw pointer |
03_02_chars_bytes_char_indices.rs | chars(), bytes(), char_indices(), suffix slicing, rev() |
03_03_string_slicing.rs | Byte-boundary slicing, is_char_boundary, get(), safe N-char prefix |
03_04_utf8_conversion.rs | as_bytes, into_bytes, from_utf8, String::from_utf8, from_utf8_lossy |
03_19_substr_range.rs | str::substr_range (1.98): get the offsets of split tokens without a search |
03_21_from_utf16_endian.rs | from_utf16le/from_utf16be and their lossy variants (1.98) |
03_23_from_utf8_lossy_owned.rs | String::from_utf8_lossy_owned and FromUtf8Error::into_utf8_lossy (1.99): lossy conversion of an owned Vec<u8>, buffer reuse for valid input |