Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

3.1 · String, &str, and the UTF-8 Contract

Domain 3 — Strings and Text Processing Duration: ~15 minutes Library components: std::string::String, primitive str, std::str::from_utf8, std::str::Utf8Error, FromUtf8Error

Introduction

String and &str are not aliases of each other. They are two different things. String is a heap-allocated, growable buffer of UTF-8 bytes. &str is a borrowed slice of UTF-8 bytes. Those bytes can be anywhere: on the heap, on the stack, or in static memory.

Each design decision in their API comes from two guarantees. The contents are always valid UTF-8. The length that len() returns is always a byte count.

This tutorial shows:

  • The internal layout of String and &str.
  • Why an index of type usize is a compile error.
  • How to iterate over chars, bytes, and (byte offset, char) pairs.
  • How to slice safely, and which slice operations panic.
  • substr_range: how to get the byte offsets of a substring without a search.
  • How to convert between byte slices and String/&str, with UTF-16LE and UTF-16BE included.

Internal Representation

String is a thin wrapper around Vec<u8> that keeps the UTF-8 invariant. It has three fields: a heap pointer, a byte length, and a byte capacity.

&str is a fat pointer: a heap (or static) address together with a byte length. For a string literal such as "hello", the compiler stores the UTF-8 bytes in the binary. The literal is a &'static str that points into that read-only segment.

Figure: String vs &str Memory Layout

let owned: String = String::from("héllo");
println!("{}", owned.len());           // 6 bytes: 'h'=1, 'é'=2, 'l'=1, 'l'=1, 'o'=1
println!("{}", owned.chars().count()); // 5 chars

let literal: &str = "world";           // a &'static str: the bytes are in the binary
println!("{}", literal.len());         // 5 bytes and 5 chars (all ASCII)

len() always counts bytes, not characters. This is the most common cause of confusion for new Rust programmers.

Why indexing by usize is not permitted

The expression owned[0] is a compile error. If the compiler accepted it, the expression would return one u8. But one u8 from the middle of a multi-byte sequence is not a valid Unicode scalar value. There is no safe way to give it to the caller. Thus str does not implement Index<usize>, and the type system rejects the expression.

03_01_string_str_representation.rs prints:

owned.len()  = 6 (bytes)
owned.chars().count() = 5 (chars)

with_capacity(16): len=0 cap=16
after push_str:    len=5 cap=16

slice &owned[0..1] = "h"
literal &str: "world", len=5

Indexing by usize is intentionally a compile error.
Use .chars(), .bytes(), .char_indices(), or explicit byte slices.

Raw pointer to first byte: 0x100e48d20   # (varies)

All assertions passed.

chars(), bytes(), and char_indices()

Three iterators supply all the access patterns:

IteratorYieldsUse when
chars()char (Unicode scalar value)You count characters or process code points.
bytes()u8You process raw bytes or ASCII-only data.
char_indices()(usize, char): byte offset and charYou slice the original string afterward.
let text = "café";   // 'é' is U+00E9: one char, two bytes in UTF-8

// chars: 4 Unicode scalar values
for ch in text.chars() {
    println!("  {ch:?}");   // 'c', 'a', 'f', 'é'
}

// bytes: 5 raw bytes (195 and 169 are the two bytes of 'é')
let byte_vec: Vec<u8> = text.bytes().collect(); // [99, 97, 102, 195, 169]

// char_indices: the byte offset where each char starts, and the char
for (byte_pos, ch) in text.char_indices() {
    println!("  byte {byte_pos}: {ch:?}");
}
// (0, 'c'), (1, 'a'), (2, 'f'), (3, 'é')

char_indices is the iterator to use when you want to slice the string afterward. The iterator guarantees that each byte offset is a valid char boundary.

Figure: UTF-8 Multi-Byte Encoding — "café"

03_02_chars_bytes_char_indices.rs prints:

chars():
  'c'
  'a'
  'f'
  'é'
  count = 4

bytes():
  [99, 97, 102, 195, 169]

char_indices():
  byte 0: 'c'
  byte 1: 'a'
  byte 2: 'f'
  byte 3: 'é'

suffix from 3rd char: "fé"

"rust" reversed: "tsur"

All assertions passed.

String Slicing and Byte Boundaries

A slice of a str with &s[start..end] is fast: it is only a pointer offset and a length adjustment. But start and end must each be on a char boundary, or the operation panics at run time.

let s = "héllo wörld";   // 'h' is byte 0, 'é' is bytes 1 and 2, the first 'l' is byte 3

let safe = &s[3..5]; // "ll": the two bounds are char boundaries
// let bad = &s[1..2]; // PANIC: byte 2 is inside 'é' (a 2-byte char)

is_char_boundary

// `s` is "héllo wörld" from the previous snippet.
// is_char_boundary(i) returns true when a char starts at byte offset i.
s.is_char_boundary(0); // true:  'h' starts here
s.is_char_boundary(1); // true:  'é' starts here
s.is_char_boundary(2); // false: the second byte of 'é'
s.is_char_boundary(3); // true:  'l' starts here

Non-panicking slicing: get()

str::get(range) returns Option<&str>. When the range is not valid, it returns None and does not panic:

// `s` is "héllo wörld" from the previous snippets.
let maybe = s.get(3..5); // Some("ll")
let bad   = s.get(1..2); // None: byte 2 is not a char boundary

substr_range (1.98) is the inverse of slicing. You give it a &str that already points into source. It returns the byte offsets as a std::range::Range. The method uses pointer arithmetic, not a search. A separate "hello" with the same text gives None. Code that splits text uses it to get the spans for diagnostics:

let source = "fn main() { hello }";
let hello = &source[12..17];   // "hello": a slice that points into `source`
// substr_range compares addresses. It does not search `source` for the text.
assert_eq!(source.substr_range(hello), Some(std::range::Range { start: 12, end: 17 }));

Safe N-char prefix using char_indices

// Returns the first `n` chars of `s`. The slice never ends inside a char.
fn first_n_chars(s: &str, n: usize) -> &str {
    // nth(n) gives the byte offset where the char with index n starts.
    match s.char_indices().nth(n) {
        Some((byte_pos, _)) => &s[..byte_pos], // the slice stops before that char
        None => s,                             // `s` has n chars or fewer: return all of it
    }
}
// first_n_chars("héllo wörld", 3) returns "hél" (4 bytes)

03_03_string_slicing.rs prints:

safe slice [3..5]: "ll"

is_char_boundary checks:
  byte 0: is_char_boundary = true
  byte 1: is_char_boundary = true
  byte 2: is_char_boundary = false
  byte 3: is_char_boundary = true
  byte 4: is_char_boundary = true

first 3 chars: "hél"
get(3..5): Some("ll")
get(1..2): None

All assertions passed.

UTF-8 Conversion API

The standard library has a full set of functions that convert between byte sequences (&[u8], Vec<u8>) and string types:

FunctionInputOutputAllocates?
str::as_bytes()&str&[u8]No
String::into_bytes()StringVec<u8>No (moves)
str::from_utf8(&[u8])&[u8]Result<&str, Utf8Error>No
String::from_utf8(Vec<u8>)Vec<u8>Result<String, FromUtf8Error>No (moves)
String::from_utf8_lossy(&[u8])&[u8]Cow<str>Only for invalid input
String::from_utf8_lossy_owned(Vec<u8>) (1.99)Vec<u8>StringOnly for invalid input
FromUtf8Error::into_utf8_lossy() (1.99)FromUtf8ErrorStringYes
String::from_utf16le / from_utf16be (1.98)&[u8]Result<String, FromUtf16Error>Yes
from_utf16le_lossy / from_utf16be_lossy (1.98)&[u8]StringAlways
// as_bytes: borrows the UTF-8 bytes of a &str (no copy, no allocation)
let bytes: &[u8] = "hello".as_bytes();

// from_utf8: validates the bytes and borrows them as a &str (no allocation)
match std::str::from_utf8(bytes) {
    Ok(s)  => { /* s: &str, the bytes are valid UTF-8 */ }
    Err(e) => { /* e: Utf8Error, e.valid_up_to() is the length of the valid prefix */ }
}

// from_utf8_lossy: always succeeds. It replaces each invalid sequence with U+FFFD.
let lossy = String::from_utf8_lossy(b"hel\xFF\xFElo");   // 0xFF and 0xFE are not valid UTF-8
// lossy == "hel\u{FFFD}\u{FFFD}lo"

The two strict conversions have one important difference. str::from_utf8 returns a &str that borrows from the input slice (no allocation). String::from_utf8 consumes a Vec<u8> and wraps it (no copy). When String::from_utf8 fails, it returns a FromUtf8Error that contains the original bytes. Call .into_bytes() on the error to recover them.

03_04_utf8_conversion.rs prints the lines below. The terminal prints each U+FFFD as the character �:

as_bytes: [104, 101, 108, 108, 111]
into_bytes: [114, 117, 115, 116]

from_utf8 valid: "héllo"
from_utf8 invalid: error at byte 0

String::from_utf8 valid: Ok("hello")
String::from_utf8 invalid: true
recovered bytes from FromUtf8Error: [255]

from_utf8_lossy: "hel��lo"
from_utf8_lossy (all valid): "clean"

All assertions passed.

Rust 1.99 adds the owned lossy conversions. String::from_utf8_lossy borrows a &[u8] and returns Cow<str>. When you own the bytes as a Vec<u8>, String::from_utf8_lossy_owned returns a String directly. If the bytes are valid UTF-8, the String reuses the buffer of the Vec, and the function copies no bytes. FromUtf8Error::into_utf8_lossy gives the same repair after a strict conversion fails:

// Valid input: the buffer of the Vec becomes the buffer of the String (no copy).
let text = String::from_utf8_lossy_owned(b"status: ok".to_vec());
assert_eq!(text, "status: ok");

// Invalid input: each bad sequence becomes U+FFFD, as with from_utf8_lossy.
let repaired = String::from_utf8_lossy_owned(b"status: \xFFok".to_vec());
assert_eq!(repaired, "status: \u{FFFD}ok");

// Strict conversion first, lossy conversion as the alternative. The error keeps the bytes.
let wire: Vec<u8> = b"id=\xF0\x90\x80;end".to_vec();   // a 4-byte char without its last byte
let recovered = match String::from_utf8(wire) {
    Ok(text) => text,
    Err(error) => error.into_utf8_lossy(),   // consumes the error, returns the repaired text
};
assert_eq!(recovered, "id=\u{FFFD};end");

UTF-16 data in a network protocol is a byte stream with an endianness. Since 1.98, String::from_utf16le(&[u8]) and from_utf16be decode it directly. An odd length or a lone surrogate gives Err. The _lossy variants replace invalid sequences with U+FFFD and always allocate a String. Unlike from_utf8_lossy, they have no borrowed fast path.

Summary

ConceptKey point
StringHeap-allocated Vec<u8> with the UTF-8 invariant
&strFat pointer (address and byte length) into valid UTF-8
len()Always the byte count, never the char count
Indexing by usizeCompile error. Use iterators or slices.
chars()Unicode scalar values
bytes()Raw UTF-8 bytes
char_indices()(byte_offset, char) pairs, safe for subsequent slicing
get(range)Slice that does not panic. It returns Option<&str>.
substr_range (1.98)Byte Range of a substring, from pointer arithmetic, not a search
from_utf8Validates bytes. It returns Result<&str, Utf8Error> (no allocation).
from_utf8_lossyAlways succeeds. It returns Cow<str>.
from_utf8_lossy_owned (1.99)Converts an owned Vec<u8> to a String. Valid input reuses the buffer.
from_utf16le / from_utf16be (1.98)Decodes UTF-16 bytes (&[u8]) of a known endianness into a String

Code Examples

FileDescription
03_01_string_str_representation.rsString layout, len() compared with chars().count(), with_capacity, raw pointer
03_02_chars_bytes_char_indices.rschars(), bytes(), char_indices(), suffix slicing, rev()
03_03_string_slicing.rsByte-boundary slicing, is_char_boundary, get(), safe N-char prefix
03_04_utf8_conversion.rsas_bytes, into_bytes, from_utf8, String::from_utf8, from_utf8_lossy
03_19_substr_range.rsstr::substr_range (1.98): get the offsets of split tokens without a search
03_21_from_utf16_endian.rsfrom_utf16le/from_utf16be and their lossy variants (1.98)
03_23_from_utf8_lossy_owned.rsString::from_utf8_lossy_owned and FromUtf8Error::into_utf8_lossy (1.99): lossy conversion of an owned Vec<u8>, buffer reuse for valid input