3.3 · Pattern Matching on Strings and FromStr Parsing
Domain 3 — Strings and Text Processing Duration: ~15 minutes Library components:
std::str::pattern::Pattern,std::str::FromStr,strmethods
Introduction
The string search methods of Rust accept a pattern, not one fixed char type or string type. A pattern can be:
- a
char - a
&str - a
&[char]slice - a
|c: char| -> boolclosure
You can pass each of these to contains, find, split, and related methods. The call site does not change.
For parsing, FromStr is the standard trait to parse a value of any type from a string. When you implement it, you get str::parse::<T>() at no cost. It also integrates with the ? operator.
This tutorial explains:
- How to use
char,&str,&[char], and closures as patterns. contains,starts_with,ends_with,find,rfind.split,splitn,split_once,rsplit,rsplitn,rsplit_once,split_whitespace.trim,trim_start,trim_end,trim_matches,strip_prefix,strip_suffix,strip_circumfix.- How to implement
FromStrand usestr::parse::<T>().
The Pattern Abstraction
The Pattern trait (in std::str::pattern) unifies the types that you can search for in a string. The trait is currently unstable as a public trait, but all the APIs that use it are stable.
Figure: Pattern Types — What Can You Search For?
let url = "https://example.com/path?key=value";
// char as pattern
url.contains('?') // true
url.find('/') // Some(6): the byte offset of the first '/'
// &str as pattern
url.starts_with("https://") // true
url.ends_with("value") // true
// &[char] as pattern: a match is any one of the chars in the slice
let seps = &['/', '?', '=']; // seps: &[char; 3]
url.contains(seps.as_ref()) // true: as_ref() gives the &[char] slice
// Closure as pattern: the method calls the predicate for each char
url.contains(|c: char| c.is_ascii_digit()) // false: the URL has no digit
"error 404: not found".contains(|c: char| c.is_ascii_digit()) // true
03_10_string_patterns.rs prints the lines below. The find and rfind lines search the sentence "the cat sat on the mat":
contains '?': true
starts_with "https://": true
ends_with "value": true
contains any separator: true
contains a digit: false
"error 404: not found" contains digit: true
find "cat": Some(4)
find ' ': Some(3)
find vowel: Some(2)
rfind ' ': Some(18)
rfind "at": Some(20)
query string: "key=value"
All assertions passed.
split and Its Variants
split returns a lazy iterator. It does not allocate until you call collect.
// Basic split
let names: Vec<&str> = "alice,bob,carol,dave".split(',').collect();
// ["alice", "bob", "carol", "dave"]
// A delimiter at an end, or two adjacent delimiters, give an empty item
",a,,b,".split(',').collect::<Vec<_>>()
// ["", "a", "", "b", ""]
// split_terminator: a delimiter at the end gives no empty item
"a,b,c,".split_terminator(',').collect::<Vec<_>>()
// ["a", "b", "c"]
// splitn: a maximum of n parts. The last part can contain the delimiter.
"key=value=extra".splitn(2, '=').collect::<Vec<_>>()
// ["key", "value=extra"]
// split_once: the idiomatic parse of key=value text. It splits at the first match.
"Content-Type: application/json".split_once(": ")
// Some(("Content-Type", "application/json"))
// rsplit / rsplitn: split from the right
"std.fmt.Display".rsplit('.').collect::<Vec<_>>()
// ["Display", "fmt", "std"]
// rsplit_once: splits at the last match
"archive.tar.gz".rsplit_once('.')
// Some(("archive.tar", "gz"))
// split_whitespace: splits on any whitespace and skips empty tokens
" one two\tthree\n".split_whitespace().collect::<Vec<_>>()
// ["one", "two", "three"]
03_11_split_and_search.rs prints:
split: ["alice", "bob", "carol", "dave"]
split on "::": ["home", "user", "documents", "file.txt"]
split with empties: ["", "a", "", "b", ""]
split_terminator: ["a", "b", "c"]
splitN(2): ["key", "value=extra"]
split_once key: "Content-Type"
split_once val: "application/json"
rsplit: ["Display", "fmt", "std"]
rsplitn(2): ["Display", "std.fmt"]
rsplit_once stem: "archive.tar"
rsplit_once ext: "gz"
split_whitespace: ["one", "two", "three"]
All assertions passed.
trim and strip
trim removes Unicode whitespace from the two ends. trim_matches does the same for any pattern. The strip_* methods return an Option, for safe removal of a prefix or a suffix.
// trim: removes whitespace from the two ends
" hello world ".trim() // "hello world"
" left".trim_start() // "left"
"right ".trim_end() // "right"
// trim_matches: removes every repeat of a char from the two ends
"\"value\"".trim_matches('"') // "value"
"###hello###".trim_matches('#') // "hello"
// trim_start_matches / trim_end_matches: one end only, with a pattern
"///path/to/file".trim_start_matches('/') // "path/to/file": the inner '/' chars stay
"3.14000".trim_end_matches('0') // "3.14"
// Closure: removes the chars that are not alphanumeric from the two ends
"...hello...".trim_matches(|c: char| !c.is_alphanumeric()) // "hello"
// strip_prefix: removes the prefix and returns an Option
"https://example.com".strip_prefix("https://") // Some("example.com")
"https://example.com".strip_prefix("ftp://") // None: the prefix is absent
// strip_suffix: removes the suffix and returns an Option
"report.pdf".strip_suffix(".pdf") // Some("report")
// Chain: remove the angle brackets at the two ends
"<value>".strip_prefix('<').and_then(|s| s.strip_suffix('>'))
// Some("value")
// strip_circumfix (1.98): the two ends in one call. None unless the prefix *and* the suffix match.
"<value>".strip_circumfix('<', '>') // Some("value")
"bar:hello:foo".strip_circumfix("bar:", ":foo") // Some("hello")
"hello".strip_circumfix('<', '>') // None
03_12_trim_strip.rs prints:
trim: "hello world"
trim_start: "left"
trim_end: "right"
trim_matches '"': "value"
trim_matches '#': "hello"
trim_start_matches '/': "path/to/file"
trim_end_matches '0': "3.14"
trim_matches closure: "hello"
strip_prefix "https://": Some("example.com")
strip_prefix "ftp://": None
strip_suffix ".pdf": Some("report")
strip prefix + suffix: Some("value")
All assertions passed.
FromStr and str::parse
FromStr is the standard interface to parse a type from its string representation. An implementation of FromStr is the idiomatic alternative to standalone parse_X(s: &str) functions.
use std::str::FromStr;
#[derive(Debug, PartialEq)]
struct Color { r: u8, g: u8, b: u8 }
// The error type of the parse. It holds a message.
#[derive(Debug)]
struct ParseColorError(String);
// Parses "r,g,b" text such as "255,127,0" into a Color.
impl FromStr for Color {
type Err = ParseColorError;
fn from_str(s: &str) -> Result<Self, Self::Err> {
let parts: Vec<&str> = s.splitn(3, ',').collect();
if parts.len() != 3 {
return Err(ParseColorError(format!("expected 3 components, got {s:?}")));
}
// Parses one component. map_err converts the ParseIntError into a ParseColorError.
let parse_u8 = |p: &str| {
p.trim() // permits spaces around the number
.parse::<u8>() // fails if the text is not an integer from 0 to 255
.map_err(|_| ParseColorError(format!("{p:?} is not 0-255")))
};
Ok(Color {
r: parse_u8(parts[0])?, // `?` returns the error of the first bad component
g: parse_u8(parts[1])?,
b: parse_u8(parts[2])?,
})
}
}
The parse() method on str calls FromStr::from_str. It is a short form that accepts turbofish syntax:
// Built-in types: the type annotation or the turbofish selects the FromStr implementation
let n: i32 = "42".parse().unwrap(); // 42
let f: f64 = "1.75".parse().unwrap(); // 1.75
let port = "8080".parse::<u16>().unwrap(); // 8080: the turbofish gives the type u16
// Custom type: parse() calls the from_str of Color
let color: Color = "255,127,0".parse().unwrap(); // Color { r: 255, g: 127, b: 0 }
The associated type Err is the error that from_str returns when the parse fails. To use it with ? in a function that returns Box<dyn Error>, the error type must implement std::error::Error. That trait requires Debug and Display.
03_13_fromstr_parse.rs prints:
"42".parse::<i32>() = 42
"1.75".parse::<f64>() = 1.75
port = 8080
bad parse: true
parsed Color: Color { r: 255, g: 127, b: 0 }
too few: Err(ParseColorError("expected 3 components, got \"255,0\""))
out of range: Err(ParseColorError("\"256\" is not 0-255"))
config key="background" color=Color { r: 30, g: 30, b: 30 }
All assertions passed.
Summary
| Tool | Description |
|---|---|
char / &str / &[char] / closure | Each one is a valid Pattern argument |
contains(pat) | Returns true if the string contains the pattern |
find(pat) / rfind(pat) | Byte offset of the first or the last match of the pattern |
split(pat) | Iterator of substrings. It consumes all the delimiters. |
splitn(n, pat) | A maximum of n parts. The last part may contain delimiters. |
split_once(pat) | Splits at the first match. It returns Option<(&str, &str)>. |
split_whitespace() | Splits on any whitespace and skips empty tokens |
trim / trim_matches | Removes whitespace or a custom pattern from the ends |
strip_prefix / strip_suffix | Safe removal of a prefix or a suffix. It returns Option<&str>. |
strip_circumfix (1.98) | Removes a prefix and a suffix in one call. It returns Option<&str>. |
FromStr::from_str | Parses a &str into a typed value |
str::parse::<T>() | Convenient call site for FromStr |
Code Examples
| File | Description |
|---|---|
03_10_string_patterns.rs | char, &str, &[char], and closures as patterns, with contains, find, rfind |
03_11_split_and_search.rs | split, splitn, split_once, rsplit, rsplitn, split_whitespace |
03_12_trim_strip.rs | trim, trim_matches, trim_start/end_matches, strip_prefix, strip_suffix |
03_20_strip_circumfix.rs | str::strip_circumfix (1.98): remove a prefix and a suffix, or get None |
03_13_fromstr_parse.rs | FromStr implementation, parse::<T>(), error handling |