3.5 · ASCII Operations and Character Classification
Domain 3 — Strings and Text Processing Duration: ~15 minutes Library components:
std::ascii,std::char, primitivecharmethods,char::from_digit,char::to_digit
Introduction
The char type of Rust represents a Unicode scalar value. Its classification methods apply the full Unicode standard. For example, is_alphabetic() returns true for 'é' and '中', not only for ASCII letters. Use these methods when you need behaviour that is correct for Unicode.
Some domains are strictly ASCII: network protocols, identifiers, numeric parsing. For these domains, the is_ascii_* family does the same checks, but only in the ASCII range. The to_ascii_* and make_ascii_* methods convert case fast and do not change non-ASCII bytes. The make_ascii_* methods convert in place.
This tutorial shows:
- Unicode-aware
charclassification:is_alphabetic,is_numeric,is_alphanumeric,is_whitespace,is_uppercase,is_lowercase. - ASCII-specific classification:
is_ascii,is_ascii_alphabetic,is_ascii_digit,is_ascii_punctuation,is_ascii_graphic,is_ascii_control. char::from_digitandto_digit, which convert digits in a given base.- Unicode case conversion:
to_lowercase,to_uppercase(these methods return iterators). - ASCII-only case conversion:
make_ascii_lowercase,make_ascii_uppercase,to_ascii_lowercase,to_ascii_uppercase. - Escape iterators:
EscapeDefault,EscapeDebug,EscapeUnicode.
Unicode Character Classification
The classification methods in this section apply the Unicode standard.
// Each method returns a bool.
'é'.is_alphabetic() // true: alphabetic in Unicode, but not ASCII
'中'.is_alphabetic() // true: a CJK ideograph
'1'.is_numeric() // true
' '.is_whitespace() // true
'\n'.is_whitespace() // true
'A'.is_uppercase() // true
'a'.is_lowercase() // true
| Method | Returns true for |
|---|---|
is_alphabetic() | Each Unicode alphabetic character |
is_numeric() | Each Unicode numeric character (fractions and superscripts included) |
is_alphanumeric() | Each character that is alphabetic or numeric |
is_whitespace() | Unicode whitespace (space, tab, newline, and others) |
is_uppercase() | Uppercase Unicode characters |
is_lowercase() | Lowercase Unicode characters |
is_control() | Control characters (C0, C1) |
Since Rust 1.97, char::is_control is a const fn, so the classification of control characters can run at compile time. For example, you can check a protocol delimiter constant while the program builds:
const DELIM: char = '\x1F'; // ASCII unit separator
// The assert! runs in const evaluation. If DELIM is not a control character,
// the build fails. The program does not start.
const _: () = assert!(DELIM.is_control(), "delimiter must be a control character");
The first part of the output of 03_16_char_classification.rs:
char alphabetic numeric alphanumeric whitespace uppercase lowercase
--------------------------------------------------------------------------------
'A' true false true false true false
'a' true false true false false true
'1' false true true false false false
' ' false false false true false false
'\n' false false false true false false
'!' false false false false false false
'é' true false true false false true
'中' true false true false false false
'\t' false false false true false false
const is_control ('\u{1f}'): true (checked at compile time)
ASCII Classification
The is_ascii_* methods examine only the 128 ASCII code points. They return false for each non-ASCII character, even if the character satisfies the Unicode criterion:
Figure: Unicode vs ASCII Classification Scope
'é'.is_ascii_alphabetic() // false: 'é' is alphabetic, but it is not ASCII
'A'.is_ascii_alphabetic() // true
'5'.is_ascii_digit() // true
'!'.is_ascii_punctuation() // true
'A'.is_ascii_graphic() // true: a graphic character is printable and is not the space
' '.is_ascii_whitespace() // true
'\x01'.is_ascii_control() // true: the control characters are 0x00–0x1F and 0x7F
is_ascii() tells you if the code point of the character is in 0..=127. Each other is_ascii_* method tests a subset of that range.
char::from_digit and to_digit
These two functions convert between the numeric value of a digit and its character in a given base:
// char::from_digit(value, radix) returns Option<char>.
char::from_digit(10, 16) // Some('a'): the hex digit for 10
char::from_digit(5, 10) // Some('5'): the decimal digit for 5
char::from_digit(16, 16) // None: 16 is not a digit in base 16
// to_digit(radix) returns Option<u32>.
'f'.to_digit(16) // Some(15)
'9'.to_digit(10) // Some(9)
'g'.to_digit(16) // None: 'g' is not a hex digit
The two functions accept a radix from 2 to 36 (inclusive). The digits 10–35 correspond to the letters 'a'–'z'.
The second part of the output of 03_16_char_classification.rs:
is_ascii:
'A'.is_ascii(): true, 'é'.is_ascii(): false
ASCII classification:
'A': alpha=true digit=false punct=false graphic=true control=false
'z': alpha=true digit=false punct=false graphic=true control=false
'5': alpha=false digit=true punct=false graphic=true control=false
' ': alpha=false digit=false punct=false graphic=false control=false
'!': alpha=false digit=false punct=true graphic=true control=false
'\u{1}': alpha=false digit=false punct=false graphic=false control=true
'_': alpha=false digit=false punct=true graphic=true control=false
'+': alpha=false digit=false punct=true graphic=true control=false
from_digit / to_digit:
from_digit(10, 16) = Some('a')
'f'.to_digit(16) = Some(15)
All assertions passed.
Unicode Case Conversion
char::to_lowercase and char::to_uppercase return iterators, not one char. The reason is that case conversion changes some Unicode characters into more than one character. The usual example is the German 'ß' (U+00DF): its uppercase form is "SS".
// Collect the iterator into a String.
let lower: String = 'A'.to_lowercase().collect(); // "a"
let upper: String = 'a'.to_uppercase().collect(); // "A"
let ss: String = 'ß'.to_uppercase().collect(); // "SS": one char becomes two
// str::to_lowercase and str::to_uppercase allocate a new String.
"Héllo, Wörld!".to_lowercase() // "héllo, wörld!"
"Héllo, Wörld!".to_uppercase() // "HÉLLO, WÖRLD!"
To make a word title case, convert the first character with to_uppercase. Then append the remainder of the word:
let word = "hello";
let mut chars = word.chars();
// chars.next() removes the first char from the iterator.
// chars.as_str() is the remainder of the word as a &str ("ello").
let titled: String = match chars.next() {
Some(c) => c.to_uppercase().collect::<String>() + chars.as_str(),
None => String::new(), // the word is empty
};
// titled is "Hello"
// The example binary applies this code to each word of "hello world".
ASCII-Only Case Conversion
Sometimes you know that your data is ASCII, or you want non-ASCII bytes to stay as they are. In these cases, use the make_ascii_* and to_ascii_* variants. They operate on u8 values, do not use Unicode tables, and can convert in place:
// The make_ascii_* methods change the value in place.
// They are methods of [u8] and str, so a Vec<u8> and a String have them too.
let mut bytes = b"Hello, World!".to_vec(); // Vec<u8>
bytes.make_ascii_lowercase(); // b"hello, world!"
bytes.make_ascii_uppercase(); // b"HELLO, WORLD!"
let mut s = String::from("Hello, World!");
s.make_ascii_lowercase(); // "hello, world!"
// The to_ascii_* methods allocate and return a new String.
"RUST".to_ascii_lowercase() // "rust"
"rust".to_ascii_uppercase() // "RUST"
// The ASCII methods do not change non-ASCII characters.
"héllo".to_ascii_lowercase() // "héllo" ('é' does not change)
03_17_case_conversion.rs prints:
'A'.to_lowercase(): "a"
'a'.to_uppercase(): "A"
'ß'.to_uppercase(): "SS"
"Héllo, Wörld!"
to_lowercase: "héllo, wörld!"
to_uppercase: "HÉLLO, WÖRLD!"
title case: ["Hello", "World"]
make_ascii_lowercase: "hello, world!"
make_ascii_uppercase: "HELLO, WORLD!"
str make_ascii_lowercase: "hello, world!"
to_ascii_lowercase: "rust"
to_ascii_uppercase: "RUST"
"héllo".to_ascii_lowercase(): "héllo" (é unchanged)
All assertions passed.
Escape Iterators
Three methods of char return iterators that produce escape sequences as text. The iterators yield char values and do not allocate:
| Method | Behaviour |
|---|---|
escape_default() | Rust literal style: \n, \t, \\, \", and \u{XXXX} for characters that are not printable or not ASCII |
escape_debug() | As escape_default(), but it does not change printable non-ASCII characters |
escape_unicode() | Always \u{XXXX} for each character, ASCII included |
// The results are written as Rust string literals: "\\n" is the 2 characters \ and n.
'\n'.escape_default().collect::<String>() // "\\n"
'"'.escape_default().collect::<String>() // "\\\"" (the 2 characters \ and ")
'é'.escape_default().collect::<String>() // "\\u{e9}" (not ASCII: a Unicode escape)
'é'.escape_debug().collect::<String>() // "é" (printable: no change)
'a'.escape_unicode().collect::<String>() // "\\u{61}" (ASCII is escaped too)
'中'.escape_unicode().collect::<String>() // "\\u{4e2d}"
str also has escape_default(), escape_debug(), and escape_unicode(). Each one escapes the full string and returns an iterator that yields char values. Collect the iterator into a String, or write it directly to a fmt::Write sink.
The iterators are lazy. Thus you can count the length of the escaped text with no allocation:
// The escaped text is \u{1f600}: 9 characters.
'\u{1F600}'.escape_unicode().count() // 9
03_18_escape_iterators.rs prints the lines below. The lines for one character show the escaped string in {:?} format, so each backslash of the result appears two times. The two str results use {} format.
escape_default per char:
'\n' → "\\n"
'\t' → "\\t"
'\\' → "\\\\"
'"' → "\\\""
'\'' → "\\'"
'a' → "a"
'é' → "\\u{e9}"
'中' → "\\u{4e2d}"
escape_debug per char:
'\n' → "\\n"
'\t' → "\\t"
'\\' → "\\\\"
'"' → "\\\""
'\'' → "\\'"
'a' → "a"
'é' → "é"
'中' → "中"
escape_unicode per char:
'a' → "\\u{61}"
'A' → "\\u{41}"
'é' → "\\u{e9}"
'中' → "\\u{4e2d}"
'\n' → "\\u{a}"
str escape_default:
tab:\there\nnewline and \"quotes\"
str escape_debug:
tab:\there\nnewline and \"quotes\"
escape_unicode char count for U+1F600: 9
All assertions passed.
Summary
| Topic | Key APIs |
|---|---|
| Unicode classification | is_alphabetic, is_numeric, is_whitespace, is_uppercase, is_lowercase |
| ASCII classification | is_ascii, is_ascii_alphabetic, is_ascii_digit, is_ascii_graphic, is_ascii_control |
| Const classification | char::is_control is a const fn since 1.97: you can classify at compile time |
| Digit conversion | char::from_digit(n, radix), char::to_digit(radix) |
| Unicode case | to_lowercase() / to_uppercase() return iterators (the result may be more than one character) |
| ASCII case | to_ascii_lowercase/uppercase, make_ascii_lowercase/uppercase (in place) |
| Escape iterators | escape_default, escape_debug, escape_unicode return lazy char iterators |
Code Examples
| File | Description |
|---|---|
03_16_char_classification.rs | Unicode and ASCII classification, from_digit, to_digit, const is_control (1.97) |
03_17_case_conversion.rs | to_lowercase/to_uppercase (Unicode), make_ascii_*/to_ascii_* |
03_18_escape_iterators.rs | EscapeDefault, EscapeDebug, EscapeUnicode on char and str |