Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

3.5 · ASCII Operations and Character Classification

Domain 3 — Strings and Text Processing Duration: ~15 minutes Library components: std::ascii, std::char, primitive char methods, char::from_digit, char::to_digit

Introduction

The char type of Rust represents a Unicode scalar value. Its classification methods apply the full Unicode standard. For example, is_alphabetic() returns true for 'é' and '中', not only for ASCII letters. Use these methods when you need behaviour that is correct for Unicode.

Some domains are strictly ASCII: network protocols, identifiers, numeric parsing. For these domains, the is_ascii_* family does the same checks, but only in the ASCII range. The to_ascii_* and make_ascii_* methods convert case fast and do not change non-ASCII bytes. The make_ascii_* methods convert in place.

This tutorial shows:

  • Unicode-aware char classification: is_alphabetic, is_numeric, is_alphanumeric, is_whitespace, is_uppercase, is_lowercase.
  • ASCII-specific classification: is_ascii, is_ascii_alphabetic, is_ascii_digit, is_ascii_punctuation, is_ascii_graphic, is_ascii_control.
  • char::from_digit and to_digit, which convert digits in a given base.
  • Unicode case conversion: to_lowercase, to_uppercase (these methods return iterators).
  • ASCII-only case conversion: make_ascii_lowercase, make_ascii_uppercase, to_ascii_lowercase, to_ascii_uppercase.
  • Escape iterators: EscapeDefault, EscapeDebug, EscapeUnicode.

Unicode Character Classification

The classification methods in this section apply the Unicode standard.

// Each method returns a bool.
'é'.is_alphabetic()    // true: alphabetic in Unicode, but not ASCII
'中'.is_alphabetic()   // true: a CJK ideograph
'1'.is_numeric()       // true
' '.is_whitespace()    // true
'\n'.is_whitespace()   // true
'A'.is_uppercase()     // true
'a'.is_lowercase()     // true
MethodReturns true for
is_alphabetic()Each Unicode alphabetic character
is_numeric()Each Unicode numeric character (fractions and superscripts included)
is_alphanumeric()Each character that is alphabetic or numeric
is_whitespace()Unicode whitespace (space, tab, newline, and others)
is_uppercase()Uppercase Unicode characters
is_lowercase()Lowercase Unicode characters
is_control()Control characters (C0, C1)

Since Rust 1.97, char::is_control is a const fn, so the classification of control characters can run at compile time. For example, you can check a protocol delimiter constant while the program builds:

const DELIM: char = '\x1F'; // ASCII unit separator
// The assert! runs in const evaluation. If DELIM is not a control character,
// the build fails. The program does not start.
const _: () = assert!(DELIM.is_control(), "delimiter must be a control character");

The first part of the output of 03_16_char_classification.rs:

char   alphabetic   numeric    alphanumeric   whitespace   uppercase    lowercase
--------------------------------------------------------------------------------
'A'    true         false      true           false        true         false
'a'    true         false      true           false        false        true
'1'    false        true       true           false        false        false
' '    false        false      false          true         false        false
'\n'   false        false      false          true         false        false
'!'    false        false      false          false        false        false
'é'    true         false      true           false        false        true
'中'    true         false      true           false        false        false
'\t'   false        false      false          true         false        false

const is_control ('\u{1f}'): true (checked at compile time)

ASCII Classification

The is_ascii_* methods examine only the 128 ASCII code points. They return false for each non-ASCII character, even if the character satisfies the Unicode criterion:

Figure: Unicode vs ASCII Classification Scope

'é'.is_ascii_alphabetic()  // false: 'é' is alphabetic, but it is not ASCII
'A'.is_ascii_alphabetic()  // true
'5'.is_ascii_digit()       // true
'!'.is_ascii_punctuation() // true
'A'.is_ascii_graphic()     // true: a graphic character is printable and is not the space
' '.is_ascii_whitespace()  // true
'\x01'.is_ascii_control()  // true: the control characters are 0x00–0x1F and 0x7F

is_ascii() tells you if the code point of the character is in 0..=127. Each other is_ascii_* method tests a subset of that range.

char::from_digit and to_digit

These two functions convert between the numeric value of a digit and its character in a given base:

// char::from_digit(value, radix) returns Option<char>.
char::from_digit(10, 16)  // Some('a'): the hex digit for 10
char::from_digit(5, 10)   // Some('5'): the decimal digit for 5
char::from_digit(16, 16)  // None: 16 is not a digit in base 16

// to_digit(radix) returns Option<u32>.
'f'.to_digit(16)          // Some(15)
'9'.to_digit(10)          // Some(9)
'g'.to_digit(16)          // None: 'g' is not a hex digit

The two functions accept a radix from 2 to 36 (inclusive). The digits 10–35 correspond to the letters 'a'–'z'.

The second part of the output of 03_16_char_classification.rs:

is_ascii:
'A'.is_ascii(): true, 'é'.is_ascii(): false

ASCII classification:
  'A': alpha=true digit=false punct=false graphic=true control=false
  'z': alpha=true digit=false punct=false graphic=true control=false
  '5': alpha=false digit=true punct=false graphic=true control=false
  ' ': alpha=false digit=false punct=false graphic=false control=false
  '!': alpha=false digit=false punct=true graphic=true control=false
  '\u{1}': alpha=false digit=false punct=false graphic=false control=true
  '_': alpha=false digit=false punct=true graphic=true control=false
  '+': alpha=false digit=false punct=true graphic=true control=false

from_digit / to_digit:
from_digit(10, 16) = Some('a')
'f'.to_digit(16) = Some(15)

All assertions passed.

Unicode Case Conversion

char::to_lowercase and char::to_uppercase return iterators, not one char. The reason is that case conversion changes some Unicode characters into more than one character. The usual example is the German 'ß' (U+00DF): its uppercase form is "SS".

// Collect the iterator into a String.
let lower: String = 'A'.to_lowercase().collect();   // "a"
let upper: String = 'a'.to_uppercase().collect();   // "A"
let ss: String    = 'ß'.to_uppercase().collect();   // "SS": one char becomes two

// str::to_lowercase and str::to_uppercase allocate a new String.
"Héllo, Wörld!".to_lowercase()  // "héllo, wörld!"
"Héllo, Wörld!".to_uppercase()  // "HÉLLO, WÖRLD!"

To make a word title case, convert the first character with to_uppercase. Then append the remainder of the word:

let word = "hello";
let mut chars = word.chars();
// chars.next() removes the first char from the iterator.
// chars.as_str() is the remainder of the word as a &str ("ello").
let titled: String = match chars.next() {
    Some(c) => c.to_uppercase().collect::<String>() + chars.as_str(),
    None    => String::new(),   // the word is empty
};
// titled is "Hello"
// The example binary applies this code to each word of "hello world".

ASCII-Only Case Conversion

Sometimes you know that your data is ASCII, or you want non-ASCII bytes to stay as they are. In these cases, use the make_ascii_* and to_ascii_* variants. They operate on u8 values, do not use Unicode tables, and can convert in place:

// The make_ascii_* methods change the value in place.
// They are methods of [u8] and str, so a Vec<u8> and a String have them too.
let mut bytes = b"Hello, World!".to_vec();   // Vec<u8>
bytes.make_ascii_lowercase();  // b"hello, world!"
bytes.make_ascii_uppercase();  // b"HELLO, WORLD!"

let mut s = String::from("Hello, World!");
s.make_ascii_lowercase();      // "hello, world!"

// The to_ascii_* methods allocate and return a new String.
"RUST".to_ascii_lowercase()    // "rust"
"rust".to_ascii_uppercase()    // "RUST"

// The ASCII methods do not change non-ASCII characters.
"héllo".to_ascii_lowercase()   // "héllo" ('é' does not change)

03_17_case_conversion.rs prints:

'A'.to_lowercase(): "a"
'a'.to_uppercase(): "A"
'ß'.to_uppercase(): "SS"

"Héllo, Wörld!"
  to_lowercase: "héllo, wörld!"
  to_uppercase: "HÉLLO, WÖRLD!"

title case: ["Hello", "World"]

make_ascii_lowercase: "hello, world!"
make_ascii_uppercase: "HELLO, WORLD!"
str make_ascii_lowercase: "hello, world!"
to_ascii_lowercase: "rust"
to_ascii_uppercase: "RUST"
"héllo".to_ascii_lowercase(): "héllo" (é unchanged)

All assertions passed.

Escape Iterators

Three methods of char return iterators that produce escape sequences as text. The iterators yield char values and do not allocate:

MethodBehaviour
escape_default()Rust literal style: \n, \t, \\, \", and \u{XXXX} for characters that are not printable or not ASCII
escape_debug()As escape_default(), but it does not change printable non-ASCII characters
escape_unicode()Always \u{XXXX} for each character, ASCII included
// The results are written as Rust string literals: "\\n" is the 2 characters \ and n.
'\n'.escape_default().collect::<String>()  // "\\n"
'"'.escape_default().collect::<String>()   // "\\\""     (the 2 characters \ and ")
'é'.escape_default().collect::<String>()   // "\\u{e9}"  (not ASCII: a Unicode escape)
'é'.escape_debug().collect::<String>()     // "é"        (printable: no change)
'a'.escape_unicode().collect::<String>()   // "\\u{61}"  (ASCII is escaped too)
'中'.escape_unicode().collect::<String>()  // "\\u{4e2d}"

str also has escape_default(), escape_debug(), and escape_unicode(). Each one escapes the full string and returns an iterator that yields char values. Collect the iterator into a String, or write it directly to a fmt::Write sink.

The iterators are lazy. Thus you can count the length of the escaped text with no allocation:

// The escaped text is \u{1f600}: 9 characters.
'\u{1F600}'.escape_unicode().count()  // 9

03_18_escape_iterators.rs prints the lines below. The lines for one character show the escaped string in {:?} format, so each backslash of the result appears two times. The two str results use {} format.

escape_default per char:
  '\n' → "\\n"
  '\t' → "\\t"
  '\\' → "\\\\"
  '"' → "\\\""
  '\'' → "\\'"
  'a' → "a"
  'é' → "\\u{e9}"
  '中' → "\\u{4e2d}"

escape_debug per char:
  '\n' → "\\n"
  '\t' → "\\t"
  '\\' → "\\\\"
  '"' → "\\\""
  '\'' → "\\'"
  'a' → "a"
  'é' → "é"
  '中' → "中"

escape_unicode per char:
  'a' → "\\u{61}"
  'A' → "\\u{41}"
  'é' → "\\u{e9}"
  '中' → "\\u{4e2d}"
  '\n' → "\\u{a}"

str escape_default:
  tab:\there\nnewline and \"quotes\"
str escape_debug:
  tab:\there\nnewline and \"quotes\"

escape_unicode char count for U+1F600: 9

All assertions passed.

Summary

TopicKey APIs
Unicode classificationis_alphabetic, is_numeric, is_whitespace, is_uppercase, is_lowercase
ASCII classificationis_ascii, is_ascii_alphabetic, is_ascii_digit, is_ascii_graphic, is_ascii_control
Const classificationchar::is_control is a const fn since 1.97: you can classify at compile time
Digit conversionchar::from_digit(n, radix), char::to_digit(radix)
Unicode caseto_lowercase() / to_uppercase() return iterators (the result may be more than one character)
ASCII caseto_ascii_lowercase/uppercase, make_ascii_lowercase/uppercase (in place)
Escape iteratorsescape_default, escape_debug, escape_unicode return lazy char iterators

Code Examples

FileDescription
03_16_char_classification.rsUnicode and ASCII classification, from_digit, to_digit, const is_control (1.97)
03_17_case_conversion.rsto_lowercase/to_uppercase (Unicode), make_ascii_*/to_ascii_*
03_18_escape_iterators.rsEscapeDefault, EscapeDebug, EscapeUnicode on char and str