Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

8.1 · Read, Write, BufRead: The I/O Trait Hierarchy

Domain 8 — I/O System Duration: ~15 minutes Library components: std::io::Read, std::io::Write, std::io::BufRead, std::io::BufReader, std::io::BufWriter, std::io::LineWriter, std::io::copy

Introduction

The Rust I/O system has three basic traits:

  • Read gets bytes from a source.
  • Write sends bytes to a destination.
  • BufRead adds line-oriented and delimiter-oriented methods to a buffered Read.

Files, sockets, standard streams, in-memory buffers, and compression wrappers all use these three traits. Thus a function that takes impl Read operates on all of them without changes.

This tutorial shows:

  • The trait hierarchy and the types that implement each trait.
  • Read: partial reads, read_to_end, read_to_string, read_exact, and why Ok(0) means EOF.
  • Write: write and write_all, flush, and the difference between io::Write and fmt::Write.
  • BufRead and BufReader: read_line, lines, split, and the fill_buf/consume primitives.
  • BufWriter and LineWriter: when the bytes that you write arrive at their destination.
  • How to compose readers with chain and take, and how to stream with io::copy.

The Trait Hierarchy

There are three small traits and one supertrait relationship: BufRead: Read. All other items in std::io are implementors or adapters of these traits.

Figure: The core I/O traits and their main implementors

(In the figure, ByteSlice is the &[u8] primitive and ByteVec is Vec<u8>. A dotted arrow means "implements".)

Two design decisions are important:

  1. Each operation can fail. Each operation returns io::Result (tutorial 8.4), because a disk can become full, a pipe can break, and a connection can reset. The lines() iterator also gives io::Result<String> items.
  2. Byte counts are contracts. read and write return the number of bytes that they moved. That number can be less than the number that you requested. The *_all and *_exact variants do the loop for you.

The write! macro operates with two different traits. io::Write moves bytes, and fmt::Write moves string slices. The write! and writeln! macros use the two traits through separate trait implementations. Tutorial 3.2 gives the details of fmt::Write.

Read: Pulling Bytes

Read has one required method: fn read(&mut self, buf: &mut [u8]) -> io::Result<usize>. This method defines the full contract:

  • It fills a maximum of buf.len() bytes.
  • It returns the number of bytes that it filled.
  • It returns Ok(0) for a non-empty buffer to signal EOF.

A short read is not an error and not EOF. A caller that obeys the contract calls read in a loop.

// These lines are in a function that returns io::Result<()>, so `?` is valid.
let mut source: &[u8] = b"HELLO WORLD";  // &[u8] implements Read

let mut chunk = [0u8; 4];                // a 4-byte buffer on the stack
let n = source.read(&mut chunk)?;        // n can be less than 4 (here n == 4)
assert_eq!(&chunk[..n], b"HELL");

let n = source.read(&mut chunk)?;        // the first read moved the slice forward
assert_eq!(&chunk[..n], b"O WO");        // the second read gives the next 4 bytes

The provided methods contain the loops that you would otherwise write manually:

  • read_to_end(&mut Vec<u8>) reads all bytes until EOF.
  • read_to_string(&mut String) does the same and also validates UTF-8. Bytes that are not valid UTF-8 cause an InvalidData error.
  • read_exact(&mut [u8]) fills the whole buffer or fails with UnexpectedEof. It is the usual method for fixed-size fields in binary formats (tutorial 8.2).
// A 5-byte source: one big-endian u32, then one more byte.
let mut source: &[u8] = &[0x12, 0x34, 0x56, 0x78, 0x9A];
let mut word = [0u8; 4];
source.read_exact(&mut word)?;                         // fills all 4 bytes
assert_eq!(u32::from_be_bytes(word), 0x1234_5678);

let err = source.read_exact(&mut word).unwrap_err();   // only 1 byte left
assert_eq!(err.kind(), std::io::ErrorKind::UnexpectedEof);

All these methods also retry internally after an ErrorKind::Interrupted error. Tutorial 8.4 tells you more about that special error.

Write: Pushing Bytes

Write is the opposite of Read and has the same type of contract. write may accept only a prefix of your data, and it returns the number of bytes that it accepted. write_all calls write in a loop until it writes all the data. Application code almost always needs write_all. If you ignore a short write, the result is silent data corruption.

let mut frame: Vec<u8> = Vec::new();     // Vec<u8> implements Write: each write appends
frame.write_all(b"LEN=11;")?;
frame.write_all(b"hello world")?;        // frame is now b"LEN=11;hello world"

// A &mut [u8] writer has a FIXED capacity. It has two failure modes:
let mut fixed = [0u8; 4];
let mut slot: &mut [u8] = &mut fixed;
let accepted = slot.write(b"toolong")?;  // Ok(4): partial write, NO error
assert_eq!(accepted, 4);                 // fixed is now b"tool"

let mut slot: &mut [u8] = &mut fixed;    // a new writer on the same 4 bytes
slot.write_all(b"cool")?;                // fits exactly, the writer is now full
let err = slot.write_all(b"x").unwrap_err();
assert_eq!(err.kind(), std::io::ErrorKind::WriteZero);   // write accepted 0 bytes

flush sends all buffered bytes to their destination. On a bare Vec, flush does nothing. Call it anyway. Then your code still operates as intended if you later replace the Vec with a BufWriter<File> or a socket. A missing flush causes problems with buffered writers (see the next slides).

BufRead and BufReader

Without a buffer, each read call on a File is a system call. If you read a File one line at a time, you call the OS for each small group of bytes. BufReader wraps any Read and keeps an internal buffer (8 KiB by default). It implements BufRead, which adds the text-oriented methods that most programs use.

Figure: How BufReader turns many small reads into few large ones

The most important methods are:

  • read_line(&mut String) appends one line, including the trailing \n. It returns 0 at EOF. Call clear() on the string between calls.
  • lines() returns an iterator of io::Result<String> items without the newline. It is the most common method to read text.
  • split(byte) is the equivalent of lines() for any delimiter byte. It gives Vec<u8> chunks, each in an io::Result.
  • read_until and skip_until do one step at a time. read_until keeps the bytes, and skip_until discards them.
  • fill_buf() and consume(n) are the two primitives that all the methods above use. fill_buf() shows you the buffered bytes. consume(n) tells the reader how many bytes you used.

In memory, &[u8] already implements BufRead, so tests do not need a wrapper. Be careful with method resolution on a bare &[u8]: the inherent split method of the slice shadows BufRead::split. Put the bytes in a Cursor (tutorial 8.2) to get the BufRead version.

// path: PathBuf, the path of a text file that exists.
// File implements Read but not BufRead. BufReader<File> implements BufRead.
for line in BufReader::new(File::open(&path)?).lines() {
    let line = line?;                    // an error can occur in the middle of the iteration
    // ... use `line`: a String without the trailing newline
}

BufWriter and LineWriter

BufWriter is the equivalent wrapper for output. It collects small writes in memory and sends them to the destination in large blocks. The example uses a counting sink as a replacement for a file, which makes the effect visible. 100 writes arrive at the destination as one call:

// CountingSink is a Write type that the example defines. It appends the size
// of each write call that it receives to its `write_calls: Vec<usize>` field.
let mut writer = BufWriter::with_capacity(4096, CountingSink::default());
for i in 0..100u32 {
    write!(writer, "{i:03},")?;                     // 4 bytes: "000,", "001,", ...
}
// get_ref() borrows the wrapped sink. All 400 bytes are still in the buffer.
assert_eq!(writer.get_ref().write_calls.len(), 0);
writer.flush()?;
assert_eq!(writer.get_ref().write_calls.len(), 1);  // one write call of 400 bytes

Remember these three behaviors:

  • The flush on drop hides errors. The Drop implementation of BufWriter tries to flush, but it cannot report a failure. End with an explicit flush() or into_inner(), so that you get the errors. into_inner() flushes and returns the wrapped writer.
  • Large writes bypass the buffer. A payload that is larger than the buffer goes directly to the destination as one call. BufWriter does not copy it into the buffer.
  • LineWriter flushes on \n. It buffers as BufWriter does. When a newline arrives, it immediately sends all bytes up to and including the newline. io::stdout() uses the same policy (tutorial 8.5).

08_03_bufwriter_linewriter.rs prints:

unbuffered: 100 writes reached the sink as 100 calls
BufWriter: 100 writes reached the sink as 1 call(s)
BufWriter: 1024-byte write bypassed the 64-byte buffer ([1024])
LineWriter: flushed through the newline, held "trailing" until into_inner

All assertions passed.

Composing Streams: chain, take, io::copy

You can compose readers as you compose iterators:

  • a.chain(b) reads all of a, then all of b. This is the concatenation of two streams.
  • r.take(n) limits a reader to n bytes. It is necessary when a length prefix gives the size of a field. It is also a low-cost protection against input that has no limit. Use r.by_ref().take(n) if you need the underlying reader after the call.
  • io::copy(&mut reader, &mut writer) moves bytes from any Read into any Write and returns the total count. It reuses one internal buffer, so memory use stays constant for a stream of any size.
// header: &[u8] = b"PREVIEW:" (8 bytes). body: a Vec<u8> of 1000 '#' bytes.
const PREVIEW_LIMIT: u64 = 64;
let mut preview = Vec::new();
let n = io::copy(
    // Read the header, then the body, and stop after 64 bytes in total.
    &mut header.chain(&body[..]).take(PREVIEW_LIMIT),
    &mut preview,
)?;
assert_eq!(n, PREVIEW_LIMIT);            // preview holds "PREVIEW:" and 56 '#' bytes

io::copy has an internal specialization for a source or a sink that is already buffered. It uses that buffer again and does not allocate a second one. As of Rust 1.99, std has no public io::copy_buf function that gives this optimization as a separate API. Use io::copy.

Summary

ConceptKey point
Read::readIt may return fewer bytes than you requested. Ok(0) means EOF.
read_to_end / read_to_stringThey read all bytes until EOF. The string version requires valid UTF-8 (InvalidData).
read_exactIt fills the whole buffer or fails. Short input gives UnexpectedEof.
Write::writeIt may accept only a prefix. It returns the count.
write_allIt loops until it writes all the bytes. It fails with WriteZero if no progress is possible.
flushIt sends buffered bytes to the destination. Call it before you rely on the data.
io::Write vs fmt::WriteBytes and str. The write! macros are the same, but the traits are different (see 3.2).
BufReadread_line, lines, split, which use fill_buf + consume
BufReaderAdds an 8 KiB read-ahead buffer to any Read
BufWriterIt batches small writes. The flush on drop hides errors, so use into_inner.
LineWriterBuffered, but it flushes at each \n (the policy of stdout)
chain / takeConcatenate readers / limit a reader
io::copyIt streams Read → Write in constant memory. Internally, it reuses a buffer that already exists.

Code Examples

FileDescription
08_01_read_write_fundamentals.rsPartial reads, read_to_end/read_to_string/read_exact, write and write_all, flush, and io::Write compared with fmt::Write
08_02_bufread_lines_split.rsread_line, lines, split, read_until/skip_until, fill_buf+consume, and BufReader on a real temporary file
08_03_bufwriter_linewriter.rsA counting sink shows the batching of BufWriter, the buffer bypass, and the newline flush of LineWriter
08_04_chain_take_copy.rschain, take, by_ref, and io::copy in one streaming pipeline with a byte limit