ESP32 Audio in Rust: A Microphone That Keeps Time on Your LAN
Field note on rusty_esp_audio: PDM microphone, PCM, ADPCM, WAV and FLAC byte-identical to ffmpeg, codec chips as register data, and PCM your laptop plays with ffplay.

Can an ESP32 stream microphone audio to my home network in Rust?
Yes. rusty_esp_audio reads the PDM microphone on a XIAO ESP32-S3 Sense and sends raw 16 kHz PCM over UDP that ffplay plays with no receiver code. A ten-minute soak read 30,000 blocks with none lost, and ten minutes over the network delivered 7,201 datagrams with none lost.
ESP32 audio with rusty_esp_audio means a small board with a microphone that captures, filters and streams sound on your own network in memory-safe Rust, with every codec path checked byte-identical against ffmpeg. On a Seeed XIAO ESP32-S3 Sense the PDM microphone read 30,000 blocks in ten minutes with none short, dropped or errored, and ten minutes of PCM over the board's own Wi-Fi delivered 7,201 datagrams with none lost.
This field note covers the pipeline, the formats, the microphone and codec chips, and streaming to a laptop, with plain labels on what has run on silicon. It belongs to the Hear pillar.
The ESP32 audio pipeline, block by block
Espressif's ESP-ADF builds audio as a graph of elements. The Janus remake keeps the idea and removes the allocation: blocks are fixed size, memory belongs to the caller, and no element allocates.
Elements that never allocate
A Pipeline holds a short list of elements and ping-pongs between two halves of one scratch buffer the caller provides. Each element can change rate, channels or encoding, and says up front how many bytes it may produce, so the scratch is sized once. The shipped elements include Gain, DcBlock, Biquad (low-pass, high-pass, peak, notch), Agc, EnergyVad, channel conversions, a mixer and a LinearResampler with exact rational phase: 480 frames at 48 kHz become exactly 160 at 16 kHz.
ESP32 audio filters here are held to outside references. Biquads land within one least significant bit of ffmpeg and SciPy. The heavy inner loops live in the shared kernel home, rusty_esp_dsp, where the ESP32-S3's vector unit can speed them up without changing a byte of output.
PCM, ADPCM, WAV and FLAC
For ESP32 audio formats, the test is simple: an outside decoder must agree exactly.
- PCM conversions between 16-bit, 24-in-32, 32-bit and float match ffmpeg's resampler byte for byte.
- IMA ADPCM is byte-identical to ffmpeg in both directions, mono and stereo. Finding that parity took the oracle: ffmpeg's encoder predicts differently from its own decoder, so there is a
ffmpeg_compatible()mode. - WAV headers write and parse for PCM, float and IMA.
- FLAC goes through rusty_flac, made
no_stdupstream for this. Mono at level 5 turned 64,000 bytes into 52,503, and ffmpeg decoded it to the exact source.
On a real XIAO ESP32-S3 the FLAC encoder produced a 203-byte stream that ffmpeg round-tripped to the source samples. It took about 91 ms per 512-sample block, below real time, so on-chip FLAC is proven correct, not yet fast.
Microphones, codec chips and the clock
Capturing sound is where a small ESP32 audio device earns trust. A microphone that drifts, drops blocks or quietly returns zeros is worse than none.
The PDM microphone soak
The XIAO's PDM ESP32 microphone runs through PdmIn, then DcBlock and EnergyVad, at 16 kHz mono in 20 ms blocks. With the radio off, a ten-minute soak gave:
- blocks: result: 30,000
- blocks per second: result: 50.000
- short reads, errors, empties: result: 0
- audio clock against system clock: result: 0.52 ppm apart
The voice activity figure taught a lesson of its own. Over ten seconds the speech ratio was 0.469; over ten minutes, same room and threshold, it was 0.235. A short capture measures the moment it was taken in, not the room. A two-second capture was also judged by ffprobe and ffmpeg, whose levels matched the firmware's arithmetic to 0.03 dB.
Codec chips as register data
ESP32 audio boards with a speaker use a codec chip such as the ES8311, which sits between the chip and the analog world and needs dozens of register writes before it makes a sound. Instead of driver code, rusty_esp_audio holds each chip's register sequences as data: ES8311, ES7210, ES8388, ES8156 and ES7243E, covering Korvo-2, S3-BOX, S3-EYE, LyraT and S3-BOX-Lite boards. They were re-derived from Espressif's ESP-ADF drivers, attributed, and reproduced register by register against a fake bus: the ES8311 bring-up is the vendor's 28 writes in the vendor's order.
The honest caveat: the tables are done, the speaker loopback on a real board is not measured. It waits for a Korvo-2 or S3-EYE.
Streaming ESP32 audio around the house
The first network format for ESP32 audio is intentionally plain: raw 16-bit little-endian PCM, one 20 ms block per UDP datagram, no header. That choice means ffplay can play a device with no receiver code at all.
From the board to ffplay
With the standalone microphone firmware sending to your laptop's address, this is the whole receiver:
ffplay -f s16le -ar 16000 -ch_layout mono -i udp://0.0.0.0:5004
To keep it as a file instead, the package ships a recorder:
cargo run -p rusty_esp_audio-esp --features std --example pcm_record -- 0.0.0.0:5004 mic.wav 600
Over the board's own access point, unicast for ten minutes, the laptop counted 7,201 datagrams at 12.002 per second with zero lost, and the board's sender reported zero dropped. Sequence numbers and timestamps arrive with the mesh transport in rusty_esp_iroh, not in this raw stream.
An intercom or baby monitor you own
A home intercom or baby monitor is the natural first build: a microphone in one room, a laptop or small hub listening in another, and nothing leaving the house. The ESP32 microphone capture, filters, voice detector and UDP stream are ready and measured on a real board. Two things are not: speaker playback through a codec chip is unmeasured, and the combined camera-and-mic sketch reads one audio block per video frame, sending 12 blocks a second from a microphone producing 50. That pacing gap is recorded and open.
To try it, pick a board with a microphone in choosing an ESP32 board, flash with espino, and read a smart home without the cloud for the bigger picture. The source is at github.com/Remade-With-Rust/rusty_esp_audio, built on ESP-IDF.
FAQ
Quick answers for builders evaluating this technology.
Which audio formats does rusty_esp_audio support on the ESP32?
Raw PCM in several sample encodings, IMA ADPCM, WAV and FLAC through rusty_flac. ADPCM and PCM conversions are byte-identical to ffmpeg, and FLAC chunks decode in ffmpeg to the exact source samples. Opus is a measured decision still to come, not a feature.
Has FLAC encoding run on a real ESP32?
Yes, once. A XIAO ESP32-S3 encoded 512 samples to a 203-byte FLAC stream that ffmpeg decoded back to the source byte for byte. It took about 91 ms per 512-sample block at level 0, which is below real time, so it is a correctness result, not a speed one.
Does it include wake words or echo cancellation?
No. The front end is deliberately light: DC blocking, biquad filters, automatic gain control and an energy voice activity detector. Wake words and echo cancellation are non-goals for v1, and speech recognition belongs on the home computer.
Which codec chips are covered?
ES8311, ES7210, ES8388, ES8156 and ES7243E, as register data re-derived from Espressif's drivers and tested register by register against a fake bus. The speaker loopback on a Korvo-2 or S3-EYE board has not been measured yet.
Is there a known limitation in the combined camera and mic sketch?
Yes. The generated camera sketch reads one audio block per video frame, so it sends 12 blocks a second from a microphone producing 50. Three quarters of the audio is discarded at the source. The standalone microphone firmware does not have this gap.