Learn/rusty_esp_audioesp-adfesp32audiorust

rusty_esp_audio Is a Pure Rust Alternative to ESP-ADF

rusty_esp_audio is a pure Rust alternative to ESP-ADF for an ESP32 microphone: fixed blocks, no per-element allocation, and codecs checked against ffmpeg.

Signed by M·
Cutaway brick living room at dusk with a microphone brick on the shelf, a soft orb above it, and golden threads staying inside the room
Is rusty_esp_audio a pure Rust alternative to ESP-ADF?

For the pipeline, yes. rusty_esp_audio is a pure Rust alternative to ESP-ADF: elements, fixed blocks the caller owns, and no allocation inside the graph. It does not include wake words or echo cancellation, and codec chips are register tables, not a second copy of the ADF drivers.

Related reading

Builder field note

ESP32 Audio in Rust: A Microphone That Keeps Time on Your LAN

APIs, composition, and home deployment for ESP32 Audio in Rust: A Microphone That Keeps Time on Your LAN live on the field note.

rusty_esp_audio is a pure Rust alternative to ESP-ADF for a microphone that has to keep time on your own network. Espressif's ESP-ADF is the market standard: a graph of elements, codec chips, and a pipeline that can allocate as it runs. Janus keeps the graph and removes the allocation. This page is the comparison. The builder field note on rusty_esp_audio is the pipeline record. It belongs to the Hear vision.

What ESP-ADF gives an ESP32 microphone

ADF won because audio is a chain. A microphone becomes samples, samples become a filtered block, a block becomes bytes on a socket. Writing that chain by hand, in a sketch, is how buffers get lost.

The idea worth keeping

A pure Rust alternative to ESP-ADF that threw the element graph away would be a different product. Pipeline still holds a short list of elements. Each element can change rate, channels, or encoding. Gain, DC block, biquad, automatic gain, an energy voice detector, channel conversion, a mixer, and a linear resampler ship in the crate. The resampler is exact: 480 frames at 48 kHz become 160 at 16 kHz.

What the allocation costs

On a small chip, a block that allocates is a block that can fail in the middle of a sentence. ADF's model allows that. Here the caller owns the scratch, and no element allocates. Biquads land within one least significant bit of ffmpeg and SciPy. The heavy loops live in rusty_esp_dsp, where an ESP32-S3 vector twin may run only if the bytes do not change.

How this pure Rust alternative to ESP-ADF is checked

The test for a format is an outside decoder. PCM conversions between 16-bit, 24-in-32, 32-bit, and float match ffmpeg's resampler byte for byte. IMA ADPCM is byte-identical to ffmpeg in both directions, mono and stereo. WAV headers write and parse for PCM, float, and IMA. FLAC chunks decode in ffmpeg to the exact source samples.

What ran on a XIAO

On a Seeed XIAO ESP32-S3 Sense the PDM microphone read 30,000 blocks in ten minutes with none short, dropped, or errored. Ten minutes of raw 16 kHz PCM over the board's own Wi-Fi delivered 7,201 datagrams with none lost, and ffplay played them with no receiver code of ours. That is the row a pure Rust alternative to ESP-ADF has to show: not a diagram, a soak.

FLAC on that board is a single correctness row, not a speed claim. 512 samples became 203 bytes, ffmpeg agreed, and the encoder took about 91 ms per block at level 0. Opus is a measured decision still to come, not a feature.

Codec chips as data

Boards with a speaker use a chip such as the ES8311 between the ESP32 and the analog jack. Instead of porting the ADF driver, the crate holds the register sequence as data: ES8311, ES7210, ES8388, ES8156, and ES7243E, covering Korvo-2, S3-BOX, S3-EYE, LyraT, and S3-BOX-Lite. They were re-derived from the ADF drivers, attributed, and replayed on a fake bus. The speaker loopback on a real Korvo-2 or S3-EYE has not been measured, so no loopback figure is claimed.

Where ESP-ADF still wins

A pure Rust alternative to ESP-ADF is the wrong tool if the product is a wake word or echo cancellation. Those are non-goals for v1. Speech recognition belongs on the home computer.

Honest limits

  • One generated camera sketch reads one audio block per video frame, so it sends about 12 blocks a second from a microphone producing 50. Three quarters of the audio is discarded at the source. The standalone microphone firmware does not have this gap. The same limit is written on the Arduino comparison.
  • Track A firmware runs on ESP-IDF. There is no C in this crate. There is C underneath, and the radio blob is Espressif's.
  • Codec-chip tables are tested on a fake bus. A listening test on a speaker board is still open.

If you are coming from an ADF example, the map is the element list, not the task API. A pure Rust alternative to ESP-ADF still has a pipeline you can draw. It does not have ADF's component registry, and it will not load an ADF element you already compiled. The gain is the scratch buffer you can see, and the ffmpeg check on the bytes that leave it. Read the field note for the element names before you port a sketch line by line. A pure Rust alternative to ESP-ADF rewards that reading, because the verbs are fewer than ADF's and they are the ones the soak actually ran. A pure Rust alternative to ESP-ADF also refuses to hide a dropped block inside a retry. The counter is the product. A pure Rust alternative to ESP-ADF is finished when that counter is the one you watch.

Where to start

Pick a board with a microphone in choosing an ESP32 board, flash with espino, and read a smart home without the cloud for why the samples stay on the LAN. The Learn index lists the rest of the family.

The crate lives in the Remade With Rust family, uses the same household identity as MATA, and the language is Rust. When a datasheet for one of those codec chips needs to become a retrieval corpus, the same house runs RAG Converter.

FAQ

Quick answers for builders evaluating this technology.

Does this pure Rust alternative to ESP-ADF allocate per audio block?

No. A pipeline ping-pongs two halves of one scratch buffer the caller provides. Each element says up front how many bytes it may produce, so the scratch is sized once. The measured microphone path did not drop a block in 30,000 reads.

Which ESP-ADF features are non-goals here?

Wake words and echo cancellation. The front end is DC blocking, biquads, automatic gain control, and an energy voice detector. Speech recognition belongs on the home computer, not on this crate.

Are the codec chips a pure Rust alternative to the ESP-ADF drivers?

They are data, re-derived from those drivers and attributed. ES8311, ES7210, ES8388, ES8156, and ES7243E replay register by register against a fake bus. The ES8311 bring-up is the vendor's 28 writes in the vendor's order. Speaker loopback on a Korvo-2 has not been measured.

Has FLAC from this pipeline run on an ESP32?

Once, as a correctness result. A XIAO ESP32-S3 encoded 512 samples to a 203-byte FLAC stream that ffmpeg decoded back to the source. It took about 91 ms per block at level 0, which is below real time.

Why would I stay on ESP-ADF?

Stay when you need a wake word, acoustic echo cancellation, or a board whose only published path is an ADF example. Switch when you want the graph, memory the caller owns, and codecs that match ffmpeg.