rusty_esp_dsp Is a Pure Rust Alternative to esp-dsp
rusty_esp_dsp is a pure Rust alternative to esp-dsp for shared ESP32 kernels: a scalar oracle, and ESP32-S3 PIE twins that must match it byte for byte.

Is rusty_esp_dsp a pure Rust alternative to esp-dsp?
Yes, as a separate kernel home, not as a wrapper. rusty_esp_dsp is a pure Rust alternative to esp-dsp for the pixel and sample loops Janus shares. It does not call Espressif's C library. A scalar kernel is the oracle, and an ESP32-S3 vector twin ships only when the bytes match.
Related reading
Builder field note
ESP32 DSP in Rust: Scalar Kernels, PIE SIMD Twins, Same BytesAPIs, composition, and home deployment for ESP32 DSP in Rust: Scalar Kernels, PIE SIMD Twins, Same Bytes live on the field note.
rusty_esp_dsp is a pure Rust alternative to esp-dsp for the loops that image, video, audio, and radio would otherwise each write for themselves. Espressif's esp-dsp is the market standard C library for that work on an ESP32. This crate does not wrap it. Every kernel has a scalar version kept forever as the oracle, and every ESP32-S3 vector twin is gated byte-identical against it. This page is the comparison. The builder field note on rusty_esp_dsp is the table. It sits in the Hear vision.
What esp-dsp is, and what this crate refuses
esp-dsp is the C library you reach for when a colour conversion or a dot product has to run on the chip's vector unit. It is the right dependency if your firmware is already C and you want Espressif's kernels.
A wrapper would not be this crate
A pure Rust alternative to esp-dsp that called the C library through FFI would still be the C library. rusty_esp_dsp is no_std, pure Rust, and allocation-free. On a XIAO ESP32-S3 the chip reported its own free memory at four stages of a run and the figure never moved: none of the eight kernels allocates. Pixel conversions that started in rusty_esp_image and PCM reductions that started in rusty_esp_audio moved here and were checked byte-identical to the copies they replaced.
The scalar path never leaves
The first rule is that the scalar kernel is the oracle, forever. An integer vector twin that is not byte-identical is a bug, not an optimisation. Float code gets an exhaustive check where the domain allows: when an RMS tail moved from f64 to the S3's native f32, a test walked all 520,093,697 possible inputs and results moved by at most 1.526e-5 dB.
PIE twins, read off the silicon
The ESP32-S3's 128-bit SIMD extension is PIE, the ee.* instructions. It has no Rust intrinsics, so the twins are hand-written through core::arch::asm! on the esp toolchain. Each instruction was run on known byte patterns, and every register it could touch was printed. The accumulator holds four lanes of 40 bits, bit-packed, not byte-aligned. The unit has no absolute-value instruction, no unsigned max or subtract, and no SAD instruction, so the SAD twins are built from signed operations that cannot saturate.
The measured table
On a XIAO ESP32-S3 Sense, per-element time moved like this:
rotate90_gray8: scalar 201,137 ps, vector 9,332 ps, down 95.4 percentdot_i16: scalar 271,354 ps, vector 11,745 ps, down 93.9 percentyuyv_to_gray8: scalar 52,875 ps, vector 6,791 ps, down 87.2 percentyuyv_to_rgb565: scalar 399,135 ps, vector 146,995 ps, down 63.2 percent
Not everything paid. A 4x4 Hadamard twin lost twice. The fused load-op family was slower than two plain instructions. RGB565 to RGB888 is not possible on this unit at all. Those results stay in the table. A pure Rust alternative to esp-dsp that published only the wins would be advertising.
Why the laptop table was the wrong map
The plan first aimed effort with a share table from a development machine, where YUYV to 24-bit colour was 63.1 percent of the work. On the ESP32-S3 that kernel was 15.4 percent, nothing exceeded 22.4 percent, and four kernels sat in a band. The laptop was between 85 and 1617 times faster depending on the kernel, mostly because of which loops the desktop compiler had vectorised. A second surprise came from the house allocator: four kernels moved by up to 8 percent with no allocator call inside them, because of where buffers landed. Only a board shows that.
Where esp-dsp still wins
Stay on esp-dsp when the firmware is C, the kernel you need is already in that library, and you do not need the scalar path to be the source of truth for a Rust caller. A pure Rust alternative to esp-dsp is the wrong import if you wanted a drop-in header.
Honest limits
- ESP32-P4 twins are a slot that delegates to scalar until a P4 is on the bench. The ESP32, S2, and C-series stay scalar on purpose.
- No battery current has been measured. Saved cycles are not a milliamp claim.
- The twins exist for kernels Janus actually shares. A kernel nobody calls is not an optimisation, which is why the field note asks whether the shipping firmware uses the loop.
Bring a kernel here only when two packages would otherwise own a copy, or when a chip-side twin is worth the asm. A pure Rust alternative to esp-dsp is not a gallery of every transform in a textbook. The scalar function stays in the tree even after the twin lands, so a reviewer can diff bytes instead of trusting a benchmark caption. That is the maintenance cost a pure Rust alternative to esp-dsp accepts, and it is why a twin that loses stays published. A pure Rust alternative to esp-dsp that deleted the slow path would make the next bug invisible. A pure Rust alternative to esp-dsp keeps the slow path on purpose.
Where to start
The audio graph that calls these kernels is the pure Rust alternative to ESP-ADF. The camera that needs the colour conversions is rusty_esp_image. Flash a board with espino after you choose one. The wider house is the Learn index.
The crate is part of Remade With Rust, the identity model is MATA, and the language is Rust. When a PIE instruction note needs to become a retrieval corpus, the same house runs RAG Converter.
FAQ
Quick answers for builders evaluating this technology.
Does rusty_esp_dsp link Espressif's esp-dsp library?
No. The scalar path is pure Rust, no_std, and allocation-free. ESP32-S3 twins are hand-written through core::arch::asm! because PIE has no Rust intrinsics. A twin that differs from the scalar kernel is a bug and does not ship.
How much faster are the ESP32-S3 twins?
On a XIAO ESP32-S3 Sense, the published table cut per-element time by 63 percent to 95 percent. rotate90_gray8 dropped 95.4 percent, dot_i16 93.9 percent, yuyv_to_gray8 87.2 percent, and yuyv_to_rgb565 63.2 percent. A 4x4 Hadamard twin lost twice and was recorded that way.
Why not benchmark this pure Rust alternative to esp-dsp on a laptop?
The laptop ranked the wrong kernel. YUYV to 24-bit colour was 63.1 percent of the work on a development machine and 15.4 percent on the ESP32-S3. The laptop mostly showed which loops the desktop compiler had vectorised.
Do the faster kernels save battery?
Fewer cycles leave the processor idle for more of each block, which is the precondition for saving power. No battery or current draw has been measured, so no battery figure is claimed.
Does the ESP32-P4 get its own twins?
Not yet. The seam has a PieP4 slot that delegates to the scalar kernels. The ESP32, S2, and C-series stay scalar by design.