ESP32 DSP in Rust: Scalar Kernels, PIE SIMD Twins, Same Bytes
Field note on rusty_esp_dsp: one scalar kernel home, ESP32-S3 PIE SIMD twins gated byte-identical, and why measuring on the chip, not the laptop, changed the plan.

What is rusty_esp_dsp and why does it matter for ESP32 devices?
rusty_esp_dsp is the shared home for the pixel and sample kernels the Janus ESP32 packages use. Each kernel has a scalar version kept as the oracle, and ESP32-S3 vector twins must produce identical bytes. On a XIAO ESP32-S3, twins cut per-element time by 63% to 95% on the measured kernels.
rusty_esp_dsp is the single home for the ESP32 DSP kernels the Janus family shares: colour conversion, scaling, block difference and the sample arithmetic that image, video, audio and radio code would otherwise each write for themselves. Every kernel has a scalar version kept forever as the oracle, and every accelerated twin for the ESP32-S3's vector unit is gated byte-identical against it. On a XIAO ESP32-S3 Sense, those twins cut per-element time by 63% to 95% across the eight ESP32 DSP kernels in the published table.
This field note explains the scalar home, the PIE SIMD seam, the byte-identity gate, and why any of it matters to a camera or microphone on your shelf. It sits beside the audio work in the Hear pillar.
One scalar home for ESP32 DSP kernels
Before this package existed, the same ESP32 DSP loops lived in several places. A kernel that two packages carry, or one a chip-side speedup would be spent on, moves here. The scalar version comes first and never leaves.
The oracle stays in the tree
The first rule of ESP32 DSP work here is that the scalar path is the oracle, forever. When pixel conversions moved in from the image package and PCM reductions moved in from audio, they were checked byte-identical to the copies they replaced. H.264 block costs were written against the house encoder's own reference code. Everything is no_std, pure Rust and allocation-free.
That last point was proven on a board, not asserted. The chip reported its own free memory at four stages of a run and the figure never moved: none of the eight kernels allocates.
cargo test --workspace
cargo check -p rusty_esp_dsp --no-default-features --target riscv32imac-unknown-none-elf
The laptop ranked the wrong kernel
The plan originally aimed its effort using a share table measured on a development machine, where YUYV to 24-bit colour was 63.1% of the work. On the ESP32-S3 that same kernel was 15.4%, nothing exceeded 22.4%, and four kernels sat in a band together. The laptop was between 85 and 1617 times faster depending on the kernel, and that spread mostly mapped which loops the desktop compiler had vectorised.
A second surprise came from adopting the house allocator: four kernels moved by up to 8% with no allocator call inside any of them. The cause was where buffers landed in memory. Placement matters on a small chip, and only a board shows it.
PIE SIMD twins for the ESP32-S3
The ESP32-S3 has a 128-bit SIMD extension Espressif calls PIE, the ee.* instructions. It has no Rust intrinsics, so the twins in rusty_esp_dsp-esp are hand-written through core::arch::asm! on the esp toolchain, behind a seam.
Read the instructions off the silicon
Rather than trust prose, each PIE SIMD instruction was run on known byte patterns and every register it could touch was printed. That settled details the obvious guess got wrong: the accumulator holds four lanes of 40 bits, bit-packed, not byte-aligned. A kernel written from the guess would still have returned plausible numbers. It also showed what the unit lacks: no absolute-value instruction, no unsigned max or subtract, no SAD instruction, so the SAD twins are built entirely from signed operations that provably cannot saturate.
rotate90_gray8: scalar (ps/element): 201,137; vector: 9,332; change: -95.4%dot_i16: scalar (ps/element): 271,354; vector: 11,745; change: -93.9%yuyv_to_gray8: scalar (ps/element): 52,875; vector: 6,791; change: -87.2%yuyv_to_rgb565: scalar (ps/element): 399,135; vector: 146,995; change: -63.2%
Not everything paid. A 4x4 Hadamard twin lost twice, the fused load-op family was slower than two plain instructions, and RGB565 to RGB888 is not possible on this unit at all. Those results are recorded with their numbers.
Byte identity, and the ESP32-P4 seam
Each ESP32 DSP twin checks 16-byte alignment and hands anything else to the scalar kernel. An integer twin that is not byte-identical is a bug. Float code gets an exhaustive check where the domain allows: when the RMS level tail moved from f64 to the S3's native f32, a test walked all 520,093,697 possible inputs and results moved by at most 1.526e-5 dB.
The seam also has a PieP4 slot for the ESP32-P4's own vector instructions. Today it delegates to the scalar kernels: P4 twins are a board row waiting for a P4. The ESP32, S2 and C-series stay scalar by design.
Why ESP32 DSP speed matters at home
For a device on your shelf, an ESP32 DSP kernel is only worth something if the shipping firmware calls it, and if saved cycles turn into something you notice.
A kernel nobody calls is not an optimisation
An audit traced every optimised kernel from the real firmware entry points and found only 4 of 39 reachable. A dependency in a manifest is not a call. After wiring the audio front end and letting the camera deliver raw formats, 20 of 39 are on shipping paths. The audio level measurement, measured at its real call site with the twin enabled, took 82.2% less time, on a microphone firmware that exists today in rusty_esp_audio. The twins sit behind an off-by-default pie-s3 feature, so the scalar kernel is always there as the reference.
Battery and latency, stated honestly
Fewer cycles per audio block or camera frame means the ESP32 finishes its work sooner and spends more of each interval idle. For a battery sensor that is the precondition for sleeping longer, and for a doorbell or intercom it is headroom for lower latency. No current draw or battery life has been measured for these kernels yet, so no battery claim is made.
The camera and radio packages that import these kernels are covered in rusty_esp_image and rusty_esp_signal. To put one on a board, use espino, or browse the Learn index. The source is at github.com/Remade-With-Rust/rusty_esp_dsp; the chip is the ESP32-S3, and the Rust-on-Espressif tooling is documented at docs.esp-rs.org.
FAQ
Quick answers for builders evaluating this technology.
What is PIE on the ESP32-S3?
PIE is Espressif's name for the ESP32-S3's 128-bit SIMD extension, the ee.* instructions. It has no Rust intrinsics, so rusty_esp_dsp reaches it through core::arch::asm! on the esp toolchain, with the scalar kernel as the reference.
Does the ESP32-P4 get vector kernels too?
Not yet. The seam has a PieP4 slot, but today it delegates to the scalar kernels. P4 twins are a board row that waits for a P4 on the bench. The ESP32, S2 and C-series stay scalar by design.
What does byte-identical mean here?
An integer vector twin must return exactly the bytes the scalar kernel returns for the same input. A twin that differs is treated as a bug, not an optimisation, and it does not ship.
Do faster kernels mean longer battery life?
Fewer cycles per block leaves the processor idle for more of each frame or audio block, which is the precondition for saving power. No battery or current draw has been measured for these kernels, so no battery figure is claimed.
Why not just trust benchmarks from a laptop?
Because they ranked the wrong kernel. On a laptop, YUYV to 24-bit colour was 63.1% of the work. On the ESP32-S3 it was 15.4%. The laptop table mostly showed which kernels the desktop compiler managed to vectorise.