Skip to content

test(qemu): measure allocator discipline and footprint vs lru - #22

Merged
ddsha441981 merged 3 commits into
mainfrom
test/embedded-footprint-evidence
Sep 9, 2026
Merged

ddsha441981 merged 3 commits into
mainfrom
test/embedded-footprint-evidence

Conversation

@ddsha441981

@ddsha441981 ddsha441981 commented Sep 5, 2026 •

Copy link
Copy Markdown
Owner

Follow-up to #21. That PR proved the crate runs on Cortex-M0. This one measures the two things an embedded reviewer would push back on, both of which were unmeasured prose until now. No new CI infrastructure — both land in the QEMU job #21 added.

1. "It needs an allocator"

True, and unavoidable — PulseMapRaw::new builds Vecs (src/raw.rs:60) and there is no const-generic or caller-storage constructor. But when it needs one is the part that decides whether firmware can use it, and that is measurable. The bump allocator in qemu-test/ now counts its calls:

Map Allocations at new() Allocations during 256 insert+get+remove
TypedPulseMap<u32, u32> (inline) 2 0
TypedPulseMap<u64, u64> (slab) 2 14 per 6 inserts

The two are buckets and slots_ttl. Inline mode never allocates again, so a fixed buckets × 128-byte arena lasts the map's whole lifetime — "needs a heap" (a dealbreaker in firmware that deliberately has none) becomes "needs a static arena at init" (a line in main). Both rows are checked assertions in the QEMU test, so they cannot silently stop being true.

The second row is the counterexample, reported rather than hidden: u64 exceeds the 6-byte key / 7-byte value inline window (src/engine/slot.rs:53), entries reach the slab, and SlabPool grows via push (src/engine/slab.rs:122,130). A static arena is not sufficient there.

2. "How big is it, really?"

Three probe binaries under qemu-test/src/bin/ plus size.sh, comparing against lru — the only other cache in the README's table that builds without std, so the whole field on bare metal. Same workload, same target, same profile, both holding 64 resident entries (len is printed, not assumed).

Emulated Cortex-M0 (thumbv6m, --features m0), figures as CI prints them:

Flash over baseline RAM Allocations
PulseMap 7,320 B 2,048 B — 32 B/entry 2
lru 7,176 B 4,520 B — 70.6 B/entry 68

Same 64 entries in 55% less RAM, touching the allocator twice instead of 68 times. Flash is a wash: 144 bytes apart, 2.0%.

Reported against interest. On Cortex-M3 the flash comparison inverts and it is not close — PulseMap 9,792 B to lru's 7,576 B. portable-atomic's spinlock AtomicU64 fallback is fatter than the critical-section route M0 takes. Where flash is the binding constraint on an M3-class part, lru is the smaller choice, and the README now says exactly that.

Why there is a third binary

footprint_none has no cache at all and is subtracted out. Without it the table is mostly measuring hprintln!, semihosting and the panic handler — 2,496 B of the thumbv6m total. Host binutils size reads ARM ELF, so no arm-none-eabi toolchain is needed; size.sh skips with a message if size is absent rather than failing the job.

Two corrections made while measuring

LTO. The first measurement had qemu-test on the default release profile, which has LTO off — the parent crate's lto = true doesn't cross a workspace boundary. That put PulseMap at 10,556 B vs lru's 6,672 B, a 58% gap that was measuring un-inlined cross-crate glue nobody flashes. With lto = true + codegen-units = 1, matching the parent and matching what real firmware ships, the M0 gap closes to 2%.

Toolchain. Flash figures are quoted from CI (stable rustc 1.98.1), not from a local build on 1.95.0, which printed ~50 bytes more per binary. Exact byte counts move between toolchain releases; the ratios don't, and the RAM and allocation counts don't move at all. Quoting CI means the source is one click from this PR.

Both are the kind of artifact that produces a confidently wrong table, so they are recorded rather than quietly fixed.

What is still not measured

No throughput or latency figure on any MCU, and the README now says so where it previously implied otherwise. QEMU is not cycle-accurate — no pipeline model, no flash wait states, no bus contention — so timing it would produce a number worth less than no number at all. Every performance figure in the docs is x86_64.

Closing that needs real silicon. An RP2040 is the honest choice, being the Cortex-M0+ part that actually needs critical-section; SysTick for cycles, same access pattern as the x86 bench, lru on the same board with the same allocator. Out of scope here.

Also unchanged from #21: riscv32imc stays compile-checked only, because qemu-system-riscv32 -machine virt has the A extension and would test a target that doesn't need the feature.

Notes

  • No changes to src/. qemu-test/ remains its own workspace, and cargo package --list confirms it stays out of the publish tarball.
  • Bump allocator extracted to qemu-test/src/bump.rs, shared by all four binaries; default-run added so cargo run is still unambiguous now that there are four.
  • Footprint numbers are printed in CI, not gated — a size gate needs a stored baseline to diff against, which is a different PR.
  • CHANGELOG appended under ## [Unreleased] — v0.6.5.

Two objections an embedded reviewer would raise about this crate had no
measurement behind them. Both now do, in the QEMU job that already exists.

1. "It needs an allocator." True, but only at construction, and that is
   the difference between unusable in firmware and one static arena in
   main(). The bump allocator in qemu-test now counts its calls:
   TypedPulseMap<u32,u32> makes 2 allocations at new() and 0 across 256
   insert+get+remove. The two are `buckets` and `slots_ttl`, so a fixed
   buckets*128 byte arena lasts the map's lifetime. Asserted, not claimed.

   The counterexample is reported rather than hidden: u64 keys/values
   exceed the 6-byte key / 7-byte value inline window, so entries reach
   the slab and allocate — 14 allocations across 6 inserts.

2. "How big is it, really?" Three probe binaries plus size.sh compare
   against lru, the only other cache in the README's table that builds
   without std. Emulated Cortex-M0, same workload, both holding 64
   resident entries:

               flash over baseline      RAM   allocations
     PulseMap              7,376 B   2,048 B             2
     lru                   7,208 B   4,520 B            68

   55% less RAM at the same entry count; flash is a wash, 2.3% apart.

   Against interest, and now in the README: on Cortex-M3 PulseMap costs
   9,864 B of flash to lru's 7,744 B. portable-atomic's spinlock
   AtomicU64 fallback is fatter than the critical-section route M0 takes,
   so where flash binds on an M3-class part, lru is the smaller choice.

The subtracted baseline is a third binary with no cache at all, so the
table measures the caches and not hprintln! and the panic handler.
qemu-test's release profile gained lto/codegen-units=1 to match the
parent and what real firmware ships — without it the numbers were
un-inlined glue nobody flashes, off by thousands of bytes.

Still not measured, and now stated in the README: no throughput or
latency figure on any MCU. QEMU is not cycle-accurate, so timing it would
produce a number worth less than none. Every perf figure in the docs is
x86_64. Closing that needs real silicon.
The committed numbers came from a local build on rustc 1.95.0; CI runs
1.98.1 and prints ~50 bytes less per binary. The conclusions are the same
either way — RAM 55% lower, M0 flash a wash, M3 flash worse — but a
README citing byte counts nobody else reproduces is a README that invites
being checked and found wrong.

Now quoting what CI prints, so the source is one click away, with the
caveat that exact bytes move between toolchain releases while the ratios
and the RAM/allocation counts do not.
The MCU facts were buried in the no_std section's prose, where someone
scanning for reasons not to use the crate would never find them. Three
entries in the list instead:

- no perf figure has ever been measured on an MCU; QEMU proves it runs,
  not that it is fast, and cannot be made to say otherwise
- lru is the smaller binary on M3-class parts, 7,576 B to our 9,792 B
- every get hit runs the on_access CAS loop (src/raw.rs:262) even in
  single-threaded builds, and on M0 that CAS is a critical section, so
  reads disable interrupts

The third is the one a firmware author would find on their own and be
annoyed we had not mentioned, since it costs interrupt latency rather
than throughput.
@ddsha441981
ddsha441981 merged commit 46412f3 into main Sep 9, 2026
23 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant