test(qemu): measure allocator discipline and footprint vs lru - #22
Merged
Merged
Conversation
Two objections an embedded reviewer would raise about this crate had no
measurement behind them. Both now do, in the QEMU job that already exists.
1. "It needs an allocator." True, but only at construction, and that is
the difference between unusable in firmware and one static arena in
main(). The bump allocator in qemu-test now counts its calls:
TypedPulseMap<u32,u32> makes 2 allocations at new() and 0 across 256
insert+get+remove. The two are `buckets` and `slots_ttl`, so a fixed
buckets*128 byte arena lasts the map's lifetime. Asserted, not claimed.
The counterexample is reported rather than hidden: u64 keys/values
exceed the 6-byte key / 7-byte value inline window, so entries reach
the slab and allocate — 14 allocations across 6 inserts.
2. "How big is it, really?" Three probe binaries plus size.sh compare
against lru, the only other cache in the README's table that builds
without std. Emulated Cortex-M0, same workload, both holding 64
resident entries:
flash over baseline RAM allocations
PulseMap 7,376 B 2,048 B 2
lru 7,208 B 4,520 B 68
55% less RAM at the same entry count; flash is a wash, 2.3% apart.
Against interest, and now in the README: on Cortex-M3 PulseMap costs
9,864 B of flash to lru's 7,744 B. portable-atomic's spinlock
AtomicU64 fallback is fatter than the critical-section route M0 takes,
so where flash binds on an M3-class part, lru is the smaller choice.
The subtracted baseline is a third binary with no cache at all, so the
table measures the caches and not hprintln! and the panic handler.
qemu-test's release profile gained lto/codegen-units=1 to match the
parent and what real firmware ships — without it the numbers were
un-inlined glue nobody flashes, off by thousands of bytes.
Still not measured, and now stated in the README: no throughput or
latency figure on any MCU. QEMU is not cycle-accurate, so timing it would
produce a number worth less than none. Every perf figure in the docs is
x86_64. Closing that needs real silicon.
The committed numbers came from a local build on rustc 1.95.0; CI runs 1.98.1 and prints ~50 bytes less per binary. The conclusions are the same either way — RAM 55% lower, M0 flash a wash, M3 flash worse — but a README citing byte counts nobody else reproduces is a README that invites being checked and found wrong. Now quoting what CI prints, so the source is one click away, with the caveat that exact bytes move between toolchain releases while the ratios and the RAM/allocation counts do not.
The MCU facts were buried in the no_std section's prose, where someone scanning for reasons not to use the crate would never find them. Three entries in the list instead: - no perf figure has ever been measured on an MCU; QEMU proves it runs, not that it is fast, and cannot be made to say otherwise - lru is the smaller binary on M3-class parts, 7,576 B to our 9,792 B - every get hit runs the on_access CAS loop (src/raw.rs:262) even in single-threaded builds, and on M0 that CAS is a critical section, so reads disable interrupts The third is the one a firmware author would find on their own and be annoyed we had not mentioned, since it costs interrupt latency rather than throughput.
2 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #21. That PR proved the crate runs on Cortex-M0. This one measures the two things an embedded reviewer would push back on, both of which were unmeasured prose until now. No new CI infrastructure — both land in the QEMU job #21 added.
1. "It needs an allocator"
True, and unavoidable —
PulseMapRaw::newbuildsVecs (src/raw.rs:60) and there is no const-generic or caller-storage constructor. But when it needs one is the part that decides whether firmware can use it, and that is measurable. The bump allocator inqemu-test/now counts its calls:new()insert+get+removeTypedPulseMap<u32, u32>(inline)TypedPulseMap<u64, u64>(slab)The two are
bucketsandslots_ttl. Inline mode never allocates again, so a fixedbuckets × 128-byte arena lasts the map's whole lifetime — "needs a heap" (a dealbreaker in firmware that deliberately has none) becomes "needs a static arena at init" (a line inmain). Both rows arechecked assertions in the QEMU test, so they cannot silently stop being true.The second row is the counterexample, reported rather than hidden:
u64exceeds the 6-byte key / 7-byte value inline window (src/engine/slot.rs:53), entries reach the slab, andSlabPoolgrows viapush(src/engine/slab.rs:122,130). A static arena is not sufficient there.2. "How big is it, really?"
Three probe binaries under
qemu-test/src/bin/plussize.sh, comparing againstlru— the only other cache in the README's table that builds withoutstd, so the whole field on bare metal. Same workload, same target, same profile, both holding 64 resident entries (lenis printed, not assumed).Emulated Cortex-M0 (
thumbv6m,--features m0), figures as CI prints them:lruSame 64 entries in 55% less RAM, touching the allocator twice instead of 68 times. Flash is a wash: 144 bytes apart, 2.0%.
Reported against interest. On Cortex-M3 the flash comparison inverts and it is not close — PulseMap 9,792 B to
lru's 7,576 B. portable-atomic's spinlockAtomicU64fallback is fatter than the critical-section route M0 takes. Where flash is the binding constraint on an M3-class part,lruis the smaller choice, and the README now says exactly that.Why there is a third binary
footprint_nonehas no cache at all and is subtracted out. Without it the table is mostly measuringhprintln!, semihosting and the panic handler — 2,496 B of the thumbv6m total. Host binutilssizereads ARM ELF, so noarm-none-eabitoolchain is needed;size.shskips with a message ifsizeis absent rather than failing the job.Two corrections made while measuring
LTO. The first measurement had
qemu-teston the default release profile, which has LTO off — the parent crate'slto = truedoesn't cross a workspace boundary. That put PulseMap at 10,556 B vslru's 6,672 B, a 58% gap that was measuring un-inlined cross-crate glue nobody flashes. Withlto = true+codegen-units = 1, matching the parent and matching what real firmware ships, the M0 gap closes to 2%.Toolchain. Flash figures are quoted from CI (stable rustc 1.98.1), not from a local build on 1.95.0, which printed ~50 bytes more per binary. Exact byte counts move between toolchain releases; the ratios don't, and the RAM and allocation counts don't move at all. Quoting CI means the source is one click from this PR.
Both are the kind of artifact that produces a confidently wrong table, so they are recorded rather than quietly fixed.
What is still not measured
No throughput or latency figure on any MCU, and the README now says so where it previously implied otherwise. QEMU is not cycle-accurate — no pipeline model, no flash wait states, no bus contention — so timing it would produce a number worth less than no number at all. Every performance figure in the docs is x86_64.
Closing that needs real silicon. An RP2040 is the honest choice, being the Cortex-M0+ part that actually needs
critical-section; SysTick for cycles, same access pattern as the x86 bench,lruon the same board with the same allocator. Out of scope here.Also unchanged from #21:
riscv32imcstays compile-checked only, becauseqemu-system-riscv32 -machine virthas the A extension and would test a target that doesn't need the feature.Notes
src/.qemu-test/remains its own workspace, andcargo package --listconfirms it stays out of the publish tarball.qemu-test/src/bump.rs, shared by all four binaries;default-runadded socargo runis still unambiguous now that there are four.## [Unreleased] — v0.6.5.