Skip to content

imu: use the process-wide cached epoch offset, like frames - #18

Closed
edgarriba wants to merge 1 commit into
mainfrom
fix/imu-epoch-offset-cached
Closed

edgarriba wants to merge 1 commit into
mainfrom
fix/imu-epoch-offset-cached

Conversation

@edgarriba

Copy link
Copy Markdown
Member

policy.rs states the invariant twice, and the IMU drain broke it.

// Frames and IMU reports both go through this so they land on ONE timeline.

/// ONE offset for the whole process for frames: recomputing per frame lets the two
/// clocks' relative jitter separate a frame stamp from an IMU stamp taken at one instant.
pub(crate) fn steady_epoch_offset_cached()

frame_epoch_ns used steady_epoch_offset_cached(). The IMU drain took a fresh
steady_epoch_offset_now() per call. The two streams were therefore converted with
independently sampled REALTIME/MONOTONIC pairs — precisely the divergence the cache exists
to prevent, and the IMU is the other half of the pair it was protecting.

The doc on frame_epoch_ns described the split as intentional ("the IMU drain takes one per
call, as the shim always did"). That was faithful to the C++ shim — oak_bridge.cpp:835 called
the uncached offset while :54 called the cached one — but the shim had the same bug. The Rust
port preserved it, and the comment then made it look deliberate. Both are corrected here.

Magnitude

Measured on a Jetson Orin (aarch64, systemd-timesyncd), 4469 samples of
CLOCK_REALTIME - CLOCK_MONOTONIC over 90 s:

min -151269 ns, max +1344 ns, spread 152613 ns  =>  ~1.68 us/s

A one-directional slew at ~1.7 ppm, not bounded jitter — so IMU stamps walk away from frame
stamps at roughly 6 ms per hour of process uptime:

uptime divergence reaches
~50 min one IMU period (5 ms @ 200 Hz)
~5.5 h one frame interval (33 ms @ 30 fps)

Plus a sharp edge: a REALTIME step — timesyncd stepping on a large offset, e.g. the first
sync after a network-less boot, network return, or an operator setting the clock — shifts IMU
stamps by the whole step instantly while frames stay put.

Practical read. Minutes-long captures accumulate tens of microseconds and are fine; no need
to redo them. A service left up for hours accrues an unbounded ms-scale IMU-vs-camera offset, and
a clock step corrupts an in-flight capture outright. That misalignment reads as a tracking or
calibration fault downstream and is unattributable after the fact, which is what makes it worth
fixing rather than tolerating.

Caveat on the evidence, stated plainly: the magnitude is measured on the clock pair the two
code paths consume, not from observed sensor-oak output — confirming it directly needs a camera
held for an hour. The code divergence itself is exact and verified by reading both call sites.

How it surfaced

Found while porting a copper OAK driver off sensor-oak onto depthai-rs directly.
Reimplementing the epoch rule from sensor-oak's own documentation surfaced that the IMU path
did not follow it.

Verification

DEPTHAI_SYS_SKIP_NATIVE=1 cargo test -p sensor-oak — 15 passed, 0 failed. 6 lines changed.

🤖 Generated with Claude Code

https://claude.ai/code/session_012PVKpnkDbtUf4Dri6tS6C5

`policy.rs` states the invariant twice and the IMU drain broke it:

    // Frames and IMU reports both go through this so they land on ONE timeline.

    /// ONE offset for the whole process for frames: recomputing per frame lets the
    /// two clocks' relative jitter separate a frame stamp from an IMU stamp taken
    /// at one instant.
    pub(crate) fn steady_epoch_offset_cached()

`frame_epoch_ns` used `steady_epoch_offset_cached()`; the IMU drain took a fresh
`steady_epoch_offset_now()` per call. So the two streams were converted with
independently sampled REALTIME/MONOTONIC pairs — exactly the divergence the cache
exists to prevent, and the IMU is the other half of the pair it was protecting.

The doc on `frame_epoch_ns` described the split as intentional ("the IMU drain
takes one per call, as the shim always did"). It was faithful to the C++ shim
(`oak_bridge.cpp:835` in the pre-depthai-rs tree called the uncached offset while
`:54` called the cached one), but the shim had the same bug — the Rust port
preserved it and the comment then made it look deliberate. Both corrected here.

MAGNITUDE, measured on a Jetson Orin (aarch64, systemd-timesyncd), 4469 samples of
CLOCK_REALTIME - CLOCK_MONOTONIC over 90 s:

    min -151269 ns, max +1344 ns, spread 152613 ns  =>  ~1.68 us/s

A one-directional slew at ~1.7 ppm, not bounded jitter. IMU stamps therefore walk
away from frame stamps at ~6 ms/hour of process uptime:

  * ~50 min  -> one IMU period   (5 ms at 200 Hz)
  * ~5.5 h   -> one frame interval (33 ms at 30 fps)

Plus a sharp edge: a REALTIME step (timesyncd stepping on a large offset — first
sync after a network-less boot, network return, an operator setting the clock)
shifts IMU stamps by the whole step instantly while frames stay put.

Practical read: minutes-long captures accumulate tens of microseconds and are
fine. A service left up for hours accrues an unbounded ms-scale IMU-vs-camera
offset, and a clock step corrupts an in-flight capture outright. That misalignment
reads as a tracking or calibration fault downstream and is unattributable after
the fact, which is what makes it worth fixing rather than tolerating.

Caveat on the evidence: the magnitude is measured on the clock pair the two code
paths consume, not from observed sensor-oak output — confirming it directly needs
a camera held for an hour. The code divergence itself is exact.

Found while porting a copper OAK driver off sensor-oak onto depthai-rs directly:
reimplementing the epoch rule from sensor-oak's own docs surfaced that the IMU
path did not follow them.

15 tests pass (`DEPTHAI_SYS_SKIP_NATIVE=1 cargo test -p sensor-oak`).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012PVKpnkDbtUf4Dri6tS6C5
@edgarriba edgarriba closed this Sep 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant