Skip to content

Composite chips run serially at ~50 s each: run sample-point chips concurrently #85

Description

@NewGraphEnvironment

Problem

drift#79's pattern for reviewing sample points is one dft_stac_composite() call per buffered point. Each chip is small and cached on its own. But each one costs about 50 s, and the calls run one after another.

Measured 2026-09-28 on the BULK floodplain (bulk_co_ff04), 300 m buffers, 2023 Jul–Aug, parallel = 4 auto, with data-raw/benchmark_composite_bulk.R chips (#79):

  • 100 chips took 84.1 min. Per chip: median 40.2 s, mean 50.5 s, max 140.9 s. Each chip is 3,660 cells, and the cache for all 100 is 3.8 MB. Peak RSS was flat at 0.48 GiB.
  • Split by stage on three chips (contended with the benchmark): stac_cube_items() took 1.4–2.5 s and returned 16 items; the whole composite took 39–95 s.
  • So the time is the COG reads, not the query: 16 scenes × 4 assets (B04/B03/B02/SCL) is about 64 remote opens and block reads for a 600 m chip. Sharing one STAC query across chips would save about 2 s of about 50.

floodplains#93 plans a pilot of roughly 30 points per stratum across about six strata, in 2–3 windows each: about 180 points × 2–3 windows, or 5–8 hours serially. The full sample would be larger.

Approach

The work is latency-bound (0.48 GiB and little CPU per chip), so chips should run concurrently:

  • PSOCK workers, not forks. parallel::mclapply() over a remote raster aborts every fork on macOS, because GDAL's curl handles do not survive a fork. The wrapper exits 0 anyway (code-check-spatial.md).
  • Budget memory per worker. GDAL reserves about 3.2 GB per process before reading a cell, and PSOCK workers outlive their master if it dies (code-check-spatial.md). Cap the worker count and clean workers up with on.exit(stopCluster()).
  • Decide the inner parallelism. Each chip currently also starts gdalcubes workers (auto min(4, cores - 1)). For a 600 m chip that is probably one chunk, so they may add process overhead without concurrency. This is not measured: measure parallel = 1 per chip against the default before combining it with an outer pool.
  • Keep the cache contract. Each chip stays its own atomic cache entry. Two workers writing the same key is last-writer-wins by design (cache_write_atomic()).

A plain wrapper would do: dft_stac_composite_points(points, buffer, years, months, workers = …) returning a list per point, or the documented lapply pattern moved onto a cluster. The API is open. Stability of point ids across calls matters, because floodplains#93 keys labels by sample id.

Acceptance

  • Wall time for 100 BULK chips measured serially and with the pool, with RSS sampled for master and workers. The same benchmark script gets a workers arm.
  • No orphaned worker processes after an interrupted run.
  • Cache entries identical to the serial path.

Relates: drift#79, floodplains#93, drift#81

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions