Problem
When gdalcubes cannot read some chunks of a cube (an expired signed URL, a refused or throttled request), it does not raise an R warning or error. It prints [WARNING] n out of m chunks have repoprted errors / incompleteness and writes the cube with those chunks NA. drift then:
- passes
cube_check_nonempty() whenever any chunk succeeded;
- caches the partial cube as complete, and serves it forever under
force = FALSE.
This affects dft_stac_cube() and dft_stac_composite() (shared stac_cube_assemble()), and probably dft_stac_fetch() (same gdalcubes write).
How it was found (drift#79, 2026-09-28)
A floodplain-wide composite over BULK (tile_size = 20000, 30 tiles) ran for 54 min. Planetary Computer SAS tokens last about 45 min, and every asset had been signed once at query time. 15 of 30 tiles reported "9 out of 9 chunks" failed. The run died later on an unrelated COG write; otherwise the holed composite would have been cached.
#79 fixes the cause it hit: features are re-signed before each read extent. With corrupted tokens, re-signing read 102,364 cells, against 0 without. Detection is still missing. A single extent that outlives a token, or a throttled request, still yields silent NA.
Why the obvious guard does not work
Capturing stderr around write_ncdf() (capture.output(type = "message")) and aborting on the report line was tried and measured on the packaged AOI with corrupted tokens:
gdalcubes parallel |
image mask |
report captured |
cells read |
| 1 |
no |
0 (1 in an earlier run of the same probe) |
0 |
| 1 |
SCL |
2 |
0 |
| 4 |
no |
0 |
0 |
| 4 |
SCL |
0 |
0 |
With worker processes (drift's default is min(4, cores - 1)) the report never reaches R. A guard that fires sometimes reads as protection it does not give, so #79 did not ship it.
Options considered
- Ask gdalcubes for a programmatic error count after compute.
- Pre-flight each extent:
HEAD one asset URL per item just before the read, and abort on 403 or 404. That catches expired tokens but not mid-read throttling.
- Post-check against an independent expectation: for each item footprint that intersects the extent, require some non-NA cells unless the SCL for that item is fully masked.
What was found (2026-10-02) — supersedes the options above
gdalcubes already records it. write_ncdf() writes a per-chunk chunk_status integer variable into every output (0 OK, 1 ERROR, 2 INCOMPLETE, 128 UNKNOWN). Worker processes carry it to the main process, so it does not depend on parallel, where the stderr line did. No upstream request was needed for this half.
A pixel post-check could not have worked. A monthly median over two scenes where one failed fills every cell from the survivor: a holed cube with no NA at all. The third option would have passed it. chunk_status flags it.
What it records is narrower than "the read failed". It records exactly one failure: a band image that fails to open (an expired, corrupted or refused signed URL, 403/404). That is the #79 failure and this issue's acceptance test. Every other I/O step in gdalcubes drops its return code. A code-check enumeration found 38 failure paths: drift now detects 25 and 13 stay silent. The silent ones include:
- a band image that opens and then fails mid-read (a throttled or refused range request): cells filled from the scenes that read;
- a mask (SCL) image that opens and then fails mid-read: the scene is used unmasked;
- a missing or unreadable worker chunk file;
- a worker killed by a signal, which hangs
write_ncdf().
These moved to #99, with an upstream gdalcubes draft (not posted). GDAL HTTP retries now absorb transient 429/5xx.
Shipped in the PR for this issue: cube_write_ncdf(), drift's single write_ncdf() call. It aborts (drift_incomplete_cube) on a non-OK chunk_status and on gdalcubes' "could not be added to output" merge warning, before anything is cached. Untiled fetch caches written before the fix are re-checked on hit and re-fetched. Details: inst/notes/gdalcubes-pc-gotchas.md, "Detecting failed chunk reads".
Acceptance
- A read with corrupted tokens, at the default
parallel, aborts and caches nothing.
- The same holds for a tiled read.
- A normal read is unaffected, and so is the cube cache key.
Relates: drift#79, drift#83
Problem
When gdalcubes cannot read some chunks of a cube (an expired signed URL, a refused or throttled request), it does not raise an R warning or error. It prints
[WARNING] n out of m chunks have repoprted errors / incompletenessand writes the cube with those chunks NA. drift then:cube_check_nonempty()whenever any chunk succeeded;force = FALSE.This affects
dft_stac_cube()anddft_stac_composite()(sharedstac_cube_assemble()), and probablydft_stac_fetch()(same gdalcubes write).How it was found (drift#79, 2026-09-28)
A floodplain-wide composite over BULK (
tile_size = 20000, 30 tiles) ran for 54 min. Planetary Computer SAS tokens last about 45 min, and every asset had been signed once at query time. 15 of 30 tiles reported "9 out of 9 chunks" failed. The run died later on an unrelated COG write; otherwise the holed composite would have been cached.#79 fixes the cause it hit: features are re-signed before each read extent. With corrupted tokens, re-signing read 102,364 cells, against 0 without. Detection is still missing. A single extent that outlives a token, or a throttled request, still yields silent NA.
Why the obvious guard does not work
Capturing stderr around
write_ncdf()(capture.output(type = "message")) and aborting on the report line was tried and measured on the packaged AOI with corrupted tokens:parallelWith worker processes (drift's default is
min(4, cores - 1)) the report never reaches R. A guard that fires sometimes reads as protection it does not give, so #79 did not ship it.Options considered
HEADone asset URL per item just before the read, and abort on 403 or 404. That catches expired tokens but not mid-read throttling.What was found (2026-10-02) — supersedes the options above
gdalcubes already records it.
write_ncdf()writes a per-chunkchunk_statusinteger variable into every output (0 OK, 1 ERROR, 2 INCOMPLETE, 128 UNKNOWN). Worker processes carry it to the main process, so it does not depend onparallel, where the stderr line did. No upstream request was needed for this half.A pixel post-check could not have worked. A monthly median over two scenes where one failed fills every cell from the survivor: a holed cube with no NA at all. The third option would have passed it.
chunk_statusflags it.What it records is narrower than "the read failed". It records exactly one failure: a band image that fails to open (an expired, corrupted or refused signed URL, 403/404). That is the #79 failure and this issue's acceptance test. Every other I/O step in gdalcubes drops its return code. A code-check enumeration found 38 failure paths: drift now detects 25 and 13 stay silent. The silent ones include:
write_ncdf().These moved to #99, with an upstream gdalcubes draft (not posted). GDAL HTTP retries now absorb transient 429/5xx.
Shipped in the PR for this issue:
cube_write_ncdf(), drift's singlewrite_ncdf()call. It aborts (drift_incomplete_cube) on a non-OKchunk_statusand on gdalcubes' "could not be added to output" merge warning, before anything is cached. Untiled fetch caches written before the fix are re-checked on hit and re-fetched. Details:inst/notes/gdalcubes-pc-gotchas.md, "Detecting failed chunk reads".Acceptance
parallel, aborts and caches nothing.Relates: drift#79, drift#83