Accuracy assessment: stratified sampling and error-adjusted area with CIs (#81) - #90
Merged
NewGraphEnvironment merged 10 commits intoSep 29, 2026
Merged
Conversation
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PRhUJsuKABLfpBGktPoiBN
Two blockers. freq() and readValues() disagree on factor strata (confirmed by probe). Sizing assumed strata equal map classes. Also: per-stratum seeds so a pilot extends, a census for tiny strata under an fpc argument, map extraction at sampled cells, and a census oracle that needs no PDFs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PRhUJsuKABLfpBGktPoiBN
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PRhUJsuKABLfpBGktPoiBN
The user chose to import mapaccuracy (CRAN, stats-only) over a second implementation. drift keeps the contract, CIs in ha, the tidy matrix and the per-stratum table the sizer needs. The FPC argument goes: the package always applies it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PRhUJsuKABLfpBGktPoiBN
Tables 8-9, section 5.1.1 and 5.2, with page and table cited. Two PA half-widths and one area SE are printed inconsistently with the paper's own equations; the equation values are kept and the printed ones recorded. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PRhUJsuKABLfpBGktPoiBN
dft_accuracy_sample() draws stratified random points from any strata raster in two chunked passes, with per-stratum seeded streams, so a larger n extends a pilot and draws do not depend on terra's sampler. dft_accuracy_estimate() wraps mapaccuracy::stehman2014() (strata may differ from the map classes) and adds CIs in ha, a long matrix and a per-stratum table. dft_accuracy_labels() is the label contract and dft_accuracy_size() is Olofsson Eq. 13 plus allocation. Tests reproduce Olofsson et al. 2014 Tables 8-9, check against a census of the bundled tiles, and pin a golden draw. Three code-check rounds, the last ended by enumerating every identity site. Relates to NewGraphEnvironment/sred#16 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PRhUJsuKABLfpBGktPoiBN
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PRhUJsuKABLfpBGktPoiBN
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PRhUJsuKABLfpBGktPoiBN
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PRhUJsuKABLfpBGktPoiBN
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PRhUJsuKABLfpBGktPoiBN
NewGraphEnvironment
deleted the
81-accuracy-assessment-for-change-maps-stra
branch
September 29, 2026 14:43
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Until now drift reported mapped area only. Mapped area is biased wherever the map is wrong, and badly so for change maps. This PR adds the standard remedy (Olofsson et al. 2014; Stehman 2014) as the
dft_accuracy_*family:dft_accuracy_sample(strata, n, seed, map =)draws stratified random points from any strata raster: map classes, a transition map, or caller-defined strata. It records the stratum weights and the map's class at each point.terra::spatSample(). The same seed with a largernextends the pilot, so pilot labels carry over.dft_accuracy_estimate(labels, strata)wraps mapaccuracy::stehman2014(), which is valid whether or not the strata are the map classes. It returns error-adjusted area in ha with CIs, overall / user's / producer's accuracy, a long error matrix, and a per-stratum table.dft_accuracy_labels()defines the label contract. NA or blank labels (nonresponse) are refused, and so are training rows.dft_accuracy_size()implements Olofsson Eq. 13 plus allocation, sized from a pilot's per-stratum SD.Related Issues
dft_rast_classify()mutates its input in place; not fixed here)Test plan
dft_accuracy_estimate()to the published precision, using absolute tolerances. Three printed values contradict the paper's own equations; the tests pin the equation values (helper-accuracy.R).accuracy_key(); every fix was mutation-checked.devtools::test()gives 1289 pass / 15 skip.R CMD checkgives 0 errors, 1 NOTE (future timestamps) and 1 WARNING: non-ASCII inR/dft_stac_fetch.R, which is already on main and untouched by this PR.@examplesrun, andpkgdown::check_pkgdown()is clean.Scale test (BULK floodplain, per CLAUDE.md)
classified_2017.tifis 14651 × 11552 = 169,248,352 cells, of which 4,108,972 hold data. Run on a 64 GB machine, with RSS sampled every 1–2 s:dft_accuracy_sample(r17, n = 100)(2 passes + draw)dft_rast_transition(), then sample the 63-stratum factor map (n = 30, 10 censuses)dft_accuracy_estimate()on 1,677 points, 63 classesThe estimator's run time grows with roughly the 2.5th power of the class count (13 s at 80 classes and 1,000 points). That comes from
stehman2014()'s class² indicator columns. It is documented and not filed upstream.Notes
planning/archive/2026-09-issue-81-accuracy-assessment/holds the plan review, the three code-check rounds and the measurement record.ce56768. One doc-only commit follows it:0a360bbqualifies bare#93references in the archive.🤖 Generated with Claude Code
https://claude.ai/code/session_01PRhUJsuKABLfpBGktPoiBN