Skip to content

About

Python tool for cross-sensor calibration

Resources

Code of conduct

Contributing

Stars

3 stars

Watchers

3 watching

Forks

Latest commit

 

History

790 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

SpectralBridge

SpectralBridge translates, validates, and compares reflectance across sensors and scales. It provides restart-safe scientific workflows for individual NEON flightlines, local drone products, production-scale cross-sensor analysis, and spectral-library reporting.

Documentation PyPI Python

Install

SpectralBridge supports Python 3.10, 3.11, and 3.12.

python -m pip install earthlab-spectralbridge

To evaluate the 2.3.0 release candidate explicitly:

python -m pip install --pre "earthlab-spectralbridge==2.3.0rc1"
python -c "import spectralbridge; print(spectralbridge.__version__)"

Contributor, documentation, and notebook dependencies are available as extras:

python -m pip install -e ".[dev]"
python -m pip install -e ".[notebooks]"

Workflows

Individual NEON flightlines

go_forth_and_multiply() is the canonical file-based NEON workflow. It can download inputs, export ENVI, build and apply topographic and BRDF corrections, convolve corrected reflectance to target sensors, extract full-scene or polygon tables, merge Parquet outputs, and create QA artifacts.

from spectralbridge import go_forth_and_multiply

go_forth_and_multiply(
    base_folder="/data/niwo",
    site_code="NIWO",
    year_month="2023-08",
    flight_lines=["NEON_D13_NIWO_DP1_L001-1_20230815_directional_reflectance"],
    extraction_mode="full",  # or "polygon"
    polygon_path=None,
)

Pipeline stages communicate through validated files. A restarted run reuses valid outputs instead of recomputing them:

NEON HDF5
  -> raw ENVI
  -> correction model + corrected ENVI
  -> target-sensor ENVI products
  -> full or polygon Parquet tables
  -> merged tables + QA

Drone processing

The drone workflow is intentionally separate from NEON acquisition. It searches local TIFF/HDF5 inputs recursively, preserves source provenance, applies the requested corrections, and retains corrected native MicaSense. An optional, wavelength-aware affine stage consumes a versioned, reviewed coefficient registry to create distinct Landsat-like translated products and spectral libraries. Translation is L = a + bM after correction; it neither fits at runtime nor performs convolution. Convolution belongs to the NEON hyperspectral branch.

from spectralbridge import run_drone_pipeline

result = run_drone_pipeline(
    "/data/drone_exports",
    output_dir="/data/drone_processed",
    apply_topo=True,
    apply_brdf=True,
    extraction_mode="polygon",
    polygon_path="/data/plots.geojson",
    apply_translation=True,
    translation_strict=False,
)

Production policy fixes weighting to site_balanced. The reviewed 18-record registry is packaged and loaded when translation_coefficients is omitted. Its coefficients retain the numeric units of the completed bulk ENVI products; the translation step refuses apparently fractional drone values instead of silently applying count-scale intercepts or converting units without evidence. See the drone translation tutorial for the import command, wavelength mapping, confidence states, and evidence limits.

Standalone translation QA needs no NEON or network access. Set landsat_qa=True to search Microsoft Planetary Computer for an overlapping Landsat Collection 2 Level 2 scene; install that optional support with python -m pip install "earthlab-spectralbridge[landsat]". Alternatively pass an analysis-ready stacked raster or a previously cached observation manifest as landsat_product. Pass comparison_neon_product only when an existing NEON-convolved product should join the common-Landsat-grid comparison. The run also writes a one-page dashboard, a self-contained PDF report, and separate publication-ready PNG/PDF translation panels. Expensive stage reuse is based on matching source/configuration fingerprints plus output validation, not file existence alone.

Production bulk translation analysis

run_bulk_pipeline() analyzes a tree of immutable, completed-flightline products. Normal bulk analysis never creates an ordinary row-level pixel cache. It reads source rasters in bounded windows and reduces observations immediately to mergeable sufficient statistics.

immutable completed-flightline ENVI + per-flight Parquet products
  -> discovery, identity, product/schema/QA, and eligibility catalog
  -> bounded raster windows
  -> one compact sufficient-statistics checkpoint per flightline
  -> pooled, flightline-balanced, and site-balanced translations
  -> per-flightline and per-site fits
  -> leave-one-site-out validation
  -> candidate coefficients + compact DuckDB/Parquet/JSON outputs
from spectralbridge import run_bulk_pipeline

result = run_bulk_pipeline(
    "/data/completed_flightlines",
    "/data/bulk_analysis",
    threads=4,
    memory_limit="8GB",
)

Use preflight_only=True first for a cheap campaign inventory. Canonical drone outputs are discovered directly through spectralbridge_flightline.json; their per-flight Parquet footers provide product, schema, row, and size summaries. Because the Landsat-like drone products are applications of an existing coefficient registry rather than independent observations, bulk labels their fits as derived_application_verification. It writes the usual compact coefficient and LOSO artifacts so operators can verify that the registry was applied consistently, but marks those artifacts diagnostic_application_verification_only. They must not be interpreted as new calibration evidence or fed back into the production registry. Mixed NEON convolution and drone application-verification flights can share a catalog, but bulk does not silently pool the two evidence classes into one fit.

The workflow is restart-safe: completed per-flightline statistics checkpoints are reused. Source observations stay in their immutable products, so the compact bulk output can be retained independently of the large staging archive. Persistent analysis storage grows mainly with flightline checkpoints and models, not with the total number of selected source pixels.

Start with the local bulk notebook for a curated tree already on disk. For source ExportPackages on CyVerse, use run_drone_bulk_production() or its thin production notebook. The package owns requested-year source resolution, manifest-aware remote inventory, one-H5-at-a-time staging, producer validation, per-year completeness gating, combined-population reporting, compact packaging, and verified upload; the notebook supplies configuration only. Expanding a run revalidates and reuses completed earlier-year flight outputs rather than recomputing them.

Results and interpretation

summarize_bulk_results() operates only on compact outputs from a completed bulk run. It does not reopen source rasters or regenerate sufficient statistics.

from spectralbridge import summarize_bulk_results

summary = summarize_bulk_results(
    "/data/bulk_analysis",
    make_figures=True,
    make_report=True,
)

The report compares pooled and balanced fits, coefficient distributions, site dependence, leave-one-site-out transferability, and correction magnitude. High R² alone is not treated as evidence that sensors are interchangeable; weak, unstable, or unusually large corrections are surfaced explicitly. Outputs are separated into a one-page summary, detailed diagnostics, three manuscript-width publication figures, and Markdown/PDF reports. All are regenerable from the completed compact result tables without the raster archive.

Spectral-library reporting

Spectral-library analysis reads a supplied Parquet library in place and writes compact summaries and optional reports. Run the inexpensive preflight before a full report to inspect schema, group counts, and estimated rendering work.

from spectralbridge import inspect_spectral_library_preflight

preflight = inspect_spectral_library_preflight(
    "/data/polygon_spectral_library.parquet"
)

The most convenient production route is to provide the spectral library to run_bulk_pipeline(). Advanced callers can use both public spectral-library surfaces directly:

from spectralbridge import (
    inspect_spectral_library_preflight,
    run_spectral_library_analysis,
)

The outputs include species summaries, quantiles, low-alpha spectral ensembles, robust and full-range views, hierarchical variability, bounded traceable extreme-spectrum diagnostics, and optional multipage PDFs. Plot bounds do not alter analytical summaries, and the workflow does not make a second copy of the input library. run_spectral_library_analysis() accepts a DuckDB connection and BulkAnalysisPaths for direct control.

Explicit row-level materialization

Most analyses should use the streaming statistics path. If a row-level harmonized dataset is genuinely required, request it explicitly:

from spectralbridge import build_harmonized_dataset

result = build_harmonized_dataset(
    "/data/completed_flightlines",
    "/data/harmonized_dataset",
)

This is an explicit build/export workflow and can be much larger than the normal compact bulk output.

Outputs and QA

On-disk outputs and naming contracts are part of the public API. Depending on the workflow, outputs include paired ENVI .img/.hdr products, Parquet tables, DuckDB catalogs, correction models, coefficient tables, stage QA JSON, figures, and reports. See the outputs contract and schema reference.

QA is evidence, not decoration. Stage reports record pass/warn/fail status and provenance; plots help diagnose corrections, convolution, masks, extraction, weighting dependence, coefficient heterogeneity, and transferability.

Command-line tools

The installed package provides focused commands including:

spectralbridge-download
spectralbridge-pipeline
spectralbridge-bulk
spectralbridge-qa
spectralbridge-qa-summary
spectralbridge-stage-qa
spectralbridge-qa-dashboard
spectralbridge-merge-duckdb
spectralbridge-validate-parquets
spectralbridge-recover-raw

Use COMMAND --help for current options. Historical cscal-* aliases remain available for compatibility.

Release-candidate validation

After installing the exact candidate in a fresh environment, external testers can download and run the small installation check without cloning the repository:

python -m pip install --pre "earthlab-spectralbridge==2.3.0rc1"
curl -O https://raw.githubusercontent.com/earthlab/spectralbridge/v2.3.0rc1/examples/release_candidate_smoke.py
python release_candidate_smoke.py --expected-version 2.3.0rc1

This checks the installed version, public imports, and primary console scripts; it is not a scientific validation. Maintainers also run a bounded exact-artifact smoke that executes the normal, drone, bulk, results, and spectral-library paths.

Documentation and examples

Development

git clone https://github.com/earthlab/spectralbridge.git
cd spectralbridge
python -m pip install -e ".[dev]"
ruff check src tests scripts
pytest -q
mkdocs build --strict

Contributions should preserve scientific assumptions, deterministic outputs, restart safety, bounded processing, and public filename contracts. See CONTRIBUTING.md.

Citation and license

Please cite the software using CITATION.cff. SpectralBridge is licensed under GPL-3.0-or-later. The supported pipeline does not require HyTools at runtime, but selected source components retain documented HyTools implementation lineage. See HYTOOLS_PROVENANCE.md and NOTICE.

About

Python tool for cross-sensor calibration

Resources

Code of conduct

Contributing

Stars

3 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages