SpectralBridge translates, validates, and compares reflectance across sensors and scales. It provides restart-safe scientific workflows for individual NEON flightlines, local drone products, production-scale cross-sensor analysis, and spectral-library reporting.
SpectralBridge supports Python 3.10, 3.11, and 3.12.
python -m pip install earthlab-spectralbridgeTo evaluate the 2.3.0 release candidate explicitly:
python -m pip install --pre "earthlab-spectralbridge==2.3.0rc1"
python -c "import spectralbridge; print(spectralbridge.__version__)"Contributor, documentation, and notebook dependencies are available as extras:
python -m pip install -e ".[dev]"
python -m pip install -e ".[notebooks]"go_forth_and_multiply() is the canonical file-based NEON workflow. It can
download inputs, export ENVI, build and apply topographic and BRDF corrections,
convolve corrected reflectance to target sensors, extract full-scene or polygon
tables, merge Parquet outputs, and create QA artifacts.
from spectralbridge import go_forth_and_multiply
go_forth_and_multiply(
base_folder="/data/niwo",
site_code="NIWO",
year_month="2023-08",
flight_lines=["NEON_D13_NIWO_DP1_L001-1_20230815_directional_reflectance"],
extraction_mode="full", # or "polygon"
polygon_path=None,
)Pipeline stages communicate through validated files. A restarted run reuses valid outputs instead of recomputing them:
NEON HDF5
-> raw ENVI
-> correction model + corrected ENVI
-> target-sensor ENVI products
-> full or polygon Parquet tables
-> merged tables + QA
The drone workflow is intentionally separate from NEON acquisition. It searches
local TIFF/HDF5 inputs recursively, preserves source provenance, applies the
requested corrections, and retains corrected native MicaSense. An optional,
wavelength-aware affine stage consumes a versioned, reviewed coefficient
registry to create distinct Landsat-like translated products and spectral
libraries. Translation is L = a + bM after correction; it neither fits at
runtime nor performs convolution. Convolution belongs to the NEON hyperspectral
branch.
from spectralbridge import run_drone_pipeline
result = run_drone_pipeline(
"/data/drone_exports",
output_dir="/data/drone_processed",
apply_topo=True,
apply_brdf=True,
extraction_mode="polygon",
polygon_path="/data/plots.geojson",
apply_translation=True,
translation_strict=False,
)Production policy fixes weighting to site_balanced. The reviewed 18-record
registry is packaged and loaded when translation_coefficients is omitted.
Its coefficients retain the numeric units of the completed bulk ENVI products;
the translation step refuses apparently fractional drone values instead of
silently applying count-scale intercepts or converting units without evidence.
See the drone translation tutorial for
the import command, wavelength mapping, confidence states, and evidence limits.
Standalone translation QA needs no NEON or network access. Set
landsat_qa=True to search Microsoft Planetary Computer for an overlapping
Landsat Collection 2 Level 2 scene; install that optional support with
python -m pip install "earthlab-spectralbridge[landsat]". Alternatively pass
an analysis-ready stacked raster or a previously cached observation manifest as
landsat_product. Pass comparison_neon_product only when an existing
NEON-convolved product should join the common-Landsat-grid comparison.
The run also writes a one-page dashboard, a self-contained PDF report, and
separate publication-ready PNG/PDF translation panels. Expensive stage reuse
is based on matching source/configuration
fingerprints plus output validation, not file existence alone.
run_bulk_pipeline() analyzes a tree of immutable, completed-flightline
products. Normal bulk analysis never creates an ordinary row-level pixel cache.
It reads source rasters in bounded windows and reduces observations immediately
to mergeable sufficient statistics.
immutable completed-flightline ENVI + per-flight Parquet products
-> discovery, identity, product/schema/QA, and eligibility catalog
-> bounded raster windows
-> one compact sufficient-statistics checkpoint per flightline
-> pooled, flightline-balanced, and site-balanced translations
-> per-flightline and per-site fits
-> leave-one-site-out validation
-> candidate coefficients + compact DuckDB/Parquet/JSON outputs
from spectralbridge import run_bulk_pipeline
result = run_bulk_pipeline(
"/data/completed_flightlines",
"/data/bulk_analysis",
threads=4,
memory_limit="8GB",
)Use preflight_only=True first for a cheap campaign inventory. Canonical drone
outputs are discovered directly through spectralbridge_flightline.json; their
per-flight Parquet footers provide product, schema, row, and size summaries.
Because the Landsat-like drone products are applications of an existing
coefficient registry rather than independent observations, bulk labels their
fits as derived_application_verification. It writes the usual compact
coefficient and LOSO artifacts so operators can verify that the registry was
applied consistently, but marks those artifacts
diagnostic_application_verification_only. They must not be interpreted as new
calibration evidence or fed back into the production registry. Mixed NEON
convolution and drone application-verification flights can share a catalog, but
bulk does not silently pool the two evidence classes into one fit.
The workflow is restart-safe: completed per-flightline statistics checkpoints are reused. Source observations stay in their immutable products, so the compact bulk output can be retained independently of the large staging archive. Persistent analysis storage grows mainly with flightline checkpoints and models, not with the total number of selected source pixels.
Start with the local bulk notebook
for a curated tree already on disk. For source ExportPackages on CyVerse, use
run_drone_bulk_production() or its
thin production notebook.
The package owns requested-year source resolution, manifest-aware remote
inventory, one-H5-at-a-time staging, producer validation, per-year completeness
gating, combined-population reporting, compact packaging, and verified upload;
the notebook supplies configuration only. Expanding a run revalidates and
reuses completed earlier-year flight outputs rather than recomputing them.
summarize_bulk_results() operates only on compact outputs from a completed
bulk run. It does not reopen source rasters or regenerate sufficient statistics.
from spectralbridge import summarize_bulk_results
summary = summarize_bulk_results(
"/data/bulk_analysis",
make_figures=True,
make_report=True,
)The report compares pooled and balanced fits, coefficient distributions, site dependence, leave-one-site-out transferability, and correction magnitude. High R² alone is not treated as evidence that sensors are interchangeable; weak, unstable, or unusually large corrections are surfaced explicitly. Outputs are separated into a one-page summary, detailed diagnostics, three manuscript-width publication figures, and Markdown/PDF reports. All are regenerable from the completed compact result tables without the raster archive.
Spectral-library analysis reads a supplied Parquet library in place and writes compact summaries and optional reports. Run the inexpensive preflight before a full report to inspect schema, group counts, and estimated rendering work.
from spectralbridge import inspect_spectral_library_preflight
preflight = inspect_spectral_library_preflight(
"/data/polygon_spectral_library.parquet"
)The most convenient production route is to provide the spectral library to
run_bulk_pipeline(). Advanced callers can use both public spectral-library
surfaces directly:
from spectralbridge import (
inspect_spectral_library_preflight,
run_spectral_library_analysis,
)The outputs include species summaries, quantiles, low-alpha spectral ensembles,
robust and full-range views, hierarchical variability, bounded traceable
extreme-spectrum diagnostics, and optional multipage PDFs. Plot bounds do not
alter analytical summaries, and the workflow does not make a second copy of the
input library. run_spectral_library_analysis() accepts a DuckDB connection and
BulkAnalysisPaths for direct control.
Most analyses should use the streaming statistics path. If a row-level harmonized dataset is genuinely required, request it explicitly:
from spectralbridge import build_harmonized_dataset
result = build_harmonized_dataset(
"/data/completed_flightlines",
"/data/harmonized_dataset",
)This is an explicit build/export workflow and can be much larger than the normal compact bulk output.
On-disk outputs and naming contracts are part of the public API. Depending on
the workflow, outputs include paired ENVI .img/.hdr products, Parquet tables,
DuckDB catalogs, correction models, coefficient tables, stage QA JSON, figures,
and reports. See the outputs contract
and schema reference.
QA is evidence, not decoration. Stage reports record pass/warn/fail status and provenance; plots help diagnose corrections, convolution, masks, extraction, weighting dependence, coefficient heterogeneity, and transferability.
The installed package provides focused commands including:
spectralbridge-download
spectralbridge-pipeline
spectralbridge-bulk
spectralbridge-qa
spectralbridge-qa-summary
spectralbridge-stage-qa
spectralbridge-qa-dashboard
spectralbridge-merge-duckdb
spectralbridge-validate-parquets
spectralbridge-recover-raw
Use COMMAND --help for current options. Historical cscal-* aliases remain
available for compatibility.
After installing the exact candidate in a fresh environment, external testers can download and run the small installation check without cloning the repository:
python -m pip install --pre "earthlab-spectralbridge==2.3.0rc1"
curl -O https://raw.githubusercontent.com/earthlab/spectralbridge/v2.3.0rc1/examples/release_candidate_smoke.py
python release_candidate_smoke.py --expected-version 2.3.0rc1This checks the installed version, public imports, and primary console scripts; it is not a scientific validation. Maintainers also run a bounded exact-artifact smoke that executes the normal, drone, bulk, results, and spectral-library paths.
- Documentation home
- Start here
- Bulk analysis vignette
- Architecture
- Notebook vignettes
- Changelog
- HyTools code provenance
git clone https://github.com/earthlab/spectralbridge.git
cd spectralbridge
python -m pip install -e ".[dev]"
ruff check src tests scripts
pytest -q
mkdocs build --strictContributions should preserve scientific assumptions, deterministic outputs, restart safety, bounded processing, and public filename contracts. See CONTRIBUTING.md.
Please cite the software using CITATION.cff. SpectralBridge is licensed under GPL-3.0-or-later. The supported pipeline does not require HyTools at runtime, but selected source components retain documented HyTools implementation lineage. See HYTOOLS_PROVENANCE.md and NOTICE.