An Apache Iceberg implementation of the Hydrofabric to disseminate continental hydrologic data
Note
To run any of the functions in this repo your AWS test account credentials need to be in your .env file and your .pyiceberg.yaml settings need to up to date with AWS_DEFAULT_REGION="us-east-1" set
The
.envfile is used for deploying to the test environment. A.prod.envfile is used in-place of that if you're deploying to the production environment. By default, when not deploying locally, the test env/catalog is the default; prod needs to be specified when the API/dashboard is launched.
This repo is managed through UV and can be installed through:
uv sync --all-extras
source .venv/bin/activateNote: Functionality is split into optional-dependencies in pyproject.toml. If you only require base functionality, install as uv sync. If you require some extras (e.g. icechunk, io), you can specify uv sync --extra icechunk --extra io as needed. For local develpoment, --all-extras is recommended for complete functionality.
To ensure that icefabric follows the specified structure, be sure to install the local dev dependencies and run pre-commit install
The following sections detail how to:
- Run the Icefabric API locally (standalone)
- Deploy the API/Dashboard through a script that pulls down an S3 catalog archive, then spins up a unified docker setup routed behind an nginx proxy
For full information on API deployment, please see the full user guide API deploy info page. For full information on Streamlit Dashboard deployment, please see the dashboard deploy info page.
To run the API locally, ensure your .env file (make sure to have your prod credentials in a .prod.env if deploying with the production env/catalog) in your project root has the right credentials, then run
python -m app.mainThis should spin up the API services at localhost:8000/.
To specify the deploy environment/iceberg catalog used (test or production (OE)), add a deploy-env flag to the command. The flag should be formatted as --deploy-env <value>:
# Test
python -m app.main --catalog glue --deploy-env test
# Prod
python -m app.main --catalog glue --deploy-env prodIf you are running the API locally (SQL) you first need to localize the Iceberg and Icechunk stores from S3. Information on this can be found in the User Guide - details here.
With the local SQL catalog created, run:
python -m app.main --catalog sqlTo run the api and dashboard together connected to a local iceberg catalog and icechunk data that has been extracted from an archive file synced from S3 please:
1. Authenticate into an AWS profile that has access to an S3 bucket with icefabric archive using the command:
aws sso login --profile your-profile-name
If you haven't created a profile linked to the AWS account please use the aws configure sso command using information associated with the AWS account. Further instructions can be found at: https://d-90678ba0c3.awsapps.com/start/#/
2. Run the following shell script to download the archived catalog, extract it, build the api, dashboard, and nginx docker images, and run docker compose up. Replace the s3 with your s3 path:
docker/deploy_local.sh s3://ngwpc-data/icefabric_catalog_archive.tar {aws_profile}This process will take a while (10-30 minutes) because we need to download ~40 GB of data, extract a large archive, build 3 docker images, and then wait for the api to spin up.
The files will be saved to your /tmp/. If the both directories are present, the shell script will not re-download the archive. Delete icefabric_local_catalog and icefabric_streamflow_obs directories to force download. Note that your tmp file system must be mounted to root and not a tmpfs file system tied to memory. You can check this by running: df -hT /tmp.
The shell script will update your .env file to have the appropriate file paths.
To build the user guide documentation for Icefabric locally, run the following commands:
uv pip install ".[docs]"
mkdocs serve -a localhost:8080Docs will be spun up at localhost:8080/
The tests folder is for all testing data so the global confest can pick it up. This allows all tests in the namespace packages to share the same scope without having to reference one another in tests
To run tests, run pytest -s from project root.
To run the subsetter tests, run pytest --run-slow as these tests take some time. Otherwise, they will be skipped
Smoke tests validate the deployed test API. These tests are skipped when the API_BASE_URL environment variable is not set, so they won't run during normal CI.
To run smoke tests against a deployed environment:
export API_BASE_URL="http://edfs.test.nextgenwaterprediction.com:8000/"
uv run pytest tests/smoke/ -vFor local (note no / following api):
export API_BASE_URL="http://localhost:8000/api"
uv run pytest tests/smoke/ -vThe smoke tests currently verify:
- The API health endpoint is reachable
- Numeric fields (
initial_value,min,max) in theparameter_metadataendpoint are never null
