Skip to content
This repository was archived by the owner on Oct 6, 2026. It is now read-only.
NGWPCPublic archive
forked from NOAA-OWP/icefabric

About

An Apache Iceberg implementation of the Hydrofabric to disseminate continental hydrologic data.

Resources

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

 
 

Repository files navigation

icefabric

icefabric

Ruff

An Apache Iceberg implementation of the Hydrofabric to disseminate continental hydrologic data

Note

To run any of the functions in this repo your AWS test account credentials need to be in your .env file and your .pyiceberg.yaml settings need to up to date with AWS_DEFAULT_REGION="us-east-1" set

The .env file is used for deploying to the test environment. A .prod.env file is used in-place of that if you're deploying to the production environment. By default, when not deploying locally, the test env/catalog is the default; prod needs to be specified when the API/dashboard is launched.

Getting Started

This repo is managed through UV and can be installed through:

uv sync --all-extras
source .venv/bin/activate

Note: Functionality is split into optional-dependencies in pyproject.toml. If you only require base functionality, install as uv sync. If you require some extras (e.g. icechunk, io), you can specify uv sync --extra icechunk --extra io as needed. For local develpoment, --all-extras is recommended for complete functionality.

Development

To ensure that icefabric follows the specified structure, be sure to install the local dev dependencies and run pre-commit install

Running/deploying the services

The following sections detail how to:

  • Run the Icefabric API locally (standalone)
  • Deploy the API/Dashboard through a script that pulls down an S3 catalog archive, then spins up a unified docker setup routed behind an nginx proxy

For full information on API deployment, please see the full user guide API deploy info page. For full information on Streamlit Dashboard deployment, please see the dashboard deploy info page.

Running the API locally standalone

To run the API locally, ensure your .env file (make sure to have your prod credentials in a .prod.env if deploying with the production env/catalog) in your project root has the right credentials, then run

python -m app.main

This should spin up the API services at localhost:8000/.

To specify the deploy environment/iceberg catalog used (test or production (OE)), add a deploy-env flag to the command. The flag should be formatted as --deploy-env <value>:

# Test
python -m app.main --catalog glue --deploy-env test
# Prod
python -m app.main --catalog glue --deploy-env prod
SQL catalog deploy

If you are running the API locally (SQL) you first need to localize the Iceberg and Icechunk stores from S3. Information on this can be found in the User Guide - details here.

With the local SQL catalog created, run:

python -m app.main --catalog sql

Full API/dashboard local deployment (from archive file)

To run the api and dashboard together connected to a local iceberg catalog and icechunk data that has been extracted from an archive file synced from S3 please:

1. Authenticate into an AWS profile that has access to an S3 bucket with icefabric archive using the command: aws sso login --profile your-profile-name

If you haven't created a profile linked to the AWS account please use the aws configure sso command using information associated with the AWS account. Further instructions can be found at: https://d-90678ba0c3.awsapps.com/start/#/

2. Run the following shell script to download the archived catalog, extract it, build the api, dashboard, and nginx docker images, and run docker compose up. Replace the s3 with your s3 path:

docker/deploy_local.sh s3://ngwpc-data/icefabric_catalog_archive.tar {aws_profile}

This process will take a while (10-30 minutes) because we need to download ~40 GB of data, extract a large archive, build 3 docker images, and then wait for the api to spin up.

The files will be saved to your /tmp/. If the both directories are present, the shell script will not re-download the archive. Delete icefabric_local_catalog and icefabric_streamflow_obs directories to force download. Note that your tmp file system must be mounted to root and not a tmpfs file system tied to memory. You can check this by running: df -hT /tmp. The shell script will update your .env file to have the appropriate file paths.

Documentation

To build the user guide documentation for Icefabric locally, run the following commands:

uv pip install ".[docs]"
mkdocs serve -a localhost:8080

Docs will be spun up at localhost:8080/

Pytests

The tests folder is for all testing data so the global confest can pick it up. This allows all tests in the namespace packages to share the same scope without having to reference one another in tests

To run tests, run pytest -s from project root.

To run the subsetter tests, run pytest --run-slow as these tests take some time. Otherwise, they will be skipped

Smoke Tests

Smoke tests validate the deployed test API. These tests are skipped when the API_BASE_URL environment variable is not set, so they won't run during normal CI.

To run smoke tests against a deployed environment:

export API_BASE_URL="http://edfs.test.nextgenwaterprediction.com:8000/"
uv run pytest tests/smoke/ -v

For local (note no / following api):

export API_BASE_URL="http://localhost:8000/api"
uv run pytest tests/smoke/ -v

The smoke tests currently verify:

  • The API health endpoint is reachable
  • Numeric fields (initial_value, min, max) in the parameter_metadata endpoint are never null

About

An Apache Iceberg implementation of the Hydrofabric to disseminate continental hydrologic data.

Resources

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages