A Firestore + Google Cloud Storage (GCS) backed implementation of a lightweight catalog interface. This package provides an opinionated catalog implementation for storing table metadata documents in Firestore and consolidated Parquet manifests in GCS.
Important: This library is modelled after Apache Iceberg but is not compatible with Iceberg; it is a separate implementation with different storage conventions and metadata layout. This library is the catalog and metastore used by opteryx.app and uses Firestore as the primary metastore and GCS for data and manifest storage.
- Firestore-backed catalog and collection storage
- GCS-based table metadata storage; export/import utilities available for artifact conversion
- Table creation, registration, listing, loading, renaming, and deletion
- Commit operations that write updated metadata to GCS and persist references in Firestore
- Simple, opinionated defaults (e.g., default GCS location derived from catalog properties)
- Lightweight schema handling (supports pyarrow schemas)
-
Ensure you have GCP credentials available to the environment. Typical approaches:
- Set
GOOGLE_APPLICATION_CREDENTIALSto a service account JSON key file, or - Use
gcloud auth application-default loginfor local development.
- Set
-
Install locally (or publish to your package repo):
python -m pip install -e ".[parquet]"rugo (the Parquet engine) is an optional dependency so it can be supplied by
whatever already provides it β the parquet extra installs it directly, and
opteryx-core bundles a matching build. Do not install the separately
published draken distribution: rugo vendors its own, and pip overwrites it.
- Create an
OpteryxCatalogand use it in your application:
from draken.interop.vector_sequence import vector_from_sequence
from draken.morsels.morsel import Morsel
from opteryx_catalog import OpteryxCatalog
catalog = OpteryxCatalog(
workspace="my_workspace",
firestore_project="my-gcp-project",
gcs_bucket="my-default-bucket",
)
# Create a collection
catalog.create_collection("example_collection", author="me")
# Schemas are described by an empty Morsel carrying the column types
schema = Morsel()
schema.append_vector("id", vector_from_sequence([], dtype="INTEGER"))
schema.append_vector("name", vector_from_sequence([], dtype="VARCHAR"))
# Create a new dataset (metadata written to a GCS path derived from the bucket property)
dataset = catalog.create_dataset("example_collection.users", schema, author="me")
# Or register a table if you already have a metadata JSON in GCS
catalog.register_table(
("example_namespace", "events"), "gs://my-bucket/path/to/events/metadata/00000001.json"
)
# Load a table
tbl = catalog.load_dataset(("example_namespace", "users"))
print(tbl.metadata)- GCP authentication: Use
GOOGLE_APPLICATION_CREDENTIALSor Application Default Credentials firestore_projectandfirestore_databasecan be supplied when creating the cataloggcs_bucketis recommended to allowcreate_datasetto write metadata automatically; otherwise passlocationexplicitly tocreate_dataset- The catalog writes consolidated Parquet manifests and does not write manifest-list artifacts in the hot path. Use the provided export/import utilities for artifact conversion when necessary.
Example environment variables:
export GOOGLE_APPLICATION_CREDENTIALS="/path/to/service-account.json"
export GOOGLE_CLOUD_PROJECT="my-gcp-project"A user-created commit on a dataset carrying triggers enqueues a refresh job
for each target materialized view: a jobs/{execution_id} document plus an
OIDC-authenticated Cloud Task to worker.opteryx (see
MATERIALIZED_VIEWS_TRIGGERS_PLAN.md). Firing environments need:
| Variable | Default | Purpose |
|---|---|---|
OPTERYX_TRIGGER_FIRING |
on | 0 disables commit-time firing entirely (local scripts, tests) |
GCP_PROJECT_ID / GCP_PROJECT / GOOGLE_CLOUD_PROJECT |
catalog's Firestore project | Project holding the jobs collection and the task queue |
OPTERYX_JOBS_DATABASE |
(default database) | Firestore database for the jobs collection β jobs.opteryx/worker.opteryx use the project default, not catalogs |
TASKS_LOCATION |
us-east1 |
Cloud Tasks queue location |
TASKS_QUEUE |
worker-dispatch |
Cloud Tasks queue name |
TASKS_TARGET_URL |
https://worker.opteryx.app/api/v1/submit |
Where the task is pushed |
TASKS_OIDC_SA |
β | Service account for the task's OIDC token. Must be the same SA jobs.opteryx enqueues as β worker.opteryx pins that OIDC subject. Without it the task is enqueued unauthenticated and the worker rejects it |
TASKS_OIDC_AUDIENCE |
the target URL | OIDC audience β the worker checks exact equality with its submit URL |
JOB_TTL_DAYS |
14 |
purge_at horizon on refresh job documents, matching jobs.opteryx |
Enqueue failures never break the commit that triggered them β they alert and
write a trigger.fire_failed audit record instead. A missed fire is a stale
materialized view, so keep alerting configured wherever firing is enabled.
A trigger fires at most once per minimum-interval-seconds. A commit inside the
interval after a firing is recorded on the trigger as last-fired-status: throttled and audited as trigger.throttled, but enqueues nothing β it is not
an error and does not alert, like a suspension. It is a throttle, not a
debounce: the first commit in a burst fires, and later ones in the interval
are dropped, so the target is refreshed as of that first commit until the next
commit after the interval fires again. Throttled task fires stay correct
because the next fire widens its window back over the skipped commits.
- New triggers get 120 seconds (
DEFAULT_MINIMUM_INTERVAL_SECONDS), written onto the record atcreate_trigger. Passminimum_interval_seconds=0for none, or any non-negative whole number of seconds. - Existing triggers are untouched. The floor is read from the record and
never defaulted at fire time, so a trigger document without the field fires
on every commit as before and pays no extra Firestore round trip. Give one a
floor with
set_trigger_minimum_interval(dataset, trigger, seconds), or in SQL withALTER TRIGGER <name> ON <table> SET MINIMUM INTERVAL TO <n> [SECONDS|MINUTES];0removes it. Re-registering a trigger (CREATE OR REPLACE MATERIALIZED VIEWrewrites every source trigger) keeps the floor the record already holds. - Two commits milliseconds apart cannot both fire. The right to fire is
claimed in a Firestore transaction on the trigger document
(
claim_trigger_fire), which stampslast-claimed-at-msin the same transaction that reads it. Of two concurrent claims one commits and the other is retried against the fresh stamp and refused. The claim is keyed on its own field rather thanlast-fired-at-ms, which records failures too, and a fire that raises after claiming hands the claim back (release_trigger_fire) so an outage does not also silence the next interval. - The floor is per trigger, not per target: a view with two sources has two triggers, each with its own floor. Coalescing what gets through stays with jobs' dedup window.
When the catalog detects a platform inconsistency β a state that should be impossible, like a
snapshot summary disagreeing with the manifest it describes β it raises or reports an exception
carrying the Alertable mixin (opteryx_catalog/exceptions.py). Those are delivered by
opteryx_catalog/alerts/.
Caller errors (DatasetNotFound, DatasetLocked, β¦) are deliberately not alertable β otherwise
every 404 files a ticket.
Delivery is by sink. stdout is the guarantee: one structured JSON line, written synchronously, so the record survives the process being killed. Everything else is an addition and is best-effort.
| Sink | What it does | Default |
|---|---|---|
stdout |
One GCP-structured JSON line per alert, routed to ops.stdout_logs by its severity |
on |
github |
Files, or folds into, one issue per distinct failure | off |
discord |
Posts to a channel webhook. Severity-gated β it interrupts people | off |
# Which channels. Comma-separated; 'both' is an alias for 'stdout,github'.
export OPTERYX_ALERTS_SINK="stdout,discord"
# Identifies the reporting job. Prefixes titles, becomes a label, and is SALTED
# INTO THE FINGERPRINT - changing it later gives every failure a new identity,
# orphaning open issues and re-alerting everything once. Pick it and leave it.
export OPTERYX_ALERTS_COMPONENT="catalog-maintenance"
export OPTERYX_ALERTS_ENVIRONMENT="production"Everything below has a working default; set it only to change that default.
| Variable | Default | Notes |
|---|---|---|
OPTERYX_ALERTS_ENABLED |
on | false silences alerting entirely |
OPTERYX_ALERTS_COOLOFF_HOURS |
24 |
How long a known failure stays quiet. Applies to every sink |
OPTERYX_ALERTS_LABELS |
β | Extra labels, comma separated |
OPTERYX_ALERTS_REPO |
β | owner/repo, or a GitHub URL. Required by the github sink |
OPTERYX_ALERTS_TOKEN_SECRET |
GITHUB_TOKEN |
Secret Manager secret holding the token. A GITHUB_TOKEN env var wins β the dev path |
OPTERYX_ALERTS_API_URL |
https://api.github.com |
For GitHub Enterprise |
OPTERYX_ALERTS_DISCORD_WEBHOOK |
β | The webhook URL directly. Skips Secret Manager, so no IAM grant needed |
OPTERYX_ALERTS_DISCORD_WEBHOOK_SECRET |
DISCORD_NOTIFICATION_WEBHOOK |
Used when the URL isn't set directly |
OPTERYX_ALERTS_DISCORD_MIN_SEVERITY |
CRITICAL |
WARNING, ERROR or CRITICAL |
OPTERYX_ALERTS_DISCORD_MENTION |
β | <@&ROLE_ID> or @here. Without this a Discord message posts silently and won't reach a phone. Get the role ID from Discord β Settings β Advanced β Developer Mode, then right-click the role β Copy ID |
The legacy PLATFORM_ISSUES_* names are still read as a fallback, with a one-time warning β the
GitHub reporter these were moved from was configured that way. Check for stale ones on a deployed
service: an old PLATFORM_ISSUES_COMPONENT silently wins over an unset OPTERYX_ALERTS_COMPONENT
and changes your fingerprints.
Reading a secret needs roles/secretmanager.secretAccessor on it for the runtime service account.
Setting OPTERYX_ALERTS_DISCORD_WEBHOOK directly avoids that entirely.
To verify delivery end to end against a real channel:
python3 scripts/send_test_alert.pyThe alerts extra installs google-cloud-secret-manager, needed only for the Secret Manager path:
pip install "opteryx-catalog[alerts]"This catalog writes consolidated Parquet manifests for fast query planning and stores table metadata in Firestore. Manifests and data files are stored in GCS. If you need different artifact formats, use the provided export/import utilities to convert manifests outside the hot path.
The package's entry point is the OpteryxCatalog class; there is no factory
helper. Alongside it, the top level exports the metastore interface (Metastore,
Dataset, View), the dataset and metadata types (SimpleDataset,
DatasetMetadata, Snapshot, DataFile, ManifestEntry), and ResourceType.
DatasetCompactor is exported from opteryx_catalog.catalog.
OpteryxCatalog(workspace=...) is a read: the workspace must already exist, or
construction raises WorkspaceNotFound. This keeps a mistyped workspace name in
a query from bringing an empty workspace into existence β in Firestore a
collection exists only because a document in it does, so writing the workspace's
$properties document is creating the workspace.
Provisioning is explicit:
OpteryxCatalog(workspace="new_workspace", create_if_missing=True)Key methods include:
create_collection(collection, properties={}, exists_ok=False)drop_namespace(namespace)list_namespaces()create_dataset(identifier, schema, location=None, partition_spec=None, sort_order=None, properties={})register_table(identifier, metadata_location)load_dataset(identifier)list_datasets(namespace)drop_dataset(identifier)rename_table(from_identifier, to_identifier)commit_table(table, requirements, updates)create_view(identifier, sql, schema=None, author=None, description=None, properties={})load_view(identifier)list_views(namespace)view_exists(identifier)drop_view(identifier)update_view_execution_metadata(identifier, row_count=None, execution_time=None)create_materialized_view(identifier, sql, source_tables, author, update_if_exists=False)get_materialized_view(identifier)/list_materialized_views(collection)drop_materialized_view(identifier, author)mark_materialized_view_refreshed(identifier, status, execution_id=None, author=None)create_trigger(dataset_identifier, name, target_view, statement_id=None, author=None)list_triggers(dataset_identifier)/drop_trigger(dataset_identifier, name, author, missing_ok=False)mark_trigger_fired(dataset_identifier, name, status)
Views are SQL queries stored in the catalog that can be referenced like tables. Each view includes:
- SQL statement: The query that defines the view
- Schema: The expected result schema (optional but recommended)
- Metadata: Author, description, creation/update timestamps
- Execution history: Last run time, row count, execution time
Example usage:
from pyiceberg.schema import Schema, NestedField
from pyiceberg.types import IntegerType, StringType
# Create a schema for the view
schema = Schema(
NestedField(field_id=1, name="user_id", field_type=IntegerType(), required=True),
NestedField(field_id=2, name="username", field_type=StringType(), required=False),
)
# Create a view
view = catalog.create_view(
identifier=("my_namespace", "active_users"),
sql="SELECT user_id, username FROM users WHERE active = true",
schema=schema,
author="data_team",
description="View of all active users in the system",
)
# Load a view
view = catalog.load_view(("my_namespace", "active_users"))
print(f"SQL: {view.sql}")
print(f"Schema: {view.metadata.schema}")
# Update execution metadata after running the view
catalog.update_view_execution_metadata(
("my_namespace", "active_users"), row_count=1250, execution_time=0.45
)A trigger is an EVENT plus a runs-as. A commit trigger is held by the dataset
whose commits fire it; a schedule or signal trigger has no dataset, so it is
held by the task it fires, under the task document's triggers subcollection.
Every trigger method takes holder_kind="dataset" (the default) or "task".
catalog.create_trigger(
"ops.rollup", "hourly", target_task="ops.rollup", kind="task", author="alice",
holder_kind="task", event_kind="schedule", schedule="0 * * * *",
time_zone="Europe/London", window_source="src.events", # OVER src.events
)
catalog.create_trigger(
"ops.rollup", "on_demand", target_task="ops.rollup", kind="task", author="alice",
holder_kind="task", event_kind="signal",
)window_source names the dataset a run is windowed OVER: at fire time the
window is that dataset's head snapshot against the task's last-window-to,
with the same superseded/gap guard a commit-fired run has. Without one the
task must be windowless, refused at arming and again at fire time
(window-unbound). A task has one trigger whichever holder it lives under.
The dispatcher is dispatch.opteryx: trigger_firing.fire_due_schedules(client)
is its once-a-minute tick (a collection-group query for next-due-at-ms <= now, each hit claimed with claim_schedule_tick so overlapping ticks cannot
double-fire), and trigger_firing.fire_signal(catalog, task, caller) is its
webhook. The caller of a signal is the event, not the context: the run assumes
the trigger's runs-as and the caller is recorded as fired_by. See
SCHEDULE_SIGNAL_TRIGGERS_DESIGN.md.
A materialized view is a normal dataset document β readable as a table, with
its own location, schema and snapshots β that additionally carries
dataset-type: "materialized_view", its defining SQL (versioned in the same
statement subcollection views use), and a source-tables list. Registration
also writes one trigger document under each source dataset:
{workspace}/{collection}/datasets/{source}/triggers/{trigger-name}
Triggers live in a subcollection, not on the dataset document, for the same
reason maintenance state does: load_dataset reads that document on every
call and must not pay for opt-in state.
The engine creates the backing table first (CREATE MATERIALIZED VIEW runs as
a CTAS), then registers it:
catalog.create_materialized_view(
"mart.daily_orders",
"SELECT customer_id, COUNT(*) FROM sales.orders GROUP BY customer_id",
source_tables=["sales.orders"], # one refresh trigger per source
author="data_team",
)
catalog.list_triggers("sales.orders")
# [{'name': 'refresh__mart__daily_orders', 'kind': 'materialized_view_refresh', ...}]Refresh is event-driven. When a source dataset takes a user-created commit
(append, overwrite, add_files, truncate_and_add_files, committed
truncate), the commit path reads that dataset's triggers and enqueues one
refresh job per target view β see Trigger firing above for the environment it
needs. Housekeeping snapshots (compaction, expiration, refresh_manifest) are
excluded via Snapshot.user_created, so maintenance never re-runs every view.
Behavior worth knowing:
- Invoker semantics: the refresh runs as the author of the commit that
fired it, with policies re-read at fire time. A committer without rights on
the view gets a denied refresh β recorded in
last-refresh-status, never silent. - Materialized views may stack. A view's source can itself be a
materialized view, and the chain refreshes a hop at a time: an inner
refresh lands a user-created commit, which fires the triggers of the views
above it. The cost is inherent to the shape β an outer view is always at
least one refresh behind the inner one, and a failed inner refresh pins
everything above it at the last good data (visible in each view's
last-refresh-status). - The trigger graph is a DAG. A node is a dataset, an edge
D -> Va refresh trigger onDtargeting viewV.create_triggerwalks forward from the target and refuses an edge that reaches back to the dataset the trigger sits on;create_materialized_viewadditionally rejects a cyclicsource-tablesgraph up front, so a bad registration writes nothing at all. Enforcement is on the trigger graph rather than only onsource-tablesbecause that is the graph that fires β a trigger created directly, or left behind by a reconciliation that failed partway, is an edge no source list mentions. Two concurrent writes can each see an acyclic graph, sofsckre-walks it and reportstrigger-cycleas the backstop. - Task triggers are outside all of this. A task records its SQL and never
what that SQL writes, so a task is not a node in the trigger graph and a loop
that runs through one β a task fired by a commit on
awhose statement writesaβ is neither rejected nor detectable here. drop_materialized_viewremoves the triggers from every source before dropping the dataset;drop_dataseton a materialized view does the same cleanup, so a raw drop cannot strand triggers.- Datasets carrying triggers, and materialized views themselves, cannot be renamed β trigger documents and source lists reference names. Drop and recreate instead.
A workspace property that refuses automated copies of that workspace's
datasets into a different workspace β materialized-view refreshes, CTAS, and
INSERT ... SELECT. Enforced in the catalog and, since the engine wiring
landed, at bind time in opteryx-core before anything is written.
On by default. A workspace is restricted from birth: absent, null, and a
workspace with no $properties document at all all read as restricted, and only
an explicit OFF clears it. Sharing a workspace's data out is a decision
someone makes, not a state a workspace drifts into. It is an ordinary settable
property, not a reserved lifecycle field:
ALTER WORKSPACE ichnos SET egress_protection TO OFF; -- opt out
ALTER WORKSPACE ichnos SET egress_protection TO ON; -- opt back indeletion_protection shares the same tri-state, default-on semantics
(_guard_is_on): a workspace is protected from birth, and deleting one is a
deliberate two-step β turn the flag off, then delete. For both guards only an
explicit falsey value clears them, so a hand-written "OFF" string keeps the
guard on rather than silently clearing it. Fail-closed is the right way for a
default-on flag to be wrong.
catalog.is_egress_restricted() # this workspace
catalog.is_egress_restricted("ichnos") # any workspace in the database
catalog.enforce_egress_policy( # the shared gate; raises EgressRestricted
source_workspaces=["ichnos"],
destination_workspace="sales",
operation="create table sales.mart.copy",
)The source workspace's flag decides, never the destination's β the property protects data leaving. A copy that stays inside the source workspace is not egress and is unaffected whatever the flag says, which is what makes a default-on setting liveable: the ordinary same-workspace view or CTAS never touches it.
It is checked at both materialized-view creation and every refresh, because a
workspace can be restricted long after a view that reads it was registered, and
each refresh writes a fresh copy. A refresh blocked this way surfaces like any
other fire failure: alert, audit record, last-fired-status: egress-blocked, and
the commit that fired it is untouched.
Clearing the lock is all-or-nothing: ALTER WORKSPACE ichnos SET egress_protection TO OFF unlocks every copy out of ichnos, for everyone,
until somebody puts it back. Where the copy that has to be allowed is one known
statement β a platform pipeline moving billing events into another workspace,
say β marking that object SECURE is the narrow form: one named object, into
named destination workspaces, withdrawable on its own, with the lock left on for
everything else.
# Run against the SOURCE workspace - the one being copied out of.
ichnos.mark_secure("ws.ops.billing_events_ingest", ["platform"], author="owner")
ichnos.list_secure() # everything ichnos has sanctioned, with who and when
ichnos.clear_secure("ws.ops.billing_events_ingest", author="owner")Only the source can sanction, and that is structural. The record lives in
the SOURCE workspace's $properties, under secure_objects, and a handle only
ever writes its own workspace's properties. A flag stored on the object would be
settable by whoever may edit the object β the party the lock protects against β
which makes it self-granting and the rule advisory. Here the destination cannot
sanction a copy into itself however hard it tries.
Both the object and the destination must match. A task's writes can be
changed by redefining it, so an exemption naming only the object would follow
that redefinition into a workspace its source never agreed to.
Anything malformed reads as NOT sanctioned. This is the permitting half of a
default-closed rule, so it has to be the conservative one about records it does
not understand β the opposite of _guard_is_on, and for the same reason.
secure_objects is a reserved property for that reason too: it is shaped, and an
unshaped one fails open, so it goes through mark_secure (which validates) or it
does not go.
What it is not: containment. Anyone with reader can still SELECT the
data and paste it wherever they like, and nothing here prevents that β it is
leaky by construction, in the way a VPC Service Controls perimeter is an egress
boundary rather than a permission. What it stops is the systematic, automated,
recurring copy: the standing view or CTAS that keeps a full mirror of someone
else's data fresh forever off the back of one read grant. Do not describe it,
or rely on it, as anything stronger.
Where this stands relative to the feature it guards. Cross-workspace MV
sources are not representable yet β _relative_identifier collapses a foreign
prefix into a relative name, and opteryx-core's register_materialized_view
rejects one outright β but they are planned, and this is the guard waiting for
them. That ordering is the point: because the flag defaults to ON, the day a view
can read another workspace it needs that workspace's owner to have said yes
first, and there is no release in which the capability ships ahead of the
boundary. Today the MV gate fires for a caller driving this library directly with
a foreign-qualified source; the load-bearing caller is CTAS, in the engine, which
is not wired up yet.
Notes about behavior:
create_datasetwill try to infer a default GCS location using the providedgcs_bucketproperty iflocationis omitted.register_tablevalidates that the providedmetadata_locationpoints to an existing GCS blob.- Views are stored as Firestore documents with complete metadata including SQL, schema, authorship, and execution history.
- Table transactions are intentionally unimplemented.
Two records answer "what was this built from?", with different lifetimes
(see PROVENANCE_DESIGN.md):
The receipt β on every snapshot document, written once and never edited:
produced-by is kind:name, and the kind decides what the name means:
task: and view: name a catalog object β that name is what the integrity
sweep checks a receipt against β while upload: names the CHANNEL data
arrived through (web, api, mesos), because there is no object to name.
Absent means nothing registered made the commit, which is what a hand-run
statement looks like. What a commit DID is not in here: operation-type
already records that, and a second field restating it is a second field that
can disagree with it. SHOW LINEAGE FOR composes the two into one readable
how column ("merged, by task ops.ingest", "appended, uploaded via web").
One entry per (catalog relation, snapshot) the statement that produced the
commit read; resolved-by is one of current, version, previous, tag,
date. Capped at 256 entries, with read-sources-truncated: true when the cap
bit. read-source-keys (the bare name, and the name at each version) makes
"which commits, anywhere, read X" and "which read version V of X" one
collection-group array_contains query each over snapshots -
find_readers in impact.py.
The standing source list β on the dataset document, maintained by the commit path from the receipts:
"sources": ["opteryx.ops.stdout_log", "opteryx.ops.stderr_log"], // most recent first, up to 64
"sources-complete": trueAn append adds to it, a rewrite (CREATE OR REPLACE, a materialized-view
refresh) replaces it, a TRUNCATE or a delete that empties the table clears
it, maintenance leaves it alone, and a rollback recomputes it from the receipts
behind the new head. The dataset is never in its own list. sources-complete
is false when a name may be missing β the cap dropped one, or a commit in the
chain carried no receipt β and true again after the next rewrite or clear.
Looking downstream β impact.find_impacted answers "if I change this,
who is affected?" from the sources lists, and impact.find_readers answers
"whose commits have read this, and which version?" from read-source-keys.
The first is about current content, so a consumer since rebuilt from
something else is correctly absent; the second remembers it. Both cross
workspace boundaries and name every consumer in full: an impact answer that
will not say who is affected has not answered the question, and knowing a
dataset exists is not being able to read it. Both cap at 64 consumers and
say when the cap bit β past that a list is not what anyone wanted.
The declaration β reads on a task document, beside writes: what the
statement is declared to read, qualified, rewritten on every registration. It
is what the integrity sweep checks a receipt against. Tasks registered before
the field existed get it from scripts/backfill_task_reads.py, which derives
it from each task's current statement with the engine's parser (dry run by
default; --apply writes only where the stored list is absent or differs).
Every commit method (add_files, truncate_and_add_files, merge_commit,
delete_rows, delete_files, append, overwrite) takes read_sources and
produced_by. None is a bug, not a state: a statement that read no
catalog relation β an upload, INSERT β¦ VALUES β passes []. A commit with
no receipt still lands (its files are already on disk) but raises a
ReceiptMissing alert, writes a missing_receipt audit record, marks the
source list incomplete, and is reported by audit_workspace as
missing-receipt for as long as it is the dataset's current snapshot. The
sweep also reports receipt-outside-declaration (the current snapshot read
something its producing task or view does not declare) and window-overrun
(a windowed task read its window source past the version it was fired for).
Receipts cannot be backfilled: every snapshot committed before this existed is permanently unrecorded, and every read surface must say "not recorded" as a distinct answer from "read nothing".
Collection-group queries need an index declared for them. Firestore
creates single-field indexes automatically, but those are COLLECTION-scoped: a
collection_group(...) query on the same field is a different query and fails
with FAILED_PRECONDITION until an index with collection-group scope exists.
Two things must both be right, and each fails differently:
- The field path must be backticked wherever it is not a bare identifier.
A path is parsed, and an unquoted segment must match
[a-zA-Z_][a-zA-Z_0-9]*, soevent-kindis refused withINVALID_ARGUMENTbefore any index is consulted - no amount of index work fixes it. Almost every field name here is hyphenated.tests/test_firestore_field_paths.pyreads the source and fails on an unquoted one, because a fake Firestore parses nothing and unit tests cannot see this. - The index must exist, with
COLLECTION_GROUPscope.
The queries this package makes, and the index each needs:
| query | collection group | field | index |
|---|---|---|---|
trigger_firing._due_schedule_triggers |
triggers |
`event-kind`, `next-due-at-ms` |
composite, ASC/ASC β exists |
impact.find_impacted |
datasets |
sources |
COLLECTION_GROUP_CONTAINS β
exists |
impact.find_readers |
snapshots |
`read-source-keys` |
COLLECTION_GROUP_CONTAINS β
exists |
list_relationships |
relationships |
`references-*` |
composite β exists |
| listener lookups | listeners |
workspace, user |
composite β exists |
Do not create these by hand. They are declared as Terraform in
opteryx-infra firestore-terraform/ - both databases, one file - and that
config is the answer to "what indexes exist". Creating one out of band is how
the previous list drifted from production: this table claimed the snapshots
index still needed creating long after it existed, and documented a
`target-view` index on triggers that was never created and that nothing
queries (every reference to that field reads it off an already-fetched
document). Both are corrected above.
Two things to know if you add one:
- The single-field ones are
google_firestore_fieldwith anindex_configblock, notgoogle_firestore_index- the gcloud spelling isindexes fields update, notindexes composite create. Anarray_containsfield takesarray_config = "CONTAINS"rather than anorder. index_configdeclares a field's entire configuration. Listing only the collection-group index strips the automatic collection-scoped ones. That has already happened todatasets.sourcesandsnapshots.``read-source-keys` `` via gcloud, which replaces the same way.
Building one backfills the whole collection group, so the snapshots
index is the expensive one - it visits every snapshot document in the
database. Create it before the engine starts writing receipts, not after, so
the backfill runs against the smaller collection.
This package includes a small Makefile target to run linting and formatting tools (ruff, isort, pycln).
Install dev tools and run linters with:
python -m pip install --upgrade pycln isort ruff
make lintRunning tests (if you add tests):
python -m pytestThis catalog supports small file compaction to improve query performance. See COMPACTION.md for detailed design documentation.
from opteryx_catalog import OpteryxCatalog
from opteryx_catalog.catalog import DatasetCompactor
catalog = OpteryxCatalog(workspace="my_workspace", gcs_bucket="my-bucket")
dataset = catalog.load_dataset("my_collection.my_dataset")
# `strategy=None` auto-detects: 'performance' when the dataset has a usable
# sort order, otherwise 'brute'.
compactor = DatasetCompactor(dataset, author="me")
# Each compact() call performs ONE read -> select -> execute -> commit pass.
# dry_run=True returns the plan dict (what would be compacted, and why);
# dry_run=False returns the committed Snapshot. Either returns None when
# nothing clears the size thresholds.
plan = compactor.compact(dry_run=True)
if plan is None:
print("nothing to do")
else:
print(plan["type"], plan["reason"])
snapshot = compactor.compact(dry_run=False)Scheduling is handled by the xb500.opteryx housekeeping service (Cloud
Scheduler β /housekeeping/trigger_compaction), which walks allowlisted
collections and audits every dataset it evaluates. See
COMPACTION.md.
Rules brute and sort_aware are independent. To attempt both in one tick,
call compact() twice in series rather than chaining them in a single call:
compactor.compact(rule="brute")
compactor.compact(rule="sort_aware")Compaction sizing is set by module constants in
opteryx_catalog.catalog.compaction, not by dataset properties β the target
output size is TARGET_SIZE_BYTES (4 GB), with MIN_SIZE_BYTES (3.5 GB) as
the lower bound of the acceptable band and MAX_SIZE_BYTES (4.1 GB) a hard
cap. The memory budget is the one runtime-tunable value, via the
OPTERYX_COMPACTION_RAM_MB environment variable (default 16384).
- No support for dataset-level transactions.
create_dataset_transactionraisesNotImplementedError. - The catalog stores metadata location references in Firestore; purging metadata files from GCS is not implemented.
- This is an opinionated implementation intended for internal or controlled environments. Review for production constraints before use in multi-tenant environments.
Contributions are welcome. Please follow these steps:
- Fork the repository and create a feature branch.
- Run and pass linting and tests locally.
- Submit a PR with a clear description of the change.
Please add unit tests and docs for new behaviors.
If you'd like, I can also add usage examples that show inserting rows using PyIceberg readers/writers, or add CI testing steps to the repository. β