Skip to content

PoC - #2

Merged
mburns merged 20 commits into
mainfrom
poc-3
Oct 6, 2026
Merged

PoC#2
mburns merged 20 commits into
mainfrom
poc-3

Conversation

@mburns

@mburns mburns commented Jun 27, 2025

Copy link
Copy Markdown
Owner

No description provided.

Michael Burns and others added 20 commits June 25, 2025 04:05
Replace the LIKE-over-HTTP search with FTS5 MATCH queries via TrailBase's query() API, with escaped user input, prefix matching, vote-weighted ranking, real pagination counts and clamped limits. A new migration rebuilds titles_fts/persons_fts as external-content tables over the real tables and drops the never-populated search_fts. Fix the Jest config so the suite runs, add query-builder tests, and update the Python smoke test and README.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Add sources/entities/events tables (instants and spans with explicit date precision, negative years, JSON detail), a v_events read model and an entities FTS index. Add scripts/ingest: a sync engine that fingerprints source files, skips unchanged sources, upserts by stable keys, removes stale rows and logs each run, with adapters for IMDB (derived), Lahman baseball, Olympic athlete events and the Wikidata age dataset. Expose the new tables as read-only record APIs, add Makefile targets, tests and docs. Stop lint-staged from passing Markdown and YAML to Biome, which does not process them.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… functions

Add a migration that flags spans versus instants, indexes [start_year, end_year] in an rtree_i32 kept in sync by triggers, exposes Julian day numbers as VIRTUAL generated columns, normalises IMDB genres and professions with json_each, switches entities_fts to the trigram tokenizer and makes the wikidata index partial. Add a /timeline handler (R*Tree overlap, window functions for per-entity order and total count, anchor offsets in years and days) and /timeline/density. The search handler returns FTS5 highlight() markup. The ingest sets the span flag, tunes pragmas for bulk loads and runs ANALYZE and PRAGMA optimize after each sync. Everything is verified against TrailBase's SQLite 3.49, Python's 3.53 and the CLI's 3.43.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…nly APIs

Replace the primary-key-less v_genre_summary view (TrailBase cannot expose it) with a genres table refreshed by the import and a v_genre_titles view keyed by titles.id; point title.html at the real v_title_episodes view and include mini-series; add indexes for every filter/order the list pages issue and drop the redundant auto-named ones. Make every record API read-only, drop the stale auto-generated email templates and the removed conflict_resolution enum from the config. Replace the template engine that stripped {% extends %} with one that implements block inheritance, build pages from the templates directory, and document serving dist via trail run --public-dir. Pin Alpine and fix the footer link.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…et v0.34

Move /search, /timeline and /timeline/density from the removed traildepot/scripts V8 runtime to a WebAssembly component built from wasm/src with the trailbase-wasm SDK: Vite bundles the TypeScript, jco componentize produces traildepot/wasm/component.wasm, and make wasm / make run wire it up. Handlers read query params through HttpRequest, return HttpResponse.json, and coerce bigint SQLite integers to numbers. Tests mock the SDK and cover the query builders. Verified end to end on TrailBase v0.34.4: component loads, all routes, pages and record APIs answer, and the v0.14 filter and count bugs are gone.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…uide

Delete the shell-era and Python-era configs (ESLint, Prettier, pre-commit), the dead build.py/test scripts (one crashed with NameError), stale helper scripts and unused types, and the committed GeoLite2 database (gitignored now). CI runs Biome, tsc, Jest, the site build, the WASM component build, ruff, the ingest tests and a sqlite3 migration check. Makefile and CONTRIBUTING describe the current stack; VS Code settings point at Biome and ruff; the unused sqlite3 npm package is gone.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ty, capped totals

Add entity_links (same person across sources: shared QID, or unique name plus birth and death years; Wikidata canonical), event_participants (a film's cast and crew from principals, a World Series roster from Lahman appearances), event_density refreshed per source on every sync, and an events.certainty column. The /timeline page query no longer carries a window count; a separate capped count supplies total/totalCapped. New routes: /timeline?participant=, /timeline/participants, /timeline/links; density reads the precomputed table. Pin the upsert join order with CROSS JOIN: with ANALYZE statistics present the planner scanned the real tables and probed the stat-less staging tables, turning the 50-second Wikidata load into hours. Mark interrupted syncs failed on the next run.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…king

scripts/wikidata_extract.py streams a Wikidata JSON dump (gz, bz2, plain or stdin) once and writes items.jsonl.gz (humans, conflicts, countries, awards: English label, classes, dates with precision and calendar, IMDb/Baseball-Reference/Olympedia ids, dated award/position/conflict/participant relations) plus labels.jsonl.gz for every item, with a regex pre-filter so most lines skip json.loads. The wikidata source turns the extract into life spans, award instants, position spans, conflict spans with combatants and participating countries, country existence spans and award establishment instants, and emits external identifiers. entity_identifiers (new migration, also refreshes v_events with span and certainty) feed an exact linker: shared identifiers link first, then shared QIDs, then name plus dates; the dump source is canonical when loaded. Entities that only take part in events are no longer dropped as orphans. Fixture: eleven real entities in dump format.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ipants

The imdb source now emits every episode of an included series as an instant on the series' lane (season, episode, rating in detail; --no-imdb-episodes to skip), stages every credited person as an entity even without a birth year, and records cast (principals) and crew roles as participants of titles and episodes. Episodes are excluded from stand-alone title events. entities_fts is now maintained by triggers (new migration) instead of a full trigram rebuild per sync, which had made every sync cost 15-20 seconds over 1.3M names. Biome no longer formats scripts/fixtures, which had pretty-printed the one-entity-per-line Wikidata fixture; the fixture is regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
scripts/ingest/csvsource.py builds a Source from a TOML spec: files to fingerprint, and {column} templates mapping rows onto entities, events (ISO, float-year, M/D/Y or year/month/day-table dates; require/when filters; min/max year to drop bad data; typed detail JSON), participants and identifiers. specs/nba.toml yields player life, career and draft events, every game since 1946 on the home team's lane with the away team as participant, and franchise eras; specs/musicbrainz.toml yields every official release on its artist's lane. Both emit the ids Wikidata carries (P3647 NBA.com, P434 MusicBrainz), and the extractor now keeps those plus Steam and Open Library ids. read_csv tolerates byte-order marks.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Add events.rank (log of votes, credits and the owner's event count, computed per source by the ingest; make rerank for all) and event_lod, the top 20 events per (bucket size 1/10/100, bucket, category, source), refreshed per source. Add indexes whose order is (start_year, id) per filter (events_by_start, events_by_source_year, events_by_entity, events_by_rank) so a window page walks one index and stops at the page size. The /timeline handler now returns events starting inside the window with keyset cursors instead of OFFSET, plus the spans active at the window start from an R*Tree point query; /timeline/overview serves the top events per bucket from event_lod for wide windows. Tests cover rank ordering, LOD contents and the query plans (no TEMP B-TREE sorts). The commit-msg hook accepts the full conventional-commit type set and both hooks drop the deprecated husky shim.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The per-event correlated subqueries made rerank take 41 minutes over 6.4M events; counting participants per event and events per owner once into keyed temp tables and joining brings it down to minutes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… order

Add span_lod: per 10- and 100-year bucket and category, the top-ranked spans overlapping the bucket (a recursive CTE expands each span over the buckets it covers), plus '*' rows across categories in both LOD tables. The timeline's 'active' list reads span_lod unless a name, kind or participant filter forces the R*Tree query, which now orders by rank. Every events query pins events (or the LOD table) as the outer loop with CROSS JOIN: with statistics present the planner started from the seven-row sources table and sorted the whole window. Index entities by exact name.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Name-filtered timeline pages, counts and active spans now start from the
entities_fts match and cross-join into events, so a filter like
q=beatles reads the few hundred matching events instead of walking the
window index and probing FTS per row (1.1 s to 130 ms on 6.4M events).

Rank: the owner-breadth term is capped at log10(n) <= 2 so an entity
with thousands of events cannot crowd every overview bucket, and the
CSV adapter gains an `unless` filter used by the MusicBrainz spec to
skip the "Various Artists" placeholder.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The WASM componentizer's native modules (@napi-rs/lzma via jco) need
Node 22.20+, so CI and the Volta pin move to Node 24 and `engines`
says so. A root ruff.toml pins the rule set and target version so the
local and CI ruff runs agree; the shebang scripts get their exec bit
and the remaining findings (import order, % formatting) are fixed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@mburns
mburns marked this pull request as ready for review October 6, 2026 04:09
@mburns
mburns merged commit 967c170 into main Oct 6, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant