Skip to content

feat: banded timeline, browse by domain, seven new sources - #6

Merged
mburns merged 6 commits into
mainfrom
feat/timeline-site
Oct 6, 2026
Merged

mburns merged 6 commits into
mainfrom
feat/timeline-site

Conversation

@mburns

@mburns mburns commented Oct 6, 2026

Copy link
Copy Markdown
Owner

Summary

Thyme becomes a general timeline site rather than an IMDB browser with a timeline API on the side.

  • Banded timeline (/timeline.html, web/timeline/): one horizontal band per data source, scrolled sideways across a world range of years, with a sticky year axis, a density minimap and zoom (ctrl+wheel, +/−). Zoomed out, each band shows its top events per century, decade or year from the LOD tables; zoomed in it reads the full event stream in ten-year tiles. Name, category, kind, source and domain filters, a detail panel with participants, and Compare from here for a second axis of years before and after an event. View state lives in the URL.
  • Browse by domain (/browse.html, wasm/src/domains.ts, /timeline/domains): Film/TV/games, Music, Sports, Books, Awards, Science, Politics, Wars, Religion, Business, Arts, Exploration, Lives and Wikipedia lists, each a set of event categories and kinds with live counts; categories without data show as stubs. Nav is Timeline · Browse · Sources · Search · About; the IMDB tables stay reachable from the Film domain.
  • Pages: sources (counts, years, licence, last sync, what is planned), entity (an entity's events across sources), search over every source's entities, and the TV page explains an empty sample import.
  • Seven new sources: Book-Crossing books, Steam games, famous painters, international cricket, MLS and Olympics from the files under data/, plus Wikipedia bibliographies and "List of …" pages through a new fetcher (make wikipedia-lists) and parser. wikidata_extract.py --all-dated keeps any Wikidata item with a calendar date as a stub.
  • Performance: a sparse source inside a dense decade timed out at 16 s and now pages on its own index in ~20 ms; multi-category overviews get category-first indexes and a rank column on the LOD tables (1 s → 70 ms). New migration U1791270000; run make rerank on existing databases.
  • README and CONTRIBUTING rewritten for the current app; the ingest CLI waits out a running server instead of failing with "database is locked".

Test plan

  • yarn lint, yarn type-check (wasm/ and web/), yarn test (51 Jest tests)
  • make test-ingest (ingest, Wikipedia list parser and Wikidata extract suites)
  • Migrations apply on a fresh database (CI) and on a 7M-event database, followed by make rerank
  • Synced all seven new sources into the large database (~590k events) and checked every page and endpoint, including headless-Chrome screenshots of the timeline at overview and detail zoom

🤖 Generated with Claude Code

Michael Burns and others added 6 commits October 5, 2026 22:13
…ries

A sparse source inside a dense decade (160 IMDB events among 500k music
releases) walked events_by_start and rejected almost every row: 16 s,
past the component's 20 s limit on the next decade. When event_density
says the requested sources are sparse in the window, the page is now
taken from events_by_source_year first and only those rows are joined
(56 ms). Source membership is checked on the events row in every query
rather than on sources.slug after the join (3 s to 27 ms for NBA 1960s).

Domain views ask the LOD tables for up to fifty categories at once. The
primary keys lead with the bucket, so that scanned every category in
range (0.7-1.2 s per tile). New category-first indexes and a rank column
on event_lod/span_lod let the overview cut the top events per bucket
across categories before touching events; `make rerank` fills the rank.
event_density gets a category index for the domain counts.

Also new: /timeline/sources (per-source totals, year range, top
categories, last sync) and density?by=source for the minimap.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The site is no longer an IMDB browser with a timeline API on the side.

Timeline (web/, a small TypeScript app bundled by Vite): one horizontal
band per data source, scrolled sideways over a world range of years; a
sticky year axis, a density minimap of the whole range, and zoom with
ctrl+wheel or +/-. Zoomed out, each band shows the top events per
century, decade or year from the LOD tables; zoomed in it reads the full
event stream in ten-year tiles. Name, category, kind and source filters,
domain chips, a detail panel with participants, and "Compare from here",
which adds a second axis of years before and after the chosen event. The
view state lives in the URL so any view can be shared. Pure layout maths
(tiers, tiles, ticks, lane packing) is unit tested under web/__tests__.

Domains (wasm/src/domains.ts, /timeline/domains): Film, TV & games,
Music, Sports, Books, Awards, Science, Politics, Wars, Religion,
Business, Arts, Exploration, Lives and Wikipedia lists, each a set of
event categories and kinds with live counts, so the menu stays short
while categories keep growing; categories without data show as stubs.

Pages: browse.html (domain grid and per-domain view with highlights by
century), sources.html (every dataset with counts, years, licence, last
sync and what is planned), entity.html (an entity's events across
sources), search.html now covers timeline entities from every source
through /search (trigram FTS, ranked by the entity's top event), and the
TV page explains an empty sample import instead of showing nothing. The
IMDB tables stay reachable from the Film domain.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…dated stubs

Five declarative sources from the files already under data/: Book-
Crossing books (271k publications with their authors as participants),
Steam games (89k releases keyed by app id, which Wikidata carries as
P1733), famous painters (lives), international cricket fixtures and MLS
matches. The CSV reader lifts the field limit for Steam's descriptions.

Wikipedia lists: scripts/wikipedia_lists_fetch.py pulls bibliographies
and "List of ..." pages through the MediaWiki API (seeded from "Lists of
books", or any page given), and the wikipedia_lists adapter turns every
dated entry into an event: cite templates and italic-title lines become
books and articles with authors, every other dated line a stub "listed"
event, and each list page an entity that takes part in its entries.

Wikidata: `wikidata_extract.py --all-dated` keeps any item with a
calendar date (inception, publication, point in time) as a "dated" stub
under its first class, so everything Wikipedia dates can be represented
before it has a parser of its own.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The rerank's big write transactions hit 'database is locked' after the default five seconds while TrailBase held readers on the database; a 300 s busy timeout lets make rerank and make sync-events run against a live server.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…gories are stubs

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@mburns
mburns merged commit f0160f3 into main Oct 6, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant