Skip to content

test: споделен фалшив D1, който хвърля при непозната заявка - #331

Merged
todorkolev merged 16 commits into
midt-bg:mainfrom
ydimitrof:test/shared-fake-d1-helper
Aug 26, 2026
Merged

todorkolev merged 16 commits into
midt-bg:mainfrom
ydimitrof:test/shared-fake-d1-helper

Conversation

@ydimitrof

@ydimitrof ydimitrof commented Aug 24, 2026 •

Copy link
Copy Markdown
Contributor

Какво и защо

Един споделен двойник за D1 в нов workspace @sigma/test-support, вместо ~29 ръчни двойника из тестовете.

Маршрут = набор от маркери и отговор; всеки маркер трябва да се среща в SQL-а, а непозната заявка хвърля и назовава както изпълнения оператор, така и всички регистрирани маркери. Празният резултат става изричен избор — { onUnmatched: 'empty' } или маршрут с празен отговор.

Три входни точки: fakeD1 за тестове на заявки, recordingD1 за тестовете на обвивка над D1 (readonlyD1), които трябва да приемат произволен SQL и да твърдят срещу дневника на извикванията, и throwingD1 за пътищата с грешка.

Свързан issue

Closes #325

Вид промяна

  • test / ci / build / chore — поддръжка

Мярката, а не твърдението

Преди да е пипнат кой да е тест: счупих всеки маркер в производствения SQL и пуснах съответния тест. 11 маркерни пътя в 9 файла останаха зелени срещу изпразнен резултат — тоест минаваха, без да проверяват нищо:

файл маркер
authorities FROM authority_totals
companies ORDER BY bidder_id
competition FROM sector_totals
contracts facet_counts
flows sector_totals
home bids_received = 1, JOIN
network FROM company_totals, FROM authority_totals WHERE authority_id
search sqlite_master
trend FROM sector_totals

Всичките отхвърлят вече. След мигрирането същата проверка дава 47 маркера → no route matched.

Гейтът scripts/check-fake-d1.mjs е написан пръв и беше червен от първия комит: 36 каста в 24 файла. Сега: 280 сканирани файла, нула каста извън @sigma/test-support.

Числата в issue-то са остарели

Issue-то е от преди #309/#313/#314/#323 и сочи файлове, които вече не съществуват (search.suggest.test.tsx, index.control-flow.test.ts) или вече нямат гол каст (agent.test.ts).

issue реално на main
файлове с ръчен двойник 29 29 ✓
от тях с каст към D1Database — 24 (36 каста)
as unknown as / гол as 55 / 33 21 / 15
връщане на { results: [] } при непозната заявка 42 14
файлове, които хвърлят 2 1

Посоката е вярна, редът на величините — не.

Какво излезе наяве, извън описаното в issue-то

  • regions.test.ts подаваше редове за области на sectorOptions — съвсем друга таблица. Стигаше до верния отговор само защото фикстурата няма поле division и филтърът изхвърляше всеки ред. Сега маршрутът казва all: [] и обяснява защо.
  • companies.test.ts регистрираше два маршрута за заявки, които никой тест в него не издава (getCompanyFacets не се упражнява там). А CSV потокът и списъчната заявка четат една и съща таблица, така че счупен ORDER BY в потока тихо падаше върху списъчния маршрут и връщаше нестранициран резултат. Разделени са по собствен маркер.
  • trend.test.ts е от същия клас като regions — старият двойник хранеше sectorOptions с редове за периоди, а [] излизаше по случайност. Новият маршрут го прави изричен избор с коментар.
  • readonly-d1.test.ts различаваше prepare: от exec: в дневника си. Сплескването им щеше да отнеме смисъла на теста — обвивка, която прати exec по пътя на prepare, издава същия текст. FakeD1Call носи via.

Три дефекта в самия помощник, намерени от употребата му

  • sql беше getter, тоест const { db, sql } = fake() хващаше празна снимка, която никога не се пълни. Точно тихата грешка, срещу която е целият PR.
  • throwingD1 хвърляше от prepare(). В D1 prepare() е мързелив — липсваща таблица излиза при изпълнение. Тест би „покривал" път за грешка, до който не стига.
  • throwingD1.bind() четеше calls.at(-1), тоест приписваше аргументите на последно подготвения оператор, а не на своя.

По-широко от искането — вашето решение

Комит 823c61f събира на едно място фасадата над истинско SQLite. Тя съществуваше в четири копия, а apps/etl стигаше до едно от тях през ../../../packages/ingest/src/test/ — релативен път през граница на workspace.

#325 се самоограничава до фалшивите двойници и оставя тестовете с истинско SQLite извън обхват. Тук е, защото „един каст навсякъде" е негово собствено условие за готовност, а два от последните четири каста бяха точно тези копия. Комитът се вади чисто, ако предпочитате.

Покритие

Пет от шестте workspace-а не мърдат. packages/ingest се качва с +0.71pp, защото d1-sqlite.ts напуска знаменателя му, напускайки workspace-а — резултатът, който #254 гони с поименен списък за изключване, постигнат по устройство.

Седми ключ в coverage-baseline.json беше неизбежен: findTestWorkspaces във check-coverage.mjs се проваля затворено за workspace със test скрипт без запис. @sigma/test-support: 100% редове, 100% клонове, 44 теста. Покриването на фасадата в новия ѝ дом опипа случай, който досега не се тестваше никъде: batch() се връща назад, когато един оператор се провали.

След прегледа

Четири комита отгоре, всеки с червен тест преди поправката (@nikimilenkov, @lyubomir-bozhinov):

  • batch() изобщо не маршрутизираше — записваше и връщаше синтетичен успех. Точно тихото зелено, срещу което е PR-ът, и то на входната точка, която производствените пътища за писане ползват изключително. exec() имаше същата дупка един метод по-горе. И двата вече търсят оператора и хвърлят, ако не е регистриран. Питат само дали е познат, не за конкретна форма на отговора — това остава разликата спрямо all()/first()/run().
  • csv-export.test.ts беше единственото място, където миграцията разхлаби проверка. Измерено: обвиване на производствения оператор оставя 34/34 зелени. Равенството се върна вътре в маршрута, през call.sql.
  • Гейтът се заобикаляше с един ред — type DBAlias = D1Database и после каст към псевдонима. Второ минаване хваща даването на второ име на типа извън allowlist-а, без да пипа обикновените анотации. И SCAN_ROOTS беше частен, тоест махането на 'apps' оставяше самотеста 12/12 зелен. Изнесен и пинат — мутирах и двете оси, за да видя, че вече гърмят.
  • results/meta по трите места, където кастът към D1Database криеше липсващ ключ.

Не пипнах пазача в eop.test.ts: местенето му от prepare() към изпълнение е по-вярно на D1, където prepare() е мързелив.

Уговорки

  • Съвпадението по маркер си остава по подниз, така че заявка може да падне от специфичен върху по-общ маршрут в същия набор. Отпада подразбиращото се пропадане към празнота, не всяко пропадане.
  • Два маркера оцеляват при чупене — FROM <rollup> в authorities и companies. Това е ограничение на измерването, не на тестовете: FROM ${src.from} се сглобява по време на изпълнение, така че литералът го няма в кода. Чупенето на самата стойност from: се хваща и от двата.

Припокриване с #254

#254 е отворен и пренаписва 19 от 29-те файла. Който влезе втори, го чака съществен rebase. Ако предпочитате #254 да мине пръв, кажете — ребейзвам този отгоре.

Как е тествано

pnpm lint · pnpm typecheck · pnpm test -- --coverage
pnpm check:coverage · pnpm check:docs · pnpm check:fake-d1

Всичко зелено локално. db 487, web 493, ingest 84, etl 20, test-support 53, shared 45, config 10.

Клонът мина и през пълния CI на форка (ydimitrof/sigma#4) — 5 от 5 проверки зелени: check, test, cacbg, semgrep, coverage-comment.

Чеклист

  • Комитите следват conventional commits и нямат Co-Authored-By: trailer към агент
  • PR-ът е с един логически обхват и е от форк към midt-bg/sigma:main
  • pnpm typecheck минава
  • pnpm test минава
  • pnpm lint е чисто
  • Няма комитнати тайни, .env* или .dev.vars
  • README.md е обновен със седмия пакет; docs/ не се нуждае от промяна

@nikimilenkov nikimilenkov left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Прегледах този PR на HEAD b6667f4 по строгия протокол: четири измерения (сигурност, коректност на помощника, вярност на миграцията по всичките 29 файла, гейт + хигиена на workspace-а), собствена мутационна батерия върху помощника и гейта, и повторение на измерванията от описанието — не на доверие, а на машина.

Първо честното: това е образцово построен PR. Гейтът е написан пръв и е бил червен от първия commit (проверих еквивалентно: върху main преброявам 37 каста в 25 файла — вашите 36/24 плюс едно споменаване в коментар, което blanking-ът на гейта правилно не брои); измерването на празните тестове е възпроизводимо; трите дефекта, които сам сте намерили в помощника, са реално поправени и всеки има именуван регресионен тест; уговорките в описанието са честни и — както се убедих лично — точни. Двата ми първи опита да „опровергая" измерването се оказаха точно двете разкрити ограничения (сглобяваният FROM ${src.from} и заявка, маршрутизирана по друг дискриминиращ маркер), което е добър знак за описанието, не лош.

Възпроизведох централното твърдение двустранно върху home/listSingleOfferContracts с мутация, която не променя семантиката на SQL-а (bids_received = 1 → bids_received=1): на main — 6/6 зелени срещу счупен маркер (тестът минава, без да проверява нищо); на този клон — fake D1: no route matched this all() query. Точно обещаното.

Миграцията е вярна: по диффа на всеки от 29-те файла нито едно твърдение не е отслабено, фикстурите не променят какво се доказва, а на няколко места (details, flows, cohort, sectors) новите маршрути са по-строги от старите произволно-отговарящи двойници. Трите разкрити семантични корекции (regions, companies, readonly-d1 via) са каквото твърдят.

Присъдата ми (съветодателна): одобрение, щом една-единствена находка се затвори — batch() трябва да мине през същия договор като всичко останало. Всичко друго по-долу е за втвърдяване по преценка.

Находката, която бих искал затворена преди сливане

batch() изобщо не маршрутизира — винаги връща синтетичен успех (fake-d1.ts:174-185). Възпроизведох директно: fakeD1 с една регистрирана run: пътека, db.batch([...]) с оператор към несъществуваща, нерегистрирана таблица — зелено, без хвърляне, синтетичен успех и за двата оператора. Това е точно тихото „минава срещу нищо", което целият PR съществува да убие, при това на входната точка, която производственият код за писане ползва изключително (staging.ts, refresh.ts, fx.ts — всичките пишат само през batch(), никога през prepare().run()). Днес е спящо — нито един тест не съчетава run: маршрут с batch(), а двата тестови файла по пътищата за писане правилно ползват истинско SQLite — но първият бъдещ тест, който посегне към fakeD1 за batch-базирана функция, ще получи фалшиво зелено от споделения договор, на който вече всички разчитат. Поправката е малка: прекарайте всеки оператор от batch() през същото matches/responder търсене (с хвърляне при пропуск, освен при lenient), плюс един тест, който пина, че непозната заявка в batch хвърля. Алтернативата — изрично да забраните batch() върху fakeD1 („ползвай d1FromSqlite") — също затваря вратата, стига да е грешка, а не мълчание.

Гейтът — две втвърдявания по преценка

  1. Заобикаля се с един ред индиректност. Пробвах емпирично: файл с import type { D1Database as DB } + as DB и с type DBAlias = D1Database + as unknown as DBAlias минава гейта чисто („280 файла сканирани" — файлът е сканиран, пропускът е в шаблона, не в обхода). Самотестът покрива границите на идентификатора и коментарите/низовете, но не и този клас. Едно второ минаване, което маркира type X = D1Database и преименуван import извън allowlist-а, го затваря — заедно с по два фикстурни случая в самотеста.
  2. Самотестът не пина обхожданите директории. Мутирах скрипта двукратно: отслабен шаблон (без голото as D1Database) — самотестът гърми с 3 провала, точно както трябва; махнато 'apps' от SCAN_ROOTS — 12/12 зелени. Тих регрес на обхвата (тестовете на web и etl отпадат от прилагането) би минал незабелязано. Един тест, който твърди какво се обхожда, пина втората ос.

Извън това гейтът е стабилен: blanking-ът на коментари/низове/regex-литерали е коректен, allowlist-ът е по име и гърми при остарял запис (проверих върху main — отказва точно както е замислено), CI стъпката е последователна със съседните („самотест първо, после гейт").

Дребни

  • run()/batch() на двойника връщат резултат без results (а фасадата и без meta), докато истинският D1Result тип винаги ги носи — маскирано от вътрешния каст. Спящо (никой производствен код не чете резултата от run()/batch() днес), но първият читател ще получи undefined от двойника и [] от истинското D1. Едно results: [] на трите места го изравнява.
  • Остарял коментар на fake-d1.ts:9 сочи packages/ingest/src/test/d1-sqlite.ts — път, който същият PR премести. Иронично предвид собствената философия на PR-а за остарелите записи, които никой не препрочита.
  • trend.test.ts е четвърта корекция от класа на regions — старият двойник е хранел sectorOptions с редове за периоди и [] е излизало по случайност; новият маршрут го прави изричен избор с коментар. Коректно поправено в кода, само липсва от разкритието в описанието — казвам го за протокола, не като недостатък.
  • Две пределно теоретични отслабвания на пазачите, за протокола: eop.test.ts мести пазача на raw_contracts от prepare() към изпълнение (подготвен, но неизпълнен оператор вече не се лови — на практика нищо не чете така); csv-export.test.ts сменя точно съвпадение с подниз върху пълния текст на оператора (надниз би минал — при маркер, който е целият оператор, е почти невъзможно).

Какво изпълних и проверих лично

  • Петте пакета на този HEAD: db 486/487 (единственият провал е познатият env артефакт — ship-domain.test.ts извиква pnpm без пълен път; файлът не е пипан от PR-а), web 493/493, ingest 84/84, etl 20/20, test-support 44/44. check:fake-d1:test + check:fake-d1 зелени; check:docs зелен; tsc --noEmit чист в новия пакет и в двата главни консуматора.
  • Двустранната мутация на home (по-горе), пробата за batch(), двете проби за заобикаляне на гейта, двете мутации на самия гейт скрипт — всичко върнато, дървото чисто след всяка стъпка.
  • Диффът на coverage-baseline.json е точно както е описан: само седмият ключ (100/100 — имайте предвид, че с глобалния толеранс това е най-стегнатият ratchet в репото; всяко бъдещо недотествано добавяне в test-support гърми веднага, което приемам за замислено); +0.71pp на ingest е предимство над базата, не промяна на базата.
  • Липсата на декларирани зависимости в packages/test-support е съществуващата конвенция на репото (shared/db/etl правят същото; node:sqlite, не външен пакет) — не е нов риск.

Координация

По въпроса от описанието за #254: това е решение на поддържащите; отбелязвам само, че тази промяна затваря условието за готовност на #325 по устройство, а #254 пренаписва 19 от същите файлове — редът на сливане определя кой поема rebase-а.

Едно затваряне на batch() и от моя страна това е одобрение — измерването-преди-твърдението, червеният-пръв гейт и честно разкритите ограничения са точно как трябва да изглежда тестова инфраструктура в слой, който публикува твърдения за реални хора.

@lyubomir-bozhinov

Copy link
Copy Markdown
Collaborator

Рефакторът е верен и всъщност вдига чувствителността на пакета — throw-on-unmatched по подразбиране, специфични маркери вместо catch-all fall-through (competition, companies, regions, integrity, details), а via/binds са по-строги от старите .args. Проверих fake-а (fake-d1.ts), guard скрипта (check-fake-d1) и high-stakes конверсиите (readonly-d1, readonly-corpus, related-persons) при HEAD b6667f42 — нито едно security/libel твърдение не е изгубено, coverage-baseline само добавя test-support (100/100), без да сваля под. Едно Minor изключение, което върви срещу тезата на самия PR:

apps/web/app/lib/csv-export.test.ts (fakeDb) — точната проверка стана маркер:

// преди: равенство върху изпълнения SQL
expect(sql).toBe('SELECT refreshed_at FROM home_totals WHERE id = 1');
// сега: същият низ е само `when` маркер → match по sql.includes(), не по ===
when: 'SELECT refreshed_at FROM home_totals WHERE id = 1',

fakeD1 мачва маркерите по съдържане, така че тук === става „съдържа". Маркерът е целият очакван SQL, затова промяна ВЪТРЕ в клаузата пак хвърля — но допълнително обвиване (префикс/суфикс около нея) вече минава, където старият toBe би паднал. Единственото място, където конверсия разхлабва проверка, вместо да я стяга.

Фикс без да въвеждаш нов handle — callback формата на маршрута получава call.sql, така че точността се връща вътре в самия route:

{
  when: 'SELECT refreshed_at FROM home_totals WHERE id = 1',
  first: (call) => {
    expect(call.sql).toBe('SELECT refreshed_at FROM home_totals WHERE id = 1');
    return refreshedAt === undefined ? null : { refreshed_at: refreshedAt };
  },
}

@ydimitrof

Copy link
Copy Markdown
Contributor Author

Благодаря и на двамата — прегледите бяха по-полезни от одобрение. Четири комита (a956606..9b1bad1), всеки с червен тест преди поправката.

batch() — блокиращата находка

@nikimilenkov е прав и възпроизведох поведението точно: batch() записваше и връщаше синтетичен успех, без изобщо да поглежда маршрутите.

exec() имаше същата дупка един метод по-горе — записваше и връщаше { count: 0, duration: 0 }, каквото и да получи. Никой не я посочи, но е от същия клас, затова се затваря в същия комит.

И двата вече търсят оператора и хвърлят при пропуск (освен при lenient). Питат само дали е регистриран изобщо — не за конкретна форма на отговора, както правят all()/first()/run() — защото batch() и exec() не искат определена форма; искат само SQL-ът да е познат. Отделен registered() до responder(), с коментар защо са различни. Шест нови теста: непозната заявка в batch хвърля; хвърля и когато по-ранен оператор в същия batch е съвпаднал; run: маршрутът се вика за всеки оператор със своите binds; all: маршрут сервира редове на batch-нат SELECT; recordingD1 продължава да пропуска всичко; непознат exec() хвърля.

Проверих и че никой съществуващ тест не разчита на старото поведение — единственият exec() върху двойник е в readonly-d1.test.ts през recordingD1, който е lenient.

csv-export.test.ts — единственото разхлабване

@lyubomir-bozhinov и @nikimilenkov стигнаха до едно и също, независимо. Измерих го, вместо да го приема: обвих производствения оператор в csv-export.ts:189 (… WHERE id = 1 LIMIT 1) — 34/34 зелени. Старият toBe би паднал.

Взех предложената форма — равенството се връща вътре в маршрута, през call.sql, без нов handle. С мутацията на място: 7 провала. Без нея: 34/34.

Гейтът — и двете оси

Възпроизведох заобикалянето и го затворих: второ минаване маркира даването на второ име на типа извън allowlist-а (type X = D1Database, преименуван импорт, interface X extends D1Database), защото няма честна причина тест да се нуждае от такова. Обикновените анотации остават недокоснати — db: D1Database, поле в Env — и това е пинато с три отделни случая, за да не стане гейт, който плаши напразно и накрая го отслабват.

SCAN_ROOTS е изнесен и пинат. Мутирах двете оси, за да видя, че вече гърмят: махането на 'apps' дава 18/19 (беше 12/12 зелени), отслабването на новия шаблон — 16/19.

Гейтът остава ok — 280 файла, нула каста извън двата allowlist-нати, а върху засадения байпас изхвърля packages/db/…:1: type DBAlias = D1Database и излиза с 1.

Дребните

  • results/meta са добавени на трите места; фасадата получи един D1Shape тип, за да не се разминат пак. Тестът е toEqual, не toMatchObject, така че бъдещо изпускане на ключ пада.
  • Остарелият коментар в fake-d1.ts:9 вече не сочи преместения път.
  • Двойният запис в batch() е с коментар на самото място: calls е дневник на входни точки, не на различни оператори.
  • all() минава през === undefined вместо falsy проверка, а коментарът казва защо взима маршрута, а не отговора: meta принадлежи на същия маршрут. (Оттам и асиметрията, която @cefothe отбеляза на форка.)
  • trend.test.ts — прав сте, четвърта корекция от класа на regions, и я нямаше в описанието. Добавена е.

Какво съзнателно не пипнах

Пазача в eop.test.ts. Местенето му от prepare() към изпълнение е по-вярно на D1, където prepare() е мързелив и липсваща таблица излиза при изпълнение — точно грешката, която сам открих в throwingD1 по-рано в този PR. Връщането му на prepare() би пинало поведение, което D1 няма.

Проверка

pnpm lint · pnpm typecheck (8/8) · pnpm test -- --coverage (7/7) · check:coverage · check:docs · check:fake-d1:test (19/19) + check:fake-d1 — всичко зелено. db 487, web 493, ingest 84, etl 20, test-support 53 (беше 44), shared 45, config 10. @sigma/test-support остава 100/100; нито един базов праг не пада.

За #254: решението кой влиза пръв е ваше. Ако предпочитате той, ребейзвам този отгоре.

@nikimilenkov nikimilenkov left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Прегледах четирите нови commit-а до 9b1bad1 със същата дисциплина: повторих всяка своя проба и мутация срещу новия HEAD, а после пуснах и втори пълен строг кръг само върху делтата (253 нови реда), за да проверя дали поправките не внасят нещо ново. Всичко от прегледа е адресирано; от моя страна това е одобрение (съветодателно; решението е на поддържащите). Вторият кръг намери остатъчни втвърдявания по преценка — нищо блокиращо; изредени са след проверките.

Повторени проби — всички дават верния резултат

Проба При прегледа На 9b1bad1
непознат оператор в batch() синтетичен успех, без хвърляне хвърля no route matched — включително когато по-ранен оператор в същия batch е съвпаднал
маршрутизиран batch n/a run: callback-ът се вика по веднъж на оператор със своите binds (проверих с два bind-а — [[1],[2]])
непознат exec() синтетичен {count: 0} (същият клас, намерен от вас) хвърля
recordingD1 + batch пропуска всичко продължава да пропуска — lenient режимът е незасегнат
файл с import type { D1Database as DB } + type DBAlias = D1Database минаваше гейта чисто гейтът гърми и назовава двата реда с действено съобщение; чистото дърво остава ok без фалшиви аларми по обикновените анотации
махнато 'apps' от SCAN_ROOTS самотестът минаваше 12/12 1 провал
отслабен alias шаблон n/a 2 провала
отслабен основен шаблон 3 провала 3 провала — остава пинат
обвит производствен SQL в csv-export.ts:189 (… LIMIT 1) 34/34 зелени (разхлабването, което открихме независимо с @lyubomir-bozhinov) 7 провала — точността се върна вътре в маршрута през call.sql, точно предложената форма

Дребните: D1Shape<T> обединява results/success/meta и тестът е toEqual, тъй че изпуснат ключ пада ✓; коментарът на fake-d1.ts:9 вече сочи фасадата до себе си ✓; разкритието за trend.test.ts е в описанието ✓. Двете решения с аргументи приемам и двете: пазачът в eop.test.ts на изпълнение е по-верен на мързеливия prepare() на D1 (същата грешка, която сам поправихте в throwingD1), а отделният registered() с коментара защо batch()/exec() не искат форма е вярното разграничение по замисъл — с една уговорка от втория кръг по-долу.

Втори кръг върху делтата — остатъчни, по преценка

  1. registered() е сляп за метода, и това е измеримо. Възпроизведох: fakeD1([{ when: 'FROM staging', all: [ред] }]) и batch-нат DELETE FROM staging WHERE id = 1 — минава и връща редовете на read маршрута като резултат на write оператора, докато същият SQL през prepare().run() правилно хвърля. Припокриване на маркер между SELECT маршрут и write оператор е реалистично (FROM staging е подниз на DELETE FROM staging). По-тясно от оригиналната дупка — иска нещастно съвпадение, не нищо — но е точно мястото, където договорът още пропуска. Стягане, което пази замисъла ви: batch да предпочита run:, да пада към all: за batch-нат SELECT, а съвпадение само по first: (или само по маркер на чужд метод) да се брои за пропуск.
  2. ALIAS шаблонът затваря три конкретни изписвания, а коментарът обещава класа. Възпроизведох: type Evade = D1Database & {} + каст минава гейта чисто (файлът се сканира — 281 файла). Същото важи за Pick<D1Database, …>, import('…').D1Database, class X implements D1Database и extends Foo, D1Database (не на първа позиция). Позицията на репото за content-mode гейтове е честен-оператор, не противник — легитимна е; но тогава коментарът „give the type a second name … and the gate treats that as the offence" надобещава. Или разширете първата алтернатива до „RHS, който споменава D1Database" + implements/списъчен extends, или смекчете коментара до best-effort.
  3. run() изпуска meta, което batch() вече резолвира — възпроизведох: meta: {changes: 5} излиза {} през prepare().run() и {changes: 5} през batch() за същия оператор. Никой не чете run-meta днес; един ред го изравнява със заглавието на собствения ви commit („същия договор").
  4. За протокола, дребни: провален строг batch оставя частични странични ефекти (истинското D1 е транзакционно; фасадата отива до rollback — за тест, който твърди състояние след провален batch, разликата ще личи); при два маршрута с застъпващи се маркери batch-нат SELECT може да отговори от run:-only маршрута с празни редове, докато .all() на същия SQL отговаря от втория; alias находките излизат с exit 1 преди cast находките, тъй че файл с двете иска два пуска.

Нищо от горното не е от класа на затвореното: оригиналът беше „зелено срещу нищо при нула маршрути", а остатъкът иска конкретно съвпадение на маркери. Затова е по преценка, не условие.

Изпълнено тук

Петте пакета: db 486/487 (единственият провал е познатият env артефакт — тестът извиква pnpm без пълен път; файлът не е пипан), web 493/493, ingest 84/84, etl 20/20, test-support 53/53. check:fake-d1:test 19/19 + check:fake-d1 зелени; typecheck чист. Дървото чисто след всяка проба. Една бележка за протокола: първата ми проба срещу новия batch() подаде run: като статичен обект и получи route.run is not a function — мой зле типизиран вход, който typecheck лови в реален тест.

Образцов цикъл: блокиращата находка затворена с шест теста, съседната дупка в exec() намерена и затворена от самия автор, разхлабването в csv-export премерено с мутация преди да се приеме, и едно съзнателно не-пипане, защитено с по-добрия аргумент. Благодаря — и на @lyubomir-bozhinov за независимото стигане до същото място.

@todorkolev

Copy link
Copy Markdown
Collaborator

Клонът изостава от main и изравняването не минава автоматично. Опитах merge на main в него, за да го
изравня вместо теб, но спрях и го отмених - единият конфликт иска решение, което е твое, не мое.

Три конфликта:

Тривиални два. apps/web/package.json - @react-router/dev 7.18.0 срещу 7.18.2, плюс новата
зависимост @sigma/test-support; решението е 7.18.2 И новата зависимост. pnpm-lock.yaml - същото, един
блок, регенерира се с pnpm install.

Същинският е packages/db/src/queries/related-persons.test.ts. Докато клонът е стоял, main разшири
ръчния двойник там през #309 и #312 с две неща, които твоят споделен fakeD1 в сегашния си вид не
изразява:

  1. contracts: именувано пространство. Страницата на дружество връзва същия ЕИК за два различни
    SQL-а - COMPANY_SQL и EIK_CONTRACTS_SQL. Затова четенията на договори са строго под ключ
    contracts:<eik> и никога не падат обратно към редовете на обхвата. Без разделянето редовете на
    връзки се четат като договори, тихо.
  2. Записване на binds. Двойникът трупа { sql, key } при всяко връзване, за да може тест да твърди
    колко четения е направило зареждането - например че дедупликацията по ЕИК наистина спестява четене.

Твоят маршрут го свежда до едно FROM interest_links il, което отговаря на всичко. Тоест изхвърля точно
предпазителя, който #309 добави, и то по начин, при който тестовете пак ще минават.

Затова го оставям на теб: как споделеният fakeD1 да изрази тези две неща е решение за неговия API, а не
конфликт за разрешаване напосоки. Ако го разреша аз и сгреша, ще изглежда зелено и няма да си личи.

Останалото по PR-а е наред - CI беше зелен преди main да мръдне, нишки за разрешаване няма.

За контекст защо main мръдна толкова: днес влязоха #313, #323, #332, #333, #312 и #329 - конвейерът за
свързани лица тръгна докрай и sigma-stage.midt.bg/conflicts вече показва 277 публикувани връзки срещу
103 преди. Твоят #323 беше този, който отпуши регистъра.

midt-bg#325: 24 test files hold 36 `as D1Database` casts, one hand-rolled double each.
Every one dispatches on `sql.includes('…')` and falls through to `{ results: [] }`
when no marker matches, so renaming a CTE or reordering a JOIN leaves the test
green against emptiness — asserting nothing. Only details.test.ts throws today.

This is the acceptance test for that work, written before the work: outside an
explicit allowlist, no file under apps/ or packages/ may type a value as a
D1Database. It is red now (24 files, 36 casts) and goes green when the last
double moves to the shared helper.

The allowlist is by name, never a directory glob — the argument the midt-bg#254 review
already made about the coverage exclusion list. A glob lets a new double leave
the gate by where it sits; a named entry means someone had to add it, which is
reviewable. A stale entry is an error rather than a no-op, so a renamed double
cannot leave the gate widened by a line nobody reads again. That fail-closed
branch is what fires right now, since the helper does not exist yet.

Matching is over blanked source — comments, strings and regex literals removed,
byte positions kept — so a comment describing the old design is not a finding.
`as unknown as D1Database` is matched before `as D1Database` because the short
spelling is a suffix of the long one and a naive pattern counts one cast twice.
`satisfies` is covered too: it is the only other operator that types a literal.

Self-test is mutation-checked — dropping the `as unknown as` alternative, the
`satisfies` alternative, the comment blanking, the trailing word boundary, or
the stale-entry check each kills exactly one named test, and no others.

scripts-test.yml needs no edit: its lane already globs scripts/*.test.mjs.
The helper midt-bg#325 asks for, as its own private workspace. `@sigma/db` exports only
`.`, and neither apps/etl nor packages/ingest depends on it, so putting the
double under db/src/test/ would have meant a subpath export plus two new
workspace deps. A separate package also sits outside all six measured
workspaces, so it cannot enter their coverage denominators by construction —
stronger than the by-name vitest.shared.ts exclusion the issue proposes, and it
leaves that file (which midt-bg#254 rewrites) untouched.

A route is a marker set and a response; every marker must appear in the SQL, and
the first matching route wins so a specific route can precede a general one.
Unmatched throws, naming the offending statement and every registered marker.
`{ onUnmatched: 'empty' }` buys emptiness back, at the call site, in writing.

Three entry points, one core: fakeD1 for query tests, recordingD1 for the tests
of a *wrapper* over D1 (readonlyD1) that must accept arbitrary SQL and assert on
a call log, throwingD1 for the error paths.

Two design notes worth keeping:

  - `first` is typed `object | null | (call) => object | null`, not `unknown`.
    A top type absorbs the union and the callback form silently loses its
    parameter type — tsc caught it. A D1 row is an object or nothing anyway.
  - No pagination feature. Keyset slicing is already `all: (call) =>
    rows.filter(r => r.id > call.binds.at(-2))`, which is what the doubles in
    companies.test.ts do by hand today.

Tests written before the code, behaviour by behaviour, and mutation-checked:
never throwing on an unmatched all() or first(), matching a route that answers a
different method, `some` for `every` over the markers, last-match instead of
first, no truncation, dropping the marker list from the message, discarding
bind() arguments, not recording prepare() or batch(), and ignoring throwingD1's
supplied error — eleven mutations, each killing a specific named test.

The `?? []` fallback in all() went away rather than getting a test: the route
lookup already guarantees the response is defined, so it was unreachable.
Returning the response instead of the route also keeps `first: null` — a route
meaning "no such row" — distinct from no route at all.

Seventh key in coverage-baseline.json: check-coverage's findTestWorkspaces fails
closed on a workspace that has a test script without a baseline entry, and this
one should be measured. 100% lines, 100% branches. The six existing workspaces
are untouched — they were already reading above their baselines before this
branch, which is pre-existing drift and not for a test refactor to ratchet.
…execution

Two defects the first migration batch walked straight into.

`sql` was a getter over `calls`, so `const { db, sql } = fake()` — the natural
way to use it, and what flows.test.ts and authorities.test.ts already wrote
against their hand-rolled spies — captured an empty snapshot that never filled
in. Every later assertion then read nothing and passed for the wrong reason,
which is the exact failure this helper exists to remove. It is a live array kept
in step with `calls` now, and a test pins the destructured form.

throwingD1 threw from prepare(). D1's prepare() is lazy and never touches the
database: a missing table surfaces on all()/first()/run(). A double that failed
earlier would let a test claim it covers an error path it never reaches — and
related-persons.test.ts, whose whole point is that an un-migrated environment
degrades instead of 500ing, hand-rolled a double that threw at execution for
exactly that reason. It now rejects from the three execution methods and records
the statement that failed, so the offending SQL stays inspectable.
Sixteen files in packages/db/src/queries, each of which built its own D1 double
that dispatched on `sql.includes('…')` and fell through to no rows. Fixtures and
assertions are unchanged — only the double moves.

Measured before touching anything, by breaking each marker in the production SQL
and running the test: ELEVEN marker paths across nine files stayed green against
an emptied result. authorities (FROM authority_totals), companies (ORDER BY
bidder_id), competition and trend and flows (FROM sector_totals), contracts
(facet_counts), home (bids_received = 1, JOIN), network (FROM company_totals,
FROM authority_totals WHERE authority_id), search (sqlite_master). Every one of
them now rejects with the marker set it was looking for.

Three things the migration turned up that were not in the issue:

  - regions.test.ts served *region* rows to sectorOptions, which asks a
    completely different table. It reached the same answer only because
    sectorOptions reads r.division, the region fixture has no such field, and
    the filter dropped every row. The route says `all: []` now, and says why.
  - companies.test.ts registered two facet routes for queries no test in it ever
    issues — getCompanyFacets is not exercised there. Dropped rather than kept
    as decoration.
  - companies' CSV stream and list query both read company_totals, so breaking
    the stream's ORDER BY quietly fell through to the list route and returned an
    unpaginated page. They are separated by their own markers now (ORDER BY
    bidder_id vs AS sort_value), and breaking either one throws.

Two markers still survive being broken — authorities' and companies' `FROM
<rollup>`. That is the harness, not the tests: `FROM ${src.from}` is composed at
runtime, so the literal never appears in the source to be mutated. Mutating the
`from:` value itself is caught by both.

Route matching is still substring-based, so a query can fall from a specific
route to a more general one in the same set. What is gone is the *default*
fall-through to emptiness — an unrouted query throws.

packages/db: 487 tests pass; coverage unmoved.
d1FromSqlite lived in packages/ingest/src/test/, and packages/db had two
byte-identical re-implementations of it (contracts-filter-sql, value-base-sql,
differing only in a local variable name) while apps/etl reached the original
through ../../../packages/ingest/src/test/d1-sqlite — a relative path across a
workspace boundary, which is what a missing shared home looks like.

It moves to @sigma/test-support beside the fake. The two are different tools and
stay different: this one runs the real SQL against a real node:sqlite database
where SQL semantics are what is under test; fakeD1 is for the TypeScript logic
around a query. Now they at least live in the same place, and the gate's
allowlist names one package instead of two.

Slightly wider than midt-bg#325 asked — the issue scopes itself to the fake doubles and
puts real-SQLite tests out of scope. It is here because "one cast everywhere" is
one of its own done-when boxes, and two of the four remaining casts were these
copies. Reviewer's call; it lifts out cleanly.

Side effect worth noting: d1-sqlite.ts leaves packages/ingest's coverage
denominator by leaving the workspace, which is the outcome midt-bg#254 wanted from a
by-name exclusion, reached by construction instead.

db 487, ingest 84, etl 20 — all pass.
readonly-d1 and readonly-corpus test a *wrapper* over D1, not a query: what
matters is which statements reach the handle underneath, not what comes back.
Marker dispatch is the wrong shape for that, so both use recordingD1 — answers
anything, records everything — with `when: []`, a route that constrains nothing.

Two things came out of it, both in the helper:

  - `when: []` matching every query was already true (every() over no markers),
    but undocumented and unpinned. Now both.
  - readonly-d1's hand-rolled log tagged its entries `prepare:` / `exec:`, and
    flattening that into plain SQL would have cost the test its point: a wrapper
    that sent an exec down the prepare path emits identical text, and the
    assertion could no longer tell. FakeD1Call carries `via` now, and the
    corpus's zero-proxy row survives as the response to a constraint-free route.

readonly-corpus also dropped a `raw()` no production path calls.

packages/db: 487 tests pass.
Three doubles. The integrity gate's fake dispatched on eleven markers and fell
through to no rows; it now names all eleven as routes and rejects anything else.
Its local builder was called `fakeD1`, which is the shared helper's name, so it
becomes `servedD1` — which is what it models anyway: a served D1 after
precompute, not any old one.

eop.test.ts also passed `{} as D1Database` twice, for paths that fail before
they reach the database. `fakeD1([])` states that instead of implying it: a
route-less double rejects any query, so if one of those paths ever did reach D1
the test would say so rather than throwing an incidental TypeError on an empty
object.

The freshness double's guard survives as a route that throws its own message —
"raw staging should not be read for planning" is a claim worth keeping in the
test, rather than degrading to the generic no-route error.

apps/etl: 20 tests pass.
…s new home

apps/web's assistant tests need meta.rows_read and meta.total_attempts: they
drive the rows-read budget that keeps a retried full scan from under-billing the
Denial-of-Wallet limit (midt-bg#122, review midt-bg#80). Flattening that to a fixed empty meta
would have quietly removed what those two tests assert, so a route can declare
its own meta. Default stays `{}`.

Moving d1-sqlite.ts here left it with no tests of its own — its callers live in
db, ingest and etl, and none of them count toward this workspace. The ratchet
caught it at 81% and it is covered directly now, including the case nothing
tested anywhere before: batch() rolls back when one statement fails. A
half-applied batch would leave a fixture in a state no production path can
reach, and whoever met it would be debugging a ghost.

While covering it, throwingD1's bind() read `calls.at(-1)` — so binding statement
A after preparing B recorded the arguments against B. Same statement-independence
bug fakeD1 already had a test against; it captures its own record now, and so
does the test.

100% lines, 100% branches, 44 tests.
assistant/tools built a double whose only real job was carrying meta; it now
declares that meta on a route. csv-export asserted the expected SQL *inside* its
fake — that assertion becomes the route's own marker, so a query that no longer
matches rejects and names both the statement and what was expected, instead of
failing an inline expect from inside a stub.

With these two the gate from the first commit goes green: 279 files scanned, no
D1Database cast outside @sigma/test-support. It opened at 36 casts in 24 files.

apps/web: 493 tests pass.
batch() recorded each statement and returned a synthetic success without ever
consulting the routes, so a batch of unregistered SQL passed against nothing —
the silent green this helper exists to kill, on the one entry point the write
paths use exclusively (staging, refresh, fx never call prepare().run()).
exec() had the identical hole one method up.

Both now look the statement up. They ask only whether it is registered at all,
not for a particular response shape the way all()/first()/run() do, and throw
naming the SQL and every marker when it is not.

Also: run() and batch() carry the `results` key a real D1Result always has —
the cast to D1Database was hiding its absence; the header no longer points at
the facade's pre-move path; and the second batch() record is documented as a
log of entry points rather than a double count.
all() returned no `meta`, run() neither `meta` nor `results`, batch() no
`results`. The cast to D1Database hid every one of them: the first caller to
read one would get `undefined` from the facade where real D1 hands back `[]`
or `{}`. One D1Shape type spells out all three keys.
The hand-rolled double asserted `expect(sql).toBe(...)` on the whole statement.
Migrating turned that string into a `when` marker, and markers match by
substring — so the one place the refactor loosened a check rather than
tightening it. Measured: wrapping the production statement leaves all 34 tests
green. The equality moves inside the route, where the callback sees `call.sql`.
Two ways past the gate, both reproduced. `type DBAlias = D1Database` and then
`as unknown as DBAlias` leaves no D1Database token for the pattern to find; a
renamed type import does the same. A second pass treats giving the type another
name outside the allowlist as the offence, while leaving ordinary annotations
(`db: D1Database`, a field on an Env type) alone.

And SCAN_ROOTS was module-private, so deleting 'apps' from it left the self-test
12/12 green while web and etl dropped out of enforcement. Exported and pinned: a
pattern applied to half the repo is a gate that passes while enforcing nothing.
@ydimitrof
ydimitrof force-pushed the test/shared-fake-d1-helper branch from 9b1bad1 to 3009dc3 Compare August 26, 2026 06:25
… markers

Markers alone are blind to the method, and it is measurable: a route declaring
only `all:` answered a batched `DELETE FROM staging` with its rows, because
`FROM staging` is a substring of the write. The same SQL through prepare().run()
threw. Narrower than an unrouted batch, but the same silent pass.

A batched write now needs a `run:` route and a batched SELECT an `all:` one;
neither settles for the other, and a SELECT no longer fires a write effect it
happens to match. A write still serves rows when it has them, for RETURNING.
exec() asks for `run:` too — it hands back no rows, so nothing else means
anything to it. That retires `registered()`: every entry point is method-aware.

Reading is decided by the leading keyword, so a `WITH … INSERT` reads as a
SELECT here. That costs a false rejection, never a false pass.

Also: run() carries the route's meta, which batch() already resolved for the
same statement, and batch() documents that it is not transactional — real D1
and the d1-sqlite.ts facade roll back, this does not.
The pattern closed three spellings while the comment promised the class. Four
more walked past it: `D1Database & {}`, `Pick<D1Database, …>`, a namespaced
`import('…').D1Database`, and a heritage list naming it off the first position.

The rule is now positional — the mention must sit right of `=`, or inside an
intersection, union, type argument or namespace, never where a parameter or a
field goes. `type Env = { DB: D1Database }` and the conditional type in
readonly-corpus.test.ts stay clean, both pinned.

`implements` is deliberately out: ReadonlyD1 implements D1Database in
production, and TypeScript forces a complete implementation there, so it is no
shortcut to a stub. The comment now says best-effort and means it.
@ydimitrof

Copy link
Copy Markdown
Contributor Author

Клонът е ребейзнат върху main (f081ad3) и трите остатъчни находки от втория кръг са затворени — 3009dc3 за ребейза, a92f769 и 8080321 за находките. Всяка проба е възпроизведена, преди да бъде приета.

@todorkolev — конфликтът е решен така, както го описа

Благодаря, че спря. Прав беше да го оставиш: и двете неща, които изброи, се губеха при небрежно разрешаване, и то тихо.

contracts: пространството е запазено. Три маршрута вместо един: EIK_CONTRACTS_SQL и LINK_CONTRACTS_SQL преди обхватния, всеки с проверка за точно съвпадение вътре в маршрута (expect(call.sql).toBe(...)) — двойникът, който заместват, диспечираше на sql === EIK_CONTRACTS_SQL, а маркерите съвпадат по подниз, тъй че равенството трябваше да отиде някъде. Точно разхлабването, което @lyubomir-bozhinov хвана в csv-export; нямаше да го върна един комит след като го поправих.

Записването на връзванията е запазено, но през calls вместо binds. Първото ми разрешаване връщаше Object.defineProperty(...) as D1Database & { binds } — и собственият ми гейт го отхвърли на related-persons.test.ts:79. Вместо да си разширя allowlist-а, изхвърлих проекцията {sql, key} и изнесох живия дневник на самия двойник през Object.assign, който се типизира като сечението без никакво твърдение. Един тестов ред се смени. Поправката е сгъната в самия миграционен комит, не отгоре, тъй че никой комит в клона не внася каст, който гейтът би отхвърлил.

Немеханичността е проверена: чупенето на FROM interest_links il в производствения SQL сваля файла от 20/20 зелени на 15 провала с no route matched.

@nikimilenkov — трите остатъчни

1. registered() беше сляп за метода — и пропускаше запис. Възпроизведох дословно: all:-маршрут отговаря на batch-нат DELETE FROM staging WHERE id = 1 с [{"results":[{"id":1}],...}], докато същият SQL през prepare().run() хвърля.

Взех по-широката поправка, а не подредбата: всяка входна точка вече е метод-осъзната. Batch-нат запис иска run:, batch-нат SELECT иска all:, и нито един не се съгласява на другия; SELECT вече не пали ефект за запис, който случайно съвпада. Запис пак сервира редове, ако маршрутът има all: — INSERT … RETURNING си остава запис. exec() също иска run:, защото не връща редове и нищо друго не значи нещо за него. Това премахва registered() изцяло — тоест обръщам решение, което защитавах в предишния отговор; ти го показа измеримо грешно.

Четенето се решава по водещата дума, тъй че WITH … INSERT минава за SELECT. Струва фалшив отказ, никога фалшиво минаване — маршрутът, който тогава ще поиска, е именно този, който няма.

2. Шаблонът обещаваше класа. Засадих твоите изписвания: и петте минаваха чисто при 281 сканирани файла. Правилото вече е позиционно — споменаването трябва да стои вдясно от =, или в сечение, обединение, типов аргумент или пространство от имена, никога там, където стои параметър или поле.

implements съзнателно го няма, и това е промяна спрямо предложението ти: първият вариант гръмна върху ReadonlyD1 implements D1Database — производствената обвивка, която прави истинското нещо. TypeScript и без това изисква пълна имплементация там, тъй че не е евтин път до заглушка. Същото важи за extends вътре в условен тип: type Params<F> = F extends (db: D1Database, …) в readonly-corpus.test.ts е истински код и гърмеше. И двата случая са пинати като негативни. Коментарът вече казва best-effort и го мисли — съгласен съм с прочита ти, че надобещаваше.

3. run() изпускаше meta. Възпроизведено: {} през run(), {changes: 5} през batch() за същия маршрут. Изравнено; run() взима маршрута, не отговора, точно както all().

4. За протокола. Непрозрачността на batch() е документирана на самото място вместо променена: истинското D1 и фасадата d1-sqlite.ts се връщат назад, този двойник не. Тест, който твърди състояние след провален batch, е тест за фасадата. Останалите две бележки (реда на изходите на гейта, празни редове от run:-only маршрут) отпадат от метод-осъзнатото маршрутизиране или са въпрос на два пуска, което приемам.

Проверка на 8080321

pnpm lint · pnpm typecheck 8/8 · pnpm test -- --coverage 7/7 — db 490, web 521, ingest 88, etl 20, test-support 60, shared 45, config 10. check:fake-d1:test 22/22 + гейтът зелен; check:coverage без падане под база; check:docs зелен. @sigma/test-support остава 100% редове / 100% клонове.

Мутирах и двете нови оси: отслабването на alias правилото дава 5 провала, а засадените изписвания излизат поименно с exit 1.

@todorkolev
todorkolev merged commit 2f72e32 into midt-bg:main Aug 26, 2026
3 checks passed
lyubomir-bozhinov added a commit to lyubomir-bozhinov/sigma that referenced this pull request Aug 26, 2026
Adopts midt-bg#331's shared fake-D1 double and its check:fake-d1 gate across every
suite this branch adds, and midt-bg#334's CACBG corpus work.

- migrate all 71 hand-rolled D1 doubles (18 files) to @sigma/test-support:
  fakeD1 for query tests, recordingD1 for the two ingest wrappers whose SQL is
  generated rather than routed, throwingD1 for the run_sql error path. An
  unmatched query now throws instead of answering with no rows, which surfaced
  two silent gaps: getCompany/getAuthority never declared the listContracts
  panel reads, and eop/etl passed `{}` as a binding that could never fail.
- keep the doubles that must record batch GROUPING (ingest staging/refresh) by
  wrapping batch() over the shared double rather than re-rolling one.
- coverage-baseline.json: union of this branch's raised floors and upstream's
  new packages/test-support entry; apps/etl branches ratcheted 97 -> 97.6.
- vitest.shared.ts: drop the now-stale src/test/d1-sqlite.ts exclusion (midt-bg#331
  moved that shim into packages/test-support, which carries its own entry).

All seven workspaces stay at or above their floors; check-fake-d1, check-docs
and check-coverage are green.
lyubomir-bozhinov added a commit to lyubomir-bozhinov/sigma that referenced this pull request Aug 26, 2026
Line-length reflow only — no assertion or fixture changes. The three files
whose fakeD1 call sites pushed a line past printWidth after the midt-bg#331 migration.
todorkolev added a commit that referenced this pull request Sep 2, 2026
* test: coverage measurement + ratchet gate in CI

Vendored from #216 (feat/coverage-ratchet) to put the coverage machinery
in place ahead of #217, while #216 is pending merge upstream. Squashes
ydimitrof's two harness commits into one; original authorship preserved
via the commit author.

@vitest/coverage-v8 through a shared vitest preset, per-workspace coverage
configs with explicit include, committed coverage-baseline.json, the
scripts/check-coverage.mjs ratchet gate (+ node:test self-test), and the CI
wiring.

Ref #93, #216.

* test(config): cover regionByName + taxonomy integrity to 100%

regionByName had zero tests (both lines uncovered). Adds its full branch
matrix (valid, whitespace-trim, null/undefined/empty/unknown, verbatim
round-trip of all 28 regions) plus edge inputs for categoryForDivision and
procedureGroup, and taxonomy-integrity invariants (unique CPV codes, single
procedure-type ownership, classified = competitive∪non-competitive,
28 unique NUTS3 regions). config: 92.85/72.22 -> 100/100 lines/branches.

* test(shared): cover format.ts edge branches (eik/unp, date fallbacks, periodRange)

Adds eik/unp passthrough, one-sided and empty periodRange, date/monthYear/
longDate no-match + datetime-prefix + out-of-range-month fallbacks, count
sign/absence, pct/signedPct dp + non-finite, entityName non-collapsing paths,
cleanName unbalanced-quote drop, ЕТ/ET latin detection. shared branches
78.1 -> 98.3; residual is signedPct's provably-unreachable defensive return.

Caught: count(-0.4) emits '−0' (missing money()'s rounded-zero sign guard);
latent only (count takes non-negative integers), logged not patched.

* test(ingest): cover staging, refresh, and base/ocds edge branches

New staging.test (0%->100%): scoped DELETE, chunked INSERT at CHUNK=100
boundaries (100/101/250 rows), null-fill of absent columns, table+column
routing per target. New refresh.test: SQL splitter (escaped '' in literals,
-- comments inside/outside literals, trailing statement), @refresh-batch
grouping, transient-table drop order, D1 orchestration. base/ocds additions:
toBool, Date.parse date fallback, annexes mapping, baseSqlLiteral numeric/
text/null branches, secured_inverse/variants coercions, party/lot/amendment
and catch-up-window branches. ingest lines 83.8->100, branches 78.9->95.1.

* test(db): close branch gaps in identity, keyset, home, regions, methodology, flows

identity slug fallbacks + undecodable name slug; keyset decodeCursor
malformed/oversized/bad-type-guard paths; home zero single-offer aggregate;
regions year + EU/national funding predicates; methodology absent-count
coalesce with positive total; flows sankey sort tiebreak on shared authority.

* test(db): cover sitemap streaming end to end

streamAuthority/Company/Contract sitemaps via a paginating fake D1: XML
escaping + C0 stripping, lastmod fallbacks (row date -> as_of -> none),
natural-person filtering, empty-chunk skip loop, CHUNK-boundary pagination,
contract page rowid windowing, and contractSitemapPages math. sitemaps
branch 34.9 -> 95.3; residual is the defensive post-close pull guard.

* test(db): cover getCompany, getAuthority, and getContract derivations

getCompany + getAuthority were entirely untested. Adds DTO assembly, share
math (won/spent denominators, zero-guards), avg-bids rounding, consortium
membership (list->participants, prose->note), hasEik, sector top6+tail rollup,
and getContract subcontractor (EUR/BGN/null/blank), framework call-off
detection, eurFromNative currency paths (EUR/BGN peg/FX/no-rate), deltaPct
suspect+zero-base guards, lot dedup/totals, and not-found. details branch
43.6 -> 88.5, lines 53 -> 100.

* test(db): cover company-centred network, defaults, and hop-2 reduction

Adds the company-centre direction, null-param default (top authority) + its
empty fallback, includeCenterOptions=false, loadCenter sample-name fallback
(authority + company), hop-2 top-1-per-neighbour dedup, centre self-skip, and
edgeless-node weighting. network branch 50 -> 88, lines 86.7 -> 100.

* test(db): cover search empty-query + trend zero-year YoY and coverage guards

search: empty/punctuation query -> empty shape, searchMoreHref unknown-kind
fallback. trend: YoY guarded against a zero prior year, coverage pct when
nothing is dated (no divide-by-zero).

* test(db): cover authorities query branches to 96%

Add coverage for the base-aggregation source (year/EU/single- vs multi-sector
primary_sector), the entity WHERE type/text filters, sort normalization, facet
label fallback and sort, page overflow, and the CSV stream body across the
CHUNK boundary. Branch 55%->96%, lines 100%.

* test(db): cover companies query branches to 95%

Add coverage for sort normalization, the base-aggregation source (year/EU/
single- vs multi-sector), the kind/text entity WHERE, facet kind mapping and
sector sort, page overflow, missing total row, and the CSV body across the
CHUNK boundary. Branch 61%->95%, lines 100%.

* test(db): cover competition query branches to 95%

Add coverage for the authority-detail wrappers (getAuthoritySingleOffer,
getAuthorityProcedureCompetition), getCompetitionSummary (both the qualifying
and null-topConcentration paths), the MAX_TOP cap, EU/national funding scope,
and a degenerate corpus exercising the zero-guard fallbacks. Branch 70%->95%.

* test(db): cover contracts query branches to 96%

Add coverage for buildFilters (every year/sector/procedure/value-bucket/EU/
bids/authority/bidder/text predicate), summary override, page overflow,
contractsSummary null row, listSingleOfferContracts modes, facet procedure
folding / sector sort / year ordering, and the streamed CSV body across the
CHUNK boundary. Branch 66%->96%, lines 100%.

* test(db): cover details query fallback branches to 98%

Add degenerate-input coverage for getCompany (absent metadata/bids/suspect
rows, null primary sector and procedure value), getAuthority (spent-nothing
authority with an unknown CPV division, zeroed tail share, null bids/suspect),
and a getContract lot defaulting to the BGN peg. Branch 88%->98%, lines 100%.

* test(db): cover network query branches to 98%

Add coverage for an unresolvable company centre (null name → empty network),
the company-kind fallback when neither rollup nor sample carries a kind, the
empty-default path with includeCenterOptions off, and deduping a hop-1
neighbour that appears twice. Branch 88%->98%, lines 100%.

* test(db): cover flows, search, and trend branches to 95%+

flows: EU/national/all funding scope + long-label truncation (branch 89->100).
search: nullish raw-query coalescing (branch ->97). trend: EU/national funding,
includeSectors=false, and an empty series with an absent coverage row
(branch 89->100).

* test(etl): cover the EOP ingest worker and bucket pipeline to 98%

index.ts: full RefreshWorkflow.run coverage (staging lifecycle, capped and
zero-ingest branches, derive-slice loop, finally-drop on success and error) and
the scheduled cron entrypoint, via mocked platform/ingest/eop seams. eop.ts:
bucket-key parse/classify, catch-up planning (uncapped/capped/default), bucket
listing status + redirect guard, OCDS/base staging, and the window walk.
Workspace 19%->98% branch, lines 100%.

* test(web): cover filters URL-state helpers to 97%

Add coverage for singleSelectFilters (unknown sector/year flags, funding/top
defaults), buildSectorGroup (category grouping, summed vs absent counts,
uncategorised skip), sortHref, and the withParams/pageNav null-override, array,
empty-result, and page-default branches. filters.ts branch 47%->97%.

* test(web): cover assistant agent, report binder, and tool registry

agent.ts: SDK-wiring coverage (model/base-URL resolution, tool-set assembly,
stream Response + onError) via mocked ai/@ai-sdk. report-schema.ts: flows block,
facts sub-line, unknown-handle, empty title, 0-row and null chart edges
(branch 78->94). tools.ts: run_sql AST-reject/error/meta-less paths, semantic
hits, eop_fetch, source_link (branch 57->95). Also exclude test fixtures and
type-only declarations from coverage (permanent 0% data files, not code).

* test(web): cover assistant format, results, eop-fetch, rag, emit-shape

render-format null-date; tool-results missing-cell + truncation flag; eop-fetch
null-date/non-array/invalid-JSON/thrown-fetch; rag embed mismatch + metadata
mapping + empty-vector guards; validateEmitShape callout/flows/timeseries and
the object/question/items/columns negatives.

* test(web): cover CSV export ranges, freshness, and multipart edges

Add coverage for the v0 freshness fallback, non-string/null q classification,
empty and empty-chunk multipart bodies, the abort-on-part-failure path, and
suffix (bytes=-N) / open-ended (bytes=A-) R2 range shapes (the fake now emits
R2-native range objects). csv-export lines 87->95, branch 76->90.

* build: exclude markdown files from coverage instrumentation

The v8 provider instruments every file matched by a workspace's include globs.
Markdown docs colocated in src (e.g. the assistant README) carry no coverable
statements, report a permanent 0%, and — being non-JS — make the reporter's
remap step throw a parse error. Exclude **/*.md alongside the existing JSON and
type-only exclusions so the ratchet total reflects executable code only.

* test(web): raise branch coverage to the 95 floor

Cover the remaining thin spots in the web workspace with real behavioural tests,
no code changes:

- cache.publicCache: default and explicit stale-while-revalidate windows
- eopSource: missing/malformed dates, DD.MM.YYYY key shape, OCDS cutoff boundary
- search.suggest: trimGroup cap + loader query/trim/headers, empty-q default
- app.ts hardening: nonce re-read path, OPTIONS short-circuit, no-Content-Type
- retry: non-Error rejection logging + default backoff past the table
- security: nonce-less headers omit the CSP outside production
- riskLogic: unknown bid count (null) and missing bidsRejected fallback
- ScrollToTop: rAF coalescing of a scroll burst

Web branch total 81.3% -> 95.47%, lines 89.3% -> 99.25%.

* style(config): wrap the curated-sector assertion per prettier

* chore: ratchet coverage floors to >=95 for every workspace

With the new tests in place, every workspace clears 95% on both lines and
branches. Regenerate the ratchet baseline from current coverage so the gate
now enforces the 95 floor going forward:

  etl      18.7/19.4  -> 100/98.5
  web      89.3/81.3  -> 99.2/95.4
  config   88.2/58.3  -> 100/100
  db       82/65.5    -> 100/97.3
  ingest   83.8/78.9  -> 100/95
  shared   94.8/78.1  -> 98.8/98.3

* test(web): assert non-null categories in buildSectorGroup tests

buildSectorGroup always returns categories, but the group type marks it
optional. vitest transpiles without type-checking so this passed locally;
tsc --noEmit under noUncheckedIndexedAccess (CI typecheck) rejected the
possibly-undefined access. Assert non-null at the three call sites.

* test(ingest): cover sparse-release OCDS nullish branches

Add branch-completion tests for releaseToContracts/Amendments/Lots on
releases with absent optional fields: missing tag/contracts keys, id-less and
identifier-less parties, a scheme-less CPV classification, an empty-string
value amount, a blank date, and an ocid/tender-id-less lot. ocds.ts branch
92.17% -> 98.26%, ingest workspace 95.06% -> 98.7%.

Residual uncovered branches are unreachable defensive code: the validDateOnly
regex reject (day is always pre-normalized to YYYY-MM-DD) and the
`rel.contracts ?? []` / `c.id ?? null` right-sides the length/id guards above
them make impossible.

* test(db): close reachable query branch gaps to 98.4%

- keyset: unsafe-direction guard, before-cursor with an ascending sort, and the
  empty before-page (both cursors null)
- flows: two-authority sankey so the authority-column sort comparator runs
- regions: empty dataset → the total==0 coverage-pct guard (no divide-by-zero)
- companies: base aggregation from a non-sector filter (no CPV predicate)
- contracts: the „Неизвестна" year bucket sinking below real years regardless
  of input order
- authorities/companies/contracts: backward pagination — page forward for a
  cursor, back for a before-cursor, then feed it back so keyset's reverse path
  runs

db branch 97.35% -> 98.39%. Residual gaps are unreachable defensive code:
CSV/sitemap `if(done)` re-entry (a stream never pulls after close), the
minContracts `?? DEFAULT` the orchestrator already normalises, split().pop()
`?? ` fallbacks (pop is always defined), the homoglyph map (every regex-matched
char is mapped), and cross-namespace network self-edges.

* test(web): cover read-only SQL guard, tool, and eop-fetch branches

- assertReadOnlySelect: empty/comment-only query, a forbidden keyword hidden
  in a single CTE-prefixed statement (cheap keyword layer, not just the AST
  guard), and the sqlite_master/sqlite_schema catalog-table rejection
- run_sql: a driver returning no results array (the results ?? [] fallback)
- eop-fetch: a non-Error thrown value → the generic fetch-error label

web branch 95.47% -> 96.0%. Residual gaps are deep AST-shape defenses
(sql-ast-guard), schema-validation guards (report-schema), and unreachable
code: regex capture-group ?? fallbacks (csp, always matched), PROD-gated
redirect/OPTIONS paths under vitest, and the module-init Date.now tag.

* chore: ratchet coverage floors up after deeper branch tests

Coverage rose across web/db/ingest with the new branch tests; raise the
ratchet floors to match (never down):

  web      99.2/95.4 -> 99.4/96
  db       100/97.3  -> 100/98.3
  ingest   100/95    -> 100/98.7

etl (100/98.5), config (100/100) and shared (98.8/98.3) unchanged. Monorepo
total 99.73% lines / 97.51% branches.

* style(db): wrap flows two-authority test per prettier

* test: make coverage-only tests mutation-sensitive

An adversarial mutation audit found six added tests that lit up a branch for
the ratchet without asserting its behaviour (each survived deleting the very
line it claimed to cover). Strengthened so the assertion fails under the
targeted mutation:

- db backward pagination (authorities/companies/contracts): pageSize 1 made
  slice+reverse a no-op; now pageSize 2 over 3 rows and asserts the page comes
  back in reversed fetch order (fails if rows.reverse() is dropped)
- db flows ribbon order: input was pre-sorted so the comparator was unguarded;
  now feeds unsorted pairs and asserts ranked toName order
- web app.harden nonce swap: unobservable under dev PROD=false; now stubs
  PROD=true and asserts the CSP nonce is replaced by a sha256 hash
- web retry backoff: asserts setTimeout was called with [50,150,150], pinning
  the BACKOFF_MS[i] ?? 150 fallback value

Coverage unchanged (db 98.39%, web 96%); same branches, real assertions.

* test(coverage): replace coverage-only assertions with mutation-sensitive ones

Adversarial audit pass over the coverage-ratchet suite: each finding was
confirmed by breaking the exact production line and watching the test stay
green, then fixed and re-verified so the mutation now fails.

- etl stagedRows: fixture now sets baseAmendments/ocdsAmendments non-zero so
  both amendment terms of the sum are guarded (were unexercised).
- db details: procedureMix assertions go from .length>0 to the folded DTO
  shape (contracts/valueEur/sharePct); numeric lot sort now fed out of order
  with multi-digit labels; consortium lot contractor name asserted through
  the row kind.
- db sitemaps: page math now guards the upper rowid bound (hi), and lastmod
  precedence proves signed_at wins over published_at when both are present.
- db network: centre-picker test asserts the full mapped shape for both the
  authority and (previously unexercised) company branch.
- db trend: sector options asserted through the includeSectors=true path.
- db keyset: oversized-cursor test uses a genuinely decodable payload so the
  length guard is exercised, not the JSON.parse catch; plus a near-limit pass.
- db contracts: listSingleOfferContracts asserts the limit reaches LIMIT ?.
- web retry: fake timers; ScrollToTop: rAF coalescing counted.

* test(etl): pin the clock in the catch-up default-today test

The 'defaults today' test only asserted the plan window was date-shaped
(/^\d{4}-\d{2}-\d{2}$/), which a mutation to any hard-coded ISO string would
survive. Pin the system clock and assert plan.to equals that exact date, so
the new Date() default is actually verified. Addresses the reviewer's
determinism note on real-clock reliance.

* test(etl): restore the data-integrity invariant weakened in the fake DB

A prior commit on this branch reshaped fakeDbFromFreshness so prepare()
ignores the SQL and always returns the freshness row, dropping the guard
that threw when raw staging was read for planning. Restore it as the stronger
positive invariant: the catch-up planner must read data_freshness, never
raw_*, so a regression that reads raw staging for the max-loaded date fails
loudly. Also fix a BG typo in the config test name (областти -> области).

* test(ingest): cover FX/OCDS edge branches; reconcile etl branch floor to merged-tree actual

Restore ingest to its line floor and hold branches after the merge:
- fx.ts: a malformed (unpadded) staged contract_date surfaces through
  findFxCoverageGaps' MIN/MAX, so loadFxRates must reject the range before
  fetching — asserted (skips the currency, warns, never fetches).
- ocds.ts: a release with no bids block nulls bids_received via the optional
  chain instead of throwing.

Reconcile apps/etl branches 98.5 -> 97 (achieved 97.46). etl LINES hold at 100.
The two uncovered branches are provably unreachable: eop.ts parseBucketKeys
`m[1] ?? ''` (a matched regex group is never undefined — the `??` is required
only by noUncheckedIndexedAccess) and integrity.ts `summary.message ?? '...'`
(dead — summarizeIntegrity guarantees a non-null message on the throw path).
97 is still far above the harness's original etl branch floor (58.2); covering
the branches would require a fake or a source edit, both disallowed.

check:coverage passes for all six workspaces.

* test(db,web): address review findings on the coverage ratchet (#254)

Four findings from ydimitrof's review, verified against the source before acting.

competition.test.ts — the „caps at MAX_TOP" test asserted top:50 → 50 and named a
clamp the source does not perform. `getCompetition` reads
`p.top === MAX_TOP ? MAX_TOP : DEFAULT_TOP`: an exact 50 selects the large size and
everything else falls back to 20, so nothing is ever reduced to 50. The old assertion
did kill a mutation (collapsing the toggle to DEFAULT_TOP fails it), but it left the
fallback — the half that stops a caller naming its own leaderboard size and its own
LIMIT — untested, under a name that overstated it. Renamed and extended to cover 999,
51, 35, 0, -1, NaN and an omitted top. Now kills three mutations, including the
`Math.min` clamp the old test would have passed.

vitest.shared.ts — dropped `**/*.json` and `**/*.md` from the coverage excludes. The
provider only reports files it can instrument as modules, so neither ever reached a
report; removing both leaves all six workspaces' numbers byte-identical. Recorded that
verification in the comment rather than the globs.

vitest.shared.ts — documented the contract behind `**/src/test/**`: the glob is wide,
so the exclusion holds only while src/test/ stays test HELPERS. Product code placed
there would leave the coverage denominator silently, which is the single way this list
can hide an untested module rather than an uncoverable file.

search.suggest.test.tsx — filled the dangling „upstream #…" placeholder with the real
reference (#225, the read-only D1 chokepoint for #199).

trend.test.ts — the local `customDb` returned the period series for every query,
including the sector_totals read. Dispatches on the SQL now, like the other fakes in
the file, so a later assertion on `sectors` cannot be fed rows of the wrong shape.

Coverage unchanged: etl 100/97.64, web 99.47/96.21, config 100/100, db 100/98.47,
ingest 100/98.69, shared 98.87/98.33. No floor moved.

* test(web): tighten the coverage denominator contract and the date-cell assertions (#254)

Second round of ydimitrof's review.

vitest.shared.ts — replaced the `**/src/test/**` directory glob with the three files it
actually covers: the cloudflare:workers/workflows stubs apps/etl aliases the real modules
to, and the ingest SQLite D1 shim. A directory glob was the one entry on this exclude
list that could hide an untested product MODULE rather than an uncoverable file, since
anything later dropped into src/test/ would leave the coverage denominator by virtue of
its location. With an explicit list, a new file there is measured until someone
deliberately adds it — a reviewable act. Stricter than a suffix convention, which would
still let a *.helper.ts product file through. Coverage is unchanged in all six
workspaces, confirming the list covers exactly what the glob did.

render-format.test.ts — the date-null case asserted `formatCell(null, 'date')` equals
`date(null)`, which only proves delegation and passes for any value the shared formatter
returns. Pinned to the literal em-dash, and added the non-null path (a formatted date and
an unparseable value that must be echoed, never rendered as a fake date). The pair now
kills two mutations the old assertion could not see — including a `date` branch
hard-coded to return the em-dash.

eop.test.ts — dropped the describe-local afterEach that duplicated the file-level
vi.unstubAllGlobals introduced when this file was union-merged with upstream.

contracts.test.ts — the two `describe('getContractFacets')` blocks now carry distinct
names for what each covers.

Coverage unchanged: etl 100/97.64, web 99.47/96.21, config 100/100, db 100/98.47,
ingest 100/98.69, shared 98.87/98.33.

* test(config): name both paths to the unknown procedure bucket (#254 review)

The comment said a whitespace-only value 'trims to ""' and grouped it with the nullish
inputs, which reads as if it hits the `if (!procedureType)` guard. It does not: '   ' is
truthy, passes the guard, and reaches the map lookup, where .get('') misses and the ??
fallback supplies PROCEDURE_UNKNOWN. Spelling out both paths so the map-miss branch is
visibly exercised here as well as by the unrecognised-type case above.

* style: run prettier on the three migrated test files

Line-length reflow only — no assertion or fixture changes. The three files
whose fakeD1 call sites pushed a line past printWidth after the #331 migration.

* test: give the fully-covered workspaces a line-floor margin

Per @todorkolev on #254: the five workspaces sitting at lines 100 drop to a
99 floor. Measured coverage is unchanged at 100% in all five — this only
widens the margin before the ratchet fires.

Why it mattered: `tolerance` is a percentage, so it is near-zero slack in a
small workspace. At 100/100 a single uncovered line red-builds
packages/config (28 lines), packages/test-support (99) and apps/etl (169) —
including on PRs that never touch tests. packages/db needed 6 and
packages/ingest 3.

Branch floors are untouched.

* test(web): pin the two ConflictDetail assertions to what they claim

Both findings are @ydimitrof's on #254, and both were real.

- the source-URL test selected `.cc-source, .conflict-detail`. There is no
  `.cc-source` in the component — the stat cells carry no per-field class — so
  it always widened to the whole card, where `toContain('—')` can be satisfied
  by any other dash. Now pinned to the „Източник" cell via its `<dt>`, with a
  positive control asserting the same cell holds the declaration link when the
  URL is present, so a broken finder cannot make the negative case pass.
- the in/out-window test asserted only that both rows render and both numbers
  appear, which holds with the split broken in either direction. Now asserts
  the per-row `contract-item-conflict` modifier and that the outside row sits
  behind the „Извън периода" disclosure.

Mutation-verified: forcing every row to carry the modifier, forcing none to,
and deleting the source-cell fallback each fail the suite.

---------

Co-authored-by: Yoan Dimitrov <ydimitrof@users.noreply.github.com>
Co-authored-by: Todor Kolev <tkolev@obecto.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

test: споделен помощник за фалшив D1 вместо ~130 локални двойника

4 participants