From b0357ae64226ee21e4bce5ba90f92c05e8ec48fc Mon Sep 17 00:00:00 2001 From: jacobvjk Date: Wed, 7 Oct 2026 16:44:43 +0100 Subject: [PATCH 1/2] docs(readme): how to add or update pathways and benchmark data A plain-language section for the people who supply pathway content, not only developers: the two parts of a pathway (description and benchmark data), why the description comes first, who does what, and the steps for each part, from the shared workbook or the preparation repository through the import scripts' dry run to a reviewed pull request. It includes a short table of the commonest benchmark import errors and where to fix them. The workbook is named, not linked: this repository is public and the workbooks live on RMI's SharePoint. Co-Authored-By: Claude Opus 5.5 --- README.md | 64 +++++++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 64 insertions(+) diff --git a/README.md b/README.md index 6dfb7b77..528f3525 100644 --- a/README.md +++ b/README.md @@ -38,6 +38,70 @@ The application is deployed using Azure Static Web Apps: Pull requests automatically deploy to preview environments with URLs provided in the PR comments. +## Adding or updating pathways + +Each pathway on the site has two parts: + +1. **A description** (the "metadata"): who published it, which sectors and regions it covers, its key features, drivers, dependencies and what data it offers. Every pathway needs one. +2. **Benchmark data** (the "timeseries"): the numbers behind the charts and the download, for example installed solar capacity in Southeast Asia, year by year. This part is optional. + +**The description always comes first.** Benchmark data may only use the region names its pathway's description lists, so a new pathway, or a new region, has to be in the description before its data can be added. + +Neither part is edited by hand. Data people fill in a shared workbook or run a preparation script, and a developer brings the result into this repository with an import script that checks the data and prints a report of every problem it finds. Every change then goes through a pull request, which builds a preview of the site (the link appears in the pull request) so the new pathway can be checked before it goes live. + +### Adding or updating a pathway description + +**Who:** someone who knows the publication fills in the workbook; a developer does the import. + +1. **Fill in the workbook.** Descriptions are written in the shared workbook `pathway_data_prepared.xlsx` (on RMI's SharePoint; ask the TPR data team for access). Each pathway gets a row in the metadata sheet and rows in the sheets for key features, core drivers, dependencies and data availability. The classification cookbook says which values are allowed and how to choose them. +2. **Tell the developer what's new.** Updating a pathway that is already on the site needs nothing extra. A **new** pathway needs one line in the importer's list of pathways (`TARGETS` in `scripts/import-pathway-data.ts`). A **new publisher** also has to be added to the list of allowed publishers (`src/schema/common/publication.v1.json`), and the importer has to learn to recognise its name (`publisherGroup` in the same script). +3. **Import (developer).** Run a dry run first. It writes nothing and prints a report of everything it could not take over cleanly, such as a value that isn't allowed or a missing field. Some problems stop the import (for example a data-availability row that is only partly "Not covered"); others are skipped and listed in the report, so read it all: + + ```bash + npx ts-node --esm scripts/import-pathway-data.ts --xlsx --dry-run + ``` + + Fix problems in the workbook (not in the files here) and repeat until the report shows nothing you want to fix. Then run the same command without `--dry-run`, followed by: + + ```bash + npx prettier --write "src/data/**/*.json" + npm run schema:check + ``` + +4. **Review and publish.** Open a pull request. Check the pathway on the preview site, then merge. + +### Adding or updating benchmark data + +**Who:** the benchmark data is prepared in a separate repository, [RMI/tpr_benchmark_data_preparation](https://github.com/RMI/tpr_benchmark_data_preparation) (RMI internal), by someone with access to the raw publication files; a developer does the import. + +1. **Prepare the data.** Run the preparation pipeline as its README describes. It produces two things: + - `benchmark_data_prepared.xlsx`, a workbook for checking the numbers by eye. Review it before going on. + - a `tpr_timeseries` folder with one file per pathway, ready to import. +2. **Import (developer).** Again, a dry run first. This import is stricter: if anything is wrong it writes nothing at all, and it lists every problem in one go: + + ```bash + npx ts-node --esm scripts/import-benchmark-data.ts --in --dry-run + ``` + + The most common problems, and where to fix them: + + | The import says | What it means | Fix it in | + | ----------------------------------------- | ------------------------------------------------------------------- | ------------------------------------------------------------------- | + | "… is not a geography pathway … declares" | The data uses a region the pathway's description doesn't list | The pathway description (see above): add the region, then re-import | + | "… is not the id of any pathway" | The pathway has no description yet | Add the description first | + | "… is not a segment of …" | A row names a part of the sector that doesn't belong to that sector | The preparation repository | + + Once the dry run is clean, run it again without `--dry-run`. It updates existing files and adds new ones next to their pathway's description; it never deletes anything. Then: + + ```bash + npm run build:timeseries + npm run schema:check + ``` + +3. **Review and publish.** Open a pull request, check the charts and the data download on the preview site, then merge. + +The file formats and every validation rule are described in [`src/data/README.md`](src/data/README.md). + ## Development ### Set-Up From 88f1f6ccb8604bd3f6745b8dc6b5884b345c32cd Mon Sep 17 00:00:00 2001 From: jacobvjk Date: Thu, 8 Oct 2026 08:54:25 +0100 Subject: [PATCH 2/2] docs(readme): match the importers' actual behaviour (#986 review) - The metadata importer only handles pathways on its TARGETS list; others are reported as unmatched and skipped. A new pathway copies its publisher, licence and links from a template file, so the README says to pick a template from the same publisher, and that a new publisher's publication details are written by hand after import. - The preparation repository writes one file per dataset, and a dataset can belong to several pathways. - An undeclared-geography or unknown-pathway error does not by itself mean the description is missing something: the data may be the one that is wrong. The table now says to check the publication and fix whichever side is wrong. Co-Authored-By: Claude Opus 5.5 --- README.md | 19 +++++++++++-------- 1 file changed, 11 insertions(+), 8 deletions(-) diff --git a/README.md b/README.md index 528f3525..b6e01e3d 100644 --- a/README.md +++ b/README.md @@ -54,7 +54,10 @@ Neither part is edited by hand. Data people fill in a shared workbook or run a p **Who:** someone who knows the publication fills in the workbook; a developer does the import. 1. **Fill in the workbook.** Descriptions are written in the shared workbook `pathway_data_prepared.xlsx` (on RMI's SharePoint; ask the TPR data team for access). Each pathway gets a row in the metadata sheet and rows in the sheets for key features, core drivers, dependencies and data availability. The classification cookbook says which values are allowed and how to choose them. -2. **Tell the developer what's new.** Updating a pathway that is already on the site needs nothing extra. A **new** pathway needs one line in the importer's list of pathways (`TARGETS` in `scripts/import-pathway-data.ts`). A **new publisher** also has to be added to the list of allowed publishers (`src/schema/common/publication.v1.json`), and the importer has to learn to recognise its name (`publisherGroup` in the same script). +2. **Tell the developer what's new.** The importer only handles pathways on its own list (`TARGETS` in `scripts/import-pathway-data.ts`); a workbook row it doesn't know is reported as "unmatched" and skipped. + - **Updating a pathway on that list** needs nothing extra. Some older pathways on the site are not on it yet. + - **A new pathway** needs a line on the list, naming the file to create and an existing pathway file to use as a template. The importer takes the publication's title and year from the workbook but copies the rest of the publication details (publisher, licence, links) from the template, so pick one from the same publisher and check those details afterwards. + - **A new publisher** has no such template. The developer adds it to the allowed publishers (`src/schema/common/publication.v1.json`), teaches the importer to recognise its name (`publisherGroup` in the same script), and writes the new file's publication details by hand after the import. 3. **Import (developer).** Run a dry run first. It writes nothing and prints a report of everything it could not take over cleanly, such as a value that isn't allowed or a missing field. Some problems stop the import (for example a data-availability row that is only partly "Not covered"); others are skipped and listed in the report, so read it all: ```bash @@ -76,20 +79,20 @@ Neither part is edited by hand. Data people fill in a shared workbook or run a p 1. **Prepare the data.** Run the preparation pipeline as its README describes. It produces two things: - `benchmark_data_prepared.xlsx`, a workbook for checking the numbers by eye. Review it before going on. - - a `tpr_timeseries` folder with one file per pathway, ready to import. + - a `tpr_timeseries` folder with one file per dataset, ready to import. A dataset usually belongs to one pathway but can be shared by several. 2. **Import (developer).** Again, a dry run first. This import is stricter: if anything is wrong it writes nothing at all, and it lists every problem in one go: ```bash npx ts-node --esm scripts/import-benchmark-data.ts --in --dry-run ``` - The most common problems, and where to fix them: + The most common problems, and what to do about them: - | The import says | What it means | Fix it in | - | ----------------------------------------- | ------------------------------------------------------------------- | ------------------------------------------------------------------- | - | "… is not a geography pathway … declares" | The data uses a region the pathway's description doesn't list | The pathway description (see above): add the region, then re-import | - | "… is not the id of any pathway" | The pathway has no description yet | Add the description first | - | "… is not a segment of …" | A row names a part of the sector that doesn't belong to that sector | The preparation repository | + | The import says | What it means | What to do | + | ----------------------------------------- | ------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | + | "… is not a geography pathway … declares" | The data names a region that is not written exactly like any region in the pathway's description | Check the publication. If the data's name is a typo or differs from the publication's wording, fix it in the preparation repository. If the publication really covers that region and the description leaves it out, add it to the description (see above) and re-import that first. | + | "… is not the id of any pathway" | The data points to a pathway id that has no description | If the id is mistyped, fix it in the preparation repository. If the pathway is new, add its description first. | + | "… is not a segment of …" | A row names a part of the sector that doesn't belong to that sector | Fix it in the preparation repository. | Once the dry run is clean, run it again without `--dry-run`. It updates existing files and adds new ones next to their pathway's description; it never deletes anything. Then: