diff --git a/README.md b/README.md index 6dfb7b77..b6e01e3d 100644 --- a/README.md +++ b/README.md @@ -38,6 +38,73 @@ The application is deployed using Azure Static Web Apps: Pull requests automatically deploy to preview environments with URLs provided in the PR comments. +## Adding or updating pathways + +Each pathway on the site has two parts: + +1. **A description** (the "metadata"): who published it, which sectors and regions it covers, its key features, drivers, dependencies and what data it offers. Every pathway needs one. +2. **Benchmark data** (the "timeseries"): the numbers behind the charts and the download, for example installed solar capacity in Southeast Asia, year by year. This part is optional. + +**The description always comes first.** Benchmark data may only use the region names its pathway's description lists, so a new pathway, or a new region, has to be in the description before its data can be added. + +Neither part is edited by hand. Data people fill in a shared workbook or run a preparation script, and a developer brings the result into this repository with an import script that checks the data and prints a report of every problem it finds. Every change then goes through a pull request, which builds a preview of the site (the link appears in the pull request) so the new pathway can be checked before it goes live. + +### Adding or updating a pathway description + +**Who:** someone who knows the publication fills in the workbook; a developer does the import. + +1. **Fill in the workbook.** Descriptions are written in the shared workbook `pathway_data_prepared.xlsx` (on RMI's SharePoint; ask the TPR data team for access). Each pathway gets a row in the metadata sheet and rows in the sheets for key features, core drivers, dependencies and data availability. The classification cookbook says which values are allowed and how to choose them. +2. **Tell the developer what's new.** The importer only handles pathways on its own list (`TARGETS` in `scripts/import-pathway-data.ts`); a workbook row it doesn't know is reported as "unmatched" and skipped. + - **Updating a pathway on that list** needs nothing extra. Some older pathways on the site are not on it yet. + - **A new pathway** needs a line on the list, naming the file to create and an existing pathway file to use as a template. The importer takes the publication's title and year from the workbook but copies the rest of the publication details (publisher, licence, links) from the template, so pick one from the same publisher and check those details afterwards. + - **A new publisher** has no such template. The developer adds it to the allowed publishers (`src/schema/common/publication.v1.json`), teaches the importer to recognise its name (`publisherGroup` in the same script), and writes the new file's publication details by hand after the import. +3. **Import (developer).** Run a dry run first. It writes nothing and prints a report of everything it could not take over cleanly, such as a value that isn't allowed or a missing field. Some problems stop the import (for example a data-availability row that is only partly "Not covered"); others are skipped and listed in the report, so read it all: + + ```bash + npx ts-node --esm scripts/import-pathway-data.ts --xlsx --dry-run + ``` + + Fix problems in the workbook (not in the files here) and repeat until the report shows nothing you want to fix. Then run the same command without `--dry-run`, followed by: + + ```bash + npx prettier --write "src/data/**/*.json" + npm run schema:check + ``` + +4. **Review and publish.** Open a pull request. Check the pathway on the preview site, then merge. + +### Adding or updating benchmark data + +**Who:** the benchmark data is prepared in a separate repository, [RMI/tpr_benchmark_data_preparation](https://github.com/RMI/tpr_benchmark_data_preparation) (RMI internal), by someone with access to the raw publication files; a developer does the import. + +1. **Prepare the data.** Run the preparation pipeline as its README describes. It produces two things: + - `benchmark_data_prepared.xlsx`, a workbook for checking the numbers by eye. Review it before going on. + - a `tpr_timeseries` folder with one file per dataset, ready to import. A dataset usually belongs to one pathway but can be shared by several. +2. **Import (developer).** Again, a dry run first. This import is stricter: if anything is wrong it writes nothing at all, and it lists every problem in one go: + + ```bash + npx ts-node --esm scripts/import-benchmark-data.ts --in --dry-run + ``` + + The most common problems, and what to do about them: + + | The import says | What it means | What to do | + | ----------------------------------------- | ------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | + | "… is not a geography pathway … declares" | The data names a region that is not written exactly like any region in the pathway's description | Check the publication. If the data's name is a typo or differs from the publication's wording, fix it in the preparation repository. If the publication really covers that region and the description leaves it out, add it to the description (see above) and re-import that first. | + | "… is not the id of any pathway" | The data points to a pathway id that has no description | If the id is mistyped, fix it in the preparation repository. If the pathway is new, add its description first. | + | "… is not a segment of …" | A row names a part of the sector that doesn't belong to that sector | Fix it in the preparation repository. | + + Once the dry run is clean, run it again without `--dry-run`. It updates existing files and adds new ones next to their pathway's description; it never deletes anything. Then: + + ```bash + npm run build:timeseries + npm run schema:check + ``` + +3. **Review and publish.** Open a pull request, check the charts and the data download on the preview site, then merge. + +The file formats and every validation rule are described in [`src/data/README.md`](src/data/README.md). + ## Development ### Set-Up