From 3d69137d7bb028f4d1857e6ce116ebf3c988e5d2 Mon Sep 17 00:00:00 2001 From: Wholeheartedly <119653204+AESMatias@users.noreply.github.com> Date: Fri, 25 Sep 2026 20:15:21 -0300 Subject: [PATCH] feat(deploy): deploy main to the server automatically after CI passes Add deploy/deploy.sh: it pulls main, builds while the old version keeps serving, restarts the containers and waits for their health checks. If the new version is not healthy within four minutes it restores the previous images and commit by itself; a failed build restarts nothing. One deploy runs at a time. A new deploy job in CI runs the script over SSH on every push to main whose backend and frontend checks pass. It stays skipped until the DEPLOY_HOST variable is set. The SSH key is pinned in authorized_keys to the deploy script, and the server's host key is verified. --- .github/workflows/ci.yml | 28 ++++++ 02-DOCS/wiki/ftd/continuous-deployment.md | 31 +++++++ 02-DOCS/wiki/sdd/decisions.md | 1 + CLAUDE.md | 1 + deploy/deploy.sh | 101 ++++++++++++++++++++++ docs/DEPLOY.md | 43 +++++++-- 6 files changed, 199 insertions(+), 6 deletions(-) create mode 100644 02-DOCS/wiki/ftd/continuous-deployment.md create mode 100755 deploy/deploy.sh diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 125267f..09bbf80 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -1,4 +1,5 @@ # Quality gates on every push and pull request, plus a weekly security scan of both images. +# A push to main that passes the backend and frontend gates is deployed to production. name: CI on: @@ -81,3 +82,30 @@ jobs: severity: CRITICAL,HIGH ignore-unfixed: true exit-code: "1" + + deploy: + name: Deploy to production + runs-on: ubuntu-latest + needs: [backend, frontend] + # Pushes to main only, and only once the server is set up (repository variable DEPLOY_HOST; + # see docs/DEPLOY.md, "Automatic deploys"). Until then the job is skipped, not failed. + if: github.event_name == 'push' && github.ref == 'refs/heads/main' && vars.DEPLOY_HOST != '' + environment: production # listed under the repository's Deployments, like Vercel + concurrency: + group: production + cancel-in-progress: false # never stop a deploy halfway; the next one waits + steps: + - name: Run deploy/deploy.sh on the server + env: + SSH_KEY: ${{ secrets.DEPLOY_SSH_KEY }} + KNOWN_HOSTS: ${{ secrets.DEPLOY_KNOWN_HOSTS }} + DEPLOY_HOST: ${{ vars.DEPLOY_HOST }} + DEPLOY_USER: ${{ vars.DEPLOY_USER || 'root' }} + run: | + install -m 700 -d ~/.ssh + printf '%s\n' "$SSH_KEY" > ~/.ssh/deploy_key + chmod 600 ~/.ssh/deploy_key + printf '%s\n' "$KNOWN_HOSTS" > ~/.ssh/known_hosts + # The server pins this key to deploy/deploy.sh in authorized_keys: it can run nothing else. + ssh -i ~/.ssh/deploy_key -o BatchMode=yes -o StrictHostKeyChecking=yes -o ServerAliveInterval=30 \ + "$DEPLOY_USER@$DEPLOY_HOST" diff --git a/02-DOCS/wiki/ftd/continuous-deployment.md b/02-DOCS/wiki/ftd/continuous-deployment.md new file mode 100644 index 0000000..6c97a74 --- /dev/null +++ b/02-DOCS/wiki/ftd/continuous-deployment.md @@ -0,0 +1,31 @@ +--- +title: Continuous deployment from main to the server +status: done +updated: 2026-09-25 +--- + +# Continuous deployment + +## Intent + +Every push to `main` that passes CI reaches production without logging in to the server, like +Vercel, and a broken release never stays online. + +## Design + +- `deploy/deploy.sh` (on the server): lock, fetch `main`, tag the running images `:previous`, + fast-forward, build while the old version serves, `up -d`, wait for Docker health checks (web, + frontend; worker running), roll back images and commit if unhealthy after 240 s, prune. +- CI job `deploy`: needs backend and frontend, runs on pushes to `main` when the `DEPLOY_HOST` + variable exists, `production` environment, no cancellation mid-deploy. It only opens SSH. +- The SSH key is pinned in `authorized_keys` with `command=".../deploy.sh"` and no forwarding or + pty, and the host key is checked (`DEPLOY_KNOWN_HOSTS`): a leaked key can only redeploy `main`. + +## Checklist + +| # | Task | Proof | Status | +|---|------|-------|--------| +| 1 | Deploy script | Simulated `docker` in a throwaway git setup: nothing new → exit 0; healthy → deployed; never healthy → rolled back to the previous commit and images, exit 1; build fails → no restart; retry → deployed; not on main → refused | ✅ | +| 2 | CI deploy job | YAML parses; needs backend and frontend; skipped without `DEPLOY_HOST` | ✅ | +| 3 | Docs | DEPLOY sections 10 and "Automatic deploys" | ✅ | +| 4 | First real run on the server | Owner sets the key, secrets and variables | ⏳ | diff --git a/02-DOCS/wiki/sdd/decisions.md b/02-DOCS/wiki/sdd/decisions.md index 01e2b0b..2c69e02 100644 --- a/02-DOCS/wiki/sdd/decisions.md +++ b/02-DOCS/wiki/sdd/decisions.md @@ -39,3 +39,4 @@ | 2026-09-25 | Count unique visitors with a keyed hash of IP and user agent in daily Redis HyperLogLogs, no cookies, totals only | Google Analytics or Plausible; a cookie; storing hashes in Postgres | No consent banner, no third party, and nothing that identifies a visitor is ever stored; about 0.8% counting error is acceptable | | 2026-09-25 | Behind the host's Nginx, the container trusts X-Forwarded-For only from REAL_IP_FROM networks | Trusting the header always; TRUSTED_PROXIES=2 | Without it every visitor shares the gateway address (shared rate limits, admin lockout); trusting it from anywhere would let clients forge their address | | 2026-09-25 | Sentences the model composes (summary, contract obligations, report findings) are written in the document's own language | Always English | Users read the summary of their own document; an English summary of a Spanish CV looked broken | +| 2026-09-25 | Deploy on push to main: CI opens SSH with a key pinned to deploy/deploy.sh, which builds on the server, health-checks and rolls back | A registry with prebuilt images; a polling timer on the server; Watchtower | No registry to run or pay for; the pinned key can do nothing but redeploy main; pushing from CI keeps deploys gated on tests and visible in GitHub's Deployments | diff --git a/CLAUDE.md b/CLAUDE.md index cd87751..5b255f2 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -13,4 +13,5 @@ | Feature: email verification, passwords, subscriptions, webhooks | `02-DOCS/wiki/ftd/account-lifecycle-billing.md` | | Feature: per-page pricing, page packs, i18n, themes, landing redesign | `02-DOCS/wiki/ftd/pages-i18n-redesign.md` | | Feature: unique visitor counter, real visitor addresses behind the host's Nginx | `02-DOCS/wiki/ftd/visitors-real-ip.md` | +| Feature: continuous deployment (push to main → server, with rollback) | `02-DOCS/wiki/ftd/continuous-deployment.md` | | Production deployment guide | `docs/DEPLOY.md` | diff --git a/deploy/deploy.sh b/deploy/deploy.sh new file mode 100755 index 0000000..17639f4 --- /dev/null +++ b/deploy/deploy.sh @@ -0,0 +1,101 @@ +#!/usr/bin/env bash +# Deploy the latest main on this server: pull, build, restart, wait until healthy, and roll back +# to the previous version if the new one does not come up. +# +# deploy/deploy.sh deploy if origin/main has new commits +# deploy/deploy.sh --force rebuild and restart even when nothing changed +# +# GitHub Actions runs it over SSH after CI passes on main (see docs/DEPLOY.md, "Automatic +# deploys"); it can also be run by hand. Only one deploy runs at a time. Database migrations run +# when the web container starts and are not undone by a rollback, so keep them backward +# compatible (add columns and tables; drop them in a later release). +set -euo pipefail + +BRANCH=main +IMAGES=(pdf-process-pipeline extracta-frontend) +HEALTH_TIMEOUT=${DEPLOY_HEALTH_TIMEOUT:-240} # seconds; Docker health checks run every 30 s +HEALTH_POLL=${DEPLOY_HEALTH_POLL:-5} + +cd "$(dirname "$0")/.." + +log() { printf '%s %s\n' "$(date -u +%H:%M:%S)" "$*"; } +fail() { + log "ERROR: $*" + exit 1 +} + +# One deploy at a time: a second push waits for the first deploy, then deploys what is newest. +if command -v flock >/dev/null 2>&1; then + exec 9>/tmp/extracta-deploy.lock + flock -w 1800 9 || fail "another deploy has been running for 30 minutes" +fi + +[[ -f .env ]] || fail ".env is missing: copy .env.sample and fill it in first" +[[ "$(git rev-parse --abbrev-ref HEAD)" == "$BRANCH" ]] || fail "this checkout is not on $BRANCH" + +git fetch --quiet origin "$BRANCH" +old=$(git rev-parse HEAD) +new=$(git rev-parse "origin/$BRANCH") +if [[ "$old" == "$new" && "${1:-}" != "--force" ]]; then + log "already at ${new:0:7}: nothing to deploy" + exit 0 +fi + +log "deploying ${old:0:7} -> ${new:0:7}" +git log --oneline "$old..$new" | sed 's/^/ /' + +# Keep the running images as :previous so a failed deploy can go back to them in seconds. +for image in "${IMAGES[@]}"; do + if docker image inspect "$image:latest" >/dev/null 2>&1; then + docker image tag "$image:latest" "$image:previous" + fi +done + +git merge --ff-only --quiet "origin/$BRANCH" || fail "cannot fast-forward to origin/$BRANCH (local changes on the server?)" + +log "building images (the site keeps running the old version meanwhile)" +if ! docker compose build; then + git reset --hard --quiet "$old" + fail "build failed: nothing was restarted, the server still runs ${old:0:7}" +fi + +log "starting the new containers" +docker compose up -d --remove-orphans + +healthy() { + local service id status + for service in web frontend; do + id=$(docker compose ps -q "$service") + [[ -n "$id" ]] || return 1 + status=$(docker inspect -f '{{if .State.Health}}{{.State.Health.Status}}{{else}}{{.State.Status}}{{end}}' "$id") + [[ "$status" == "healthy" ]] || return 1 + done + id=$(docker compose ps -q worker) # no health check: it must at least be running + [[ -n "$id" && "$(docker inspect -f '{{.State.Running}}' "$id")" == "true" ]] +} + +log "waiting for web, frontend and worker to be healthy (up to ${HEALTH_TIMEOUT}s)" +waited=0 +until healthy; do + if ((waited >= HEALTH_TIMEOUT)); then + log "the new version did not become healthy; last web logs:" + docker compose logs --tail 30 web | sed 's/^/ /' || true + log "rolling back to ${old:0:7}" + git reset --hard --quiet "$old" + for image in "${IMAGES[@]}"; do + if docker image inspect "$image:previous" >/dev/null 2>&1; then + docker image tag "$image:previous" "$image:latest" + fi + done + docker compose up -d --remove-orphans + fail "deploy of ${new:0:7} failed and was rolled back to ${old:0:7}" + fi + sleep "$HEALTH_POLL" + waited=$((waited + HEALTH_POLL)) +done + +# Free disk: untagged images from older builds and week-old build cache. The :previous images stay +# for the next rollback. +docker image prune --force >/dev/null +docker builder prune --force --filter until=168h >/dev/null +log "deployed ${new:0:7} in production" diff --git a/docs/DEPLOY.md b/docs/DEPLOY.md index 1dd3b76..384449b 100644 --- a/docs/DEPLOY.md +++ b/docs/DEPLOY.md @@ -254,14 +254,45 @@ an event was lost. A refund made from PayPal's dashboard ends the plan time it p ```bash cd /opt/extracta -git pull -docker compose build -docker compose up -d -./scripts/docker-cleanup.sh +./deploy/deploy.sh ``` -Migrations run automatically on start. Rebuilding also installs the latest Debian security -patches into the images. +It pulls `main`, builds the images while the old version keeps serving, restarts the containers +and waits until `web`, `frontend` and `worker` are healthy. If the new version does not come up +within 4 minutes it goes back to the previous images and commit on its own. A build failure +restarts nothing. `./deploy/deploy.sh --force` rebuilds even when `main` has not changed (for +example, to pick up new Debian security patches). Migrations run automatically on start and are +not undone by a rollback, so keep them backward compatible. + +### Automatic deploys (push to main → production) + +The `deploy` job in `.github/workflows/ci.yml` runs `deploy/deploy.sh` on the server over SSH +after every push to `main` whose backend and frontend checks pass. It stays skipped until these +steps are done once. + +1. **A key that can only deploy.** On the server: + + ```bash + ssh-keygen -t ed25519 -N "" -C github-actions-deploy -f ~/.ssh/extracta_deploy + echo "command=\"$PWD/deploy/deploy.sh\",no-port-forwarding,no-X11-forwarding,no-agent-forwarding,no-pty $(cat ~/.ssh/extracta_deploy.pub)" >> ~/.ssh/authorized_keys + ``` + + The `command=` prefix pins the key to the deploy script: whoever holds it can trigger a deploy + of what is already on `main`, and nothing else. Run the `echo` from the project folder. +2. **The server's identity**, so GitHub refuses to talk to an impostor: + + ```bash + echo "YOUR_SERVER_IP $(cut -d' ' -f1,2 /etc/ssh/ssh_host_ed25519_key.pub)" + ``` + +3. **In GitHub** → the repository → **Settings → Secrets and variables → Actions**: + - Secrets: `DEPLOY_SSH_KEY` = the whole output of `cat ~/.ssh/extracta_deploy` (then delete + that file from the server: `rm ~/.ssh/extracta_deploy`); `DEPLOY_KNOWN_HOSTS` = the line + from step 2. + - Variables: `DEPLOY_HOST` = the server's IP; `DEPLOY_USER` = `root` (or the user that owns the + project folder). +4. Push to `main`. **Actions** shows the run, and **Deployments → production** keeps the history. + A red deploy job means the server kept (or went back to) the previous version: its log says why. ## 11. Operations