From a939f7daee7f88a0d74e256f73f27230ebde0ecc Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?=EA=B9=80=EA=B8=B0=EB=AF=BC?= Date: Tue, 29 Sep 2026 18:40:12 +0900 Subject: [PATCH 1/3] [Test] Reproduce stale admin role through OpenProxy replica lag Add guarded Patroni-managed standby delay, HTTP permission controls, and redacted two-run evidence. Relates to #362. --- docs/design/opensql-permission-replica-lag.md | 88 ++++ ...opensql-permission-replica-lag-20260929.md | 119 +++++ ...l-permission-routing-preflight-20260929.md | 65 +++ scripts/opensql/permission_fixture.sh | 73 +++ scripts/opensql/permission_node_probe.sh | 69 +++ scripts/opensql/permission_replica_lag.py | 449 ++++++++++++++++++ .../run_permission_replica_lag_local.py | 131 +++++ scripts/opensql/standby_apply_delay_guard.sh | 188 ++++++++ 8 files changed, 1182 insertions(+) create mode 100644 docs/design/opensql-permission-replica-lag.md create mode 100644 docs/test-results/opensql-permission-replica-lag-20260929.md create mode 100644 docs/test-results/opensql-permission-routing-preflight-20260929.md create mode 100755 scripts/opensql/permission_fixture.sh create mode 100755 scripts/opensql/permission_node_probe.sh create mode 100644 scripts/opensql/permission_replica_lag.py create mode 100644 scripts/opensql/run_permission_replica_lag_local.py create mode 100755 scripts/opensql/standby_apply_delay_guard.sh diff --git a/docs/design/opensql-permission-replica-lag.md b/docs/design/opensql-permission-replica-lag.md new file mode 100644 index 00000000..21238721 --- /dev/null +++ b/docs/design/opensql-permission-replica-lag.md @@ -0,0 +1,88 @@ +# OpenProxy 권한 조회와 복제 지연 재현 설계 + +## 목적과 경계 + +관리자 역할 회수 API가 성공한 뒤 시작한 `/admin/**` 요청이 OpenProxy를 통해 +뒤처진 standby의 옛 ADMIN 역할을 읽고 허용되는지 판정한다. 이 변경은 운영 인가 +코드를 수정하지 않는다. Redis 캐시 부활 경쟁과 WebSocket 세션 권한은 별도 과제다. + +`JwtAuthenticationFilter`는 JWT의 `userId`로 `RoleAuthorityService`를 호출한다. +Redis `auth:roles:{userId}`가 없으면 `UserRoleRepository.findRoleCodesByUserId`가 +DB에서 역할을 읽고 30초간 캐시한다. `UserRoleCommandService.revokeRole`은 DB +커밋 뒤 역할 캐시를 지운다. 권한 SQL이 standby에 도착하는지는 추측하지 않고 +실제 HTTP 경로와 노드별 통계로 판정한다. + +## 사전 조건과 중단 기준 + +- 승인된 GCP 계정·프로젝트의 기존 세 VM만 사용한다. 실험용 앱은 기존 앱과 + 별도 프로세스로 실행하고, Flyway·Worker·Dispatcher를 끈다. +- primary 1대와 streaming standby 2대를 확인한다. 이미 다른 복제 지연이나 + 장애·부하 시험이 실행 중이면 시작하지 않는다. +- 역할 SQL의 standby 도착을 먼저 입증하지 못하면 지연을 적용하지 않는다. +- 두 standby 호스트에 시험 JVM·SSH 연결과 독립된 자동 원복 타이머, 원본 + Patroni YAML 백업을 준비한다. 단독 standby에서 자동 원복을 확인한 뒤에만 + 두 노드에 지연을 적용한다. `finally`와 수동 원복은 추가 방어선이다. +- 지연 동안 수신·재생 LSN 차이와 디스크 여유를 감시한다. 미리 정한 상한, + 역할 변경, 타이머 오작동, 관측 누락 중 하나라도 발생하면 즉시 재개한다. +- 자동 원복 장치가 아직 검증되지 않았다면 클러스터 실험은 실행하지 않는다. + +### 2026-09-29 실제 사전 시험에서 발견한 제약 + +이 클러스터의 Patroni 4.0.5는 SQL로 일시정지한 standby WAL replay를 다음 +HA 루프에서 직접 재개했다. 두 standby의 로그 모두 `Resuming paused WAL replay for +PostgreSQL 14+`를 기록했고, 실제 재생 중단은 약 2~3초만 유지됐다. 따라서 아래 +`pg_wal_replay_pause()` 기반 M·C2 순서는 **현재 구성에서 실행 불가**다. 독립 +systemd 자동 재개가 정상이라는 사실은 이 제약을 해결하지 않는다. 현재 러너는 +라우팅 기준선 R을 측정한 뒤 기본적으로 M 진입을 거부한다. + +직접 `ALTER SYSTEM`으로 준 `recovery_min_apply_delay` 역시 Patroni가 약 +1초 뒤 제거했다. 이 값은 standby의 Patroni 로컬 +`postgresql.recovery_conf`에서 관리하고, Patroni REST `POST /reload`를 +사용해야 유지된다. 호스트 타이머가 원본 YAML을 정확히 되돌리고 reload한다. +지연된 standby가 failover 후보가 되는 위험을 최소화하도록 짧은 관측 구간을 +두며, 이 시험 중 primary 장애를 의도적으로 주입하지 않는다. `patronictl +pause`나 Patroni 중지로 HA 제어 자체를 우회하지 않는다. + +## 독립 시나리오 + +| 이름 | 조건 | 유효성 및 판정 | +| --- | --- | --- | +| R | 정상 복제, 앱 OpenProxy URL | 대상 사용자 Redis 키를 매 요청 전에 비우고 역할 SQL의 노드별 호출 수 차이를 기록한다. | +| M | 두 standby의 `recovery_min_apply_delay` 적용, 앱 OpenProxy URL | 회수 API 2xx, primary에서 ADMIN 없음, 양쪽 standby에 ADMIN 있음, 대상 Redis 키 없음이 선행 조건이다. 그다음 첫 `/admin/workers` 상태, 실제 라우팅, Redis 값·PTTL을 기록한다. | +| C1 | standby가 회수를 반영한 뒤, 앱 OpenProxy URL | 별도 시험 사용자와 비어 있는 Redis 키로 403인지 확인한다. | +| C2 | 두 standby는 지연, 앱은 검증된 primary 직결 URL | 별도 시험 사용자와 비어 있는 Redis 키로 403인지 확인한다. | + +각 실행의 관리자와 대상 사용자 이메일은 서로 다른 무작위 run ID를 포함한다. +JWT는 요청을 받는 시험 앱과 같은 서명 설정으로 발급하며 토큰을 기록하지 않는다. +대상 사용자의 회수 전 `/admin/workers` 200을 확인한다. 관리자 API의 401이나 +5xx는 권한 차단으로 해석하지 않는다. + +## 본 실험 순서 + +1. 두 시험 사용자를 primary에 만들고 양쪽 standby에서도 ADMIN이 보이는지 확인한다. +2. 독립 자동 원복 타이머가 무장된 것을 확인한 후 양쪽 standby의 Patroni 로컬 + 설정에 apply delay를 추가하고 Patroni를 reload한다. 양쪽의 실제 + `pg_settings` 값이 동일하며 타이머가 예정되어 있을 때만 다음 단계로 간다. +3. 관리자 HTTP API로 대상 ADMIN을 회수한다. 회수 응답과 세 노드의 역할 행, + Redis 키 부재를 확인한다. 회수 뒤 읽은 primary LSN은 정확한 commit LSN이 + 아닌 참조 시점으로만 사용한다. +4. 대상 사용자로 `/admin/workers`를 한 번 호출하고 상태 코드, 역할 SQL의 + 노드별 증가량, Redis의 ADMIN 여부와 PTTL을 기록한다. +5. 즉시 양쪽 standby의 원본 Patroni 설정을 복원한다. 회수가 반영된 이후의 첫 403 시각을 + 1초 간격으로 관측해 복제 지연 구간과 캐시 잔류 구간을 구분한다. +6. 성공·실패와 관계없이 설정 원복·복제·단일 primary를 확인하고 이번 실행의 + 사용자·역할·Redis 키만 정리한다. 정상 원복 뒤 자동 원복 예약을 해제한다. + +## 증거와 완료 조건 + +기존 `ha_evidence.py`의 run ID와 요청·fault 이벤트를 재사용한다. 추가 +비식별 증거에는 코드 커밋, 설정 해시, UTC 시각, 노드 별칭, 실제 역할 상태, +실효 지연값, receive/replay LSN, SQL 호출 수 차이, HTTP 상태, Redis ADMIN 여부와 +PTTL, 자동 원복 발동 여부를 기록한다. 비밀번호·JWT·응답 본문·주소·실제 +사용자 정보는 공개 파일에 넣지 않는다. + +결과는 `STALE_ADMIN_ALLOWED`, `DENIED`, `INVALID`로 나눈다. `DENIED`는 +이번 환경에서 재현되지 않았다는 뜻이며 향후 설정 변경까지 안전하다는 증명은 +아니다. `pg_stat_statements`는 노드별 누적 통계이므로 다른 요청과 분리할 수 +없으면 `INVALID`로 처리한다. 이 PR의 성공은 취약점 재현 자체가 아니라 네 +시나리오의 유효한 판정, 원본 증거, 정상 원복 및 한계 문서화다. diff --git a/docs/test-results/opensql-permission-replica-lag-20260929.md b/docs/test-results/opensql-permission-replica-lag-20260929.md new file mode 100644 index 00000000..4c6188bf --- /dev/null +++ b/docs/test-results/opensql-permission-replica-lag-20260929.md @@ -0,0 +1,119 @@ +# OpenProxy 권한 조회의 복제 지연 재현 + +## 결론 + +2026-09-29 UTC, 기존 OpenSQL 3노드에서 DocGrid의 실제 HTTP 관리자 인가 경로를 +두 번 실행했다. 두 번 모두 **관리자가 ADMIN 역할을 회수한 뒤**, primary에는 +ADMIN이 없고 두 standby에는 아직 ADMIN이 남아 있는 동안 OpenProxy 경유 +`GET /admin/workers`가 **200**을 반환했다. 해당 요청의 역할 SQL은 standby에 +도착했고 Redis에 옛 ADMIN이 다시 캐시됐다. 같은 시점 primary 직결 대조군은 +**403**, 복제 정상화 후 OpenProxy 대조군도 **403**이었다. + +이 결과는 OpenProxy의 오작동이라는 주장이 아니다. 읽기 분산이 가능한 경로에서 +최신 역할 판정을 수행한 **애플리케이션의 일관성 경계 문제**를 재현한 것이다. +Redis의 조회·DB 조회·재저장 사이 경쟁과 WebSocket 장기 세션은 이번 시험의 +원인이거나 해결책으로 판정하지 않았다. 수정은 별도 작업에서 해야 한다. + +## 범위·실행 기준 + +- OpenSQL `v3.17.8.7`, PostgreSQL 17.8, Patroni 4.0.5의 기존 VM 3대. + 제품·설정 기준 해시: + `0a0b86eff71c4c24740a99ea0f334fca70e5c11e1a813c3a3bc646a825549794`. + 앱 JAR의 backend 코드 기준 커밋은 `2d4d56619280`이다. +- 앱 두 인스턴스와 Redis는 **시험을 실행한 Mac의 localhost**에서만 기동했다. + 하나는 OpenProxy A/B JDBC URL, 다른 하나는 primary 직결 URL을 사용했다. + 둘의 Hikari 풀이 분리되며 Flyway·Worker·Dispatcher는 껐다. SSH 터널은 + 접속 수단이므로 이 시험은 GCP 내부 부하·장애 전환 성능 측정이 아니다. +- 시험 계정 네 개는 실행별 무작위 식별자를 가진 일회용 계정이다. 인증 토큰은 + 두 시험 JVM과 동일한 비공개 키로 메모리에서만 서명했다. 공개 증거에는 + 계정·프로젝트 ID·IP·JDBC URL·비밀번호·토큰·응답 본문을 넣지 않았다. +- 원본 요청 원장, 설정 manifest, 앱 로그와 상세 JSON은 Git 밖의 소유자 전용 + 디렉터리에 남겼다. 아래 SHA-256은 두 번째 실행의 로컬 원본 확인용이다. + `permission-scenarios.json`: + `1f83f13a518e9b8ea6cc83e5afd43559a4a81d5535c5fc2e5b10628864d5c9a5`; + `events.jsonl`: + `0a028ff133533e4507f1e6d0034233d4588967b9d6871f3dedb5dbb76fbff9aa`. + +## 왜 복제 지연을 이렇게 만들었나 + +첫 사전 시험의 `pg_wal_replay_pause()`는 Patroni가 다음 HA 루프에서 재개했다. +`ALTER SYSTEM SET recovery_min_apply_delay`도 Patroni가 약 1초 뒤 제거했다. +둘 다 유효한 지연 관문이 아니다. +[사전 시험 상세](opensql-permission-routing-preflight-20260929.md)에 실패 사례를 +남겼다. Patroni가 관리하는 **각 standby의 로컬 +`postgresql.recovery_conf.recovery_min_apply_delay`**에 한시적으로 120초를 +넣고 Patroni `POST /reload`로 적용했다. PostgreSQL은 이 값으로 commit WAL +적용을 늦춘다. [PostgreSQL 설정 설명](https://www.postgresql.org/docs/17/runtime-config-replication.html), +[Patroni 로컬 설정·reload](https://patroni.readthedocs.io/en/latest/patronictl.html). + +적용 전 각 VM 호스트에 원본 Patroni YAML의 root-only 백업과 독립된 +`systemd-run` 240초 원복 타이머를 만들었다. 시험 프로세스의 `finally`도 +원본 파일 복원·Patroni reload·실효값 0 확인·타이머 해제를 수행한다. +단독 node2에서는 30초 지연 중 새 역할이 node3에만 먼저 보였다가 node2에도 +반영됐다. 단독 node3에서는 45초 타이머 서비스가 **종료된 뒤** +`delay=0,default`, streaming, backlog 0을 확인했다. 타이머 만료 직후 +서비스가 아직 `active/running`일 때 지연이 보인 것은 원복 실패로 판정하지 +않았다. VM 자체가 재부팅되는 경우 `/run` 백업과 transient timer가 +보존되는지는 검증하지 않았으므로 이 가드는 운영용 자동화가 아니다. + +## 실행 명령·관측 + +실제 계정·프로젝트·키·개인 `.env` 위치는 실행 시 비공개 환경변수로 전달한다. +명령은 시험용 코드의 위치를 보여주기 위한 축약형이며 비밀값을 포함하지 않는다. + +| 실행 위치 | 명령 또는 코드 경로 | 이유 | 두 번째 실행 결과 요약 | +| --- | --- | --- | --- | +| 이 Mac | `python3 scripts/opensql/run_permission_replica_lag_local.py --env-file --jar --output --apply-delay` | OpenProxy 경유·primary 직결 앱을 독립 JVM으로 기동하고 한 번의 원장에 연결 | 두 JVM 기동, 종료 시 둘 다 정지. 총 결과 `STALE_ADMIN_ALLOWED` | +| node1~3 VM 호스트 | `sudo /usr/local/sbin/docgrid-permission-node-probe health ` | 한 primary·두 streaming standby와 `postgres` DB의 `pg_stat_statements` 확인 | node1 `f,t,2,0`; node2·3 `t,t,0,1` (recovery, extension, sender, receiver) | +| node1 VM 호스트 | `sudo /usr/local/sbin/docgrid-permission-fixture create ` | 실사용자와 분리된 ADMIN 시험 계정 네 개 생성 | 회수 전 각 사용자로 관리자 API 200; 양 standby에도 ADMIN 복제 확인 | +| 이 Mac → 프록시 앱 | Redis `DEL auth:roles:{id}` 후 `GET /admin/workers` 6회 | 캐시 미스를 강제해 역할 SQL의 실제 라우팅 확인 | R: HTTP 200×6, 역할 SQL 호출 증가 node1 +0, node2 +5, node3 +1 | +| node2·3 VM 호스트 | `standby-apply-delay-guard arm 240 120` → `apply ...` → `status ...` | 원본 백업과 독립 원복을 먼저 건 뒤 양 standby에 120초 apply delay 적용 | 두 노드 모두 `armed=true`, `delay=120000,configuration file`, streaming | +| 이 Mac → 프록시 앱 | 관리자 `DELETE /admin/users/{id}/roles/ADMIN` | 앱 트랜잭션·afterCommit 캐시 무효화를 실제 HTTP로 통과 | M·C2 회수 API 모두 200; primary ADMIN 없음, 양 standby ADMIN 있음, 대상 Redis 키 없음 | +| 이 Mac → 프록시 앱 | M 사용자 `GET /admin/workers`, 전후 `role-stats`, Redis `GET/PTTL` | 뒤처진 읽기의 보안 영향 확인 | HTTP **200**; 역할 SQL node3 +1·나머지 +0; Redis ADMIN 재저장, PTTL 24,900ms | +| 이 Mac → primary 직결 앱 | C2 사용자 `GET /admin/workers`, 전후 `role-stats` | 같은 지연 상태에서 최신 primary 판정과 대조 | HTTP **403**; 역할 SQL node1 +1·standby +0; Redis ADMIN 없음 | +| node1~3 VM 호스트 | `permission-node-probe lsn `, 가드 `status` | 수신은 진행됐지만 standby 적용이 뒤처졌는지 확인 | 회수 뒤 primary/receive `0/5A45BE8`, 양 standby replay `0/5A45B28`; 관측 뒤 backlog 각 456B | +| node2·3 VM 호스트 | `standby-apply-delay-guard restore 240 120` → `cancel ...` | 원본 Patroni 설정·정상 복제 복원 | 두 노드 `delay=0,default`, streaming, backlog 0, 타이머 비활성 | +| 이 Mac → 프록시 앱 | 복제 반영 후 C1의 관리자 역할 회수·`GET /admin/workers` | 지연이 사라지면 같은 프록시 경로가 403인지 대조 | 회수 API 200, C1 요청 **403**, 세 노드 모두 ADMIN 없음 | +| node1 VM 호스트 | `permission-fixture remove/status ` | 이번 실행의 계정·역할만 삭제 | `status` 조회 행 0개. 임시 Redis·앱·터널도 종료 | + +두 번째 실행에서는 09:30:09 UTC에 M 회수 200, 09:30:36에 M 요청 200, +09:30:41에 C2 요청 403이었다. 그 시점 양 standby의 WAL receive LSN은 +`0/5A45CF0`, replay LSN은 `0/5A45B28`이었다. 최종 양 노드의 receive와 +replay는 모두 `0/5A479E0`으로 같아졌고 원본 지연값 0을 유지했다. 이 LSN은 +커밋의 정확한 LSN이 아니라 각 관측 시점의 위치다. + +시험 종료 후에도 `postgres` DB에 앞선 사전 시험에서 만든 +`pg_stat_statements` 확장과 VM의 root 전용 probe/fixture/guard 스크립트는 +남아 있다. 일회용 계정·Redis 프로세스·시험 JVM·SSH 터널·Patroni 지연 설정은 +남기지 않았다. 공개 스크립트의 최종 호스트 볼륨 자동 탐색 변경은 지연을 +재적용하지 않고 두 standby의 read-only `status`로 확인했다. + +## 두 실행의 교차 확인 + +| 항목 | 첫 유효 실행 | LSN 수집을 보강한 재실행 | +| --- | --- | --- | +| 정상 상태 역할 SQL 도착 | primary 0 / standby 2·4 | primary 0 / standby 5·1 | +| 권한 회수 후 M | 200, standby node2 +1, Redis ADMIN, PTTL 24,586ms | 200, standby node3 +1, Redis ADMIN, PTTL 24,900ms | +| 같은 지연 중 C2(primary 직결) | 403, primary +1 | 403, primary +1 | +| 복제 복원 뒤 C1(OpenProxy) | 403 | 403 | +| 지연 중 receive-replay 차이 | 두 standby 각 192B | 두 standby 각 456B | +| 종료 상태 | 양 standby 지연 0·streaming·backlog 0, 시험 계정 0 | 동일; 최종 receive/replay LSN 일치 | + +LSN 수집을 처음 추가한 실행은 사전 검사에서 `docgrid` DB에도 +`pg_stat_statements` 확장이 있어야 한다고 잘못 요구해 중단됐다. 실제 통계 +조회는 확장을 설치한 `postgres` DB에서 다른 DB의 `dbid`로 수행한다. +검사만 수정했고, 이 중단된 실행은 fixture 생성·지연 적용 전에 끝났으므로 +위의 유효 실행 횟수에 포함하지 않는다. + +## 판정의 한계와 다음 조치 + +`pg_stat_statements`는 누적 통계라 단일 SQL과 단일 HTTP 요청을 직접 연결하는 +추적 ID는 아니다. 다만 시험 JVM 외의 Worker·Dispatcher를 끄고 캐시를 매번 +비웠으며, M/C2 직전·직후 카운터를 비교했고 DB 계정별 통계를 사용했다. +이 인과관계와 두 대조군은 이번 환경에서 **복제 지연 중 stale ADMIN 허용**을 +지지한다. 모든 부하·모든 라우팅·장애 상황에서 같은 빈도라는 주장은 아니다. + +이 시험은 HTTP 권한 수정이 아니다. 안전한 읽기 분산 경계는 권한 판정을 +primary로 보내는 것과, Redis cache resurrection 경쟁을 별도로 막는 것을 +함께 검토해야 한다. WebSocket 기존 구독, 실제 OpenProxy 프로세스 장애, +Patroni failover, 앱 VM의 GCP 내부 부하 시험은 여기서 검증하지 않았다. diff --git a/docs/test-results/opensql-permission-routing-preflight-20260929.md b/docs/test-results/opensql-permission-routing-preflight-20260929.md new file mode 100644 index 00000000..0cb9e875 --- /dev/null +++ b/docs/test-results/opensql-permission-routing-preflight-20260929.md @@ -0,0 +1,65 @@ +# OpenSQL 권한 조회 라우팅 사전 시험과 Patroni 재생 제약 + +## 판정 + +2026-09-29 UTC, 기존 3노드 OpenSQL·OpenProxy 환경에서 실제 DocGrid HTTP +인가 경로의 **권한 SQL이 standby에 도착함**을 확인했다. Redis 키를 비우고 +`GET /admin/workers`를 6회 호출했을 때 모두 200이었고, 노드별 +`pg_stat_statements`의 해당 SQL 호출 수는 primary 0, standby A 5, +standby B 1 증가했다. 이것은 정상 복제 상태의 라우팅 증거일 뿐, 권한 회수 +후 stale ADMIN 허용의 증거는 아니다. + +두 standby의 WAL replay를 동시에 멈추려는 다음 관문은 통과하지 못했다. +각 standby에서 `pg_wal_replay_pause()`는 잠시 `paused`가 되었지만, +Patroni 4.0.5가 다음 HA 루프에서 직접 재개했다. 노드 로그에는 각각 +08:15:37.302/08:15:38.291 `recovery has paused`, 08:15:40.503/08:15:40.504 +`Resuming paused WAL replay for PostgreSQL 14+`가 남았다. **회수 API, +stale-role 판정, C1·C2 대조군은 실행하지 않았다.** 이 실행의 결과는 +`INVALID`이며 보안 취약점이나 안전성을 주장할 수 없다. + +## 실행 경로와 증거 + +공개 문서에는 계정·프로젝트 ID·IP·JDBC URL·비밀번호·JWT·사용자 이메일을 +기록하지 않는다. 원본 개인 실행 원장과 두 앱 로그는 Git 밖의 권한 제한된 +로컬 출력 디렉터리에 보관한다. 이 실행은 아직 커밋되지 않은 작업 트리에서 +수행했으므로 특정 Git SHA로 완전 재현했다고 주장하지 않는다. + +| 수행 위치 | 명령·동작 | 이유 | 결과 요약 | +| --- | --- | --- | --- | +| VM host, node1~3 | `sudo /usr/local/sbin/docgrid-permission-node-probe health docgrid-nodeN` | primary 1대·streaming standby 2대와 노드별 통계 확장을 확인 | node1 `f,t,2,0`; node2·3 `t,t,0,1` (순서: recovery, extension, sender 수, receiver 수) | +| VM host, node2·3 | `sudo /usr/local/sbin/docgrid-standby-replay-guard arm ...` 후 `status`·`cancel` | JVM·SSH와 독립된 자동 재개 예약 검증 | 양쪽 모두 예정 타이머를 확인했고 정상 상태에서 취소 가능 | +| VM host, node1 | `sudo /usr/local/sbin/docgrid-permission-fixture create ...` → `remove` → `status` | 실행별 격리 계정의 생성·정리 검증 | 시험 계정 4개 생성, 삭제 후 조회 행 0개 | +| 이 Mac | 프록시 경유 앱과 primary 직결 앱을 별도 JVM·Hikari 풀로 localhost에 기동 | HTTP 인가 경로를 통과하면서 C2 경로를 분리 | 두 앱 기동; 본 실행에서는 C2 요청 전 중단 | +| 이 Mac → 앱 | 매 요청 전 Redis `DEL auth:roles:{testUserId}` 후 `GET /admin/workers` 6회 | DB 조회가 실제 발생하게 강제 | HTTP 200 6회, Redis 역할 캐시 생성 확인 | +| VM host, node1~3 | `sudo /usr/local/sbin/docgrid-permission-node-probe role-stats docgrid-nodeN` 전후 비교 | 권한 SQL의 물리 도착 노드 판정 | primary +0, standby A +5, standby B +1 | +| VM host, node2·3 | 자동 재개 예약 후 `pg_wal_replay_pause()` 및 `status` | M 시나리오 선행 조건인 양쪽 `paused` 확인 | Patroni가 두 노드를 약 2~3초 안에 재개하여 관문 실패 | +| VM host, node1~3 | fixture `status`, standby guard `status` | 실패 후 원복 검증 | 시험 계정 행 0개; node2·3 모두 `streaming=true`, `replay=not paused`, `armed=false`, backlog 0B | + +`pg_stat_statements`는 누적 통계다. 시험 계정의 Redis 키를 각 요청 직전에 +비우고 백그라운드 Worker·Dispatcher를 끈 별도 앱을 사용했지만, 호출 수 +증가만으로 단일 요청과 SQL을 1:1 대응시킬 수는 없다. 이 한계는 후속 실험에 +남겨둔다. + +## 원인과 다음 관문 + +PostgreSQL의 수동 WAL replay pause를 Patroni가 정상 HA 루프에서 해제한다. +Patroni 로그와 [공식 릴리스 노트](https://patroni.readthedocs.io/en/latest/releases.html)의 +PostgreSQL 14 이상 동작 설명이 일치한다. 타이머는 연결이 끊겼을 때 복구하기 +위한 안전장치였고, 이번 조기 재개의 원인은 타이머가 아니라 Patroni였다. + +따라서 SQL pause를 계속 반복하거나 Patroni를 중지하는 방식으로 결과를 +만들지 않는다. 이 사전 시험 시점에는 라우팅 기준선 R만 **검증**했고 +권한 회수 후 stale-role M·C1·C2는 **미검증**이었다. + +### 후속 실험 기록 + +같은 날 `ALTER SYSTEM SET recovery_min_apply_delay`를 node2 단독으로 시험했지만 +Patroni가 약 1초 뒤 해당 recovery 파라미터를 제거해 실효값이 기본 0으로 +돌아갔다. 이것 역시 타이머 원복이 아니라 Patroni의 설정 소유권 때문이었다. +이 시도에서 M·C1·C2는 실행하지 않았고 일회용 계정은 삭제했다. + +이후 각 standby의 Patroni **로컬** `postgresql.recovery_conf`를 원본 백업과 +독립 원복 타이머 아래에서 잠시 변경하는 방법을 단독 노드에서 검증했다. +그 방법으로 두 번의 유효한 HTTP 권한 시험을 마쳤으며 결과와 원복 증거는 +[복제 지연 재현 결과](opensql-permission-replica-lag-20260929.md)에 별도로 +기록했다. 이 문서의 `INVALID`는 첫 사전 시험에만 해당한다. diff --git a/scripts/opensql/permission_fixture.sh b/scripts/opensql/permission_fixture.sh new file mode 100755 index 00000000..272c16c8 --- /dev/null +++ b/scripts/opensql/permission_fixture.sh @@ -0,0 +1,73 @@ +#!/usr/bin/env bash +# Create or remove only this run's four permission-test users on the current primary. +set -euo pipefail + +action="${1:?action required: create|remove|status}" +container="${2:?primary container required}" +run_id="${3:?run ID required}" + +if [[ "$container" != 'docgrid-node1' ]] || (( EUID != 0 )); then + echo 'Run as root on the known primary container only.' >&2 + exit 2 +fi +if [[ ! "$run_id" =~ ^[a-z0-9][a-z0-9-]{0,23}$ ]]; then + echo 'Invalid run ID.' >&2 + exit 2 +fi + +query() { + /usr/bin/docker exec "$container" sh -c ' + . /var/lib/docgrid/opensql/etc/credentials.env + export PGPASSWORD="${PG_SUPERUSER_PASSWORD}" + exec /var/lib/docgrid/opensql/bin/psql -h 127.0.0.1 -U postgres -d docgrid \ + -X -v ON_ERROR_STOP=1 -At -F , -c "$1" + ' sh "$1" +} + +if [[ "$(query 'SELECT pg_is_in_recovery()')" != 'f' ]]; then + echo 'Known fixture host is not primary; stop instead of writing elsewhere.' >&2 + exit 1 +fi + +emails="('ha-${run_id}-admin@invalid.example', 'ha-${run_id}-m@invalid.example', 'ha-${run_id}-c1@invalid.example', 'ha-${run_id}-c2@invalid.example')" +case "$action" in + create) + # 1. A unique run ID confines writes to four disposable accounts and the existing ADMIN role. + if [[ "$(query "SELECT count(*) FROM users WHERE email IN $emails")" != '0' ]]; then + echo 'Fixture already exists; choose a new run ID.' >&2 + exit 1 + fi + query "BEGIN; + INSERT INTO users (email, password_hash, name, status) VALUES + ('ha-${run_id}-admin@invalid.example', '!disabled!', 'HA test admin', 'ACTIVE'), + ('ha-${run_id}-m@invalid.example', '!disabled!', 'HA test M', 'ACTIVE'), + ('ha-${run_id}-c1@invalid.example', '!disabled!', 'HA test C1', 'ACTIVE'), + ('ha-${run_id}-c2@invalid.example', '!disabled!', 'HA test C2', 'ACTIVE'); + INSERT INTO user_roles (user_id, role_id, assigned_by, assigned_at) + SELECT u.id, r.id, a.id, CURRENT_TIMESTAMP + FROM users u CROSS JOIN roles r + JOIN users a ON a.email = 'ha-${run_id}-admin@invalid.example' + WHERE u.email IN $emails AND r.code = 'ADMIN'; + COMMIT;" >/dev/null + query "SELECT split_part(email, '@', 1), id FROM users WHERE email IN $emails ORDER BY email" + ;; + status) + # 2. Only the run's IDs and role presence leave the database container. + query "SELECT split_part(u.email, '@', 1), u.id, + EXISTS (SELECT 1 FROM user_roles ur JOIN roles r ON r.id = ur.role_id + WHERE ur.user_id = u.id AND r.code = 'ADMIN') + FROM users u WHERE u.email IN $emails ORDER BY u.email" + ;; + remove) + # 3. Remove only mappings and users identified by the exact run-scoped email set. + query "BEGIN; + DELETE FROM user_roles WHERE user_id IN (SELECT id FROM users WHERE email IN $emails); + DELETE FROM users WHERE email IN $emails; + COMMIT;" >/dev/null + echo 'fixture-removed' + ;; + *) + echo 'Unknown fixture action.' >&2 + exit 2 + ;; +esac diff --git a/scripts/opensql/permission_node_probe.sh b/scripts/opensql/permission_node_probe.sh new file mode 100755 index 00000000..0cbd0870 --- /dev/null +++ b/scripts/opensql/permission_node_probe.sh @@ -0,0 +1,69 @@ +#!/usr/bin/env bash +# Read-only, fixed SQL probes for the existing three-node permission experiment. +set -euo pipefail + +action="${1:?action required: health|role|role-stats|lsn|delay-config}" +container="${2:?container required}" +user_id="${3:-}" + +if [[ ! "$container" =~ ^docgrid-node[123]$ ]] || (( EUID != 0 )); then + echo 'Run as root for a known DocGrid container only.' >&2 + exit 2 +fi + +query() { + local database="$1" + local sql="$2" + # 1. The superuser password remains in the container's root-only file. + /usr/bin/docker exec "$container" sh -c ' + . /var/lib/docgrid/opensql/etc/credentials.env + export PGPASSWORD="${PG_SUPERUSER_PASSWORD}" + exec /var/lib/docgrid/opensql/bin/psql -h 127.0.0.1 -U postgres \ + -d "$1" -X -v ON_ERROR_STOP=1 -At -F , -c "$2" + ' sh "$database" "$sql" +} + +case "$action" in + health) + # 2. No role names, addresses or credential values are exported. + query postgres "SELECT pg_is_in_recovery(), + EXISTS (SELECT 1 FROM pg_extension WHERE extname = 'pg_stat_statements'), + (SELECT count(*) FROM pg_stat_replication WHERE state = 'streaming'), + (SELECT count(*) FROM pg_stat_wal_receiver WHERE status = 'streaming')" + query docgrid "SELECT EXISTS ( + SELECT 1 FROM pg_extension WHERE extname = 'pg_stat_statements')" + ;; + role) + if [[ ! "$user_id" =~ ^[1-9][0-9]*$ ]]; then + echo 'A positive numeric test user ID is required.' >&2 + exit 2 + fi + query docgrid "SELECT EXISTS ( + SELECT 1 FROM user_roles ur JOIN roles r ON r.id = ur.role_id + WHERE ur.user_id = $user_id AND r.code = 'ADMIN')" + ;; + role-stats) + # 3. Query text is used only inside the server filter; publish IDs and counts. + query postgres "SELECT queryid, calls FROM pg_stat_statements + WHERE dbid = (SELECT oid FROM pg_database WHERE datname = 'docgrid') + AND userid = (SELECT oid FROM pg_roles WHERE rolname = 'docgrid_app') + AND query ILIKE '%user_roles%' + ORDER BY queryid" + ;; + lsn) + query postgres "SELECT pg_is_in_recovery(), + CASE WHEN pg_is_in_recovery() THEN pg_last_wal_receive_lsn()::text + ELSE pg_current_wal_lsn()::text END, + CASE WHEN pg_is_in_recovery() THEN pg_last_wal_replay_lsn()::text + ELSE pg_current_wal_lsn()::text END" + ;; + delay-config) + query postgres "SELECT setting, unit, context, source, pending_restart, + COALESCE(sourcefile, '') FROM pg_settings + WHERE name = 'recovery_min_apply_delay'" + ;; + *) + echo 'Unknown read-only probe.' >&2 + exit 2 + ;; +esac diff --git a/scripts/opensql/permission_replica_lag.py b/scripts/opensql/permission_replica_lag.py new file mode 100644 index 00000000..56ec1bdf --- /dev/null +++ b/scripts/opensql/permission_replica_lag.py @@ -0,0 +1,449 @@ +#!/usr/bin/env python3 +"""Run the real HTTP permission path against the three-node OpenSQL cluster. + +This is an opt-in, one-shot experiment, not a production service or an automated +CI test. The VM-host rescue timers are independent of this process and SSH. +""" + +from __future__ import annotations + +import argparse +import base64 +import hashlib +import hmac +import json +import os +import re +import socket +import subprocess +import sys +import time +import urllib.error +import urllib.request +import uuid +from concurrent.futures import ThreadPoolExecutor +from datetime import datetime, timezone +from pathlib import Path + + +ROOT = Path(__file__).resolve().parents[2] +EVIDENCE = ROOT / "scripts/opensql/ha_evidence.py" +CONTRACT = ROOT / "docs/test-results/opensql-contract-evidence/contract-manifest.json" +NODES = ("docgrid-node1", "docgrid-node2", "docgrid-node3") +STANDBYS = NODES[1:] +LABEL = re.compile(r"[a-z0-9][a-z0-9-]{0,23}\Z") +DELAY_SECONDS = 120 +GUARD_SECONDS = 240 + + +class InvalidExperiment(RuntimeError): + """Stop a run whose preconditions or evidence cannot support a conclusion.""" + + +def utc_now() -> str: + return datetime.now(timezone.utc).isoformat(timespec="milliseconds").replace("+00:00", "Z") + + +def required(name: str) -> str: + value = os.environ.get(name, "") + if not value: + raise InvalidExperiment(f"Missing {name}") + return value + + +def command(argv: list[str], *, timeout: int = 30) -> str: + result = subprocess.run(argv, cwd=ROOT, capture_output=True, text=True, timeout=timeout) + if result.returncode: + # Do not copy remote stderr: client exceptions may contain endpoint details. + raise InvalidExperiment(f"Command failed: {argv[0]} {argv[1]} (exit {result.returncode})") + return result.stdout.strip() + + +def remote(node: str, script: str, *args: str) -> str: + if node not in NODES or script not in {"docgrid-permission-node-probe", + "docgrid-standby-apply-delay-guard", "docgrid-permission-fixture"}: + raise InvalidExperiment("Unexpected remote target") + for value in args: + if not re.fullmatch(r"[a-z0-9-]+", value): + raise InvalidExperiment("Unsafe remote argument") + return command(["gcloud", "compute", "ssh", node, + f"--zone={required('OPENSQL_GCP_ZONE')}", + f"--project={required('OPENSQL_EXPECTED_PROJECT')}", + f"--ssh-key-file={required('OPENSQL_SSH_KEY')}", "--quiet", + f"--command=sudo /usr/local/sbin/{script} {' '.join(args)}"], timeout=35) + + +def check_target() -> None: + account = command(["gcloud", "auth", "list", "--filter=status:ACTIVE", + "--format=value(account)"]) + project = command(["gcloud", "config", "get-value", "project"]) + if account != required("OPENSQL_EXPECTED_ACCOUNT") or project != required("OPENSQL_EXPECTED_PROJECT"): + raise InvalidExperiment("Active GCP account or project differs from approved target") + + +def redis(*parts: str) -> object: + host = os.environ.get("REDIS_HOST", "127.0.0.1") + port = int(os.environ.get("REDIS_PORT", "6379")) + if host not in {"localhost", "127.0.0.1"}: + raise InvalidExperiment("This local experiment requires loopback Redis") + commands = [] + password = os.environ.get("REDIS_PASSWORD", "") + if password: + commands.append(("AUTH", password)) + commands.append(parts) + with socket.create_connection((host, port), timeout=2) as connection: + connection.settimeout(2) + stream = connection.makefile("rb") + result = None + for values in commands: + payload = f"*{len(values)}\r\n".encode() + for value in values: + data = str(value).encode() + payload += f"${len(data)}\r\n".encode() + data + b"\r\n" + connection.sendall(payload) + prefix = stream.read(1) + line = stream.readline().removesuffix(b"\r\n") + if prefix == b"-": + raise InvalidExperiment("Redis command failed") + if prefix == b"$": + length = int(line) + result = None if length == -1 else stream.read(length).decode() + if length != -1: + stream.read(2) + elif prefix == b":": + result = int(line) + elif prefix == b"+": + result = line.decode() + else: + raise InvalidExperiment("Unexpected Redis response") + return result + + +def cache(user_id: int) -> dict[str, object]: + key = f"auth:roles:{user_id}" + value = redis("GET", key) + return {"present": value is not None, "admin": value is not None and + "ADMIN" in value.split(","), "pttl_ms": redis("PTTL", key)} + + +def token(user_id: int, run_id: str) -> str: + secret = required("JWT_SECRET").encode() + if len(secret) < 32: + raise InvalidExperiment("JWT secret is too short") + algorithm, digest = (("HS512", hashlib.sha512) if len(secret) >= 64 else + ("HS384", hashlib.sha384) if len(secret) >= 48 else + ("HS256", hashlib.sha256)) + + def encoded(value: dict[str, object]) -> str: + return base64.urlsafe_b64encode(json.dumps(value, separators=(",", ":")).encode()).decode().rstrip("=") + + now = int(time.time()) + parts = [encoded({"alg": algorithm, "typ": "JWT"}), + encoded({"sub": f"ha-{run_id}@invalid.example", "userId": user_id, + "jti": str(uuid.uuid4()), "iat": now, "exp": now + 600})] + signature = hmac.new(secret, ".".join(parts).encode(), digest).digest() + return ".".join(parts + [base64.urlsafe_b64encode(signature).decode().rstrip("=")]) + + +def ledger(run_dir: Path, action: str, *args: str) -> str: + return command([sys.executable, str(EVIDENCE), action, "--run-dir", str(run_dir), *args]) + + +def http(run_dir: Path, base_url: str, path: str, jwt: str, operation: str, + method: str = "GET") -> int: + request_id = ledger(run_dir, "sent", "--operation", operation) + request = urllib.request.Request(base_url + path, method=method, + headers={"Authorization": "Bearer " + jwt}) + try: + with urllib.request.urlopen(request, timeout=5) as response: + status = response.status + except urllib.error.HTTPError as error: + status = error.code + except (urllib.error.URLError, TimeoutError): + ledger(run_dir, "unknown", "--request-id", request_id, "--reason", "connection_lost") + raise InvalidExperiment("HTTP outcome unknown") from None + ledger(run_dir, "ack" if 200 <= status < 300 else "fail", + "--request-id", request_id, "--http-status", str(status)) + return status + + +def role_stats() -> dict[str, int]: + with ThreadPoolExecutor(max_workers=3) as pool: + rows_by_node = dict(zip(NODES, pool.map( + lambda node: remote(node, "docgrid-permission-node-probe", "role-stats", node), NODES))) + return {node: sum(int(row.split(",")[1]) for row in rows.splitlines() if row) + for node, rows in rows_by_node.items()} + + +def deltas(before: dict[str, int], after: dict[str, int]) -> dict[str, int]: + delta = {node: after[node] - before[node] for node in NODES} + if any(value < 0 for value in delta.values()): + raise InvalidExperiment("pg_stat_statements counters reset during the run") + return delta + + +def roles(user_id: int) -> dict[str, bool]: + with ThreadPoolExecutor(max_workers=3) as pool: + values = pool.map(lambda node: remote(node, "docgrid-permission-node-probe", + "role", node, str(user_id)), NODES) + return {node: value == "t" for node, value in zip(NODES, values)} + + +def lsn() -> dict[str, dict[str, str]]: + with ThreadPoolExecutor(max_workers=3) as pool: + values = pool.map(lambda node: remote(node, "docgrid-permission-node-probe", + "lsn", node), NODES) + return {node: dict(zip(("in_recovery", "receive_lsn", "replay_lsn"), + value.split(","), strict=True)) + for node, value in zip(NODES, values)} + + +def status(run_id: str) -> dict[str, str]: + with ThreadPoolExecutor(max_workers=2) as pool: + values = pool.map(lambda node: remote(node, "docgrid-standby-apply-delay-guard", + "status", node, run_id), STANDBYS) + return dict(zip(STANDBYS, values)) + + +def restore_delay(run_id: str) -> None: + def attempt(node: str) -> str | None: + try: + remote(node, "docgrid-standby-apply-delay-guard", "restore", node, run_id, + str(GUARD_SECONDS), str(DELAY_SECONDS)) + return None + except Exception: + return node + + with ThreadPoolExecutor(max_workers=2) as pool: + errors = [node for node in pool.map(attempt, STANDBYS) if node] + if errors: + raise InvalidExperiment("Delay reset failed on " + ", ".join(errors)) + + +def fixture(run_id: str, action: str) -> dict[str, int]: + output = remote(NODES[0], "docgrid-permission-fixture", action, NODES[0], run_id) + if action != "create": + return {} + ids = {} + for row in output.splitlines(): + alias, user_id = row.split(",") + ids[alias.removeprefix(f"ha-{run_id}-")] = int(user_id) + if set(ids) != {"admin", "m", "c1", "c2"}: + raise InvalidExperiment("Fixture returned an incomplete user set") + return ids + + +def require_status(actual: int, expected: int, label: str) -> None: + if actual != expected: + raise InvalidExperiment(f"{label}: expected HTTP {expected}, received {actual}") + + +def run(args: argparse.Namespace) -> None: + check_target() + redis("PING") + if not args.proxy_url.startswith("http://127.0.0.1:") or not args.primary_url.startswith("http://127.0.0.1:"): + raise InvalidExperiment("Both test app instances must be on loopback") + if args.proxy_url == args.primary_url: + raise InvalidExperiment("Proxy and direct-primary app URLs must differ") + contract = json.loads(CONTRACT.read_text()) + versions = contract["snapshot"]["nodes"][0]["versions"] + run_id = uuid.uuid4().hex[:12] + if not LABEL.fullmatch(run_id): + raise InvalidExperiment("Invalid generated run ID") + for node in NODES: + lines = remote(node, "docgrid-permission-node-probe", "health", node).splitlines() + expected = "f,t,2,0" if node == NODES[0] else "t,t,0,1" + if len(lines) != 2 or lines[0] != expected: + raise InvalidExperiment("Expected one primary, two streaming standbys, " + "and postgres-DB statistics extension on " + node) + initial = status(run_id) + if any("streaming=true" not in line or "delay=0,default" not in line for line in initial.values()): + raise InvalidExperiment("Standby not streaming or already delayed") + + # 1. This manifest refers to the redacted contract snapshot; no URL or secret is stored. + config_hash = contract["evidence_sha256"] + run_dir = args.output / f"permission-{run_id}" + ledger(run_dir, "init", "--scenario", "permission-replica-lag", + "--config-sha256", config_hash, + "--opensql-version", versions["opensql"], + "--openproxy-version", "recorded-in-contract", + "--patroni-version", versions["patroni"], + "--etcd-version", versions["etcd"]) + evidence: dict[str, object] = {"run_id": run_id, "started_at": utc_now(), + "contract_sha256": config_hash, + "initial_standbys": initial, "initial_lsn": lsn()} + ids: dict[str, int] = {} + fixture_attempted = False + armed = False + delayed = False + fault_open = False + try: + fixture_attempted = True + ids = fixture(run_id, "create") + jwt = {alias: token(user_id, run_id) for alias, user_id in ids.items()} + for alias in ("admin", "m", "c1", "c2"): + require_status(http(run_dir, args.proxy_url, "/admin/workers", jwt[alias], + f"baseline-{alias}"), 200, f"baseline {alias}") + for alias in ("m", "c1", "c2"): + if not all(roles(ids[alias]).values()): + raise InvalidExperiment("Fixture ADMIN role has not reached every standby") + + # 2. Each R request is a genuine cache miss; cumulative node statistics show arrival. + baseline = role_stats() + for index in range(6): + redis("DEL", f"auth:roles:{ids['m']}") + require_status(http(run_dir, args.proxy_url, "/admin/workers", jwt["m"], + f"routing-{index}"), 200, "routing baseline") + route = deltas(baseline, role_stats()) + evidence["R"] = {"role_sql_delta": route, "cache": cache(ids["m"])} + if route[NODES[1]] + route[NODES[2]] == 0: + raise InvalidExperiment("Role SQL did not reach standby; replay pause is prohibited") + if not args.apply_delay: + raise InvalidExperiment("R recorded; pass --apply-delay only after the independent " + "reset timer and one-standby delay have been verified") + + # 3. Arm both independent VM-host reset timers before delaying either standby. + armed = True + with ThreadPoolExecutor(max_workers=2) as pool: + list(pool.map(lambda node: remote(node, "docgrid-standby-apply-delay-guard", + "arm", node, run_id, str(GUARD_SECONDS), + str(DELAY_SECONDS)), STANDBYS)) + guard_deadline = time.monotonic() + GUARD_SECONDS - 30 + if any("armed=true" not in line for line in status(run_id).values()): + raise InvalidExperiment("An independent rescue timer is not pending") + with ThreadPoolExecutor(max_workers=2) as pool: + list(pool.map(lambda node: remote(node, "docgrid-standby-apply-delay-guard", + "apply", node, run_id, str(GUARD_SECONDS), + str(DELAY_SECONDS)), STANDBYS)) + delayed = True + before_revoke = status(run_id) + if any(f"delay={DELAY_SECONDS * 1000},configuration file" not in line + for line in before_revoke.values()): + raise InvalidExperiment("Both standbys did not retain the effective apply delay") + ledger(run_dir, "fault", "--name", "standby-apply-delay", "--phase", "start") + fault_open = True + + # 4. Commit role revocation through the actual administrator HTTP API. + for alias in ("m", "c2"): + require_status(http(run_dir, args.proxy_url, f"/admin/users/{ids[alias]}/roles/ADMIN", + jwt["admin"], f"revoke-{alias}", "DELETE"), 200, + f"revoke {alias}") + m_roles, c2_roles = roles(ids["m"]), roles(ids["c2"]) + if m_roles != {NODES[0]: False, NODES[1]: True, NODES[2]: True} or \ + c2_roles != {NODES[0]: False, NODES[1]: True, NODES[2]: True}: + raise InvalidExperiment("Revocation did not produce the required primary/standby split") + if cache(ids["m"])["present"] or cache(ids["c2"])["present"]: + raise InvalidExperiment("Revoked users' Redis keys were not invalidated") + evidence["after_revocation_lsn"] = lsn() + if any(f"delay={DELAY_SECONDS * 1000},configuration file" not in line + for line in status(run_id).values()): + raise InvalidExperiment("A standby lost its delay before stale-read observation") + + if time.monotonic() >= guard_deadline: + raise InvalidExperiment("Independent rescue timer is near its deadline") + m_before = role_stats() + m_http = http(run_dir, args.proxy_url, "/admin/workers", jwt["m"], "stale-M") + m_after = role_stats() + m_delta = deltas(m_before, m_after) + m_cache = cache(ids["m"]) + c2_http = http(run_dir, args.primary_url, "/admin/workers", jwt["c2"], "primary-C2") + c2_delta = deltas(m_after, role_stats()) + c2_cache = cache(ids["c2"]) + evidence["M"] = {"http_status": m_http, "role_sql_delta": m_delta, + "cache": m_cache, "roles": m_roles} + evidence["C2"] = {"http_status": c2_http, "cache": c2_cache, + "roles": c2_roles, "role_sql_delta": c2_delta} + evidence["delayed_standbys"] = status(run_id) + evidence["delayed_lsn"] = lsn() + if any(f"delay={DELAY_SECONDS * 1000},configuration file" not in line + for line in evidence["delayed_standbys"].values()): + raise InvalidExperiment("Standby apply delay ended during observation") + + # 5. Restore the normal apply policy immediately after the stale-read observation. + restore_delay(run_id) + delayed = False + ledger(run_dir, "fault", "--name", "standby-apply-delay", "--phase", "end") + fault_open = False + for _ in range(20): + if all(not value for value in roles(ids["m"]).values()): + break + time.sleep(1) + if any(roles(ids["m"]).values()): + raise InvalidExperiment("Replication did not catch up after replay resumed") + redis("DEL", f"auth:roles:{ids['c1']}") + require_status(http(run_dir, args.proxy_url, f"/admin/users/{ids['c1']}/roles/ADMIN", + jwt["admin"], "revoke-C1", "DELETE"), 200, "revoke C1") + for _ in range(20): + if not any(roles(ids["c1"]).values()): + break + time.sleep(1) + if any(roles(ids["c1"]).values()): + raise InvalidExperiment("C1 did not replicate before its fresh-read check") + c1_http = http(run_dir, args.proxy_url, "/admin/workers", jwt["c1"], "fresh-C1") + evidence["C1"] = {"http_status": c1_http, "cache": cache(ids["c1"]), + "roles": roles(ids["c1"])} + if c2_delta[NODES[0]] == 0 or c2_delta[NODES[1]] + c2_delta[NODES[2]] != 0: + raise InvalidExperiment("C2 did not prove direct-primary role lookup") + evidence["result"] = ("STALE_ADMIN_ALLOWED" if m_http == 200 and m_cache["admin"] + and m_delta[NODES[1]] + m_delta[NODES[2]] > 0 and c1_http == 403 + and c2_http == 403 else "DENIED" if m_http == 403 and + c1_http == 403 and c2_http == 403 else "INVALID") + except Exception as error: + evidence["result"] = "INVALID" + evidence["stop_reason"] = str(error) if isinstance(error, InvalidExperiment) else type(error).__name__ + raise + finally: + # 6. The independent timers remain armed until every standby is visibly normal. + if armed: + try: + restore_delay(run_id) + delayed = False + if fault_open: + ledger(run_dir, "fault", "--name", "standby-apply-delay", "--phase", "end") + fault_open = False + for node in STANDBYS: + remote(node, "docgrid-standby-apply-delay-guard", "cancel", node, run_id) + except Exception: + evidence["cleanup_warning"] = "Manual standby inspection required" + if fixture_attempted: + try: + for user_id in ids.values(): + redis("DEL", f"auth:roles:{user_id}") + fixture(run_id, "remove") + except Exception: + evidence["fixture_warning"] = "Manual fixture cleanup required" + evidence["finished_at"] = utc_now() + try: + evidence["final_standbys"] = status(run_id) + evidence["final_lsn"] = lsn() + except Exception: + evidence["final_standbys"] = "unverified; manual inspection required" + evidence["result"] = "INVALID" + if delayed: + evidence["result"] = "INVALID" + evidence["cleanup_warning"] = "Apply delay may still be active; inspect both standbys" + path = run_dir / "permission-scenarios.json" + path.write_text(json.dumps(evidence, ensure_ascii=False, sort_keys=True, indent=2) + "\n") + path.chmod(0o600) + ledger(run_dir, "export" if fault_open else "finish") + print(f"run_dir={run_dir} result={evidence['result']}") + + +def main() -> int: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--proxy-url", required=True) + parser.add_argument("--primary-url", required=True) + parser.add_argument("--output", type=Path, required=True) + parser.add_argument("--apply-delay", action="store_true", + help="Apply a temporary standby delay after independent reset verification") + args = parser.parse_args() + try: + run(args) + except (InvalidExperiment, OSError, subprocess.TimeoutExpired, ValueError) as error: + print(f"Experiment stopped: {error}", file=sys.stderr) + return 1 + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/opensql/run_permission_replica_lag_local.py b/scripts/opensql/run_permission_replica_lag_local.py new file mode 100644 index 00000000..d7444ed8 --- /dev/null +++ b/scripts/opensql/run_permission_replica_lag_local.py @@ -0,0 +1,131 @@ +#!/usr/bin/env python3 +"""Start two isolated local DocGrid apps and run the permission lag experiment. + +The private env file is read as Java properties, never sourced by a shell. +Application logs remain owner-only in the private output directory. +""" + +from __future__ import annotations + +import argparse +import os +import signal +import socket +import subprocess +import sys +import time +from pathlib import Path + + +ROOT = Path(__file__).resolve().parents[2] +RUNNER = ROOT / "scripts/opensql/permission_replica_lag.py" + + +def load_env(path: Path) -> dict[str, str]: + values = {} + for line in path.read_text().splitlines(): + line = line.strip() + if not line or line.startswith("#"): + continue + key, separator, value = line.partition("=") + if not separator or not key.isidentifier(): + raise ValueError("Invalid key in private env file") + values[key] = value.strip().strip('"').strip("'") + return values + + +def port_listening(port: int) -> bool: + try: + with socket.create_connection(("127.0.0.1", port), timeout=0.5): + return True + except OSError: + return False + + +def start_app(jar: Path, env: dict[str, str], port: int, log_path: Path) -> subprocess.Popen: + # 1. Each JVM has its own Hikari pool; only the direct instance overrides the URL. + log = log_path.open("x") + log_path.chmod(0o600) + try: + process = subprocess.Popen([ + "java", "-jar", str(jar), f"--server.port={port}", + "--server.address=127.0.0.1", "--management.server.address=127.0.0.1", + "--spring.profiles.active=opensql-ha", "--spring.flyway.enabled=false", + "--management.server.port=0", "--indexing.worker.enabled=false", + "--sync.dispatcher.enabled=false", "--sync.reconciliation.enabled=false", + ], cwd=ROOT, env=env, stdout=log, stderr=subprocess.STDOUT, + start_new_session=True) + finally: + log.close() + return process + + +def await_port(process: subprocess.Popen, port: int) -> None: + for _ in range(90): + if process.poll() is not None: + raise RuntimeError("A test app exited during startup; inspect its private log") + if port_listening(port): + return + time.sleep(1) + raise RuntimeError("A test app did not open its port; inspect its private log") + + +def stop_app(process: subprocess.Popen) -> None: + # 2. Stop only the two child JVMs started by this wrapper. + if process.poll() is not None: + return + os.killpg(process.pid, signal.SIGTERM) + try: + process.wait(timeout=15) + except subprocess.TimeoutExpired: + os.killpg(process.pid, signal.SIGKILL) + process.wait(timeout=5) + + +def main() -> int: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--env-file", type=Path, required=True) + parser.add_argument("--jar", type=Path, required=True) + parser.add_argument("--output", type=Path, required=True) + parser.add_argument("--apply-delay", action="store_true", + help="Run the guarded standby apply-delay phase after local apps start") + args = parser.parse_args() + if args.output.exists() or any(port_listening(port) for port in (18080, 18081)): + print("Output directory or test app port already exists", file=sys.stderr) + return 2 + values = load_env(args.env_file) + required = ("OPENSQL_APP_JDBC_URL", "OPENSQL_APP_DIRECT_JDBC_URL", + "OPENSQL_APP_USER", "OPENSQL_APP_PASSWORD", "JWT_SECRET") + if any(not values.get(key) for key in required): + print("Private env file is missing a required setting", file=sys.stderr) + return 2 + args.output.mkdir(parents=True, mode=0o700) + base_env = os.environ | values | {"REDIS_HOST": "127.0.0.1", "REDIS_PORT": "6379"} + children = [] + try: + proxy = start_app(args.jar, base_env, 18080, args.output / "proxy-app.log") + children.append(proxy) + await_port(proxy, 18080) + direct_env = base_env | {"OPENSQL_APP_JDBC_URL": values["OPENSQL_APP_DIRECT_JDBC_URL"]} + direct = start_app(args.jar, direct_env, 18081, args.output / "primary-app.log") + children.append(direct) + await_port(direct, 18081) + # 3. The runner owns the remote guard, HTTP requests, evidence, and cleanup. + runner_args = [ + sys.executable, str(RUNNER), "--proxy-url", "http://127.0.0.1:18080", + "--primary-url", "http://127.0.0.1:18081", "--output", str(args.output), + ] + if args.apply_delay: + runner_args.append("--apply-delay") + result = subprocess.run(runner_args, cwd=ROOT, env=base_env, check=False) + return result.returncode + except (OSError, RuntimeError) as error: + print(f"Local app setup stopped: {error}", file=sys.stderr) + return 1 + finally: + for child in reversed(children): + stop_app(child) + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/opensql/standby_apply_delay_guard.sh b/scripts/opensql/standby_apply_delay_guard.sh new file mode 100755 index 00000000..c220f96d --- /dev/null +++ b/scripts/opensql/standby_apply_delay_guard.sh @@ -0,0 +1,188 @@ +#!/usr/bin/env bash +# Run as root on a standby VM host; Patroni owns this recovery parameter. +set -euo pipefail + +action="${1:?action required}" +container="${2:?container required}" +run_id="${3:?run ID required}" +duration="${4:-180}" +delay_seconds="${5:-90}" +if [[ ! "$container" =~ ^docgrid-node[23]$ ]] || (( EUID != 0 )) || + [[ ! "$run_id" =~ ^[a-z0-9][a-z0-9-]{0,63}$ ]] || + [[ ! "$duration" =~ ^[0-9]+$ ]] || (( duration < 30 || duration > 300 )) || + [[ ! "$delay_seconds" =~ ^[0-9]+$ ]] || (( delay_seconds < 1 || delay_seconds > 180 )); then + echo 'Invalid standby guard target or duration.' >&2 + exit 2 +fi + +unit="docgrid-delay-reset-${run_id}-${container}" +timer="${unit}.timer" +mount_source="$(/usr/bin/docker inspect "$container" --format '{{range .Mounts}}{{if eq .Destination "/var/lib/docgrid"}}{{.Source}}{{end}}{{end}}')" +if [[ "$mount_source" != /* ]] || [[ "$(basename "$mount_source")" != "$container" ]]; then + echo 'The expected dedicated container volume is missing.' >&2 + exit 1 +fi +config="${mount_source}/opensql/etc/patroni/patroni.yml" +backup_dir='/run/docgrid-permission-delay' +backup="${backup_dir}/${container}-${run_id}.yml" +psql_command='. /var/lib/docgrid/opensql/etc/credentials.env; export PGPASSWORD="${PG_SUPERUSER_PASSWORD}"; exec /var/lib/docgrid/opensql/bin/psql -h 127.0.0.1 -U postgres -d postgres -X -v ON_ERROR_STOP=1 -At -F , -c "$1"' + +query() { + # 1. Keep the root-only DB password inside the existing container. + /usr/bin/docker exec "$container" sh -c "$psql_command" sh "$1" +} +setting() { query "SELECT setting, source FROM pg_settings WHERE name = 'recovery_min_apply_delay'"; } +is_standby() { [[ "$(query 'SELECT pg_is_in_recovery()')" == 't' ]]; } +is_streaming() { + [[ "$(query "SELECT COALESCE((SELECT status FROM pg_stat_wal_receiver LIMIT 1), 'none')")" == 'streaming' ]] +} +timer_pending() { + systemctl is-active --quiet "$timer" && + [[ "$(systemctl show "$timer" -p NextElapseUSecMonotonic --value)" != 'infinity' ]] +} +reload_patroni() { + /usr/bin/docker exec "$container" curl --silent --show-error --fail --max-time 5 \ + --output /dev/null --request POST http://127.0.0.1:8008/reload +} + +edit_config() { + # 2. Insert/remove only a local recovery_conf block; preserve exact original bytes. + python3 - "$1" "$config" "$backup" "$delay_seconds" <<'PY' +import os +import shutil +import sys +import tempfile + +action, config, backup, seconds = sys.argv[1:] +original = open(backup, 'rb').read() +current = open(config, 'rb').read() +lines = original.splitlines(keepends=True) +start = [i for i, line in enumerate(lines) if line == b'postgresql:\n'] +if len(start) != 1: + raise SystemExit('Expected exactly one top-level postgresql block') +end = next((i for i in range(start[0] + 1, len(lines)) + if lines[i] and lines[i][:1] not in (b' ', b'\t', b'\n', b'#')), + len(lines)) +if any(line.lstrip().startswith(b'recovery_conf:') for line in lines[start[0] + 1:end]): + raise SystemExit('Pre-existing recovery_conf must not be overwritten') +if b'recovery_min_apply_delay' in original: + raise SystemExit('Pre-existing apply delay must not be overwritten') +changed = b''.join(lines[:end]) + (f' recovery_conf:\n recovery_min_apply_delay: {seconds}s\n').encode() + b''.join(lines[end:]) +expected, replacement = (original, changed) if action == 'apply' else (changed, original) +if current == replacement: + raise SystemExit(0) +if current != expected: + raise SystemExit('Patroni config changed independently; refusing to overwrite it') +metadata = os.stat(config) +fd, temp = tempfile.mkstemp(prefix='.docgrid-delay-', dir=os.path.dirname(config)) +try: + with os.fdopen(fd, 'wb') as stream: + stream.write(replacement) + stream.flush() + os.fsync(stream.fileno()) + os.chown(temp, metadata.st_uid, metadata.st_gid) + shutil.copystat(config, temp) + os.replace(temp, config) +finally: + if os.path.exists(temp): + os.unlink(temp) +PY +} + +restore_once() { + if [[ ! -f "$backup" ]]; then + [[ "$(setting)" == '0,default' ]] + return + fi + edit_config restore || return 1 + reload_patroni || return 1 + for (( check = 0; check < 10; check++ )); do + [[ "$(setting)" == '0,default' ]] && return 0 + sleep 1 + done + return 1 +} +restore() { + # 3. The VM-host timer retries even if the client process or SSH disappears. + for (( attempt = 0; attempt < 12; attempt++ )); do + if restore_once; then + echo 'delay-reset' + return 0 + fi + sleep 5 + done + echo 'Patroni delay reset failed; manual cluster inspection required.' >&2 + return 1 +} + +case "$action" in + arm) + command -v systemd-run >/dev/null + if ! is_standby || ! is_streaming || [[ "$(setting)" != '0,default' ]] || + timer_pending || [[ ! -f "$config" ]] || [[ -L "$config" ]]; then + echo 'Standby baseline, Patroni config, or timer precondition failed.' >&2 + exit 1 + fi + install -d -m 0700 "$backup_dir" + [[ ! -e "$backup" ]] || { echo 'Guard backup already exists.' >&2; exit 1; } + cp --preserve=all "$config" "$backup" + chmod 0600 "$backup" + systemd-run --quiet --collect --unit="$unit" --on-active="${duration}s" \ + /usr/local/sbin/docgrid-standby-apply-delay-guard restore "$container" "$run_id" \ + "$duration" "$delay_seconds" >/dev/null + timer_pending || { echo 'Independent reset timer did not become pending.' >&2; exit 1; } + echo 'armed' + ;; + apply) + timer_pending && [[ -f "$backup" ]] || + { echo 'Independent timer or backup missing.' >&2; exit 1; } + if ! is_standby || ! is_streaming || [[ "$(setting)" != '0,default' ]]; then + echo 'Standby baseline changed before apply.' >&2 + exit 1 + fi + free_kb="$(/usr/bin/docker exec "$container" df -Pk /var/lib/docgrid/opensql/data/pgsql | awk 'NR == 2 {print $4}')" + if [[ ! "$free_kb" =~ ^[0-9]+$ ]] || (( free_kb < 1048576 )); then + echo 'Less than 1 GiB free on standby data filesystem.' >&2 + exit 1 + fi + if ! edit_config apply || ! reload_patroni; then + restore || true + echo 'Could not apply Patroni-managed standby delay.' >&2 + exit 1 + fi + for (( check = 0; check < 15; check++ )); do + if [[ "$(setting)" == "$(( delay_seconds * 1000 )),configuration file" ]]; then + timer_pending && echo 'delay-applied' && exit 0 + break + fi + sleep 1 + done + restore || true + echo 'Effective delay or independent reset timer was not verified.' >&2 + exit 1 + ;; + restore) + restore + ;; + cancel) + if ! is_standby || ! is_streaming || [[ "$(setting)" != '0,default' ]]; then + echo 'Do not cancel until default delay and streaming return.' >&2 + exit 1 + fi + if systemctl is-active --quiet "$timer"; then systemctl stop "$timer"; fi + if [[ -f "$backup" ]] && cmp --silent "$config" "$backup"; then rm -- "$backup"; fi + echo 'cancelled' + ;; + status) + if timer_pending; then armed='true'; else armed='false'; fi + printf 'standby=%s streaming=%s armed=%s delay=%s backlog_bytes=%s free_kb=%s\n' \ + "$(is_standby && echo true || echo false)" \ + "$(is_streaming && echo true || echo false)" "$armed" "$(setting)" \ + "$(query 'SELECT COALESCE(pg_wal_lsn_diff(pg_last_wal_receive_lsn(), pg_last_wal_replay_lsn()), 0)::bigint')" \ + "$(/usr/bin/docker exec "$container" df -Pk /var/lib/docgrid/opensql/data/pgsql | awk 'NR == 2 {print $4}')" + ;; + *) + echo 'Unknown guard action.' >&2 + exit 2 + ;; +esac From 5fba43e10901eb72f1b13b57f7cc7fb704eced19 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?=EA=B9=80=EA=B8=B0=EB=AF=BC?= Date: Thu, 1 Oct 2026 01:26:39 +0900 Subject: [PATCH 2/3] [Test] Harden OpenSQL permission replica-lag evidence and cleanup --- scripts/opensql/ha_evidence.py | 6 +- scripts/opensql/permission_fixture.sh | 56 +++++- scripts/opensql/permission_replica_lag.py | 171 +++++++++++++++--- .../run_permission_replica_lag_local.py | 20 +- scripts/opensql/standby_apply_delay_guard.sh | 80 +++++++- scripts/opensql/test_ha_evidence.py | 11 ++ .../opensql/test_permission_replica_lag.py | 63 +++++++ .../opensql/test_standby_apply_delay_guard.py | 66 +++++++ 8 files changed, 428 insertions(+), 45 deletions(-) create mode 100644 scripts/opensql/test_permission_replica_lag.py create mode 100644 scripts/opensql/test_standby_apply_delay_guard.py diff --git a/scripts/opensql/ha_evidence.py b/scripts/opensql/ha_evidence.py index ef81cda9..22b994f2 100644 --- a/scripts/opensql/ha_evidence.py +++ b/scripts/opensql/ha_evidence.py @@ -219,6 +219,9 @@ def initialize(args): raise EvidenceError("redacted config의 SHA-256 64자리가 필요합니다") if not LABEL.fullmatch(args.scenario): raise EvidenceError("scenario는 100자 이내의 영문·숫자·점·하이픈·밑줄만 허용합니다") + run_id = getattr(args, "run_id", None) or str(uuid.uuid4()) + if not LABEL.fullmatch(run_id): + raise EvidenceError("run_id는 100자 이내의 영문·숫자·점·하이픈·밑줄만 허용합니다") versions = (args.opensql_version, args.openproxy_version, args.patroni_version, args.etcd_version) if any(not value.strip() or len(value) > 128 or "\n" in value for value in versions): @@ -228,7 +231,7 @@ def initialize(args): directory = args.run_dir directory.mkdir(mode=0o700, parents=True, exist_ok=False) manifest = { - "schema_version": 1, "run_id": str(uuid.uuid4()), + "schema_version": 1, "run_id": run_id, "scenario": args.scenario, "started_at": utc_now(), "git_sha": git_sha, "config_sha256": args.config_sha256.lower(), "versions": {"opensql": args.opensql_version, "openproxy": args.openproxy_version, @@ -267,6 +270,7 @@ def parser(): sub = commands.add_subparsers(dest="command", required=True) init = sub.add_parser("init", help="새 실험과 환경 지문 생성") init.add_argument("--run-dir", type=Path, required=True) + init.add_argument("--run-id", help="외부 시험과 동일한 실행 ID를 사용할 때 지정") init.add_argument("--scenario", required=True) init.add_argument("--config-sha256", required=True, help="비밀 제거된 설정 스냅샷의 SHA-256") for name in ("opensql", "openproxy", "patroni", "etcd"): diff --git a/scripts/opensql/permission_fixture.sh b/scripts/opensql/permission_fixture.sh index 272c16c8..4a641fbc 100755 --- a/scripts/opensql/permission_fixture.sh +++ b/scripts/opensql/permission_fixture.sh @@ -1,10 +1,11 @@ #!/usr/bin/env bash -# Create or remove only this run's four permission-test users on the current primary. +# Guard and remove only this run's four permission-test users on the known primary. set -euo pipefail action="${1:?action required: create|remove|status}" container="${2:?primary container required}" run_id="${3:?run ID required}" +guard_seconds="${4:-900}" if [[ "$container" != 'docgrid-node1' ]] || (( EUID != 0 )); then echo 'Run as root on the known primary container only.' >&2 @@ -14,6 +15,18 @@ if [[ ! "$run_id" =~ ^[a-z0-9][a-z0-9-]{0,23}$ ]]; then echo 'Invalid run ID.' >&2 exit 2 fi +if [[ ! "$guard_seconds" =~ ^[0-9]+$ ]] || + (( guard_seconds < 300 || guard_seconds > 1800 )); then + echo 'Fixture guard duration must be 300-1800 seconds.' >&2 + exit 2 +fi + +unit="docgrid-fixture-reset-${run_id}" +timer="${unit}.timer" +timer_pending() { + systemctl is-active --quiet "$timer" && + [[ "$(systemctl show "$timer" -p NextElapseUSecMonotonic --value)" != 'infinity' ]] +} query() { /usr/bin/docker exec "$container" sh -c ' @@ -31,9 +44,28 @@ fi emails="('ha-${run_id}-admin@invalid.example', 'ha-${run_id}-m@invalid.example', 'ha-${run_id}-c1@invalid.example', 'ha-${run_id}-c2@invalid.example')" case "$action" in + stale-count) + # Refuse another run while an earlier short-lived ADMIN fixture remains. + query "SELECT count(*) FROM users + WHERE email ~ '^ha-[a-z0-9]{12}-(admin|m|c1|c2)@invalid[.]example$'" + ;; + arm) + # 1. The VM-host timer outlives the client JVM and retries while this host is up. + command -v systemd-run >/dev/null + if timer_pending || [[ "$(query "SELECT count(*) FROM users WHERE email IN $emails")" != '0' ]]; then + echo 'Fixture timer or run-scoped users already exist.' >&2 + exit 1 + fi + systemd-run --quiet --collect --unit="$unit" --on-active="${guard_seconds}s" \ + --on-unit-active=60s \ + /usr/local/sbin/docgrid-permission-fixture remove "$container" "$run_id" \ + "$guard_seconds" >/dev/null + timer_pending || { echo 'Independent fixture timer did not become pending.' >&2; exit 1; } + echo 'fixture-armed' + ;; create) - # 1. A unique run ID confines writes to four disposable accounts and the existing ADMIN role. - if [[ "$(query "SELECT count(*) FROM users WHERE email IN $emails")" != '0' ]]; then + # 2. Never create an ADMIN fixture without its independent removal timer. + if ! timer_pending || [[ "$(query "SELECT count(*) FROM users WHERE email IN $emails")" != '0' ]]; then echo 'Fixture already exists; choose a new run ID.' >&2 exit 1 fi @@ -52,20 +84,34 @@ case "$action" in query "SELECT split_part(email, '@', 1), id FROM users WHERE email IN $emails ORDER BY email" ;; status) - # 2. Only the run's IDs and role presence leave the database container. + # 3. Only the run's IDs and role presence leave the database container. query "SELECT split_part(u.email, '@', 1), u.id, EXISTS (SELECT 1 FROM user_roles ur JOIN roles r ON r.id = ur.role_id WHERE ur.user_id = u.id AND r.code = 'ADMIN') FROM users u WHERE u.email IN $emails ORDER BY u.email" ;; remove) - # 3. Remove only mappings and users identified by the exact run-scoped email set. + # 4. An expired timer may retry this exact, idempotent removal after DB recovery. query "BEGIN; DELETE FROM user_roles WHERE user_id IN (SELECT id FROM users WHERE email IN $emails); DELETE FROM users WHERE email IN $emails; COMMIT;" >/dev/null echo 'fixture-removed' ;; + cancel) + # 5. Keep the timer armed unless removal is visible on the current primary. + if [[ "$(query "SELECT count(*) FROM users WHERE email IN $emails")" != '0' ]]; then + echo 'Run-scoped users remain; refusing to cancel cleanup.' >&2 + exit 1 + fi + if systemctl is-active --quiet "$timer"; then systemctl stop "$timer"; fi + echo 'fixture-cancelled' + ;; + guard-status) + if timer_pending; then armed=true; else armed=false; fi + printf 'armed=%s fixture_count=%s\n' "$armed" \ + "$(query "SELECT count(*) FROM users WHERE email IN $emails")" + ;; *) echo 'Unknown fixture action.' >&2 exit 2 diff --git a/scripts/opensql/permission_replica_lag.py b/scripts/opensql/permission_replica_lag.py index 56ec1bdf..3cea8dc8 100644 --- a/scripts/opensql/permission_replica_lag.py +++ b/scripts/opensql/permission_replica_lag.py @@ -34,6 +34,7 @@ LABEL = re.compile(r"[a-z0-9][a-z0-9-]{0,23}\Z") DELAY_SECONDS = 120 GUARD_SECONDS = 240 +FIXTURE_GUARD_SECONDS = 900 class InvalidExperiment(RuntimeError): @@ -44,6 +45,19 @@ def utc_now() -> str: return datetime.now(timezone.utc).isoformat(timespec="milliseconds").replace("+00:00", "Z") +def write_private(path: Path, contents: bytes) -> None: + # The private evidence must never be briefly created with a permissive umask. + descriptor = os.open(path, os.O_WRONLY | os.O_CREAT | os.O_EXCL, 0o600) + with os.fdopen(descriptor, "wb") as output: + output.write(contents) + output.flush() + os.fsync(output.fileno()) + + +def file_sha256(path: Path) -> str: + return hashlib.sha256(path.read_bytes()).hexdigest() + + def required(name: str) -> str: value = os.environ.get(name, "") if not value: @@ -66,11 +80,20 @@ def remote(node: str, script: str, *args: str) -> str: for value in args: if not re.fullmatch(r"[a-z0-9-]+", value): raise InvalidExperiment("Unsafe remote argument") + remote_command = (f"sudo sha256sum /usr/local/sbin/{script}" if args == ("hash",) else + f"sudo /usr/local/sbin/{script} {' '.join(args)}") return command(["gcloud", "compute", "ssh", node, f"--zone={required('OPENSQL_GCP_ZONE')}", f"--project={required('OPENSQL_EXPECTED_PROJECT')}", f"--ssh-key-file={required('OPENSQL_SSH_KEY')}", "--quiet", - f"--command=sudo /usr/local/sbin/{script} {' '.join(args)}"], timeout=35) + f"--command={remote_command}"], timeout=35) + + +def deployed_sha256(node: str, script: str) -> str: + value = remote(node, script, "hash").split()[0] + if not re.fullmatch(r"[0-9a-f]{64}", value): + raise InvalidExperiment("Remote helper fingerprint is malformed") + return value def check_target() -> None: @@ -221,7 +244,8 @@ def attempt(node: str) -> str | None: def fixture(run_id: str, action: str) -> dict[str, int]: - output = remote(NODES[0], "docgrid-permission-fixture", action, NODES[0], run_id) + output = remote(NODES[0], "docgrid-permission-fixture", action, NODES[0], run_id, + str(FIXTURE_GUARD_SECONDS)) if action != "create": return {} ids = {} @@ -233,6 +257,40 @@ def fixture(run_id: str, action: str) -> dict[str, int]: return ids +def fixture_guard_status(run_id: str) -> str: + return remote(NODES[0], "docgrid-permission-fixture", "guard-status", NODES[0], run_id, + str(FIXTURE_GUARD_SECONDS)) + + +def delayed_status_is_safe(values: dict[str, str]) -> bool: + required_fields = ("standby=true", "streaming=true", "armed=true", + f"delay={DELAY_SECONDS * 1000},configuration file", + "nofailover=true", "nofailover_cleared=false") + return set(values) == set(STANDBYS) and all(all(field in line for field in required_fields) + for line in values.values()) + + +def restored_status_is_safe(values: dict[str, str]) -> bool: + required_fields = ("standby=true", "streaming=true", "armed=false", + "delay=0,default", "nofailover=false", "nofailover_cleared=true") + return set(values) == set(STANDBYS) and all(all(field in line for field in required_fields) + for line in values.values()) + + +def classify_result(m_http: int, m_admin_cached: bool, m_delta: dict[str, int], + c1_http: int, c2_http: int) -> str: + # A 403 without an observed standby role lookup is not evidence of safe routing. + standby_lookup = (m_delta[NODES[0]] == 0 and + m_delta[NODES[1]] + m_delta[NODES[2]] > 0) + if not standby_lookup or c1_http != 403 or c2_http != 403: + return "INVALID" + if m_http == 200 and m_admin_cached: + return "STALE_ADMIN_ALLOWED" + if m_http == 403 and not m_admin_cached: + return "DENIED" + return "INVALID" + + def require_status(actual: int, expected: int, label: str) -> None: if actual != expected: raise InvalidExperiment(f"{label}: expected HTTP {expected}, received {actual}") @@ -241,6 +299,9 @@ def require_status(actual: int, expected: int, label: str) -> None: def run(args: argparse.Namespace) -> None: check_target() redis("PING") + if remote(NODES[0], "docgrid-permission-fixture", "stale-count", NODES[0], + "preflight", str(FIXTURE_GUARD_SECONDS)) != "0": + raise InvalidExperiment("Earlier run-scoped ADMIN fixtures require manual inspection") if not args.proxy_url.startswith("http://127.0.0.1:") or not args.primary_url.startswith("http://127.0.0.1:"): raise InvalidExperiment("Both test app instances must be on loopback") if args.proxy_url == args.primary_url: @@ -250,35 +311,71 @@ def run(args: argparse.Namespace) -> None: run_id = uuid.uuid4().hex[:12] if not LABEL.fullmatch(run_id): raise InvalidExperiment("Invalid generated run ID") + health = {} for node in NODES: lines = remote(node, "docgrid-permission-node-probe", "health", node).splitlines() expected = "f,t,2,0" if node == NODES[0] else "t,t,0,1" if len(lines) != 2 or lines[0] != expected: raise InvalidExperiment("Expected one primary, two streaming standbys, " "and postgres-DB statistics extension on " + node) + health[node] = lines initial = status(run_id) - if any("streaming=true" not in line or "delay=0,default" not in line for line in initial.values()): - raise InvalidExperiment("Standby not streaming or already delayed") - - # 1. This manifest refers to the redacted contract snapshot; no URL or secret is stored. - config_hash = contract["evidence_sha256"] + if not restored_status_is_safe(initial): + raise InvalidExperiment("Standby not streaming, eligible, or at its original policy") + + # Exact source and installed-helper hashes make a dirty worktree traceable. + source_files = ("permission_replica_lag.py", "run_permission_replica_lag_local.py", + "ha_evidence.py", "permission_fixture.sh", + "standby_apply_delay_guard.sh", "permission_node_probe.sh") + source_hashes = {name: file_sha256(ROOT / "scripts/opensql" / name) + for name in source_files} + installed_hashes = { + NODES[0]: {"permission_fixture.sh": deployed_sha256(NODES[0], "docgrid-permission-fixture")}, + NODES[1]: {"standby_apply_delay_guard.sh": deployed_sha256( + NODES[1], "docgrid-standby-apply-delay-guard")}, + NODES[2]: {"standby_apply_delay_guard.sh": deployed_sha256( + NODES[2], "docgrid-standby-apply-delay-guard")}, + } + if any(digest != source_hashes[name] for installed in installed_hashes.values() + for name, digest in installed.items()): + raise InvalidExperiment("Installed VM helper differs from the checked-out source") + + # 1. Hash the exact, secret-free live preflight used for this run. + live_preflight = {"health": health, "standbys": initial, + "prior_contract_sha256": contract["evidence_sha256"], + "app_jar_sha256": args.app_jar_sha256, + "source_sha256": source_hashes, "installed_helper_sha256": installed_hashes, + "proxy_jdbc_url_sha256": hashlib.sha256( + required("OPENSQL_APP_JDBC_URL").encode()).hexdigest(), + "primary_jdbc_url_sha256": hashlib.sha256( + required("OPENSQL_APP_DIRECT_JDBC_URL").encode()).hexdigest()} + snapshot = json.dumps(live_preflight, ensure_ascii=False, sort_keys=True, + separators=(",", ":")).encode() + b"\n" + config_hash = hashlib.sha256(snapshot).hexdigest() run_dir = args.output / f"permission-{run_id}" ledger(run_dir, "init", "--scenario", "permission-replica-lag", + "--run-id", run_id, "--config-sha256", config_hash, "--opensql-version", versions["opensql"], "--openproxy-version", "recorded-in-contract", "--patroni-version", versions["patroni"], "--etcd-version", versions["etcd"]) + write_private(run_dir / "live-preflight.json", snapshot) evidence: dict[str, object] = {"run_id": run_id, "started_at": utc_now(), - "contract_sha256": config_hash, + "contract_sha256": contract["evidence_sha256"], + "live_preflight_sha256": config_hash, + "app_jar_sha256": args.app_jar_sha256, "initial_standbys": initial, "initial_lsn": lsn()} ids: dict[str, int] = {} - fixture_attempted = False + fixture_guard_attempted = False armed = False delayed = False fault_open = False try: - fixture_attempted = True + fixture(run_id, "arm") + fixture_guard_attempted = True + if fixture_guard_status(run_id) != "armed=true fixture_count=0": + raise InvalidExperiment("Independent fixture cleanup timer is not pending") ids = fixture(run_id, "create") jwt = {alias: token(user_id, run_id) for alias, user_id in ids.items()} for alias in ("admin", "m", "c1", "c2"): @@ -317,9 +414,8 @@ def run(args: argparse.Namespace) -> None: str(DELAY_SECONDS)), STANDBYS)) delayed = True before_revoke = status(run_id) - if any(f"delay={DELAY_SECONDS * 1000},configuration file" not in line - for line in before_revoke.values()): - raise InvalidExperiment("Both standbys did not retain the effective apply delay") + if not delayed_status_is_safe(before_revoke): + raise InvalidExperiment("Both standbys must be delayed and excluded from promotion") ledger(run_dir, "fault", "--name", "standby-apply-delay", "--phase", "start") fault_open = True @@ -335,9 +431,8 @@ def run(args: argparse.Namespace) -> None: if cache(ids["m"])["present"] or cache(ids["c2"])["present"]: raise InvalidExperiment("Revoked users' Redis keys were not invalidated") evidence["after_revocation_lsn"] = lsn() - if any(f"delay={DELAY_SECONDS * 1000},configuration file" not in line - for line in status(run_id).values()): - raise InvalidExperiment("A standby lost its delay before stale-read observation") + if not delayed_status_is_safe(status(run_id)): + raise InvalidExperiment("A standby lost its delay or promotion exclusion") if time.monotonic() >= guard_deadline: raise InvalidExperiment("Independent rescue timer is near its deadline") @@ -355,9 +450,8 @@ def run(args: argparse.Namespace) -> None: "roles": c2_roles, "role_sql_delta": c2_delta} evidence["delayed_standbys"] = status(run_id) evidence["delayed_lsn"] = lsn() - if any(f"delay={DELAY_SECONDS * 1000},configuration file" not in line - for line in evidence["delayed_standbys"].values()): - raise InvalidExperiment("Standby apply delay ended during observation") + if not delayed_status_is_safe(evidence["delayed_standbys"]): + raise InvalidExperiment("Standby delay or promotion exclusion ended during observation") # 5. Restore the normal apply policy immediately after the stale-read observation. restore_delay(run_id) @@ -384,10 +478,9 @@ def run(args: argparse.Namespace) -> None: "roles": roles(ids["c1"])} if c2_delta[NODES[0]] == 0 or c2_delta[NODES[1]] + c2_delta[NODES[2]] != 0: raise InvalidExperiment("C2 did not prove direct-primary role lookup") - evidence["result"] = ("STALE_ADMIN_ALLOWED" if m_http == 200 and m_cache["admin"] - and m_delta[NODES[1]] + m_delta[NODES[2]] > 0 and c1_http == 403 - and c2_http == 403 else "DENIED" if m_http == 403 and - c1_http == 403 and c2_http == 403 else "INVALID") + evidence["result"] = classify_result(m_http, m_cache["admin"], m_delta, + c1_http, c2_http) + evidence["observation"] = evidence["result"] except Exception as error: evidence["result"] = "INVALID" evidence["stop_reason"] = str(error) if isinstance(error, InvalidExperiment) else type(error).__name__ @@ -405,28 +498,47 @@ def run(args: argparse.Namespace) -> None: remote(node, "docgrid-standby-apply-delay-guard", "cancel", node, run_id) except Exception: evidence["cleanup_warning"] = "Manual standby inspection required" - if fixture_attempted: + if fixture_guard_attempted: try: + cache_failed = False for user_id in ids.values(): - redis("DEL", f"auth:roles:{user_id}") + try: + redis("DEL", f"auth:roles:{user_id}") + except Exception: + cache_failed = True + # Remove DB ADMIN rows even when the local Redis cleanup fails. fixture(run_id, "remove") + if fixture_guard_status(run_id) != "armed=true fixture_count=0": + raise InvalidExperiment("Run-scoped fixture removal is not visible") + fixture(run_id, "cancel") + if fixture_guard_status(run_id) != "armed=false fixture_count=0": + raise InvalidExperiment("Fixture cleanup timer did not stop") + if cache_failed: + raise InvalidExperiment("Local Redis keys require manual inspection") except Exception: evidence["fixture_warning"] = "Manual fixture cleanup required" evidence["finished_at"] = utc_now() try: evidence["final_standbys"] = status(run_id) evidence["final_lsn"] = lsn() + if not restored_status_is_safe(evidence["final_standbys"]): + evidence["result"] = "INVALID" + evidence["cleanup_warning"] = "Standby policy or timer did not return to baseline" except Exception: evidence["final_standbys"] = "unverified; manual inspection required" evidence["result"] = "INVALID" if delayed: evidence["result"] = "INVALID" evidence["cleanup_warning"] = "Apply delay may still be active; inspect both standbys" + if "cleanup_warning" in evidence or "fixture_warning" in evidence: + evidence["result"] = "INVALID" path = run_dir / "permission-scenarios.json" - path.write_text(json.dumps(evidence, ensure_ascii=False, sort_keys=True, indent=2) + "\n") - path.chmod(0o600) - ledger(run_dir, "export" if fault_open else "finish") + write_private(path, (json.dumps(evidence, ensure_ascii=False, sort_keys=True, + indent=2) + "\n").encode()) + ledger(run_dir, "export" if fault_open or evidence["result"] == "INVALID" else "finish") print(f"run_dir={run_dir} result={evidence['result']}") + if evidence["result"] == "INVALID": + raise InvalidExperiment("Run result or cleanup invalid; inspect private evidence") def main() -> int: @@ -434,10 +546,13 @@ def main() -> int: parser.add_argument("--proxy-url", required=True) parser.add_argument("--primary-url", required=True) parser.add_argument("--output", type=Path, required=True) + parser.add_argument("--app-jar-sha256", required=True) parser.add_argument("--apply-delay", action="store_true", help="Apply a temporary standby delay after independent reset verification") args = parser.parse_args() try: + if not re.fullmatch(r"[0-9a-f]{64}", args.app_jar_sha256): + raise InvalidExperiment("A SHA-256 of the exact app JAR is required") run(args) except (InvalidExperiment, OSError, subprocess.TimeoutExpired, ValueError) as error: print(f"Experiment stopped: {error}", file=sys.stderr) diff --git a/scripts/opensql/run_permission_replica_lag_local.py b/scripts/opensql/run_permission_replica_lag_local.py index d7444ed8..3462bf14 100644 --- a/scripts/opensql/run_permission_replica_lag_local.py +++ b/scripts/opensql/run_permission_replica_lag_local.py @@ -8,6 +8,7 @@ from __future__ import annotations import argparse +import hashlib import os import signal import socket @@ -42,6 +43,14 @@ def port_listening(port: int) -> bool: return False +def file_sha256(path: Path) -> str: + digest = hashlib.sha256() + with path.open("rb") as source: + for block in iter(lambda: source.read(1024 * 1024), b""): + digest.update(block) + return digest.hexdigest() + + def start_app(jar: Path, env: dict[str, str], port: int, log_path: Path) -> subprocess.Popen: # 1. Each JVM has its own Hikari pool; only the direct instance overrides the URL. log = log_path.open("x") @@ -87,20 +96,26 @@ def main() -> int: parser.add_argument("--env-file", type=Path, required=True) parser.add_argument("--jar", type=Path, required=True) parser.add_argument("--output", type=Path, required=True) + parser.add_argument("--redis-port", type=int, default=16379, + help="Port of a disposable, isolated loopback Redis instance") parser.add_argument("--apply-delay", action="store_true", help="Run the guarded standby apply-delay phase after local apps start") args = parser.parse_args() if args.output.exists() or any(port_listening(port) for port in (18080, 18081)): print("Output directory or test app port already exists", file=sys.stderr) return 2 - values = load_env(args.env_file) + # Non-secret route overrides may come from the process; the private file wins on conflicts. + values = os.environ | load_env(args.env_file) required = ("OPENSQL_APP_JDBC_URL", "OPENSQL_APP_DIRECT_JDBC_URL", "OPENSQL_APP_USER", "OPENSQL_APP_PASSWORD", "JWT_SECRET") if any(not values.get(key) for key in required): print("Private env file is missing a required setting", file=sys.stderr) return 2 args.output.mkdir(parents=True, mode=0o700) - base_env = os.environ | values | {"REDIS_HOST": "127.0.0.1", "REDIS_PORT": "6379"} + if not 1 <= args.redis_port <= 65535: + print("Invalid isolated Redis port", file=sys.stderr) + return 2 + base_env = values | {"REDIS_HOST": "127.0.0.1", "REDIS_PORT": str(args.redis_port)} children = [] try: proxy = start_app(args.jar, base_env, 18080, args.output / "proxy-app.log") @@ -114,6 +129,7 @@ def main() -> int: runner_args = [ sys.executable, str(RUNNER), "--proxy-url", "http://127.0.0.1:18080", "--primary-url", "http://127.0.0.1:18081", "--output", str(args.output), + "--app-jar-sha256", file_sha256(args.jar), ] if args.apply_delay: runner_args.append("--apply-delay") diff --git a/scripts/opensql/standby_apply_delay_guard.sh b/scripts/opensql/standby_apply_delay_guard.sh index c220f96d..6ccc3965 100755 --- a/scripts/opensql/standby_apply_delay_guard.sh +++ b/scripts/opensql/standby_apply_delay_guard.sh @@ -1,5 +1,5 @@ #!/usr/bin/env bash -# Run as root on a standby VM host; Patroni owns this recovery parameter. +# Run as root on a standby VM host; Patroni owns recovery and failover tags. set -euo pipefail action="${1:?action required}" @@ -23,7 +23,8 @@ if [[ "$mount_source" != /* ]] || [[ "$(basename "$mount_source")" != "$containe exit 1 fi config="${mount_source}/opensql/etc/patroni/patroni.yml" -backup_dir='/run/docgrid-permission-delay' +# Persist the original bytes across a VM reboot; transient timers do not survive it. +backup_dir='/var/lib/docgrid-permission-delay' backup="${backup_dir}/${container}-${run_id}.yml" psql_command='. /var/lib/docgrid/opensql/etc/credentials.env; export PGPASSWORD="${PG_SUPERUSER_PASSWORD}"; exec /var/lib/docgrid/opensql/bin/psql -h 127.0.0.1 -U postgres -d postgres -X -v ON_ERROR_STOP=1 -At -F , -c "$1"' @@ -32,6 +33,47 @@ query() { /usr/bin/docker exec "$container" sh -c "$psql_command" sh "$1" } setting() { query "SELECT setting, source FROM pg_settings WHERE name = 'recovery_min_apply_delay'"; } +failover_excluded() { + # Both the local REST state and the cluster view must observe the safety tag. + local local_state cluster_state + local_state="$(/usr/bin/docker exec "$container" curl --silent --show-error --fail \ + --max-time 5 http://127.0.0.1:8008/patroni)" + cluster_state="$(/usr/bin/docker exec "$container" curl --silent --show-error --fail \ + --max-time 5 http://127.0.0.1:8008/cluster)" + python3 - "$local_state" "$cluster_state" <<'PY' +import json +import sys + +local, cluster = (json.loads(value) for value in sys.argv[1:]) +name = local.get('patroni', {}).get('name') +member = next((item for item in cluster.get('members', []) if item.get('name') == name), {}) +safe = (local.get('role') == 'replica' + and local.get('tags', {}).get('nofailover') is True + and member.get('role') == 'replica' + and member.get('tags', {}).get('nofailover') is True) +print('true' if safe else 'false') +PY +} +nofailover_cleared() { + local local_state cluster_state + local_state="$(/usr/bin/docker exec "$container" curl --silent --show-error --fail \ + --max-time 5 http://127.0.0.1:8008/patroni)" + cluster_state="$(/usr/bin/docker exec "$container" curl --silent --show-error --fail \ + --max-time 5 http://127.0.0.1:8008/cluster)" + python3 - "$local_state" "$cluster_state" <<'PY' +import json +import sys + +local, cluster = (json.loads(value) for value in sys.argv[1:]) +name = local.get('patroni', {}).get('name') +member = next((item for item in cluster.get('members', []) if item.get('name') == name), {}) +safe = (local.get('role') == 'replica' + and local.get('tags', {}).get('nofailover') is not True + and member.get('role') == 'replica' + and member.get('tags', {}).get('nofailover') is not True) +print('true' if safe else 'false') +PY +} is_standby() { [[ "$(query 'SELECT pg_is_in_recovery()')" == 't' ]]; } is_streaming() { [[ "$(query "SELECT COALESCE((SELECT status FROM pg_stat_wal_receiver LIMIT 1), 'none')")" == 'streaming' ]] @@ -46,7 +88,7 @@ reload_patroni() { } edit_config() { - # 2. Insert/remove only a local recovery_conf block; preserve exact original bytes. + # 2. Change the local delay and nofailover tag together; preserve exact original bytes. python3 - "$1" "$config" "$backup" "$delay_seconds" <<'PY' import os import shutil @@ -68,6 +110,20 @@ if any(line.lstrip().startswith(b'recovery_conf:') for line in lines[start[0] + if b'recovery_min_apply_delay' in original: raise SystemExit('Pre-existing apply delay must not be overwritten') changed = b''.join(lines[:end]) + (f' recovery_conf:\n recovery_min_apply_delay: {seconds}s\n').encode() + b''.join(lines[end:]) +tag_lines = changed.splitlines(keepends=True) +tag_start = [i for i, line in enumerate(tag_lines) if line == b'tags:\n'] +if len(tag_start) > 1 or any(line.startswith(b'tags:') and line != b'tags:\n' for line in tag_lines): + raise SystemExit('Unexpected Patroni tags block') +if tag_start: + tag_end = next((i for i in range(tag_start[0] + 1, len(tag_lines)) + if tag_lines[i] and tag_lines[i][:1] not in (b' ', b'\t', b'\n', b'#')), + len(tag_lines)) + tags = tag_lines[tag_start[0] + 1:tag_end] + if any(line.lstrip().startswith((b'nofailover:', b'failover_priority:')) for line in tags): + raise SystemExit('Pre-existing failover policy must not be overwritten') + changed = b''.join(tag_lines[:tag_end]) + b' nofailover: true\n' + b''.join(tag_lines[tag_end:]) +else: + changed += (b'' if changed.endswith(b'\n') else b'\n') + b'tags:\n nofailover: true\n' expected, replacement = (original, changed) if action == 'apply' else (changed, original) if current == replacement: raise SystemExit(0) @@ -91,13 +147,14 @@ PY restore_once() { if [[ ! -f "$backup" ]]; then - [[ "$(setting)" == '0,default' ]] + [[ "$(setting)" == '0,default' ]] && [[ "$(nofailover_cleared)" == 'true' ]] return fi edit_config restore || return 1 reload_patroni || return 1 for (( check = 0; check < 10; check++ )); do - [[ "$(setting)" == '0,default' ]] && return 0 + [[ "$(setting)" == '0,default' ]] && + [[ "$(nofailover_cleared)" == 'true' ]] && return 0 sleep 1 done return 1 @@ -119,6 +176,7 @@ case "$action" in arm) command -v systemd-run >/dev/null if ! is_standby || ! is_streaming || [[ "$(setting)" != '0,default' ]] || + [[ "$(nofailover_cleared)" != 'true' ]] || timer_pending || [[ ! -f "$config" ]] || [[ -L "$config" ]]; then echo 'Standby baseline, Patroni config, or timer precondition failed.' >&2 exit 1 @@ -136,7 +194,8 @@ case "$action" in apply) timer_pending && [[ -f "$backup" ]] || { echo 'Independent timer or backup missing.' >&2; exit 1; } - if ! is_standby || ! is_streaming || [[ "$(setting)" != '0,default' ]]; then + if ! is_standby || ! is_streaming || [[ "$(setting)" != '0,default' ]] || + [[ "$(nofailover_cleared)" != 'true' ]]; then echo 'Standby baseline changed before apply.' >&2 exit 1 fi @@ -151,7 +210,8 @@ case "$action" in exit 1 fi for (( check = 0; check < 15; check++ )); do - if [[ "$(setting)" == "$(( delay_seconds * 1000 )),configuration file" ]]; then + if [[ "$(setting)" == "$(( delay_seconds * 1000 )),configuration file" ]] && + [[ "$(failover_excluded)" == 'true' ]]; then timer_pending && echo 'delay-applied' && exit 0 break fi @@ -165,7 +225,8 @@ case "$action" in restore ;; cancel) - if ! is_standby || ! is_streaming || [[ "$(setting)" != '0,default' ]]; then + if ! is_standby || ! is_streaming || [[ "$(setting)" != '0,default' ]] || + [[ "$(nofailover_cleared)" != 'true' ]]; then echo 'Do not cancel until default delay and streaming return.' >&2 exit 1 fi @@ -175,9 +236,10 @@ case "$action" in ;; status) if timer_pending; then armed='true'; else armed='false'; fi - printf 'standby=%s streaming=%s armed=%s delay=%s backlog_bytes=%s free_kb=%s\n' \ + printf 'standby=%s streaming=%s armed=%s delay=%s nofailover=%s nofailover_cleared=%s backlog_bytes=%s free_kb=%s\n' \ "$(is_standby && echo true || echo false)" \ "$(is_streaming && echo true || echo false)" "$armed" "$(setting)" \ + "$(failover_excluded)" "$(nofailover_cleared)" \ "$(query 'SELECT COALESCE(pg_wal_lsn_diff(pg_last_wal_receive_lsn(), pg_last_wal_replay_lsn()), 0)::bigint')" \ "$(/usr/bin/docker exec "$container" df -Pk /var/lib/docgrid/opensql/data/pgsql | awk 'NR == 2 {print $4}')" ;; diff --git a/scripts/opensql/test_ha_evidence.py b/scripts/opensql/test_ha_evidence.py index 997d876d..dc01007c 100644 --- a/scripts/opensql/test_ha_evidence.py +++ b/scripts/opensql/test_ha_evidence.py @@ -122,6 +122,17 @@ def test_invalid_run_metadata_is_rejected_before_directory_creation(self): })()) self.assertFalse(other.exists()) + def test_external_run_id_matches_the_scenario_id(self): + """A caller-supplied ID keeps the fixture, run directory, and ledger joinable.""" + other = Path(self.temporary.name) / "same-run" + self.assertEqual(0, EVIDENCE.main([ + "init", "--run-dir", str(other), "--run-id", "scenario123", + "--scenario", "permission-replica-lag", "--config-sha256", "b" * 64, + "--opensql-version", "test", "--openproxy-version", "test", + "--patroni-version", "test", "--etcd-version", "test", + ])) + self.assertEqual("scenario123", json.loads((other / "manifest.json").read_text())["run_id"]) + if __name__ == "__main__": unittest.main() diff --git a/scripts/opensql/test_permission_replica_lag.py b/scripts/opensql/test_permission_replica_lag.py new file mode 100644 index 00000000..21ad624a --- /dev/null +++ b/scripts/opensql/test_permission_replica_lag.py @@ -0,0 +1,63 @@ +"""Regression checks for permission-lag evidence classification and safety gates.""" + +from __future__ import annotations + +import importlib.util +import unittest +from pathlib import Path + + +SCRIPT = Path(__file__).with_name("permission_replica_lag.py") +SPEC = importlib.util.spec_from_file_location("permission_replica_lag", SCRIPT) +if SPEC is None or SPEC.loader is None: + raise RuntimeError("Permission experiment module cannot be loaded") +EXPERIMENT = importlib.util.module_from_spec(SPEC) +SPEC.loader.exec_module(EXPERIMENT) + + +class PermissionReplicaLagTest(unittest.TestCase): + """Keep a denied HTTP response distinct from a proven standby denial.""" + + def test_403_without_standby_query_is_invalid(self): + """An authentication failure or primary read must not appear as a safe denial.""" + no_query = dict.fromkeys(EXPERIMENT.NODES, 0) + self.assertEqual("INVALID", EXPERIMENT.classify_result(403, False, no_query, 403, 403)) + primary_query = no_query | {EXPERIMENT.NODES[0]: 1} + self.assertEqual("INVALID", EXPERIMENT.classify_result(403, False, primary_query, 403, 403)) + + def test_standby_read_separates_stale_and_denied(self): + """Both outcomes require a real standby lookup and two 403 controls.""" + standby_query = {EXPERIMENT.NODES[0]: 0, EXPERIMENT.NODES[1]: 1, + EXPERIMENT.NODES[2]: 0} + self.assertEqual("STALE_ADMIN_ALLOWED", EXPERIMENT.classify_result( + 200, True, standby_query, 403, 403)) + self.assertEqual("DENIED", EXPERIMENT.classify_result( + 403, False, standby_query, 403, 403)) + self.assertEqual("INVALID", EXPERIMENT.classify_result( + 403, True, standby_query, 403, 403)) + self.assertEqual("INVALID", EXPERIMENT.classify_result( + 200, True, standby_query, 200, 403)) + + def test_standby_status_requires_both_tag_and_timer(self): + """A delay alone cannot authorize a revocation while replicas remain promotable.""" + delayed = ("standby=true streaming=true armed=true delay=120000,configuration file " + "nofailover=true nofailover_cleared=false backlog_bytes=456 free_kb=2000000") + both = dict.fromkeys(EXPERIMENT.STANDBYS, delayed) + self.assertTrue(EXPERIMENT.delayed_status_is_safe(both)) + self.assertFalse(EXPERIMENT.delayed_status_is_safe( + both | {EXPERIMENT.NODES[1]: delayed.replace("nofailover=true", "nofailover=false")})) + self.assertFalse(EXPERIMENT.delayed_status_is_safe({EXPERIMENT.NODES[1]: delayed})) + + def test_restored_status_requires_promotion_eligibility(self): + """Cleanup is not complete while either standby remains excluded or delayed.""" + restored = ("standby=true streaming=true armed=false delay=0,default " + "nofailover=false nofailover_cleared=true backlog_bytes=0 free_kb=2000000") + both = dict.fromkeys(EXPERIMENT.STANDBYS, restored) + self.assertTrue(EXPERIMENT.restored_status_is_safe(both)) + self.assertFalse(EXPERIMENT.restored_status_is_safe( + both | {EXPERIMENT.NODES[2]: restored.replace("nofailover_cleared=true", + "nofailover_cleared=false")})) + + +if __name__ == "__main__": + unittest.main() diff --git a/scripts/opensql/test_standby_apply_delay_guard.py b/scripts/opensql/test_standby_apply_delay_guard.py new file mode 100644 index 00000000..f47bb042 --- /dev/null +++ b/scripts/opensql/test_standby_apply_delay_guard.py @@ -0,0 +1,66 @@ +"""Exercise the exact embedded Patroni YAML editor without a live VM.""" + +from __future__ import annotations + +import subprocess +import sys +import tempfile +import unittest +from pathlib import Path + + +GUARD = Path(__file__).with_name("standby_apply_delay_guard.sh") +MARKER = "python3 - \"$1\" \"$config\" \"$backup\" \"$delay_seconds\" <<'PY'\n" + + +class StandbyGuardConfigTest(unittest.TestCase): + """Preserve original YAML bytes while toggling delay and promotion exclusion.""" + + def setUp(self): + source = GUARD.read_text() + self.editor = source.split(MARKER, 1)[1].split("\nPY\n}", 1)[0] + self.temporary = tempfile.TemporaryDirectory() + self.addCleanup(self.temporary.cleanup) + self.config = Path(self.temporary.name) / "patroni.yml" + self.backup = Path(self.temporary.name) / "original.yml" + + def edit(self, action): + return subprocess.run([sys.executable, "-", action, str(self.config), + str(self.backup), "120"], input=self.editor, + text=True, capture_output=True, check=False) + + def test_apply_and_restore_preserve_original(self): + """A real existing tags block gains only nofailover and then returns byte-for-byte.""" + original = (b"scope: docgrid\npostgresql:\n data_dir: /data\ntags:\n" + b" noloadbalance: false\n clonefrom: false\n") + self.config.write_bytes(original) + self.backup.write_bytes(original) + self.assertEqual(0, self.edit("apply").returncode) + changed = self.config.read_bytes() + self.assertIn(b"recovery_min_apply_delay: 120s", changed) + self.assertIn(b"nofailover: true", changed) + self.assertEqual(0, self.edit("restore").returncode) + self.assertEqual(original, self.config.read_bytes()) + + def test_preexisting_failover_policy_is_not_overwritten(self): + """An owner-defined tag must stop the experiment before editing config.""" + original = b"postgresql:\n data_dir: /data\ntags:\n nofailover: false\n" + self.config.write_bytes(original) + self.backup.write_bytes(original) + self.assertNotEqual(0, self.edit("apply").returncode) + self.assertEqual(original, self.config.read_bytes()) + + def test_concurrent_change_is_not_destroyed_on_restore(self): + """Do not overwrite a config another administrator changed during the run.""" + original = b"postgresql:\n data_dir: /data\n" + self.config.write_bytes(original) + self.backup.write_bytes(original) + self.assertEqual(0, self.edit("apply").returncode) + changed = self.config.read_bytes() + b"other: value\n" + self.config.write_bytes(changed) + self.assertNotEqual(0, self.edit("restore").returncode) + self.assertEqual(changed, self.config.read_bytes()) + + +if __name__ == "__main__": + unittest.main() From d154dd180a83d9e77acbebbb1becb1c4f0caf6c1 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?=EA=B9=80=EA=B8=B0=EB=AF=BC?= Date: Thu, 1 Oct 2026 01:47:52 +0900 Subject: [PATCH 3/3] [Docs] Record OpenSQL permission-delay safety rerun --- docs/design/opensql-permission-replica-lag.md | 57 +++++++++---- ...opensql-permission-replica-lag-20260929.md | 5 ++ ...-permission-replica-lag-safety-20261001.md | 81 +++++++++++++++++++ 3 files changed, 126 insertions(+), 17 deletions(-) create mode 100644 docs/test-results/opensql-permission-replica-lag-safety-20261001.md diff --git a/docs/design/opensql-permission-replica-lag.md b/docs/design/opensql-permission-replica-lag.md index 21238721..30771308 100644 --- a/docs/design/opensql-permission-replica-lag.md +++ b/docs/design/opensql-permission-replica-lag.md @@ -19,11 +19,18 @@ DB에서 역할을 읽고 30초간 캐시한다. `UserRoleCommandService.revokeR - primary 1대와 streaming standby 2대를 확인한다. 이미 다른 복제 지연이나 장애·부하 시험이 실행 중이면 시작하지 않는다. - 역할 SQL의 standby 도착을 먼저 입증하지 못하면 지연을 적용하지 않는다. +- 이전 실행의 시험 ADMIN 계정이 0개인지 먼저 확인한다. 이번 실행의 네 계정은 + node1 VM 호스트에 독립 정리 타이머가 무장된 뒤에만 생성한다. - 두 standby 호스트에 시험 JVM·SSH 연결과 독립된 자동 원복 타이머, 원본 - Patroni YAML 백업을 준비한다. 단독 standby에서 자동 원복을 확인한 뒤에만 + Patroni YAML의 root 전용 영속 백업을 준비한다. 단독 standby에서 원복을 확인한 뒤에만 두 노드에 지연을 적용한다. `finally`와 수동 원복은 추가 방어선이다. -- 지연 동안 수신·재생 LSN 차이와 디스크 여유를 감시한다. 미리 정한 상한, - 역할 변경, 타이머 오작동, 관측 누락 중 하나라도 발생하면 즉시 재개한다. +- 지연과 함께 각 standby의 로컬 `tags.nofailover: true`를 임시 적용한다. + Patroni 로컬 REST와 클러스터 뷰 양쪽에서 승격 제외가 확인되기 전에는 권한을 + 회수하지 않는다. 두 standby를 동시에 제외한 동안 primary가 상실되면 자동 + failover가 멈추므로, 이 실험은 승인된 짧은 구간에서만 실행한다. +- 적용 전 디스크 여유가 1 GiB 이상인지 검사하고, 관측 시점마다 수신·재생 LSN + 차이, 역할, 타이머 상태를 기록한다. 역할 변경, 타이머 오작동, 관측 누락 또는 + 자동 원복 시한 접근 시 즉시 중단한다. 연속적인 backlog 감시는 구현하지 않았다. - 자동 원복 장치가 아직 검증되지 않았다면 클러스터 실험은 실행하지 않는다. ### 2026-09-29 실제 사전 시험에서 발견한 제약 @@ -39,10 +46,17 @@ systemd 자동 재개가 정상이라는 사실은 이 제약을 해결하지 1초 뒤 제거했다. 이 값은 standby의 Patroni 로컬 `postgresql.recovery_conf`에서 관리하고, Patroni REST `POST /reload`를 사용해야 유지된다. 호스트 타이머가 원본 YAML을 정확히 되돌리고 reload한다. -지연된 standby가 failover 후보가 되는 위험을 최소화하도록 짧은 관측 구간을 -두며, 이 시험 중 primary 장애를 의도적으로 주입하지 않는다. `patronictl +지연된 standby는 적용 기간에 승격 후보에서 제외하고, 원복 뒤 승격 제외 태그가 +양쪽 Patroni 뷰에서 사라졌는지 확인한다. 이 시험 중 primary 장애를 의도적으로 +주입하지 않는다. `patronictl pause`나 Patroni 중지로 HA 제어 자체를 우회하지 않는다. +VM 호스트의 `systemd-run` 타이머는 Mac 프로세스·SSH가 죽어도 작동하지만 VM +재부팅까지 지속되는 영구 타이머는 아니다. Patroni 원본 YAML은 영속 디렉터리에 +보관하므로 재부팅 후에는 수동 `restore`가 가능하다. node1 자체가 사라지면 +시험 계정 자동 삭제도 보장하지 못한다. 이 경우 시험을 `INVALID`로 처리하고 +관리자가 복제·태그·계정을 확인해야 한다. + ## 독립 시나리오 | 이름 | 조건 | 유효성 및 판정 | @@ -59,30 +73,39 @@ JWT는 요청을 받는 시험 앱과 같은 서명 설정으로 발급하며 ## 본 실험 순서 -1. 두 시험 사용자를 primary에 만들고 양쪽 standby에서도 ADMIN이 보이는지 확인한다. -2. 독립 자동 원복 타이머가 무장된 것을 확인한 후 양쪽 standby의 Patroni 로컬 - 설정에 apply delay를 추가하고 Patroni를 reload한다. 양쪽의 실제 - `pg_settings` 값이 동일하며 타이머가 예정되어 있을 때만 다음 단계로 간다. +1. node1 호스트의 독립 계정 정리 타이머를 무장한 뒤 네 시험 사용자를 만들고, + 양쪽 standby에서도 ADMIN이 보이는지 확인한다. +2. 각 standby의 독립 설정 원복 타이머를 확인한 후 Patroni 로컬 설정에 apply + delay와 `nofailover`를 함께 추가하고 reload한다. 실효 `pg_settings`와 + 로컬·클러스터 태그 모두 일치하며 타이머가 예정되어 있을 때만 다음 단계로 간다. 3. 관리자 HTTP API로 대상 ADMIN을 회수한다. 회수 응답과 세 노드의 역할 행, Redis 키 부재를 확인한다. 회수 뒤 읽은 primary LSN은 정확한 commit LSN이 아닌 참조 시점으로만 사용한다. 4. 대상 사용자로 `/admin/workers`를 한 번 호출하고 상태 코드, 역할 SQL의 노드별 증가량, Redis의 ADMIN 여부와 PTTL을 기록한다. -5. 즉시 양쪽 standby의 원본 Patroni 설정을 복원한다. 회수가 반영된 이후의 첫 403 시각을 - 1초 간격으로 관측해 복제 지연 구간과 캐시 잔류 구간을 구분한다. -6. 성공·실패와 관계없이 설정 원복·복제·단일 primary를 확인하고 이번 실행의 - 사용자·역할·Redis 키만 정리한다. 정상 원복 뒤 자동 원복 예약을 해제한다. +5. 즉시 양쪽 standby의 원본 Patroni 설정을 복원한다. 회수가 양쪽 standby에 + 반영됐는지 확인한 뒤, 별도 사용자와 비어 있는 캐시로 C1의 403을 확인한다. + 현재 러너는 복구 중 첫 403 시각이나 캐시 잔류 기간을 측정하지 않는다. +6. 성공·실패와 관계없이 설정 원복·복제·승격 제외 해제를 확인한다. 이번 실행의 + Redis 키를 지운 뒤 사용자·역할을 삭제하고 계정 0개를 확인한다. 정상 원복 뒤 + 각 자동 원복 예약을 해제한다. 삭제나 원복 확인이 실패하면 관측값은 보존하되 + 전체 실행 결과와 종료 코드를 실패로 둔다. ## 증거와 완료 조건 -기존 `ha_evidence.py`의 run ID와 요청·fault 이벤트를 재사용한다. 추가 -비식별 증거에는 코드 커밋, 설정 해시, UTC 시각, 노드 별칭, 실제 역할 상태, +기존 `ha_evidence.py`의 run ID를 시험 계정·결과 JSON·원장 manifest와 동일하게 +사용한다. 실행 직전의 비밀 제거 스냅샷에는 설치된 호스트 스크립트·로컬 소스·앱 +JAR 해시, JDBC URL의 해시, 세 노드의 health와 standby 상태, 이전 계약 해시를 +담고 해당 바이트의 SHA-256을 원장에 기록한다. 추가 비식별 증거에는 UTC 시각, +노드 별칭, 실제 역할 상태, 실효 지연값, receive/replay LSN, SQL 호출 수 차이, HTTP 상태, Redis ADMIN 여부와 PTTL, 자동 원복 발동 여부를 기록한다. 비밀번호·JWT·응답 본문·주소·실제 사용자 정보는 공개 파일에 넣지 않는다. -결과는 `STALE_ADMIN_ALLOWED`, `DENIED`, `INVALID`로 나눈다. `DENIED`는 +결과는 `STALE_ADMIN_ALLOWED`, `DENIED`, `INVALID`로 나눈다. `DENIED`도 +M의 역할 SQL이 standby에 도착했고 primary에는 도착하지 않았다는 증거와 +두 대조군 403을 요구한다. 단순 HTTP 403만으로는 `INVALID`다. `DENIED`는 이번 환경에서 재현되지 않았다는 뜻이며 향후 설정 변경까지 안전하다는 증명은 아니다. `pg_stat_statements`는 노드별 누적 통계이므로 다른 요청과 분리할 수 -없으면 `INVALID`로 처리한다. 이 PR의 성공은 취약점 재현 자체가 아니라 네 +없으면 `INVALID`로 처리한다. 이 시험의 성공은 취약점 재현 자체가 아니라 네 시나리오의 유효한 판정, 원본 증거, 정상 원복 및 한계 문서화다. diff --git a/docs/test-results/opensql-permission-replica-lag-20260929.md b/docs/test-results/opensql-permission-replica-lag-20260929.md index 4c6188bf..dd43cc9c 100644 --- a/docs/test-results/opensql-permission-replica-lag-20260929.md +++ b/docs/test-results/opensql-permission-replica-lag-20260929.md @@ -117,3 +117,8 @@ LSN 수집을 처음 추가한 실행은 사전 검사에서 `docgrid` DB에도 primary로 보내는 것과, Redis cache resurrection 경쟁을 별도로 막는 것을 함께 검토해야 한다. WebSocket 기존 구독, 실제 OpenProxy 프로세스 장애, Patroni failover, 앱 VM의 GCP 내부 부하 시험은 여기서 검증하지 않았다. + +이후 독립 계정 정리, standby 승격 제외, 엄격한 판정 및 실행 출처 연결을 +보강해 두 standby에서 재검증했다. +[안전장치 보강 후 실행 결과](opensql-permission-replica-lag-safety-20261001.md)를 +참고한다. 위의 2026-09-29 관측값과 당시 `/run` 백업 설명은 역사적 결과로 유지한다. diff --git a/docs/test-results/opensql-permission-replica-lag-safety-20261001.md b/docs/test-results/opensql-permission-replica-lag-safety-20261001.md new file mode 100644 index 00000000..57cb4e12 --- /dev/null +++ b/docs/test-results/opensql-permission-replica-lag-safety-20261001.md @@ -0,0 +1,81 @@ +# OpenProxy 권한 복제 지연 시험의 안전장치·증거 보강 결과 + +## 결론과 범위 + +2026-10-01 KST, 기존 OpenSQL 3노드에서 두 standby에 **동시에** 임시 +`recovery_min_apply_delay=120s`와 `tags.nofailover: true`를 적용해 실제 +DocGrid HTTP 권한 경로를 재검증했다. ADMIN 회수 후 OpenProxy 경유 요청은 +**200**이었고, 역할 SQL은 지연된 standby에 도착했으며 Redis에 옛 ADMIN이 +재저장됐다. 같은 시점 primary 직결 대조군과 복제 정상화 후 OpenProxy 대조군은 +모두 **403**이었다. 이 문서는 [기존 재현 결과](opensql-permission-replica-lag-20260929.md)에 +시험 안전성·판정·출처 검증을 더한 후속 기록이다. **인가 결함 자체를 고친 +결과는 아니다.** + +동시 적용 중 두 standby는 승격 후보에서 제외됐다. 따라서 그 짧은 구간에 +primary까지 상실했다면 자동 failover가 불가능했을 수 있다. 이 위험을 승인받은 +기존 시험 클러스터에서만 수행했고, primary 장애는 주입하지 않았다. +[Patroni의 `nofailover` 설정](https://patroni.readthedocs.io/en/latest/yaml_configuration.html)은 +해당 멤버가 리더 경쟁에 참여하지 않도록 한다. + +## 네 가지 보강과 검증 + +아래 명령의 ``·``·``는 실제 값을 공개하지 않는 +자리표시자다. 호스트 스크립트는 기존 VM의 root 전용 경로에 설치했다. +계정·프로젝트 ID·IP·JDBC URL·암호·JWT는 Git에 기록하지 않았다. + +| 보강 | 실행 위치와 명령·코드 | 왜 실행했나 | 결과 요약과 해석 | +| --- | --- | --- | --- | +| 시험 ADMIN 계정 독립 정리 | node1 VM 호스트: `sudo /usr/local/sbin/docgrid-permission-fixture arm docgrid-node1 300` → `create` → `systemctl start docgrid-fixture-reset-.service` → `guard-status` → `cancel` | 클라이언트 JVM·SSH가 비정상 종료돼도 VM 호스트의 별도 타이머가 이번 실행의 네 계정을 지우도록 한다. 실제 DB에 계정을 만든 뒤 서비스의 정리 동작을 연습했다. | 무장 확인 후에만 계정이 생성됐다. 정리 서비스 실행 후 `armed=true fixture_count=0`, 취소 후 `armed=false fixture_count=0`; 이전 실행 계정 검사도 0. **300초를 기다린 자동 발화 자체는 시험하지 않았고**, VM 재부팅·node1 상실에는 transient timer가 지속되지 않는다. | +| 지연 standby 승격 방지 | node2·3 VM 호스트 각각: `standby-apply-delay-guard arm 90 15` → `apply` → `status` → `restore` → `cancel` | 지연된 복제본을 새 primary로 뽑지 않도록 Patroni 로컬 설정의 apply delay와 `nofailover`를 함께 적용·원복한다. 원본 YAML은 root-only 영속 백업에 둔다. | 두 호스트 모두 적용 중 `streaming=true`, `delay=15000,configuration file`, `nofailover=true`가 로컬·클러스터 Patroni 뷰에서 확인됐다. 원복 뒤 `delay=0,default`, `nofailover=false`, `nofailover_cleared=true`, backlog 0. 그 뒤 본시험에서도 두 standby에 동일한 관문을 통과했다. | +| 판정·정리 실패 처리 | 로컬: `python3 -m unittest scripts.opensql.test_permission_replica_lag scripts.opensql.test_standby_apply_delay_guard scripts.opensql.test_ha_evidence` 및 실제 `permission_replica_lag.py` 실행 | 단순 403을 안전한 라우팅으로 오판하지 않고, standby SQL 도착·두 대조군·원복 상태를 모두 확인한다. cleanup 실패는 관측값과 별개로 전체 실행을 무효화한다. | 해당 단위 테스트 **14개** 통과. 본시험에서 M의 역할 SQL은 primary +0, node2 +1, node3 +0; C2는 primary +1, standby +0. C1·C2 모두 403. 최종 원복 경고가 없어서 결과 `STALE_ADMIN_ALLOWED`를 유지했다. cleanup 실패를 실제 클러스터에 주입한 시험은 아니다. | +| 동일 실행의 출처 연결 | 로컬: `run_permission_replica_lag_local.py --env-file --jar --redis-port 16379 --output --apply-delay`; VM: 각 helper의 `sha256sum`; 종료 후 `ha_evidence.py verify --run-dir ` | 소스·설치 스크립트·실행 JAR·라이브 사전 상태가 서로 다른 시점의 것이면 결과를 혼동할 수 있다. 하나의 run ID와 해시로 원장을 묶는다. | 실행 ID `53688fbe7475`가 manifest·시나리오 JSON·시험 계정에 공통으로 쓰였다. 실행 코드 커밋 `5fba43e10901eb72f1b13b57f7cc7fb704eced19`; live preflight SHA-256 `84955dc776a71391f99c359f1e1d6e1d4e798044a1eec94dd014ba3cb020cab5`; JAR SHA-256 `a885e4740e9eeb9615452a84a3b07979a76410f2d3644c4e61075161613cac00`. 설치 helper와 체크아웃 소스의 해시가 일치했고 원장 검증은 `complete=true`였다. | + +실행 JAR은 기존 DocGrid 코드의 빌드 산출물이다. 이 후속 변경은 서버의 +인가 로직을 수정하지 않으며, 두 시험 JVM은 로컬 loopback에서만 실행했다. +Redis는 기존 인스턴스와 격리한 임시 컨테이너의 `127.0.0.1:16379`를 사용했다. +SSH 터널을 사용한 기능·정합성 시험이지 GCP 내부 부하 또는 장애 전환 성능 +시험이 아니다. + +## 최종 실제 실행의 관측값 + +최초 안전장치 재현은 코드 커밋 전의 예비 실행이었고, 아래 표는 **코드 커밋 뒤** +재실행한 `53688fbe7475`만을 기준으로 한다. UTC 실행 구간은 +`2026-09-30T16:28:32.947Z`부터 `16:32:02.963Z`까지다. 원본 원장은 +Git 밖의 소유자 전용 디렉터리에 보존했다. + +| 관측 지점 | 실제 결과 | 해석 | +| --- | --- | --- | +| 시작 전 | primary 1대, streaming standby 2대. 두 standby 모두 지연 0, `nofailover=false`, backlog 0 | 이전 실험 설정 없이 시작했다. | +| R 정상 라우팅 | 6회 cache miss 중 역할 SQL 증가: primary +0, node2 +3, node3 +3 | 실제 앱의 OpenProxy 경로에서 권한 조회가 standby에 도착했다. | +| 두 standby 지연 구간 | 각각 `delay=120000,configuration file`, `nofailover=true`, `armed=true`, backlog 192B | 재생 지연과 승격 제외가 동시에 유효했다. 타이머 무장도 관측됐다. | +| M 회수 뒤 요청 | primary ADMIN 없음·양 standby ADMIN 있음. OpenProxy `GET /admin/workers` **200**; 역할 SQL primary +0, node2 +1, node3 +0; Redis ADMIN 재저장, PTTL 24,515ms | 뒤처진 권한을 읽은 앱이 관리 요청을 허용했다. 보안 정합성 결함은 미해결이다. | +| C2 지연 중 primary 직결 | **403**; 역할 SQL primary +1, standby +0 | 같은 DB 변경을 최신 primary에서 읽으면 거부된다. | +| C1 복제 복구 후 OpenProxy | **403**; 양 standby에서도 ADMIN 없음 | 복제 정상화 뒤 새 사용자의 권한 판단은 거부됐다. 첫 403 도달 시간은 측정하지 않았다. | +| 종료·원장 | 양 standby streaming, 지연 0, `nofailover=false`, backlog 0, timer 비활성. 시험 계정 0; `ha_evidence verify` complete true, 요청 16건 중 14건 2xx·2건 예상된 403·결과 불명 0 | 판정은 `STALE_ADMIN_ALLOWED`이고 cleanup 경고는 없다. `403` 2건은 대조군의 정상적인 거부다. | + +비공개 원본 파일 확인용 SHA-256: `permission-scenarios.json` = +`72fc51c833e037ff49dcd83a241983498c24fb71b8d74ca8e6b05c5f77a2c85c`, +`events.jsonl` = +`42e93ed529be86154b046e833fa3cf9616fe620d85044f00034e8315efb4f453`. +라이브 preflight에는 당시 노드 health·standby 상태·이전 계약 증거 해시·앱 JAR· +로컬 소스 및 설치 helper 해시·JDBC URL **해시**를 넣었다. 이는 전 제품 설정을 +다시 수집한 스냅샷은 아니므로 그 수준의 재현성까지 주장하지 않는다. + +## 원복 확인과 남은 한계 + +시험 후 별도 조회로 양 standby의 지연 0·`nofailover` 해제·streaming·backlog +0을 다시 확인했다. node1의 이전 실행 시험 계정도 0이었다. 임시 Redis, 두 +시험 JVM, SSH 터널과 로컬 임시 자격증명 사본은 종료·삭제했다. 기존 VM의 +root 전용 helper 및 그 설치 전 백업은 남겨 두었다. 원래 DB 자격증명은 변경하지 +않았다. + +- `systemd-run` transient timer와 node1 계정 정리는 VM 재부팅·node1 상실까지 + 보장하지 않는다. 영속 백업을 통한 수동 원복 경로는 있지만, 이 고장 유형은 + 이번에 실제로 주입하지 않았다. +- cleanup 실패·이중 standby 중 primary 상실을 실제로 주입하지 않았다. 코드의 + `INVALID` 판정과 `nofailover` 관문은 단위 테스트·상태 조회로 검증했지만, + 장애 전체 조합의 안전성을 증명한 것은 아니다. +- `pg_stat_statements` 차이는 격리된 시험 구간의 누적 카운터 증거다. SQL + 단건의 분산 추적 ID나 높은 부하에서의 보안 위반율을 뜻하지 않는다. +- 이 결과는 OpenProxy 자체 결함, 권한 수정 완료, Redis cache resurrection + 해결, WebSocket 권한 회수, OpenProxy/primary 장애 시 HA 성능을 뜻하지 않는다.