Skip to content

Multiple External Services Outage (502 Bad Gateway): revela, nextrole, intel, web3-capital #45

Description

@yannvr

Incident Description

A wide outage is affecting multiple external-facing applications in the fleet, all returning HTTP 502 Bad Gateway with Nginx headers.

Impacted Apps:

  1. revela (https://revela.club)
    • HTTP 502 Bad Gateway
    • Latency: 71 ms
  2. nextrole (https://nextrole.site)
    • HTTP 502 Bad Gateway
    • Latency: 68 ms
  3. intel (https://intel.hyperdrift.io)
    • HTTP 502 Bad Gateway
    • Latency: 82 ms
  4. web3-capital (https://web3.hyperdrift.io)
    • HTTP 502 Bad Gateway
    • Latency: 111 ms

Healthy Apps (Internal/Sandbox):

  • sandbox (HTTP 200, 28ms)
  • cargo (HTTP 200, 28ms)

Attempted Remediation:

  • Attempted to run the healing runbook (heal_service) on revela and nextrole.
  • Result: Failed. No automated runbooks exist for these services (no automated runbook for 'revela' — escalate to a human).

Diagnosis & Action Required:

This is a coordinated outage across multiple external services returning Nginx 502 errors, indicating a routing layer, reverse proxy, or upstream backend connection issue. Since these services are outside automated operational scope and lack wired runbooks, human intervention is required to investigate the routing layer or nginx reverse proxy configuration/backends.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions