Skip to content

DUNDER-227: harden the opslevel app workloads with a read-only root filesystem - #92

Merged
jasonopslevel merged 2 commits into
mainfrom
jason/dunder-227-openshift-readonly-rootfs
Aug 28, 2026
Merged

jasonopslevel merged 2 commits into
mainfrom
jason/dunder-227-openshift-readonly-rootfs

Conversation

@jasonopslevel

Copy link
Copy Markdown
Contributor

Companion to the app-image change in
OpsLevel!19844 (DUNDER-227).

That MR makes the app image ship its code, vendored gems, and bootsnap cache
root-owned and unwritable at runtime, leaving tmp/, log/, and /tmp as the
only paths the app writes. That is what makes a read-only root filesystem viable
here.

What this does

readOnlyRootFilesystem is a container-level field, and the chart only
templated pod-level securityContext, so this adds a container-level block to
the six app workloads — web, the four workers, and scheduler:

containerSecurityContext:
  readOnlyRootFilesystem: true
  allowPrivilegeEscalation: false
  capabilities:
    drop:
      - ALL

Three emptyDir mounts, via a shared helper, cover everything the app still
writes:

path why
/home/opslevel/tmp config.cache_store is unset, so Rails 7.2 defaults to a file store at tmp/cache
/home/opslevel/log Rails creates it on demand; gitignored, so it never arrives via COPY
/tmp Tempfile backs ActiveStorage uploads and app/models/document.rb, and a read-only rootfs takes /tmp with it

Action required for OpenShift installs

Set opslevel.securityContext: {}.

emptyDir volumes are created root:root 0755, so something has to grant the app
group write — hence the new fsGroup: 1000 default. OpenShift's restricted-v2
SCC assigns fsGroup from the namespace's range and rejects a hardcoded value
at admission
, so OpenShift users must clear the pod-level block and let the SCC
inject it. This is called out in values.yaml and in the changie entry.

Notes

  • scheduler had no pod-level securityContext block at all, so it was rendering
    without fsGroup and could not have written to its own volumes. Added.
  • The init-certs init containers are deliberately left unhardened: they run
    update-ca-certificates, which writes /etc/ssl/certs on their own root
    filesystem. Same for the scheduler's migrations init container.
  • OpsSight is a separate image and is untouched.

Verification

helm lint is clean, and all six workloads render as expected:

opslevel-scheduler       fsGroup=1000  rorfs=true caps=["ALL"]
opslevel-web             fsGroup=1000  rorfs=true caps=["ALL"]
opslevel-worker-faktory  fsGroup=1000  rorfs=true caps=["ALL"]
opslevel-worker-high     fsGroup=1000  rorfs=true caps=["ALL"]
opslevel-worker-low      fsGroup=1000  rorfs=true caps=["ALL"]
opslevel-worker-search   fsGroup=1000  rorfs=true caps=["ALL"]

Rendered output was also parsed to confirm every hardened container mounts all
three writable paths, with certificate.enabled=true and opslevel.tls.enabled=true.

Not cluster-validated. The permission model is verified at the image level and
the templates are verified by rendering, but "does the app actually boot with a
read-only rootfs" needs a real deploy. Highest-risk unknowns are a write path not
found by grepping, and the fsGroup/emptyDir interaction.

Follow-ups (not in this PR)

  • opslevel-kubernetes needs the same container hardening for its raw
    Deployments (clusters/new-prod-runners/opssight/*, the etl/* fleet). The
    HelmRelease paths inherit these defaults; the raw manifests won't.
  • config.cache_store being unset means each replica keeps its own file cache in
    its own emptyDir, with no shared invalidation across web pods. Pointing it at
    the Redis already deployed here would fix that and drop /home/opslevel/tmp
    from the writable set.

@wesleyjellis wesleyjellis left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this all makes sense

# OpenShift: set `opslevel.securityContext: {}`. The restricted SCC assigns fsGroup
# from the namespace's range and rejects a hardcoded value at admission.
securityContext:
fsGroup: 1000

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

might be worth a comment to explain the lack of runAsGroup means that the user default runs with group 0 which is how we can 'read' the app code and such in /home/opslevel

@jasonopslevel

Copy link
Copy Markdown
Contributor Author

Pushed c138910 after the approval — flagging because it fixes a real bug in the approved version, not a cosmetic follow-up.

The container securityContext and the three emptyDir mounts were inserted inside the {{- if .Values.certificate.enabled }} guard that wraps the container volumeMounts and pod volumes keys in the worker and scheduler templates. With certificate.enabled at its default of false — every install not using a custom CA — none of the hardening rendered: no readOnlyRootFilesystem, no dropped capabilities, no writable volumes. Only web was unaffected, since its volumeMounts/volumes keys are unconditional.

The fix moves the guard below the unconditional blocks, so the hardening always renders and the certificate mounts stay conditional.

My original verification only ever passed --set certificate.enabled=true, which is exactly why this slipped through. It now renders correctly in both states of certificate.enabled and opslevel.tls.enabled, all six workloads, with the certificate/ca/ssl mounts intact when enabled.

…ilesystem

The app image now ships its code, vendored gems, and bootsnap cache
root-owned and unwritable at runtime, which makes a read-only root
filesystem viable for the app containers.

readOnlyRootFilesystem is a container-level field and the chart only
templated pod-level securityContext, so add a container-level block to
web, the four workers, and the scheduler, along with dropping all
capabilities and disallowing privilege escalation.

Three emptyDir mounts cover everything the app still writes:
/home/opslevel/tmp (Rails' default file cache lives at tmp/cache),
/home/opslevel/log, and /tmp -- Tempfile backs ActiveStorage uploads and
app/models/document.rb, and a read-only rootfs takes /tmp with it.

fsGroup: 1000 is what makes those volumes writable, since emptyDir is
created root:root 0755. OpenShift installs must clear
opslevel.securityContext: the restricted SCC assigns fsGroup from the
namespace range and rejects a hardcoded value at admission. Noted in
values.yaml.

The scheduler had no pod-level securityContext block at all, so it was
rendering without fsGroup and would not have been able to write to its
own volumes.

The init-certs init containers are deliberately left unhardened: they run
update-ca-certificates, which writes /etc/ssl/certs on their own root
filesystem.
The container securityContext and the tmp//log//tmp mounts were inserted
inside the `{{- if .Values.certificate.enabled }}` guard that wraps the
container's volumeMounts and the pod's volumes keys in the worker and
scheduler templates. With certificate.enabled at its default of false --
which is every install that does not use a custom CA -- none of it
rendered: no readOnlyRootFilesystem, no dropped capabilities, no writable
volumes.

Move the guard below the unconditional blocks so the app hardening always
renders and the certificate mounts stay conditional. web.yaml was already
correct; its volumeMounts and volumes keys are unconditional.

Verified across certificate.enabled and opslevel.tls.enabled in both
states: all six app workloads render readOnlyRootFilesystem with the three
emptyDir mounts, and the certificate/ca/ssl mounts still appear when
enabled. The earlier verification only exercised certificate.enabled=true,
which is why this was missed.
@jasonopslevel
jasonopslevel force-pushed the jason/dunder-227-openshift-readonly-rootfs branch from c138910 to 6153f9a Compare August 24, 2026 19:00
@jasonopslevel
jasonopslevel merged commit d4b07ee into main Aug 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants