Skip to content

fix: bound sandbox resource use on host and in container - #6

Merged
andybbruno merged 1 commit into
masterfrom
fix/sandbox-resource-exhaustion
Sep 14, 2026
Merged

andybbruno merged 1 commit into
masterfrom
fix/sandbox-resource-exhaustion

Conversation

@andybbruno

Copy link
Copy Markdown
Owner

Addresses the Corridor finding about unbounded output capture in _run_docker, plus related resource-exhaustion issues found while fixing it. All verified against real Docker.

The reported finding (valid)

run_docker used subprocess.run(capture_output=True), which buffers the entire output of docker exec in host memory. The max_output_bytes cap only ran afterwards, and the container's --memory limit does not apply to the host-side pipe.

Output is now streamed through Popen with a reader thread per stream, read in 64 KB chunks, and the process is killed as soon as a stream reaches the cap. DockerRunResult carries a truncated flag so the backend can tell a capped command from a completed one.

Reproduction, dd if=/dev/zero bs=1M count=100000 | cat in a real container:

elapsed 0.1s   truncated True   exit -13   output 100116 bytes
host peak RSS 110 MB  (baseline interpreter)

Before, that streamed 100 GB into the host process. The sandbox stays usable afterwards.

Note

The finding's alternative remediation — "enable the existing capture-offload path by default, it already uses head -c" — does not apply. There is no such path in this repo; _docker.py had only the single subprocess.run call.

Related issues found while fixing it

Timed-out commands kept running in the container. The timeout killed the host-side docker exec client only. Verified before the fix — after a 2s timeout on sh -c 'while true; do :; done' & sleep 30:

  PID  COMMAND
    7  sh -c sh -c 'while true; do :; done' & echo started; sleep 30
   13  sh -c while true; do :; done      ← still burning the container's CPU quota
   14  sleep 30

Same shape as the reported finding, bounded to the container rather than the host: a few timed-out loops permanently consume the 0.5 CPU, and repeats exhaust --pids-limit 128 until the sandbox is unusable.

The fix relies on a verified property of docker exec: each exec is its own session and process-group leader, and descendants keep that PGID even when reparented to PID 1. Each exec now records its PID, and timeouts (and cap kills) reap the group with kill -9 -<pgid>. Only sh builtins are used, and the pid file lives in the container's /tmp, never the shared dir.

Killed processes became zombies. Only visible after the fix above: PID 1 was sleep infinity, which never wait()s, so each reaped orphan held a PID slot forever — a slower version of the same exhaustion. The container now runs with --init. Post-fix ps shows no residue after either a timeout or a cap kill.

Unbounded startup hangs. docker info and docker run had no timeout, so a wedged daemon hung the DockerSandbox() constructor indefinitely. Both are bounded now, and a failed or timed-out start does a best-effort docker rm -f — a timed-out docker run could previously leak the container it had just created.

atexit.register(self.close) leaked every sandbox. The bound method kept a strong reference, so an abandoned sandbox was never collected and __del__ never fired; its container and temp dir survived until process exit. The hook now holds a weakref and is unregistered on close(). Verified: del sandbox; gc.collect() removes the container.

Minor: max_output_bytes is validated like the other limits (0 previously produced silently empty output), and a dead if ...: pass branch in close() is gone.

Checked, not a problem

  • Symlink escape via the shared dir. The container can ln -s /etc/passwd /shared/x, and the host-side file tools read that same directory — but FilesystemBackend._resolve_path resolves then enforces relative_to(root), and reads use O_NOFOLLOW.
  • image / extra_run_args injection. Developer-controlled, not agent-controlled, and passed as argv without a shell.

Not changed, worth knowing

The container runs as root with default capabilities. Defaulting to --user, --cap-drop=ALL or --security-opt=no-new-privileges would break legitimate workloads (apt install, pip install into system paths), and all three are reachable through extra_run_args today. On Linux this also means container-written files in the shared dir are root-owned, so shutil.rmtree on close can silently fail (ignore_errors=True) and leave temp dirs behind. An opt-in hardening flag would be a reasonable follow-up.

Testing

48 tests pass (7 new, covering the capped read path, process-group reaping on both timeout and cap, --init, failed-start cleanup and the atexit weakref); ruff check and format clean. The new helper tests exercise the real Popen path against local sh commands, so they need no Docker.

The uv.lock change is an unrelated version sync (0.0.2 → 0.1.1) that was already in the working tree.

🤖 Generated with Claude Code

Commands run through `execute()` are agent-controlled, so their output
and lifetime have to be bounded. Several paths were unbounded:

- `run_docker` buffered all stdout/stderr in host memory via
  `capture_output=True`; `max_output_bytes` was only applied afterwards,
  so `dd if=/dev/zero | cat` could exhaust host RAM. Output is now
  streamed in chunks and the process killed at the cap.
- The per-command timeout only killed the host-side `docker exec`
  client; the command kept running in the container, burning its CPU
  and PID budget. Each exec now records its process group (docker exec
  is a session leader, so descendants keep the pgid) and timeouts and
  cap kills reap the whole group.
- Reaped children became zombies because PID 1 was `sleep infinity`;
  the container now runs with `--init`.
- `docker info` and `docker run` had no timeout, so a wedged daemon hung
  the constructor forever, and a timed-out `docker run` leaked its
  container. Both are bounded and failed starts are cleaned up.
- `atexit.register(self.close)` held a strong reference, so an abandoned
  sandbox was never collected and its container and temp dir survived
  until process exit. The hook now holds a weakref and is unregistered
  on close.
- `max_output_bytes` is validated like the other limits.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@andybbruno
andybbruno merged commit 2706f86 into master Sep 14, 2026
3 checks passed
@andybbruno
andybbruno deleted the fix/sandbox-resource-exhaustion branch September 14, 2026 12:55
andybbruno added a commit that referenced this pull request Sep 14, 2026
fix: bound sandbox resource use on host and in container
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants