Skip to content

Vagrant: cut VM boot time and add storage tuning knobs - #1381

Merged
gusthoff merged 4 commits into
AdaCore:mainfrom
gusthoff:dev/topic/infrastructure/vagrant/boot-time/2026-09-18
Sep 18, 2026
Merged

gusthoff merged 4 commits into
AdaCore:mainfrom
gusthoff:dev/topic/infrastructure/vagrant/boot-time/2026-09-18

Conversation

@gusthoff

Copy link
Copy Markdown
Collaborator

Boot was dominated by two avoidable costs: a stale eth0 stanza in the
base box that blocked startup for a fixed two minutes, and uncached guest
disk reads that made concurrent VMs contend for the host disk. Boot falls
from 3:11 to ~1:45 by default, and to 10-20 s with caching enabled.

  • Remove the box's stale /etc/netplan/01-netcfg.yaml. It declares an
    eth0 that does not exist here, and netplan feeds it into
    systemd-networkd-wait-online, which is invoked without --any and so
    blocks network-online.target -- and ssh.service -- for its full
    120 s timeout.
  • Add LEARN_VM_HOST_IO_CACHE (on/off, default off) and
    LEARN_VM_STORAGE_CONTROLLER to control host I/O caching on the storage
    controller. Off by default, matching VirtualBox: with caching on, guest
    writes may sit in host RAM, so a host crash can corrupt a guest file
    system.
  • Create machines as linked clones, sharing one base disk image instead of
    a full copy each.
  • Mask the snapd units, which sit on the critical chain to ssh.service
    and are unused in these build VMs.

Only Vagrantfile changes. Validated on both machines with four VMs
running: boot times as above, toolchain verified by compiling and running
a program, and make test_content_courses_ada-in-practice byte-identical
against an unchanged control VM.

Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com

gusthoff and others added 4 commits September 18, 2026 16:48
The box ships `/etc/netplan/01-netcfg.yaml` declaring an `eth0` that does
not exist on this hardware; the NIC is `enp0s3`, configured by
`00-installer-config.yaml`. netplan feeds every declared interface into
the generated `systemd-networkd-wait-online` drop-in, and that unit is
invoked without `--any`, so it waits for `eth0` until its 120 s timeout
expires and then fails. `network-online.target`, and therefore
`ssh.service`, stays blocked for that whole time.

Remove the file during provisioning, with `run: "always"` so existing
VMs are corrected on their next `vagrant up` rather than only VMs
created after a destroy.

Measured on `bento/ubuntu-26.04` 202606.01.0: boot 3:11 before, 1:45
after, with `systemd-networkd-wait-online` down from a 120 s timeout to
178 ms.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Several machines run at once, one `web`/`epub` pair per worktree, and all
of them are created from the same box. `vb.linked_clone` makes each
machine share a single base disk image instead of taking a full private
copy, so the host page cache can serve the blocks they read in common
instead of reading each copy separately.

Creation-time only: existing machines keep their full copies until they
are destroyed and recreated, so this has no effect until then and is not
yet exercised.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
VirtualBox creates the storage controller with host I/O caching off, so
every guest read goes to the physical disk. With several machines booting
near-identical images that is the dominant cost of `vagrant up`: the guest
waits tens of seconds for its own virtual disk to appear, and everything
else, including `ssh.service`, queues behind it.

`LEARN_VM_HOST_IO_CACHE` turns caching on, and
`LEARN_VM_STORAGE_CONTROLLER` names the controller to act on, since that
name is box-specific. Both are read on every `vagrant up`, so they apply
to existing machines. The default is off, matching VirtualBox: with
caching on, guest writes may sit in host RAM, so a host crash can corrupt
a guest file system.

An unrecognized cache value is refused rather than ignored, so a typo
cannot silently leave the controller in a state the caller did not ask
for.

Validation: `useHostIOCache` confirmed to flip in the VM's own
configuration; boot time falls by roughly 9x with caching on, and the
source-code example build is unchanged either way, as it reads from the
synced folders and is otherwise CPU-bound.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`snapd.apparmor.service` and `snapd.socket` sit on the critical chain to
`basic.target`, and so delay `ssh.service` -- the unit Vagrant waits for
while it prints "Connection reset. Retrying...". Nothing in these build
VMs uses snap.

The units are masked rather than removed: masking is idempotent, so the
provisioner is safe to re-run, and `systemctl unmask` restores them.
`run: "always"` so existing machines are covered too, since they report
"Machine already provisioned" and would otherwise never run it.

Validation: both machines report `masked` for all four units, a second
run is a no-op, and `systemctl is-system-running` still reports
`running` with no failed units.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@gusthoff
gusthoff merged commit 59d7dff into AdaCore:main Sep 18, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant