Skip to content

H.264 passthrough: black-region, colour and NVENC fixes, plus HiDPI, back-pressure and the 4:4:4 cost work - #238

Open
pletch wants to merge 15 commits into
sol1:mainfrom
pletch:pr/h264-stack
Open

pletch wants to merge 15 commits into
sol1:mainfrom
pletch:pr/h264-stack

Conversation

@pletch

@pletch pletch commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Follows up pletch#1. This is the H.264 set from my fork, rebased onto
v1.10.2 and reordered the way you asked: the fixes come before the combine-cost
work and the auxiliary-view drop, so you can take a prefix of the branch and
stop wherever you like. The diagnostic instrumentation is not in it at all --
it stays in my fork, where the troubleshooting happens. Every commit builds and
passes cargo test, the JS tests and clippy on its own.

Commits, in order

Foundations -- everything after these builds on them:

  1. feat(hidpi) -- per-connection Native Resolution: the framebuffer at
    the browser's physical pixels, and exact RDP desktop scaling (e.g. 200%)
    through the display-control channel. Patch 011.
  2. feat(h264): combine AVC444 into full 4:4:4 -- reworks the 4:4:4
    combine: guacd marks paired views (004 gains a trailing <paired> flag),
    frames are copied in the decoder's own pixel format, and the decoder's
    diagnostics go to the console through one diagnostic() call. Brings
    tests/bench and tests/h264-instruction-format.mjs.
  3. feat(rdp): back-pressure -- patch 012 holds the RDPGFX frame ack
    while the client is behind, capped at 400ms, crediting the spacing the
    server already provided, so it does not fight a server that paces itself.
  4. feat(transport): binary blobs -- h264 and audio blobs as binary
    WebSocket frames instead of base64, about 25% off the wire. Opt-in per
    client (binaryBlobs=1), so older clients and recordings are unchanged.

Fixes:

  1. SPS colour signalling -- Windows declares full range with no colour
    description, which Chrome's hardware decoder reads as limited (crushed
    blacks); stock xrdp declares nothing. rustguac completes the SPS on the
    wire, checked byte-for-byte against ffmpeg's h264_metadata, and logs what
    the host's SPS said once per session.
  2. Black and green regions on Windows (the Bug: Black "ghost" regions after dynamic RDP resize on Windows RDPGFX hosts #118 / Bug: Black regions return without a resize: Windows recreates the RDPGFX surface at the same size, and 005 (#118) does not fire #223 family) -- a
    keyframe with zero region rects is no longer painted whole, and a black
    keyframe after Windows recreates its surface at the same size is withheld
    so the picture already on screen survives.
  3. Patch 015 -- guacd tells the client when a surface was recreated at
    the size it already had, so the withholding needs both signals and never
    hides a screen that has genuinely gone black (lock screen, blanking).
    Followed by a correction: surface sizes tracked per surface ID, since a
    session tears several surfaces down together.
  4. POC type 2 / NVENC -- Chrome's hardware decoder holds a whole DPB of
    pictures without bitstream_restriction, which NVENC at its defaults
    omits; every picture arrived five late, past the decode watchdog, and the
    screen froze or stayed white. max_num_reorder_frames=0 is added where POC
    type 2 already guarantees it, and nowhere else.
  5. Tunnel exceptions -- an exception in an instruction handler closes the
    tunnel, which read as an ordinary disconnect at both ends. The client now
    logs it with its stack, and the overlay keeps the real message instead of
    "Connection lost".

Features that build on the fixes:

  1. Combine cost -- copyTo()'s synchronous read-back turned out to be
    almost all of what 4:4:4 costs. Copying only the damaged rows took
    main-thread time in it from ~77% to ~16% on a 2992x1648 Windows session.
    A gate gives 4:4:4 up when the copy share, sync timeouts or flush time say
    the client cannot keep up, and probes back with an exponential backoff.
  2. Dropping the auxiliary view in transit -- the largest piece (~5k
    lines, most of it h264_refs.rs and tests), and the one to leave out if
    you would rather not carry it. Windows must be offered AVC444 to use its
    hardware encoder, and FreeRDP only advertises RDPGFX 10.x alongside AVC444,
    so the chroma view cannot be declined at the source. This removes it
    between guacd and the browser where the slice headers prove nothing
    surviving refers to it: 13% of H.264 bytes on Windows, 43% on xrdp, and one
    decode per frame instead of two. Streams that cannot be proved pass
    through unchanged. RUSTGUAC_H264_AUX_DROP=0 switches it off. Adds a small
    src/instruction.rs for reading the instruction stream.
  3. Full Colour (4:4:4) -- the one per-entry choice: off (default) drops
    the chroma view where it can and never combines; on keeps it and combines.
    Stored as h264_chroma444.
  4. docs -- the guacd trace-level instruction points at guacd.env.

Deploying

  • guacd must be rebuilt: 004 changes, and 011, 012 and 015 are new.
    Built against FreeRDP 3 on Debian 13 with exactly this patch set.
  • One new entry field, h264_chroma444. Nothing else changes on disk.
  • New environment variable: RUSTGUAC_H264_AUX_DROP (0 off, unproven for
    experiments).
  • Client overrides for testing, as window global, query parameter or
    localStorage: h264Chroma444, h264CombineMaxPixels, h264CombineLog,
    h264FullRange, h264KeepBlackKeyframes.

Testing

  • Per commit: cargo test, clippy and the tests/*.mjs suites.
  • The same features run on my deployment -- my fork's main-fork, which
    carries the diagnostic instrumentation on top -- against Windows 11
    (hardware encoding, AVC444) and xrdp. Happy for you to test the parts you
    take; docs/rdp-h264.md has the verification steps and what each log line
    means.
  • Harnesses for the riskier pieces: tests/aux-drop-replay.mjs <recording>
    decodes a recording with and without the auxiliary views and compares, and
    RUSTGUAC_AUX_DROP_RECORDING=<recording> cargo test aux_drop_over_a_recording -- --ignored runs the real dropper over one.

Not in this PR

  • The diagnostic instrumentation -- an RDPGFX surface-operation trace in
    guacd, per-session frame telemetry, a browser-to-server diagnostic endpoint,
    and the black-region probes. It is what found the causes fixed above, and it
    stays in my fork. The decoder's Guacamole.H264Decoder.onDiagnostic hook is
    the seam it plugs into, if you ever want it.
  • The guacd warning in 004 that suggests GUAC_RDP_H264_CAPS_FILTER=1, a
    variable that was never implemented. Worth a one-line fix whenever
    convenient.

pletch added 14 commits October 3, 2026 11:47
Requests the framebuffer in the browser's physical pixels rather than its CSS
pixels, per connection, so text stays sharp on a HiDPI display. The browser
reports its devicePixelRatio on connect and the entry decides, since only the
entry knows whether the target scales its own UI.

The framebuffer factor and the desktop scale are separate numbers. The
framebuffer takes the browser's true ratio, capped at MAX_NATIVE_FACTOR,
because the client fits whatever framebuffer arrives into the available CSS
area: one framebuffer pixel lands on one physical pixel only when the two
agree. Snapping it to 1.8 on a 2.0 display leaves the client stretching by
1.111, which measures as a uniformly soft picture with almost no single-pixel
edges anywhere -- 0.01% of adjacent pixels differing by more than 100 levels,
against a native render's 0.82%. The cost of keeping them apart is a desktop
scaled 180% inside a 200% framebuffer drawing its UI about 10% small, which is
legible and adjustable on the host where a resample is neither.

The desktop scale goes out on two channels that are not equally capable
(patch 011). At connection time only 100/140/180 survive: MS-RDPBCGR restricts
deviceScaleFactor to those three, and FreeRDP transposes the pair when it
synthesises the single-monitor definition, so only equal values reach the
server intact. The display-control layout has no such problem -- disp.c builds
it directly, nothing transposes it, and MS-RDPEDISP allows desktopScaleFactor
anywhere in 100-500 -- so that layout carries the exact percentage beside the
nearest legal device factor, and since the client fits the display shortly
after connecting it is what the session ends up scaled by. The scale is
re-sent on every display update, because a MONITOR_LAYOUT carrying zeroes
resets the session to 100%.

X11 behind xrdp has no per-connection DPI negotiation, so the session has to
scale itself. docs/xrdp-dpi-scaling.md records that the exact scale already
arrives there and is already stored, so a patch would only have to act on it.
A WebGL2 shader unpacks the auxiliary view's packed chroma and converts to RGB
in one pass, inverting the encoder's chroma filter. Frames are copied in the
decoder's own pixel format, since copyTo() will not convert NV12 to I420, and a
lost context falls back to 4:2:0. The cost is trimmed by skipping the paint of a
paired main view, transferring an OffscreenCanvas rather than reading it back,
overlapping the two copyTo() calls, and uploading only the rows the damage rects
touch.

tests/bench times each stage of the combine with gl.finish() forcing
completion, and tests/h264-instruction-format.mjs round-trips the h264
instruction, trailing <paired> flag included, through the real
Guacamole.Parser.
Patch 012 holds the RDPGFX frame acknowledgement by the amount the client's
processing lag exceeds its target, minus the spacing the server has already
provided since the previous frame. That subtraction is what makes it a floor
rather than a second controller: a server that steers its own capture interval
from the same round trip reads an additive hold as client latency and answers
it with a longer interval, which shortens the hold, which shortens the
interval -- two controllers driving one frame rate through each other's
sensor. Crediting the elapsed gap leaves a self-pacing server in sole
ownership of the loop while still throttling one that floods, so no
per-connection target is needed and GUAC_RDP_H264_LAG_TARGET stays a
deployment-wide default.
Base64 sends four bytes for every three, so a quarter of everything on the
wire was encoding overhead. Round-tripped against two real Windows
captures: 21.2MB -> 16.1MB on an idle desktop (24.3%) and 145.9MB ->
109.5MB on a video session (24.9%), both reconstructing byte for byte.

Only `h264` and `audio` streams convert. Both reach
Guacamole.ArrayBufferReader, which was decoding base64 to an ArrayBuffer
only to hand over the bytes, so the reader now takes either form and the
decode disappears. `img` is excluded on purpose: its blobs feed
DataURIReader, which concatenates base64 straight into a data: URI and
wants the encoded form -- converting it would only have to be undone.

The conversion is in rustguac rather than a guacd patch because rustguac
tees the raw guacd stream to disk as the session recording. Converting
upstream of that tee would turn every recording binary and drag the
recording format, SessionRecording.js and the playback page along with
it. The guacd -> rustguac hop is loopback, so leaving it as text there
costs nothing. Order in guacd_to_ws is record, then convert, then send:
recordings see exactly what they saw before.

Clients opt in with binaryBlobs=1 on the WebSocket query. A client that
does not send it -- an older cached client.html, a third-party
integration, the recording player -- gets base64 exactly as today, which
matters because client.html is cached in memory at startup and a browser
can be holding an older one than the server.

WebSocket delivers text and binary frames in one order, so a blob lifted
out of the run arrives between the same neighbours it had. Nothing here
can reorder the stream, which is what keeps this a change of encoding
rather than one of transport.

Use is visible from both ends, since the failure mode is a silent fall
back to base64: `__guac_tunnel.binaryFrames` and `.binaryBytes` in the
browser, and `binary_blobs=true|false` on the Starting proxy log line.

tests/binary-blob-format.mjs pins the 8-byte header across the Rust/JS
boundary by lifting the constants out of both sources rather than
restating them. A disagreement there fails silently -- the client drops
frames it cannot recognise, so video stops while both ends look healthy.
MS-RDPEGFX defines full-range BT.709 and hosts encode to it, but the SPS
does not always say so in a form the browser acts on. Windows declares a
range and no description, which Chrome's hardware decoder reads as limited
and paints with crushed blacks; stock xrdp declares nothing at all, since it
passes x264 no VUI parameters. src/h264_rewrite.rs completes the signalling
in either shape -- adding the description beside a declared range, or
writing the whole block where there is none -- once per session, checked
byte-for-byte against ffmpeg's h264_metadata filter. It is the only fix that
reaches a stock-xrdp host at all, since AVC420 never calls setColorSpace.
A keyframe with zero region rects was painted
whole, as decoder green, because "no rects" was read as "whole picture valid";
in MS-RDPEGFX the rects are what changed, so an empty list changes nothing. An
absent list (older guacd) still means the whole picture; an empty one is now
decoded for its references and not painted (h264_undisplayed).

The main cause: Windows recreating its surface at the same size, mid-session,
with no resize. The new surface is empty, so its first keyframe decodes black
and covers the whole screen, and Windows then repaints only what it thinks
changed -- the behaviour behind sol1#118, where the reporter proved
that SuppressOutput off/on and RefreshRect do not make Windows re-stream the
surface. What fixed that was re-sending pixels the client side already had.
Under passthrough guacd has none, but the browser does: it is still showing
the right picture when the black keyframe arrives. So a keyframe decoded >=98%
black over a framebuffer that has kept its size for 5s is decoded for its
references and withheld (h264_black_keyframe_kept), and the partial repaint
lands on the old picture. h264KeepBlackKeyframes=off disables it.

Black is judged from a 64x32 downscale of the decoded keyframe -- one small
drawImage() and an 8KB read-back, and only for keyframes over a settled
framebuffer, which are rare.
…face

Windows occasionally deletes and recreates its RDPGFX surface at the dimensions
it already had, mid-session, with no resize and no ResetGraphics. The new
surface is empty, so its first picture is black over the whole surface, and
Windows then repaints only what it believes changed -- trusting the client to
hold the rest, which under passthrough it does and guacd does not.

Withholding that picture already worked, on a guess: a keyframe decoding at
least 98% black over a framebuffer that had kept its size for five seconds. The
guess is the problem. A screen can go black on its own -- a blank screensaver,
a display blanking on lock, a fade to black -- and suppressing one of those
leaves the previous desktop on display, which in a remote-access product is the
wrong way to be wrong.

guacd can tell the difference and the browser cannot. DeleteSurface carries
only a surface ID, so the dimensions are recorded at CreateSurface; a creation
that directly follows a deletion of the same surface at the same size is the
shape after which Windows does not repaint the screen. That reaches the client
as a trailing <recreated> flag on the h264 instruction, and the decoder now
requires it as well as the black picture.

Each half covers the other's false positive, and neither is redundant yet. A
legitimately black screen produces no recreation, so the flag declines it. A
same-size recreation carrying real content is not black, so the sample declines
that -- which matters because the trigger for the recreation is still
unidentified, and the leading candidates are secure-desktop switches. A
flag-only test would withhold a lock screen sight unseen, whatever it showed.

A black keyframe with no signal behind it is logged to the console and
painted, never withheld, so a gap in the signal shows itself rather than being
covered.

The detection is narrow on purpose. The last-deleted record is consumed by the
next creation whether or not it matches, so "recreated" can only mean a
creation that directly followed a deletion, with no timer to tune. A delete of
a surface whose size was never recorded leaves the record alone, because
Windows sends the delete twice and the second finds nothing to look up. Only an
H.264 command clears the flag: after a recreation followed by a progressive or
planar command, guacd still holds the pixels and there is nothing to withhold.

Asking the server to repaint is not an alternative. SuppressOutput and
RefreshRect are both ignored by the RDPGFX surface cache, and marking the layer
dirty was shipped as a fix for the resize case and retracted a day later,
having cured nothing, with RefreshRect suspected of making it worse. Every
attempt so far has tried to make something repaint; this one declines to paint
over a picture that is already correct.

Separately, this closes a hazard in the rect handling it sits next to. Three
conditions collapsed to a rect count of zero -- the server genuinely declaring
zero, a malformed metablock, and a failed allocation -- and zero reaches the
client as "this picture changes nothing", so the last two would blank the
display. The unreadable cases now report -1 and are sent as one rect covering
the surface, leaving zero to mean only what the server said.

The instruction test gains the new flag. It trails <paired>, so one position
short reads <paired> and two short reads a rect coordinate, which is non-zero
for nearly every rect -- and a picture wrongly marked recreated is withheld
from the screen rather than merely drawn oddly.
…d record

A session holds more than one RDPGFX surface and tears them down together. A
single record of the last deletion is therefore overwritten by whichever
surface happens to be deleted last, and the recreation that follows is compared
against the wrong surface's size -- or against nothing at all.

The sequence that shows it is not exotic. An ordinary resize on a Windows host
reads:

    DeleteSurface: surface=1
    DeleteSurface: surface=0
    CreateSurface: surface=0, 2000x1360

Here surface 0 is deleted last and the comparison happens to be against the
right surface. Reverse those two deletions, which nothing in the protocol
forbids and which the ordering within a batch does not guarantee, and a
same-size recreation of surface 0 is missed entirely: the record holds surface
1, the IDs do not match, and the flag never fires.

So the size lives in the surface's own slot, marked deleted rather than freed,
and survives until that ID is created again. The double deletion Windows sends
in practice still resolves: the second finds the slot already marked and leaves
the recorded size alone.

Giving up the single record also gives up the property that "recreated" meant a
creation immediately following a deletion. That was never the thing worth
bounding -- a surface that is deleted and comes back at the same size is empty
whether that took a millisecond or a minute, and the client's picture is
equally valid in both cases. The comparison is against the size, and the size
does not go stale.

The slot type is now named rather than anonymous, so the helper and the array
refer to the same type rather than two identical layouts that happen to agree.
Chrome's hardware H.264 decoder takes its reorder depth from the VUI's
bitstream_restriction when present, as zero for the High-family profiles with
constraint_set3_flag, and as the whole DPB otherwise. It has no shortcut for
pic_order_cnt_type 2, even though that POC type is the stream stating that
output order is decode order.

NVENC at its defaults writes exactly the shape that falls through: Main
profile, constraint_set1 only, POC type 2, a VUI with timing info and no
restriction. At 2992x1648 and level 5.0 the DPB is five pictures, so every
picture emerged five late -- on an idle desktop past the client's 1000ms
decode watchdog, which discards it on arrival -- and nothing was ever painted.
The session sat white or frozen on a stream that mstsc and ffmpeg decode
without delay. The tell is frames_abandoned lines in the browser console with
no decode error beside them.

The SPS rewriter now adds bitstream_restriction with max_num_reorder_frames 0
and max_dec_frame_buffering equal to max_num_ref_frames, which is what the
xrdp fork's VA-API encoder declares; the other fields are the values the
standard infers when the block is absent. Only where POC type 2 already
guarantees the answer: types 0 and 1 can reorder, and saying otherwise would
be a guess about the host. An SPS with no VUI gets an empty one carrying only
the restriction. A VUI that cannot be walked to its end costs this edit and
leaves the colour work alone.

It is decided at the first SPS beside the colour state, applied to every
keyframe's SPS, and reported once as its own edit -- the rewriter now says
which edit it made rather than that it made one, so an NVENC host, whose
colour needs nothing, is not logged as having a description spliced in. An encoder that declares
the restriction itself, as the NVENC path now does, finds nothing to do.

Checked against ffmpeg's reading of the rewritten SPS field by field, with
and without a VUI, using NVENC's and the VA-API encoder's real SPS as
fixtures.
…nnel

WebSocketTunnel closes itself on any exception thrown while an instruction
is handled, and passes on only the exception's message. client.html then
showed even that briefly: the CLOSED transition that always follows put
"Connection lost" over it. And the server saw a clean close from the
browser, indistinguishable from the user shutting the tab. So a client-side
bug that ended a session left nothing behind at either end.

The tunnel now keeps the exception, client.html logs it to the console
with its stack, and the specific message stays on the overlay.
…t costs

AVC444's cost is not the combine. VideoFrame.copyTo()'s *synchronous* half --
the driver's texture copy and staging map, before the promise exists -- is
about 10ms per call plus 5ms per megapixel, against under 2ms for the plane
uploads, the shader and the blit together. It hid because the obvious
instrument times the promise, which reads 0.0ms: by then the blocking work is
already done. The ImageBitmap handoff, the shader, software decode and GPU
bandwidth were each ruled out with a measurement rather than an argument.

So the copy reads only the damaged rows, on both views, rounded outward to 16
-- the grid the v1 chroma layout's 16-row tiling needs, and a superset of the
v2 layout's one-to-one rows and both layouts' chroma at y >> 1, so one band
serves both views. Until this the uploads and the shader were banded to the
damage while the copy read the whole frame every picture: the widest stage of
the pipeline feeding the narrowest. On a Windows host at 2992x1648 with light
typing, main-thread time inside copyTo() fell from ~77% to ~16%, and decode
behind it from 28-38ms to 1-4ms.

Several bands rather than one bounding span, because anything scattered
defeats a single span: a clock in one corner and a caret in the other span the
whole screen between them, and a desktop reliably has both. Each extra call is
a fixed stall, so minWorthwhileGap() derives from the plane width what a gap
must save to be worth splitting -- twice what it costs -- and
COPY_BAND_MAX_BANDS caps the count by closing the cheapest gaps first. At 10ms
a call that leaves most desktop damage in a single band; what it removes is
the marginal second and third copy, not the crop that pays. The test asserts
that something still splits, since a threshold change can otherwise take the
whole path out of reach while every assertion continues to pass.

The gate then follows the measured cost rather than the resolution.

COMBINE_MAX_PIXELS (4K) is a prior only -- a ceiling on what a session may
open with before anything has been measured. Keyed on framebuffer area,
deliberately: the desktop scale only says whether HiDPI scaling was applied,
so a 4K display at devicePixelRatio 1 slips past it, and the cost is not a
property of the host at all, which is why this was taken for an xrdp problem
until a Windows session was run at native resolution. Re-checked at every main
view while combining, since the framebuffer is resized after connecting --
but never between a paired main view and its auxiliary view, because the main
view is uploaded unpainted for the auxiliary one to paint and stopping in
between discards the picture, which when it is the connect-time keyframe
leaves the session looking hung until a resize. And not started at all until
the size has settled, because the connect-time fit passes through a smaller
framebuffer on its way to the final one and an auxiliary view landing inside
that window is a race a reload can change.

COMBINE_COPY_TRIP_SHARE gives up when 30% of wall clock over a busy window
goes inside copyTo(). Share alone, and not also a mean per picture: many cheap
copies is the shape that hurts and per picture is blind to it -- xrdp at
1920x1080 dragging a scrollbar ran 39-47 pictures a second at 12ms each,
46-60% of the main thread, while the per-picture figure sat under any sane
threshold and vetoed the trip. Measuring is legitimate here where it is not
for the GPU work: this is a wall-clock delta across a synchronous call on a
path already paying it, not execution needing a gl.finish() that would stall
the pipeline the gate protects. A window is never discarded for having too few
pictures, only held open until it has them, or a session below three pictures
a second never reaches a verdict.

Sync-timeout and slow-flush latches sit beneath it for the shapes the copy
share cannot see, gating on the symptom because GPU cost cannot be measured
cheaply. Measured at 2992x2000, 4:4:4 held 10% of syncs for a mean of 271ms
with 25 timeouts a minute where 4:2:0 had none in thousands; at 1920x1072
neither held at all, and 4:4:4 ran at 33-41 syncs/s with a 16-22ms mean flush
against 54.7/s and 0.6ms. The flush latch needs a busy window so that a static
desktop keeps full chroma, which is where it is worth having.

What the gate protects is input as much as frame rate. sendMouseState() runs
synchronously in the DOM handler, on the thread copyTo() blocks, and the
browser coalesces the mousemove events piling up behind it -- so the
intermediate positions of a drag are lost, not merely delayed. Across one
suspension the picture rate was unchanged, 211 in 5s against 207, while
copyTo() went from 52% of the main thread to nothing and decoded-to-painted
from 12.9ms to 0.3ms: the same decoder doing the same work, behind a thread
that was no longer blocked.

It is a latch with hysteresis rather than a controller. It gives up, waits for
quiet, tries again, and doubles the wait per trip to a cap of eight minutes,
easing it back by one doubling per five minutes of combining without a trip so
a video at lunchtime does not leave an eight-minute wait in front of an
unrelated trip that evening. There is no attempt limit: a permanent latch
condemns the rest of a session for a workload that has passed, and bounds
re-probing no better than backing off does. Probing is the only signal there
is -- a suspended session paints 4:2:0, which never times out and never
flushes slowly -- so the first window after a resume is short, 2s and 8
pictures, because the whole of a probe is spent combining at a price the
client cannot afford and ten seconds of that on sustained video is a visible
stutter on a timer. The answer is not a close one: a session that cannot
sustain the combine copies whole planes at ~42ms a picture against a 20ms
line.

An explicit h264Chroma444 override disables all of it. An override is an
instruction, and a latch that fought it would make the A/B it exists for
impossible.

sync_hold reports the holds and the display flush per chroma mode, and only
for a window containing a sync timeout unless the combine log asks for all of
them. Read the flush next to the sync rate: the hold cannot see a slow display
queue, because the sync is acked only after flush() completes, so a stuck
queue shows as fewer syncs with no holds at all.
… allows

The combine gate in H264Decoder.js gives up 4:4:4 by discarding the
auxiliary view after decoding it, so it removes main-thread cost and
nothing else: both pictures have already crossed the link and both have
been decoded. This removes the second one from the wire instead -- the
same picture for less bandwidth and one decode per frame rather than
two. Measured at 13% of the H.264 payload against a Windows host and 43%
against the xrdp fork, the difference being how often each sends chroma.
Half is the easy assumption and is wrong on both.

It matters most where AVC420 cannot be asked for. FreeRDP emits the
RDPGFX v10 capability sets only when AVC444 is requested, so a Windows
host that wants H.264 at all must be offered AVC444 -- and AVC444 is
also what puts it on its hardware encoder, so the auxiliary view cannot
be declined at the source. Native Resolution has meanwhile made what
that view carries redundant: a 4:2:0 chroma block covers one logical
pixel at a devicePixelRatio of 2. So the host is obliged to send a
picture the client does not need and cannot refuse, and dropping it in
transit is the only point in the path where refusing is possible.

Gated per stream, because the two views are one H.264 sequence sharing
one decoded picture buffer, and dropping access units is safe only if
two things hold.

Nothing surviving may predict from a dropped picture. h264_refs reads
that out of the slice headers: which view each picture belongs to,
whether the auxiliary ones are reference pictures, whether the two views
name disjoint long-term indices or reach each other through the default
reference list -- and whether an auxiliary view ever *claims* an index
main reads, which the reference lists do not show and the marking
operations do. Reading a different index is not sufficient: a dropped
picture that had marked itself with main's index would take that index's
contents with it.

And the decoder must have room for the pictures a gap makes it invent.
Dropping leaves holes in frame_num, and 8.2.5.2 obliges a decoder to
fill each with an inferred non-existing frame held as a SHORT-TERM
reference; the sliding window can only evict short-term pictures. A
stream whose max_num_ref_frames is entirely consumed by long-term
references has nowhere to put one and fails outright -- which freezes a
session whose reference chains are perfectly separate, and is the reason
asking only about references is not sufficient.

Neither question is answerable from the connect-time keyframe burst: an
IDR names no reference and marks nothing. So the thresholds count inter
slices, ten main and three auxiliary, with none on pictures -- excluding
IDRs is the stronger guard, and a count of pictures is a bad proxy for
elapsed evidence on a desktop that sends few. An auxiliary view that has
named nothing is not evidence either, since an empty set is disjoint
from everything. From the first SPS to the decision: 2.4s on Windows,
2.5s on xrdp, almost all of it the burst.

What it does: sets gaps_in_frame_num_value_allowed_flag on every SPS
from the first instruction, since an SPS only rides a keyframe and
setting it on a stream never dropped from is inert; waits; drops the
h264, blob and end instructions of each non-IDR auxiliary view once the
stream proves itself, clearing the trailing <paired> flag on main views
so the client paints them rather than holding them for a view that is no
longer coming; and keeps watching, stopping if a later slice contradicts
the verdict. A session that never gets past the waiting stage reports
what it was short of, since the gate is otherwise silent and a stream
that never qualifies would look identical to one that is merely slow.

That summary is written from Drop, and has to be. Writing it after the read
loop reaches it only when guacd closes first; a session normally ends the
other way round, with the browser going and the guacd-side future dropped
where it stands, so everything the summary carries -- what the drop saved,
what an undecided stream was short of, and the warning for leaked in-flight
streams -- would be missing in practice. Measured against a Windows host: 159
auxiliary pictures, 491 KiB, 0 in flight, no overflow.

Auxiliary IDRs are kept, and that settles more than it appears to. guacd
flags a keyframe on nal_type == 5, so its flag is IDR and every
auxiliary view carrying one passes through. A main slice naming
long-term 0 straight after an auxiliary IDR is naming that auxiliary
picture -- and gets the identical picture whether or not anything is
dropped, because it is kept, while the auxiliary views that are dropped
claim long-term 1. A veto for this costs the drop on every Windows host
and protects nothing.

An AVC420-only host never reaches a verdict and is passed through
borrowed rather than rebuilt, tested at both corroboration levels since
every such host runs this path.

It sits in guacd_to_ws after the recording tee, so recordings keep the
full 4:4:4 stream, and before the colour rewrite and the binary blob
splitter. guacd frees each stream as soon as it has written the blobs,
so nothing upstream waits on one that is swallowed here.

The client has to be told, because it cannot tell. Combining is switched on
by an auxiliary view and off by nothing in particular, and auxiliary IDRs are
deliberately kept -- so one of those arms the combiner and every main view
after it goes through the combine path, paying a plane read-back, six texture
uploads and a shader pass to produce the ordinary 4:2:0 picture drawImage()
produces almost free, waiting for a view that has been removed upstream.
Measured on a Windows session with the drop active and the gate never
contradicted: 163 pictures in 10s at 19.0ms of copying each, 31% of the main
thread, until the copy gate gave up 4:4:4 for a reason that was true and
beside the point -- saving bandwidth and a decode while adding client cost,
which is the opposite of what this is for.

So rustguac says so, as an h264-aux instruction carrying 1 or 0, sent when the
state changes and at most twice a session. Not a guacd opcode: it originates
in the proxy, and Guacamole.Client ignores opcodes it has no handler for, so a
cached client from before this behaves as it always did. Sent after the
recording tee, so recordings keep the full 4:4:4 stream and replay combining
normally. It gates the combine ahead of the h264Chroma444 override, which
everything else there defers to: the override exists so a session can be
compared against the other setting, and with the auxiliary view off the wire
there is nothing to compare -- combining can produce the cost of 4:4:4 but not
the result. tests/h264-aux-instruction.mjs pins the shape across the Rust/JS
boundary, because both ends fail silently: a wrong opcode is ignored and
leaves exactly that cost with nothing logged anywhere.

The SPS rewriter gains the gaps flag as a third edit, reported like the other
two: once a session, for what actually changed. The droppable verdict prints
the claims that decide it, naming the marking operation that assigns them, and
gives the reads afterwards as the aside they are -- on a Windows host the
reads are empty, which on the one line a person uses to judge whether a drop
is safe looks exactly like the vacuous disjointness the gate guards against
separately.

There is no per-connection setting here: every session waits for the
corroborated sample. RUSTGUAC_H264_AUX_DROP=0 is the deployment-wide kill
switch and =unproven relaxes the gate for experiments. docs/rdp-h264.md
describes it from the user's side.

Two harnesses come with it, because the failures this guards against are
found by measurement rather than by reading. tests/aux-drop-replay.mjs
strips the auxiliary views from a recording, sets the gaps flag
independently of the Rust that does it in production, decodes both
streams with ffmpeg and compares the main pictures -- a mismatch in
picture count is the cleaner signal and is checked first, since missing
pictures also break the positional alignment. An in-crate test runs the
real dropper over a real recording at several chunk sizes and checks the
result is a stream a client could follow, which is the half a picture
comparison cannot see: an h264 whose blobs never arrive hangs a browser
without corrupting anything. The SPS edit is checked against ffmpeg's
own reading of the result, and the slice-header walk against a real
libx264 stream whose expected values come from trace_headers.

What would retire it is a VideoFrame API yielding decoded planes as GPU
textures without the colour conversion. One that only adds zero-copy lands on
RGB again, and the 4:2:0 upsampling inside that transform is irreversible on a
packed auxiliary view. w3c/webcodecs#37 is the live thread.
AVC444 is always offered, and has to be: a Windows host engages its hardware
encoder only in AVC444 mode, and FreeRDP advertises the RDPGFX 10.x capability
sets only alongside it -- offered AVC420 alone it offers 8.1, and Windows at
8.1 sends no H.264 at all. So whether the chroma is used is decided on this
side, and the one question left for a user is whether they want it.

Full Colour (4:4:4), per RDP entry, off by default. Off is Standard colour: the
auxiliary view is dropped in transit wherever the stream proves it can be
spared -- from the least evidence that can answer, since an admin has vouched
for the target -- and the browser is told never to combine, so a stream the
gate refuses still paints 4:2:0 rather than paying for 4:4:4 nobody asked for.
On switches the dropper off and lets the browser combine, under the same cost
gates as before. Ad-hoc sessions have no entry and get Standard colour from
the corroborated sample.

Stored as h264_chroma444. The no-combine half reaches client.html through
SessionInfo as h264_no_combine and sets the h264Chroma444 override before the
decoder exists; the mapping onto the dropper is aux_drop_setting() in
src/session.rs, pinned by colour_setting_maps_onto_drop_and_combine.
The installer reads GUACD_LOG_LEVEL from /opt/rustguac/guacd.env, which it
creates once and never overwrites, so that is the place to raise guacd's
logging for the verification steps -- simpler than overriding ExecStart.
@pletch pletch mentioned this pull request Oct 3, 2026
Comments only, plus one log message reworded; no code changes.

The aux-drop modules opened with long notes written while the work was in
progress -- what each capture said, which theory came first, dated sessions
-- with the facts a maintainer needs spread through them, and three that
were no longer true: a never-combine setting that does not exist, an entry
option that is now Standard colour, and an auxiliary-IDR question the code
settles by never dropping one. Both headers now say what the module does,
what it reads and why, and how the hosts measured behave.

Elsewhere, dates and history are taken out of comments and the reasoning
kept: the cost model's fit, the copy-share and flush thresholds, the
recovery backoff, the black-keyframe evidence and the SPS and reference-
chain test fixtures now read as the current rules and the measurements
behind them, rather than as an account of what they replaced.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant