Skip to content

feat: encode H.264 with VA-API on Intel and AMD GPUs on Linux - #313

Merged
devopvoid merged 1 commit into
feat/nvenc-encoderfrom
feat/vaapi-encoder
Oct 1, 2026
Merged

devopvoid merged 1 commit into
feat/nvenc-encoderfrom
feat/vaapi-encoder

Conversation

@devopvoid

Copy link
Copy Markdown
Owner

Stacked on #312 (then #310, #303): this PR targets its branch and should be merged after it.

Important

Not yet run on hardware. It was developed without a VA-API device: the WSL environment available has Mesa's d3d12 VA driver, but no /dev/dri render node, and the driver fails to initialize through WSLg's display as well. The code has only been syntax-checked. With VA-API, far more of the encoder's behavior is ours than with NVENC or Media Foundation. See Needs testing below.

On Linux, HardwareVideoEncoderFactory now uses the VA-API encoder of the GPU driver after NVENC: Intel's iHD driver, and Mesa's radeonsi for AMD. The order on Linux becomes NVENC, then VA-API, then OpenH264. The Java API is unchanged.

Loading

  • libva.so.2 and libva-drm.so.2 are loaded at run time. Machines without libva, or without a driver that can encode, are unaffected.
  • The build needs only the libva headers, vendored from libva 2.17.0 (MIT). They're unchanged except va_version.h, which the libva build generates. dependencies/libva/README.md records the details. This makes the cross-compiled ARM targets independent of what their sysroots contain. The license goes into the platform jar under META-INF/licenses/libva/.
  • The encoder uses the first /dev/dri/renderD* node whose driver can encode H.264 Constrained Baseline with CBR. It takes the full encode entrypoint where there is one, otherwise the low-power one. The display is shared by all encoders.

Encoder

VaapiH264Encoder encodes synchronously on the encoder thread:

  • one reference frame and P-frames only, with the reconstructed surfaces swapped each frame
  • picture order count type 2, frame_num counting modulo 256, and idr_pic_id incremented per IDR frame
  • CAVLC, with deblocking on
  • IDR frames only when WebRTC requests a key frame
  • frame cropping for sizes that aren't a multiple of 16, and a level chosen from frame size and macroblock rate
  • rate control, frame rate and HRD parameters (CBR, 1.5 s buffer), sent again when the rates change
  • frames written as NV12, directly into the surface where the driver allows vaDeriveImage, otherwise through vaCreateImage/vaPutImage

Headers: the driver writes the parameter sets and slice headers; no packed headers are passed. If a key frame comes out without both SPS and PPS (as it would from a driver that expects packed headers), the encoder gives up, so the stream falls back to software instead of sending something nothing can decode.

Testing

  • The VA-API, Linux platform, NVENC and hardware-factory sources pass g++ -std=c++20 -Wall -Wextra -fsyntax-only in WSL (Debian 12) against the vendored headers. The only warning comes from WebRTC's own ScalingSettings class.
  • Windows is unaffected, since none of this code builds there. HardwareVideoEncoderIntegrationTest and VideoCodecFactoryTests still pass.
  • CI (clang) will be the first full Linux build.
  • HardwareVideoEncoderIntegrationTest now also accepts VA-API (...) as a hardware encoder.

Needs testing

On an Intel iGPU and an AMD GPU, on Linux:

mvn -pl webrtc test -Dtest=HardwareVideoEncoderIntegrationTest -Dwebrtc.test.hardwareEncoder=true

This fails unless H.264 is actually encoded on the GPU. A call's encoderImplementation stat should read VA-API (<driver>). Most likely to need fixing:

  • Missing SPS/PPS: a driver that doesn't write the parameter sets itself, most likely newer iHD versions with the low-power entrypoint, would make every stream fall back to OpenH264. The fix would be to write SPS/PPS as packed headers ourselves.
  • Strict drivers: some drivers are strict about the reference and rate-control parameters. Image quality, and whether the bitrate follows the target over a longer call, are worth checking.

HardwareVideoEncoderFactory now uses the VA-API encoder of the GPU driver
on Linux, after NVENC: Intel's iHD driver and Mesa's radeonsi for AMD. libva
is loaded at run time, so building needs only its headers, vendored from
libva 2.17.0 (MIT, its license goes into the platform jar), and machines
without libva or an encoding driver are unaffected. The first DRM render
node whose driver encodes H.264 Constrained Baseline with CBR is used,
with either the full or the low-power encode entrypoint.

The encoder is synchronous, with one reference frame, P-frames only,
picture order count type 2 and CAVLC; it fills the sequence, picture and
slice parameters and the rate control, frame rate and HRD parameters, and
leaves writing the parameter sets and slice headers to the driver. A key
frame without SPS and PPS, which a driver that expects packed headers
would produce, makes the encoder give up, so that the software encoder
takes over rather than sending a stream nothing can decode. Frames are
written into the surface as NV12, derived where the driver allows it.
@devopvoid
devopvoid added this pull request to stack #311 October 1, 2026 22:04
@devopvoid
devopvoid merged commit 9e0b06e into main Oct 1, 2026
17 checks passed
@devopvoid
devopvoid deleted the feat/vaapi-encoder branch October 1, 2026 22:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant