Skip to content

feat: decode VP9 with VideoToolbox on macOS - #318

Open
devopvoid wants to merge 2 commits into
mainfrom
feat/macos-vp9-decoder
Open

devopvoid wants to merge 2 commits into
mainfrom
feat/macos-vp9-decoder

Conversation

@devopvoid

@devopvoid devopvoid commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

HardwareVideoDecoderFactory now decodes VP9 in hardware on macOS, through the VP9 decoder VideoToolbox offers. The Java API is unchanged, and so is what gets negotiated.

PeerConnectionFactory factory = PeerConnectionFactory.builder()
        .setVideoDecoderFactory(new HardwareVideoDecoderFactory())
        .build();
Platform HardwareVideoDecoderFactory
Windows H.264 and AV1 on the GPU, through Media Foundation (unchanged)
Linux Same as DefaultVideoDecoderFactory, all software (unchanged)
macOS H.264 through VideoToolbox, as before; VP9 profile 0 through VideoToolbox's VP9 decoder, on Macs that have one

Behavior

  • Availability: VideoToolbox has its VP9 decoder only after VTRegisterSupplementalVideoDecoderIfAvailable, which the factory calls once per process, followed by VTIsHardwareDecodeSupported. A Mac without a VP9 decoder gets no VP9 factory, and decodes with libvpx as before.
  • Negotiation: the hardware decoder offers VP9 profile 0, a format the software factory offers too, so negotiation does not change. Profile 2 stays with libvpx.
  • Fallback: FallbackVideoDecoder hands the stream to libvpx when:
    • the session cannot be created, or does not run in hardware (the session requires it),
    • the stream is not profile 0 / 8 bit 4:2:0, or the size is outside 64x64 to 4096x4096 (Chromium's limits),
    • a frame carries spatial layers (see below),
    • a key frame's header cannot be parsed, so there is nothing to create a session from,
    • more than 8 sessions exist in the process,
    • the decoder fails three times in a row.
      After a single failure the decoder asks for a key frame. An inter frame the header parser has no size for is decoded all the same: WebRTC's parser reports a header only for a frame that states its size, and most inter frames take it from a reference.
  • Diagnostics: the decoder in use shows in the decoderImplementation stat of inbound-rtp as VideoToolbox (VP9).

Native side

  • VTVp9Decoder creates its VTDecompressionSession on the first key frame, from the header parsed with WebRTC's ParseUncompressedVp9Header. The format description carries the vpcC box (it has to start with the 4-byte version/flags header, or VideoToolbox refuses it) and the colour properties.
    • The output pixel format follows the stream's range, VideoRange for studio and FullRange for full. VideoToolbox rescales when they differ, which would make frames differ from what libvpx gives.
    • The whole encoded image is one sample. A frame that is not shown, given on its own, comes out as a picture, so splitting a superframe along its index gives extra frames.
    • A key frame of another size or range replaces the session. VideoToolbox does not take a new format in a running session (kVTFormatDescriptionChangeNotSupportedErr); VTDecompressionSessionCanAcceptFormatDescription is asked first, so a change of colour space alone may stay.
    • Decoding is synchronous, one frame in and one out, which suits WebRTC's decoder thread.
    • Frames go on as the CVPixelBuffer VideoToolbox made, wrapped in ObjCFrameBuffer, without a copy. The Java sink converts to I420 where it needs it, as it does for the H.264 decoder.
  • VTVideoDecoderFactory offers VP9 profile 0 when the probe says the Mac decodes it.
  • MacHardwareVideoCodecFactories defines the platform hooks for macOS, as Windows and Linux do. The #ifdef __APPLE__ shortcuts in DefaultVideoCodecFactories.cpp are gone. No encoders: VideoToolbox has none for VP9 or AV1.

Spatial layers

A frame with spatial layers reaches a decoder with its layers back to back and no superframe index. In an experiment, VideoToolbox decoded a key picture with the index, and refused it without. The decoder therefore returns WEBRTC_VIDEO_CODEC_FALLBACK_SOFTWARE for a frame with spatial layers, and libvpx takes the stream, as it does in Chrome by default. Temporal layers alone are fine.

This is checked with streams libwebrtc makes itself, through scalabilityMode: at 640x480 L2T2, L3T3 and S2T1 end up in libvpx and keep playing, while L1T2 and L1T3 stay in VideoToolbox. WebRTC encodes spatial layers only above a certain size, so the test call can send other sizes than 320x240 now. A stream from a browser or an SFU is still to be tried.

Measurements

Done before the implementation, on an Apple M2 with libvpx's own encoder, 8 bit 4:2:0, with synthetic content and without the later NV12 to I420 conversion:

Size VideoToolbox, process CPU per frame libvpx 1 thread libvpx 4 threads
1080p 0.18 ms 5.56 ms 5.98 ms
720p 0.12 ms 2.62 ms 2.94 ms

The gain is CPU. On the clock, multi-threaded libvpx is about as fast (1080p: 1.93 ms against 1.72 ms), so this is not a latency improvement. The guide says "less CPU", not "faster".

Testing

  • HardwareVideoDecoderIntegrationTest, six new tests on macOS:
    • macDecodesVp9WithVideoToolbox: a call in VP9, the receiver reports VideoToolbox. Skipped without a hardware decoder, fails with -Dwebrtc.test.hardwareDecoder=true.
    • macDecodesVp9InSoftwareByDefault: the shared factory does not use VideoToolbox.
    • macFollowsResolutionChange: the sender switches to half the size in the middle of a call (scaleResolutionDownBy), frames arrive at the new size, and the decoder is still VideoToolbox.
    • macVp9NeedsNoKeyFrames: the receiver asks for at most one key frame, and decodes at most two. A decoder that fails inter frames makes it ask for one per frame, and only that key frame decodes; this is what the first version of the decoder did, and the other tests did not notice.
    • macVp9TemporalLayers: L1T3 stays in VideoToolbox.
    • macVp9SpatialLayers: L2T2 at 640x480 is decoded by libvpx.
  • Decoded frames are now checked to carry a picture (the luma ramp of the test source), not only to have the right size.
  • On an Apple M2 (macOS 14.5), with -Dwebrtc.test.hardwareDecoder=true: mvn -pl webrtc test 223 tests pass, -Pjni-check 223 tests pass, with no FATAL ERROR in native method.

Not verified yet:

  • Other Macs: an M1, an M3 and Intel Macs, where VP9 hardware decoding depends on the GPU.
  • The x86_64 target, which CI cross compiles; only arm64 was built here.
  • A stream with spatial layers from a browser or an SFU (libwebrtc's own is covered above).
  • The switch to libvpx in the middle of a stream (it cannot be forced from a test).
  • Long calls with packet loss.
  • Whether the two numbers that are guesses fit: the session limit of 8, and the size limits taken from Chromium.
  • CI runners are virtualized and likely have no VP9 decoder, so macDecodesVp9WithVideoToolbox is probably skipped there.

Needs testing

On other Macs, with the decoder required:

mvn -pl webrtc test -Dtest=HardwareVideoDecoderIntegrationTest -Dwebrtc.test.hardwareDecoder=true

This fails unless VP9 is decoded in hardware. The decoderImplementation stat of a call should read VideoToolbox (VP9). Worth checking, too:

  • a call with a browser that sends VP9, ideally one with spatial layers (Chrome with a scalability mode such as L2T1),
  • CPU use of a long 1080p receive, hardware against libvpx.

Docs

The video codecs guide (docs/guide/advanced/video-codecs.md) has the new macOS row and says what the hardware decoder does and does not take. The Javadoc of HardwareVideoDecoderFactory has the same.

HardwareVideoDecoderFactory now decodes VP9 profile 0 in hardware on
macOS, through the VP9 decoder VideoToolbox offers once it is registered.
H.264 stays with the default decoders, which use VideoToolbox already.

The decoder creates its session on the first key frame, from the header
parsed with WebRTC's VP9 header parser: the format description carries the
vpcC box and the colour properties, and the output pixel format follows the
range of the stream. The whole encoded image goes in as one sample, so a
hidden frame is not decoded on its own. A key frame of another size or
range replaces the session. Frames come out as the CVPixelBuffer
VideoToolbox made, without a copy.

Anything it cannot decode goes to libvpx through FallbackVideoDecoder: other
profiles, sizes outside 64x64 to 4096x4096, frames with spatial layers, a
session that cannot be created or is not in hardware, and a decoder that
keeps failing.

The macOS shortcuts in DefaultVideoCodecFactories are gone; macOS now
defines the platform hooks like Windows and Linux do.

Tests: VP9 through VideoToolbox, VP9 in software by default, and a
resolution change in the middle of a call. A new check makes sure decoded
frames carry a picture.
… size for

WebRTC's VP9 header parser gives a header only for a frame that states its
size: key frames, and inter frames that do not take the size from a
reference. Most inter frames take it, and the parser returns nothing for
them. The decoder treated that as a failure, so in a stream without temporal
layers every inter frame failed, the receiver asked for a key frame, and
only that key frame decoded: a few frames a second, and a key frame
request for nearly every frame.

A frame without a header is decoded now. A key frame without one is handed
to the software decoder, since there is nothing to create a session from.

The tests did not notice because they waited for ten frames and checked the
picture and the size. New ones read the statistics of the receiver, that a
stream asks for no more than one key frame, and cover temporal layers and
spatial layers, which a real libwebrtc stream now makes: L2T2 at 640x480
ends up in libvpx. The test call can send another size than 320x240 for
that, since WebRTC encodes no spatial layers below a size.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant