Skip to content

feat: decode VP9 on the GPU with Media Foundation on Windows - #319

Open
devopvoid wants to merge 3 commits into
mainfrom
feat/windows-vp9-decoder
Open

devopvoid wants to merge 3 commits into
mainfrom
feat/windows-vp9-decoder

Conversation

@devopvoid

Copy link
Copy Markdown
Owner

Important

Not built or run on Windows yet. It was written on a Mac, which has no Windows compiler, by reading the AV1 path this copies. The Windows CI jobs compile it for the first time, and it needs a run on a Windows machine with a VP9 decoder: see Needs testing below.

Stacked on #318: this PR targets main, so until #318 is merged its two commits (408c544, 6479395) show up here as well. Only 3c5776b is new; the diff of that commit is the one to review.

HardwareVideoDecoderFactory now decodes VP9 on Windows in hardware, through the VP9 decoder Media Foundation has on Direct3D 11, the way it already does H.264 and AV1. The Java API is unchanged, and so is what gets negotiated.

Platform HardwareVideoDecoderFactory
Windows H.264, AV1 and VP9 (profile 0) on the GPU, through Media Foundation on Direct3D 11 (DXVA)
Linux Same as DefaultVideoDecoderFactory, all software (unchanged)
macOS H.264 and VP9 through VideoToolbox (#318)

Behavior

  • Availability: the factory offers VP9 where the GPU has the DXVA decoder profile D3D11_DECODER_PROFILE_VP9_VLD_PROFILE0 with NV12 output, and Windows has a VP9 decoder that uses Direct3D 11. That decoder is the VP9 Video Extensions of the Microsoft Store, which a system may not have. Without either, VP9 stays with libvpx.
  • Negotiation: VP9 profile 0, a format the software factory offers too, so negotiation does not change.
  • Fallback: FallbackVideoDecoder hands the stream to libvpx when the hardware decoder fails to configure or fails while decoding, as for H.264 and AV1.
  • Diagnostics: the decoder shows in the decoderImplementation stat of inbound-rtp as MediaFoundation (...).

Native side

MFVideoDecoder takes VP9 as a third codec (MFVideoFormat_VP90), and MFVideoDecoderFactory probes it like AV1. Two things are particular to VP9, both in MFVideoDecoder:

  • A frame with spatial layers goes to libvpx. The layers of such a frame reach a decoder back to back, without the superframe index that says where one ends. VideoToolbox does not decode that (see feat: decode VP9 with VideoToolbox on macOS #318); a Media Foundation decoder is not known to either, and libvpx does, so the safe choice is the same.
  • No output type is not a failure to configure. A VP9 stream states its size in the key frames only, so the decoder may offer no output type before it has seen one. For VP9 without a known size, Configure goes on, and the output type is set when the transform asks for it with MF_E_TRANSFORM_STREAM_CHANGE, which DrainOutput already answers for a change of size. This is a guess: whether the VP9 decoder needs it is unknown. If it does not, nothing changes; if it fails, the stream ends up in libvpx.

Testing

  • HardwareVideoDecoderIntegrationTest: the VP9 tests are not for macOS only any more. They run on both platforms: hardwareDecodesVp9, defaultDecodesVp9InSoftware, hardwareFollowsResolutionChange, vp9NeedsNoKeyFrames, vp9TemporalLayers and vp9SpatialLayers. The names lost their mac prefix.
  • A VP9 decoder is required with -Dwebrtc.test.hardwareVp9Decoder=true on Windows, the way an AV1 decoder is with its own property; on macOS the existing webrtc.test.hardwareDecoder applies.
  • On an Apple M2 (macOS 14.5) the tests run as before, with -Dwebrtc.test.hardwareDecoder=true; the Windows paths are skipped there. The run was of the test classes touched; this PR's code is Windows only.

Not verified yet:

  • Everything on Windows: the build, the factory on a GPU with and without the extension, decoding, the fallback.
  • The guess above.
  • Whether the Media Foundation decoder takes frames without a superframe index; spatial layers go to libvpx in any case.
  • ARM64 Windows, which CI builds but cannot test with a GPU decoder.

Needs testing

On a Windows machine whose GPU decodes VP9 (the AMD RX 9070 XT the other PRs were tried on should, but that was not checked):

mvn -pl webrtc test -Dtest=HardwareVideoDecoderIntegrationTest -Dwebrtc.test.hardwareDecoder=true -Dwebrtc.test.hardwareVp9Decoder=true

This fails unless VP9 is decoded in hardware. The log of the factory says Media Foundation hardware decoders, H.264: … AV1: … VP9: …; a 0 for VP9 means the Store extension or the GPU profile is missing. Worth looking at:

  • vp9NeedsNoKeyFrames: how many key frames the receiver asks for. A decoder that fails inter frames makes it ask for one per frame; that is what feat: decode VP9 with VideoToolbox on macOS #318 fixed on macOS.
  • offers no output type yet in the log, which would mean the guess was needed.
  • hardwareFollowsResolutionChange: the stream settles a new size through MF_E_TRANSFORM_STREAM_CHANGE.

Docs

The video codecs guide (docs/guide/advanced/video-codecs.md) has VP9 in the Windows row and says what it needs. The Javadoc of HardwareVideoDecoderFactory has the same.

HardwareVideoDecoderFactory now decodes VP9 profile 0 in hardware on
macOS, through the VP9 decoder VideoToolbox offers once it is registered.
H.264 stays with the default decoders, which use VideoToolbox already.

The decoder creates its session on the first key frame, from the header
parsed with WebRTC's VP9 header parser: the format description carries the
vpcC box and the colour properties, and the output pixel format follows the
range of the stream. The whole encoded image goes in as one sample, so a
hidden frame is not decoded on its own. A key frame of another size or
range replaces the session. Frames come out as the CVPixelBuffer
VideoToolbox made, without a copy.

Anything it cannot decode goes to libvpx through FallbackVideoDecoder: other
profiles, sizes outside 64x64 to 4096x4096, frames with spatial layers, a
session that cannot be created or is not in hardware, and a decoder that
keeps failing.

The macOS shortcuts in DefaultVideoCodecFactories are gone; macOS now
defines the platform hooks like Windows and Linux do.

Tests: VP9 through VideoToolbox, VP9 in software by default, and a
resolution change in the middle of a call. A new check makes sure decoded
frames carry a picture.
… size for

WebRTC's VP9 header parser gives a header only for a frame that states its
size: key frames, and inter frames that do not take the size from a
reference. Most inter frames take it, and the parser returns nothing for
them. The decoder treated that as a failure, so in a stream without temporal
layers every inter frame failed, the receiver asked for a key frame, and
only that key frame decoded: a few frames a second, and a key frame
request for nearly every frame.

A frame without a header is decoded now. A key frame without one is handed
to the software decoder, since there is nothing to create a session from.

The tests did not notice because they waited for ten frames and checked the
picture and the size. New ones read the statistics of the receiver, that a
stream asks for no more than one key frame, and cover temporal layers and
spatial layers, which a real libwebrtc stream now makes: L2T2 at 640x480
ends up in libvpx. The test call can send another size than 320x240 for
that, since WebRTC encodes no spatial layers below a size.
HardwareVideoDecoderFactory now decodes VP9 profile 0 on Windows, through
the VP9 decoder Media Foundation has on Direct3D 11, the way it does H.264
and AV1. The factory offers VP9 where the GPU has the DXVA decoder profile
(D3D11_DECODER_PROFILE_VP9_VLD_PROFILE0, with NV12 output) and Windows has
a VP9 decoder that uses Direct3D 11: the VP9 Video Extensions of the
Microsoft Store. Without either, VP9 stays with libvpx, and negotiation
does not change.

MFVideoDecoder takes VP9 as a third codec. Two things are particular to it:

- A frame with spatial layers goes to libvpx. Its layers reach a decoder
  back to back without a superframe index; VideoToolbox does not decode
  that, and a Media Foundation decoder is not known to.
- A VP9 stream states its size in the key frames only, so the decoder may
  have no output type to offer before it has seen one. That is no longer a
  failure to configure; the output type is set when the transform asks for
  it, as it is for a change of size.

The VP9 tests of the decoder test class run on Windows too, not on macOS
only. A VP9 decoder is required with -Dwebrtc.test.hardwareVp9Decoder=true
on Windows, as an AV1 decoder is with its own property; on macOS the
existing property applies.

This has not been built or run on Windows.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant