Conversation
HardwareVideoDecoderFactory now decodes VP9 profile 0 in hardware on macOS, through the VP9 decoder VideoToolbox offers once it is registered. H.264 stays with the default decoders, which use VideoToolbox already. The decoder creates its session on the first key frame, from the header parsed with WebRTC's VP9 header parser: the format description carries the vpcC box and the colour properties, and the output pixel format follows the range of the stream. The whole encoded image goes in as one sample, so a hidden frame is not decoded on its own. A key frame of another size or range replaces the session. Frames come out as the CVPixelBuffer VideoToolbox made, without a copy. Anything it cannot decode goes to libvpx through FallbackVideoDecoder: other profiles, sizes outside 64x64 to 4096x4096, frames with spatial layers, a session that cannot be created or is not in hardware, and a decoder that keeps failing. The macOS shortcuts in DefaultVideoCodecFactories are gone; macOS now defines the platform hooks like Windows and Linux do. Tests: VP9 through VideoToolbox, VP9 in software by default, and a resolution change in the middle of a call. A new check makes sure decoded frames carry a picture.
… size for WebRTC's VP9 header parser gives a header only for a frame that states its size: key frames, and inter frames that do not take the size from a reference. Most inter frames take it, and the parser returns nothing for them. The decoder treated that as a failure, so in a stream without temporal layers every inter frame failed, the receiver asked for a key frame, and only that key frame decoded: a few frames a second, and a key frame request for nearly every frame. A frame without a header is decoded now. A key frame without one is handed to the software decoder, since there is nothing to create a session from. The tests did not notice because they waited for ten frames and checked the picture and the size. New ones read the statistics of the receiver, that a stream asks for no more than one key frame, and cover temporal layers and spatial layers, which a real libwebrtc stream now makes: L2T2 at 640x480 ends up in libvpx. The test call can send another size than 320x240 for that, since WebRTC encodes no spatial layers below a size.
HardwareVideoDecoderFactory now decodes VP9 profile 0 on Windows, through the VP9 decoder Media Foundation has on Direct3D 11, the way it does H.264 and AV1. The factory offers VP9 where the GPU has the DXVA decoder profile (D3D11_DECODER_PROFILE_VP9_VLD_PROFILE0, with NV12 output) and Windows has a VP9 decoder that uses Direct3D 11: the VP9 Video Extensions of the Microsoft Store. Without either, VP9 stays with libvpx, and negotiation does not change. MFVideoDecoder takes VP9 as a third codec. Two things are particular to it: - A frame with spatial layers goes to libvpx. Its layers reach a decoder back to back without a superframe index; VideoToolbox does not decode that, and a Media Foundation decoder is not known to. - A VP9 stream states its size in the key frames only, so the decoder may have no output type to offer before it has seen one. That is no longer a failure to configure; the output type is set when the transform asks for it, as it is for a change of size. The VP9 tests of the decoder test class run on Windows too, not on macOS only. A VP9 decoder is required with -Dwebrtc.test.hardwareVp9Decoder=true on Windows, as an AV1 decoder is with its own property; on macOS the existing property applies. This has not been built or run on Windows.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Important
Not built or run on Windows yet. It was written on a Mac, which has no Windows compiler, by reading the AV1 path this copies. The Windows CI jobs compile it for the first time, and it needs a run on a Windows machine with a VP9 decoder: see Needs testing below.
Stacked on #318: this PR targets
main, so until #318 is merged its two commits (408c544, 6479395) show up here as well. Only 3c5776b is new; the diff of that commit is the one to review.HardwareVideoDecoderFactorynow decodes VP9 on Windows in hardware, through the VP9 decoder Media Foundation has on Direct3D 11, the way it already does H.264 and AV1. The Java API is unchanged, and so is what gets negotiated.HardwareVideoDecoderFactoryDefaultVideoDecoderFactory, all software (unchanged)Behavior
D3D11_DECODER_PROFILE_VP9_VLD_PROFILE0with NV12 output, and Windows has a VP9 decoder that uses Direct3D 11. That decoder is the VP9 Video Extensions of the Microsoft Store, which a system may not have. Without either, VP9 stays with libvpx.FallbackVideoDecoderhands the stream to libvpx when the hardware decoder fails to configure or fails while decoding, as for H.264 and AV1.decoderImplementationstat ofinbound-rtpasMediaFoundation (...).Native side
MFVideoDecodertakes VP9 as a third codec (MFVideoFormat_VP90), andMFVideoDecoderFactoryprobes it like AV1. Two things are particular to VP9, both inMFVideoDecoder:Configuregoes on, and the output type is set when the transform asks for it withMF_E_TRANSFORM_STREAM_CHANGE, whichDrainOutputalready answers for a change of size. This is a guess: whether the VP9 decoder needs it is unknown. If it does not, nothing changes; if it fails, the stream ends up in libvpx.Testing
HardwareVideoDecoderIntegrationTest: the VP9 tests are not for macOS only any more. They run on both platforms:hardwareDecodesVp9,defaultDecodesVp9InSoftware,hardwareFollowsResolutionChange,vp9NeedsNoKeyFrames,vp9TemporalLayersandvp9SpatialLayers. The names lost theirmacprefix.-Dwebrtc.test.hardwareVp9Decoder=trueon Windows, the way an AV1 decoder is with its own property; on macOS the existingwebrtc.test.hardwareDecoderapplies.-Dwebrtc.test.hardwareDecoder=true; the Windows paths are skipped there. The run was of the test classes touched; this PR's code is Windows only.Not verified yet:
Needs testing
On a Windows machine whose GPU decodes VP9 (the AMD RX 9070 XT the other PRs were tried on should, but that was not checked):
This fails unless VP9 is decoded in hardware. The log of the factory says
Media Foundation hardware decoders, H.264: … AV1: … VP9: …; a0for VP9 means the Store extension or the GPU profile is missing. Worth looking at:vp9NeedsNoKeyFrames: how many key frames the receiver asks for. A decoder that fails inter frames makes it ask for one per frame; that is what feat: decode VP9 with VideoToolbox on macOS #318 fixed on macOS.offers no output type yetin the log, which would mean the guess was needed.hardwareFollowsResolutionChange: the stream settles a new size throughMF_E_TRANSFORM_STREAM_CHANGE.Docs
The video codecs guide (
docs/guide/advanced/video-codecs.md) has VP9 in the Windows row and says what it needs. The Javadoc ofHardwareVideoDecoderFactoryhas the same.