Conversation
HardwareVideoDecoderFactory now decodes VP9 profile 0 in hardware on macOS, through the VP9 decoder VideoToolbox offers once it is registered. H.264 stays with the default decoders, which use VideoToolbox already. The decoder creates its session on the first key frame, from the header parsed with WebRTC's VP9 header parser: the format description carries the vpcC box and the colour properties, and the output pixel format follows the range of the stream. The whole encoded image goes in as one sample, so a hidden frame is not decoded on its own. A key frame of another size or range replaces the session. Frames come out as the CVPixelBuffer VideoToolbox made, without a copy. Anything it cannot decode goes to libvpx through FallbackVideoDecoder: other profiles, sizes outside 64x64 to 4096x4096, frames with spatial layers, a session that cannot be created or is not in hardware, and a decoder that keeps failing. The macOS shortcuts in DefaultVideoCodecFactories are gone; macOS now defines the platform hooks like Windows and Linux do. Tests: VP9 through VideoToolbox, VP9 in software by default, and a resolution change in the middle of a call. A new check makes sure decoded frames carry a picture.
… size for WebRTC's VP9 header parser gives a header only for a frame that states its size: key frames, and inter frames that do not take the size from a reference. Most inter frames take it, and the parser returns nothing for them. The decoder treated that as a failure, so in a stream without temporal layers every inter frame failed, the receiver asked for a key frame, and only that key frame decoded: a few frames a second, and a key frame request for nearly every frame. A frame without a header is decoded now. A key frame without one is handed to the software decoder, since there is nothing to create a session from. The tests did not notice because they waited for ten frames and checked the picture and the size. New ones read the statistics of the receiver, that a stream asks for no more than one key frame, and cover temporal layers and spatial layers, which a real libwebrtc stream now makes: L2T2 at 640x480 ends up in libvpx. The test call can send another size than 320x240 for that, since WebRTC encodes no spatial layers below a size.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
HardwareVideoDecoderFactorynow decodes VP9 in hardware on macOS, through the VP9 decoder VideoToolbox offers. The Java API is unchanged, and so is what gets negotiated.HardwareVideoDecoderFactoryDefaultVideoDecoderFactory, all software (unchanged)Behavior
VTRegisterSupplementalVideoDecoderIfAvailable, which the factory calls once per process, followed byVTIsHardwareDecodeSupported. A Mac without a VP9 decoder gets no VP9 factory, and decodes with libvpx as before.FallbackVideoDecoderhands the stream to libvpx when:After a single failure the decoder asks for a key frame. An inter frame the header parser has no size for is decoded all the same: WebRTC's parser reports a header only for a frame that states its size, and most inter frames take it from a reference.
decoderImplementationstat ofinbound-rtpasVideoToolbox (VP9).Native side
VTVp9Decodercreates itsVTDecompressionSessionon the first key frame, from the header parsed with WebRTC'sParseUncompressedVp9Header. The format description carries thevpcCbox (it has to start with the 4-byte version/flags header, or VideoToolbox refuses it) and the colour properties.VideoRangefor studio andFullRangefor full. VideoToolbox rescales when they differ, which would make frames differ from what libvpx gives.kVTFormatDescriptionChangeNotSupportedErr);VTDecompressionSessionCanAcceptFormatDescriptionis asked first, so a change of colour space alone may stay.CVPixelBufferVideoToolbox made, wrapped inObjCFrameBuffer, without a copy. The Java sink converts to I420 where it needs it, as it does for the H.264 decoder.VTVideoDecoderFactoryoffers VP9 profile 0 when the probe says the Mac decodes it.MacHardwareVideoCodecFactoriesdefines the platform hooks for macOS, as Windows and Linux do. The#ifdef __APPLE__shortcuts inDefaultVideoCodecFactories.cppare gone. No encoders: VideoToolbox has none for VP9 or AV1.Spatial layers
A frame with spatial layers reaches a decoder with its layers back to back and no superframe index. In an experiment, VideoToolbox decoded a key picture with the index, and refused it without. The decoder therefore returns
WEBRTC_VIDEO_CODEC_FALLBACK_SOFTWAREfor a frame with spatial layers, and libvpx takes the stream, as it does in Chrome by default. Temporal layers alone are fine.This is checked with streams libwebrtc makes itself, through
scalabilityMode: at 640x480L2T2,L3T3andS2T1end up in libvpx and keep playing, whileL1T2andL1T3stay in VideoToolbox. WebRTC encodes spatial layers only above a certain size, so the test call can send other sizes than 320x240 now. A stream from a browser or an SFU is still to be tried.Measurements
Done before the implementation, on an Apple M2 with libvpx's own encoder, 8 bit 4:2:0, with synthetic content and without the later NV12 to I420 conversion:
The gain is CPU. On the clock, multi-threaded libvpx is about as fast (1080p: 1.93 ms against 1.72 ms), so this is not a latency improvement. The guide says "less CPU", not "faster".
Testing
HardwareVideoDecoderIntegrationTest, six new tests on macOS:macDecodesVp9WithVideoToolbox: a call in VP9, the receiver reportsVideoToolbox. Skipped without a hardware decoder, fails with-Dwebrtc.test.hardwareDecoder=true.macDecodesVp9InSoftwareByDefault: the shared factory does not use VideoToolbox.macFollowsResolutionChange: the sender switches to half the size in the middle of a call (scaleResolutionDownBy), frames arrive at the new size, and the decoder is still VideoToolbox.macVp9NeedsNoKeyFrames: the receiver asks for at most one key frame, and decodes at most two. A decoder that fails inter frames makes it ask for one per frame, and only that key frame decodes; this is what the first version of the decoder did, and the other tests did not notice.macVp9TemporalLayers:L1T3stays in VideoToolbox.macVp9SpatialLayers:L2T2at 640x480 is decoded by libvpx.-Dwebrtc.test.hardwareDecoder=true:mvn -pl webrtc test223 tests pass,-Pjni-check223 tests pass, with noFATAL ERROR in native method.Not verified yet:
macDecodesVp9WithVideoToolboxis probably skipped there.Needs testing
On other Macs, with the decoder required:
This fails unless VP9 is decoded in hardware. The
decoderImplementationstat of a call should readVideoToolbox (VP9). Worth checking, too:L2T1),Docs
The video codecs guide (
docs/guide/advanced/video-codecs.md) has the new macOS row and says what the hardware decoder does and does not take. The Javadoc ofHardwareVideoDecoderFactoryhas the same.