Conversation
HardwareVideoDecoderFactory now decodes on Linux where there is an NVIDIA GPU: H.264 in all its profiles and VP9 in profile 0, through the NVDEC engine. It is the decoder counterpart of the NVENC encoder, and loads the same way: libcuda.so.1 and libnvcuvid.so.1 are opened at run time, so a machine without the NVIDIA driver is not affected, and falls back to the software decoders as before. What gets negotiated does not change. NvdecVideoDecoder uses NVDEC's own parser, which finds the pictures in an encoded image and calls back for the sequence (which creates the decoder), each picture to decode, and each to show. With no display delay all of that happens inside Decode. A decoded picture is NV12 in GPU memory; it is copied to system memory, cropped to the picture within the surface, and converted to I420 with libyuv, as the Media Foundation decoder does on Windows. Whatever NVDEC cannot take returns WEBRTC_VIDEO_CODEC_FALLBACK_SOFTWARE and FallbackVideoDecoder hands the stream to the software decoder: another profile or bit depth, a size the GPU does not decode, a failure of the decoder, and VP9 frames with spatial layers, whose layers reach a decoder without the index that says where each ends. A first frame that is not a key frame asks WebRTC for one. The headers are the three of nv-codec-headers (tag n12.0.16.1, the version of the NVENC header) that NVDEC needs, unchanged, in dependencies/nvdec with their licenses; they are compile-time only. They are the same files webrtc-java-media carries for FFmpeg. The decoder tests of the webrtc module run on Linux too: H.264 and VP9 are checked there as on Windows, and the implementation name NVDEC counts as hardware. This has not been built or run on Linux, and never on an NVIDIA GPU. The sources compile with clang against WebRTC's headers on a Mac (syntax only, with Linux defines); how NVDEC behaves on real streams could not be seen.
devopvoid
force-pushed
the
feat/linux-nvdec-decoder
branch
from
October 4, 2026 09:42
7ef5a99 to
d84ba77
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Important
Not built or run on Linux, and never on an NVIDIA GPU. It was written on a Mac. The sources compile against WebRTC's headers there (syntax only, with Linux defines); the Linux CI jobs compile and link it for the first time, and it needs a run on a machine with an NVIDIA GPU: see Needs testing below.
HardwareVideoDecoderFactorynow decodes on Linux where there is an NVIDIA GPU: H.264 in all its profiles and VP9 in profile 0, through the NVDEC engine. It is the decoder counterpart of the NVENC encoder (#312), and loads the same way. The Java API is unchanged, and so is what gets negotiated.HardwareVideoDecoderFactoryBehavior
libcuda.so.1andlibnvcuvid.so.1are opened withdlopenwhen the factory is made, so a machine without the NVIDIA driver is not affected and decodes in software as before. The factory is offered only where the first CUDA device decodes at least one of the codecs (cuvidGetDecoderCaps, 8 bit 4:2:0 into NV12).FallbackVideoDecoderhands the stream to the software decoder when NVDEC returnsWEBRTC_VIDEO_CODEC_FALLBACK_SOFTWARE. That happens for another profile or bit depth, a size the GPU does not decode (checked against the caps of the device), a decoder that fails, and VP9 frames with spatial layers: their layers reach a decoder without the superframe index that says where each ends, and NVDEC is not known to take that, where libvpx does. A first frame that is not a key frame asks WebRTC for one.decoderImplementationstat ofinbound-rtpasNVDEC (<GPU name>).Native side
NvdecVideoDecoderuses NVDEC's own parser, which finds the pictures in an encoded image and calls back for the sequence (which creates the decoder), each picture to decode, and each to show. With no display delay all of that happens insideDecode, on the caller's thread. A decoded picture is NV12 in GPU memory; it is copied to system memory withcuMemcpy2D, cropped to the picture within the surface (1088 rows for 1080), and converted to I420 with libyuv, asMFVideoDecoderdoes.A frame is matched to its input by a timestamp counted per image, so the decoder does not depend on frames coming out in input order. A resolution change at a key frame makes the decoder again.
NvdecLibraryloads the two libraries once, probes the device, and keeps the primary CUDA context, the wayNvencLibrarydoes.NvdecContextScopemakes it current for the calls.NvdecVideoDecoderFactoryand the Linux platform hook (LinuxHardwareVideoCodecFactories.cpp), which had no decoders.dependencies/nvdechasdynlink_cuda.h,dynlink_cuviddec.handdynlink_nvcuvid.hofnv-codec-headersat tagn12.0.16.1, unchanged: the version of the NVENC header, and the same fileswebrtc-java-mediacarries for FFmpeg (feat: decode video in hardware in the media player #320). They define the structures NVDEC's parser and decoder take, which would be error-prone to declare by hand. They are compile-time only. Their licenses are gathered inLICENSEand installed into the Linux platform jars underMETA-INF/licenses/nvdec.Testing
HardwareVideoDecoderIntegrationTestruns on Linux too: the H.264 test and the VP9 tests (hardwareDecodesVp9,hardwareFollowsResolutionChange,vp9NeedsNoKeyFrames,vp9TemporalLayers,vp9SpatialLayers) check there what they check on Windows, and the implementation nameNVDECcounts as hardware. A decoder is required with-Dwebrtc.test.hardwareDecoder=true(H.264) and-Dwebrtc.test.hardwareVp9Decoder=true(VP9); without them the tests are skipped where there is no decoder.clang++ -std=c++20 -Wall -Wextra -fsyntax-onlyagainst the project's WebRTC headers with Linux defines, with a negative control to show that the check reports errors. The only warning is the unusedenvparameter that the NVENC factory has too.Not verified yet:
pitch x aligned heightand cropped on the copy, following NVIDIA's sample; it could not be checked.cuvidCtxLockis not used.Needs testing
On Linux with an NVIDIA GPU and its driver (
libcuda.so.1andlibnvcuvid.so.1present):This fails unless H.264 and VP9 are decoded in hardware. The
decoderImplementationstat of a call should readNVDEC (<GPU name>). The log lineNVDEC available on <GPU>, H.264: .., VP9: ..tells what the device offers. Worth looking at:vp9NeedsNoKeyFramesandhardwareFollowsResolutionChange: how many key frames the receiver asks for, and whether a new size makes the decoder again.Docs
The video codecs guide (
docs/guide/advanced/video-codecs.md) has NVDEC in the Linux row and says what it needs. The Javadoc ofHardwareVideoDecoderFactoryhas the same.