Skip to content

feat: add encoded frame transforms and a media recorder - #300

Merged
devopvoid merged 3 commits into
mainfrom
feat/encoded-frame-transforms
Sep 27, 2026
Merged

devopvoid merged 3 commits into
mainfrom
feat/encoded-frame-transforms

Conversation

@devopvoid

Copy link
Copy Markdown
Owner

Summary

Two features that build on each other:

Encoded frame transforms (insertable streams). RTCRtpSender and RTCRtpReceiver get setTransform(RTCEncodedFrameTransformer), which sees every encoded frame between encoder and packetizer, or depacketizer and decoder, and may read, replace or drop its payload. Frames expose their metadata (RTCEncodedVideoFrame: key frame, size, RID, layers; RTCEncodedAudioFrame: sequence number, audio level, CSRCs). Senders gain generateKeyFrame(), receivers requestKeyFrame().

MediaRecorder (media module). Records senders and receivers into .mkv, .webm or .mp4 without re-encoding: encoded frames go into the file as they are.

Design

  • Transforms never run on media threads. Each sender/receiver with a transform gets its own worker thread (attached as a JVM daemon, never joined), so a slow transform cannot stall the call, and a transform may call back into the peer connection, even close it, without deadlocking. The queue is bounded; a transform that falls seconds behind has frames dropped.
  • No copy unless asked for. The payload is copied into Java only when getData() is called, into a per-thread direct buffer reused across frames. A frame used after its transform returned, or from another thread, throws IllegalStateException; nothing can reach a native frame that has moved on.
  • Fail closed. A transform that throws drops the frame, so a broken encryption transform never sends plaintext.
  • One native transformer per sender/receiver, installed lazily and kept, found through a weak, self-cleaning registry (hand-rolled ref counting so a lookup can never resurrect a dying transformer). The Java transform and native observers share it, and a video sender is not restarted on every change. Senders/receivers without a transform are untouched.
  • Recorder frames stay native. The extension API (webrtc_java_api.h) gains encoded frame observers (appended members, same version). The recorder copies frames on the WebRTC thread into a bounded queue and muxes on its own thread. It starts video at a key frame and requests one (throttled), writes the header once all tracks are described or after 3 s, maps RTP timestamps onto a common timeline, and uses fragmented MP4 so a cut-short file still plays. A codec the container cannot hold (e.g. VP8 in MP4) leaves that track out with a warning instead of failing the file.
  • FFmpeg is additionally built with the matroska/webm/mp4/mov muxers, extract_extradata and the AV1 parser. The CI FFmpeg cache key already hashes the component list, so no workflow change is needed.

Also in this PR

  • EncryptedRecordingExample: AES-GCM end-to-end encryption through transforms (codec header kept in the clear and authenticated), with the decrypted side recorded.
  • Guides: Encoded Transforms and Media Recording; README and docs landing page updated.
  • The media module's API version check moved to a shared ApiCheck helper used by the player and the recorder.

Testing

  • webrtc: full suite 184/184, including 9 new transform tests (encrypted round trip, metadata, stale-frame access, drop and clear, exceptions, audio, key frame requests, shared transforms). Clean under -Pjni-check (no FATAL ERROR in native method).
  • webrtc-java-media: full suite 47/47, including 7 new recorder tests that read the recording back (codec, size, duration), cover the left-out-codec path, empty-file deletion and stopping from the listener.
  • Ran EncryptedRecordingExample end to end: 808 frames encrypted/decrypted, 0 rejected, 10 s recording written.
  • Tested on Windows x86_64 only.

Encoded frame transforms (insertable streams): RTCRtpSender and
RTCRtpReceiver get setTransform(RTCEncodedFrameTransformer), which sees
every encoded frame between encoder and packetizer, or depacketizer and
decoder, and may read, replace or drop its payload. The transform runs on
a thread of its own per sender or receiver, never on a media thread, and
frames reach Java without a copy unless the payload is read. Senders gain
generateKeyFrame(), receivers requestKeyFrame().

One native transformer is installed per sender or receiver, found through
a weak registry, so that a Java transform and native observers share it
and a video sender is not restarted on every change.

The native extension API gains encoded frame observers, through which
the media module's new MediaRecorder writes the frames of senders and
receivers into Matroska, WebM or MP4 files without re-encoding. FFmpeg
is built with the matroska, webm and mp4 muxers, extract_extradata and
the AV1 parser for it.
macOS prefers H.264 through VideoToolbox by default. A transform that
changes the whole payload breaks H.264, whose packetizer splits frames at
start codes in it, so nothing reached the receiver there; the recorder
tests check for VP8 as well. VideoToolbox on CI runners also encodes far
fewer frames, so the tests wait for fewer.
@devopvoid
devopvoid merged commit a61e232 into main Sep 27, 2026
11 checks passed
@devopvoid
devopvoid deleted the feat/encoded-frame-transforms branch September 27, 2026 14:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant