Skip to content

feat(control): verified application transfers over 0x50, and non-blocking UDP sockets - #136

Merged
igorls merged 3 commits into
mainfrom
feat/app-streams
Sep 25, 2026
Merged

igorls merged 3 commits into
mainfrom
feat/app-streams

Conversation

@igorls

@igorls igorls commented Sep 25, 2026

Copy link
Copy Markdown
Owner

What

Verified application transfers. Files that don't fit one application message (screenshots, logs, patches) can now be sent as a verified bulk transfer. Frames are MGXF1 binary plaintext inside the authenticated 0x50 envelope, so they take the same direct, hole-punched, and relayed paths as APPSEND and work in --gossip-only mode without TUN or admin rights.

  • Frames are OFFER / DATA (960-byte chunks) / ACK / DONE / CANCEL. An ACK carries a cumulative point plus a selective bitmap covering the next 512 chunks.
  • The sender uses a congestion window: slow start, additive increase, and halving on loss. It estimates round trips with Karn's rule and retransmits on timeout, or early when three later chunks overtake a missing one.
  • The receiver checks the SHA-256 before sending DONE. The sender checks its staged bytes against the same hash before sending anything.
  • A receiver only accepts channels a local app registered with XFERLISTEN. Limits: 32 MiB per file, 8 transfers each way, 96 MiB of memory each way. Offers older than one hour are refused, and a re-sent offer for a recently finished transfer is answered with DONE rather than delivered twice.
  • If a receiver loses its state (restart or expiry), it answers CANCEL(unknown) and the sender offers the file again once, not once per in-flight chunk.
  • New control commands: XFERINFO, XFERLISTEN, XFEROFFER, XFERPUT, XFERSTART, XFERSTATUS, XFERRECV, XFERGET, XFERDONE, XFERCANCEL. On Unix they require the daemon owner (same uid or root), because the control socket is world-writable and these commands carry file contents. APPINFO now reports "transfers":1.
  • Reference doc: docs/reference/app-transfers.md.

Fixes

  • Non-blocking UDP on every platform (the reviewed pilot fix, patch c2362fe0 against a0fa2f3). On macOS/BSD and Windows the gossip socket was blocking, and SWIM empties it by calling recvFrom until it returns null. So once any peer was live, the event loop blocked inside recvfrom after the last datagram and the control socket stopped answering (status, APPSEND, APPRECV). Linux keeps SOCK_NONBLOCK; macOS/BSD use fcntl and Windows uses FIONBIO. IPv4 and IPv6 regression tests cover the empty queue.
  • The gossip-only and macOS loops also wake on the control socket (and TUN). Previously each control command could wait out the 200 ms gossip poll: staging 10 MiB took 43 s before and 41 ms after.
  • 0x50 replay state is now updated only after a packet authenticates, so forged datagrams can't push real nonces out of the 128-entry ring. Transfer frames stay out of that ring, so bulk traffic can't flush it either.

Validation

  • zig build test -Dno-sodium=true: macOS arm64 382/384 (2 skipped); Windows 190/191 (1 skipped).
  • Cross-compiles cleanly for x86_64/aarch64 Linux and x86_64 Windows.
  • Deterministic simulation tests: 3 MiB delivered clean; 1 MiB with 15% loss, 5% duplication and 40 ms jitter; a receiver restart mid-transfer; rejections (no listener, too large, stale offer); forged ACK/DONE from another peer ignored; a replayed offer not delivered again. A control-socket integration test moves a file between two ControlSockets over real 0x50 encryption and checks that non-owners are refused.
  • Two gossip-only daemons on one Apple Silicon Mac over loopback: 10 MiB delivered and verified in 369 ms (about 27 MB/s) with no retransmits.
  • Mac to Windows over the LAN (Meshrooms paired room, 2026-09-25): 41 KB PNG + 8 MiB in 4.2 s to the application receipt; a 116 KB Windows screenshot + 5 MiB back in about 3.1 s. Every hash matched and each message arrived exactly once.

Remaining limits / follow-up

  • Real WAN paths, NAT-to-NAT through relays, and sustained loss are covered only by simulation; the cross-machine run was LAN only.
  • Transfer state is in memory. A sender restart abandons the transfer, and the application (Meshrooms) retries it.
  • Windows named-pipe throughput for large XFERPUT/XFERGET wasn't measured separately; the Windows run above used the pipe and was fine.
  • The mobile FFI still has its own 0x50 path, which doesn't know about transfers.

Event loops drain the gossip socket until recvFrom returns null, as
Linux sockets (SOCK_NONBLOCK) do. On macOS/BSD the socket was blocking,
so once any peer was live the loop parked in recvfrom after the last
datagram and the control socket (status, APPSEND, APPRECV) stopped
answering, in --gossip-only and macOS TUN modes alike.
Add MGXF1 transfer frames inside the authenticated 0x50 plaintext, so
files follow the same direct, hole-punched, and relayed paths as
application messages and work in --gossip-only mode.

- OFFER/DATA/ACK/DONE/CANCEL with 960-byte chunks, cumulative plus
  512-chunk selective acknowledgements, a congestion window (slow start,
  additive increase, halving on loss), Karn RTT estimation, RTO and fast
  retransmit.
- The receiver verifies SHA-256 before DONE; the sender verifies staged
  bytes against the declared hash before sending.
- Receivers accept only channels registered with XFERLISTEN, within
  per-file, slot, and memory limits (32 MiB, 8+8, 96 MiB per direction).
  A receiver that lost state answers CANCEL(unknown) and the sender
  re-offers once for all in-flight chunks.
- Control commands XFERINFO/LISTEN/OFFER/PUT/START/STATUS/RECV/GET/
  DONE/CANCEL; owner uid only on Unix, since the socket is 0666.
  APPINFO advertises "transfers":1.
- 0x50 replay state is updated only after authentication; transfer
  frames stay out of the 128-entry nonce ring.
- Gossip-only and macOS loops wake on the control socket (and TUN), so
  commands no longer wait out the 200 ms gossip poll.

Two gossip-only daemons on Apple Silicon moved 10 MiB in 369 ms over
loopback with the hash verified.
Replace the macOS-only fcntl change with the pilot's reviewed control
fix (patch c2362fe0 against a0fa2f3): Winsock sockets use FIONBIO,
macOS/BSD use F_GETFL/F_SETFL preserving flags, Linux keeps atomic
SOCK_NONBLOCK, setup covers IPv4, IPv6, and connected sockets and
closes the socket on failure, and IPv4/IPv6 empty-drain regression
tests guard against the event loop parking in recvfrom again.
Copilot AI lite review requested due to automatic review settings September 25, 2026 05:41
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@igorls

igorls commented Sep 25, 2026

Copy link
Copy Markdown
Owner Author

Windows validation (x86_64, Windows 11 Pro)

  • Built from a clean detached worktree at the PR head a807b093ff68c9a36a3876a480a79689411def79 with Zig 0.16.0: zig build -Doptimize=ReleaseFast -Dno-sodium=true. meshguard.exe SHA-256 08e4b65714f5b971e486df9f344a0c1905bbaa7e85da71d9fc9488338678b31d.
  • zig build test -Dno-sodium=true: 190 pass, 1 skip (191), 4/4 steps. The failed command: banner in the output is Zig echoing the test runner's stderr (the SWIM revocation warnings), not a failure.
  • Runtime: replaced the Sept 22 pilot transport and kept its identity, trust set, gossip port, named-pipe control path and --gossip-only. Config file hashes were unchanged after startup. APPINFO returned {"protocol":1,"maxPayload":952,"transfers":1}, and the Mac peer was alive within seconds.
  • Both directions used the Windows named pipe for XFER* (Meshrooms staging outbound files and receiving inbound ones): 41 KB + 8 MiB in, 116 KB + 5 MiB out. Every hash matched and no transfer got stuck.
  • Windows Firewall: the new binary path had no inbound allow rule (earlier builds got one from the OS prompt), and both directions still worked. The Windows side seeds the Mac, which likely keeps that UDP flow open. A Windows node that only receives may need the rule, so the Windows run instructions could mention it.

@igorls
igorls merged commit 25b20fb into main Sep 25, 2026
6 checks passed

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unresolved critical replay, transfer-state, arithmetic, and lifecycle issues block approval.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 5 High severity

Open (5)
What changed in this PR

Adds verified bulk application transfers over authenticated 0x50 datagrams and improves cross-platform UDP responsiveness.

Changes:

  • Adds MGXF1 framing, selective ACKs, congestion control, integrity verification, and bounded transfer state.
  • Adds owner-authorized XFER* control commands and documentation.
  • Makes UDP sockets non-blocking across platforms and improves event-loop wakeups.
File Summary
src/​services/​transfers.zig Transfer lifecycle, reliability, verification, and tests
src/​services/​control.zig Transfer commands and authorization
src/​protocol/​transfer.zig Binary transfer framing
src/​net/​udp.zig Cross-platform non-blocking UDP
src/​main.zig Transfer dispatch, replay handling, and event loops
src/​lib.zig Public module exports
src/​discovery/​swim.zig Configurable polling timeout
docs/​reference/​app-transfers.md Transfer API and protocol documentation
docs/​reference/​app-channels.md Transfer documentation link
CHANGELOG.md Release notes

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/main.zig
// rules; keeping them out of this small ring stops bulk traffic from flushing
// the nonces that protect application messages.
const plaintext_slice = plaintext[0..payload_len];
if (!lib.protocol.Transfer.looksLikeTransferFrame(plaintext_slice) and isDuplicateAppNonce(nonce)) return;
Comment on lines +262 to +269
if (offset + bytes.len <= out.staged) {
if (!std.mem.eql(u8, out.data[@intCast(offset)..][0..bytes.len], bytes)) return error.InvalidOffset;
return out.staged;
}
if (offset != out.staged or offset + bytes.len > out.data.len) return error.InvalidOffset;
@memcpy(out.data[@intCast(offset)..][0..bytes.len], bytes);
out.staged += bytes.len;
return out.staged;
Comment on lines +374 to +377
if (out.state != .sending and out.state != .staging and now_ns - out.finished_at > self.limits.retention_ns) self.freeOutgoing(slot, out);
if (out.state == .staging) {
// Staging has no clock of its own; start one at the first tick and drop abandoned uploads.
if (out.last_progress == 0) out.last_progress = now_ns else if (now_ns - out.last_progress > self.limits.retention_ns) self.freeOutgoing(slot, out);
return;
}
const reason: ?wire.Reason = blk: {
if (now_unix - o.created_unix > self.limits.max_offer_age_secs) break :blk .expired;
return;
}
// Acknowledge promptly on gaps and periodically otherwise.
if (index != in.cumulative - 1 or in.unacked >= ACK_EVERY) {
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants