p2p.kiwi: WebRTC Screen Sharing Without the Zoom Tax
Hook
Every Zoom call you've joined has sent your video through servers that could record it. p2p.kiwi claims to eliminate that middleman entirely—but peer-to-peer is only half the truth.
Context
Developer screen sharing has always been a compromise between convenience and control. Corporate tools like Zoom or Google Meet offer one-click simplicity but force you through centralized servers that can (and do) record sessions, inject watermarks, and require accounts tied to email addresses. Self-hosted alternatives like Jitsi require infrastructure expertise and ongoing maintenance. Terminal multiplexers like tmux with tmate offer true peer-to-peer for CLI work but leave GUI applications and browsers completely inaccessible.
The gap is obvious: developers doing quick pair programming sessions—reviewing a pull request, debugging a layout issue, walking through a local build—don't need enterprise features, recording capabilities, or persistent rooms. They need to share their screen right now, with minimal friction, and without wondering if their proprietary code is being scraped for AI training data. p2p.kiwi targets exactly this workflow: generate a link, share it, start streaming. No accounts, no downloads for the viewer (it's web-based for receivers), and theoretically no servers in the middle of your media stream.
Technical Insight
Under the hood, p2p.kiwi follows the standard WebRTC peer-to-peer topology, but the implementation details reveal where 'simple' becomes 'clever.' The architecture relies on three components: an Electron shell for the sender (to access native screen capture APIs), a signaling server for connection brokerage, and WebRTC's peer connection infrastructure for the actual media streaming.
The Electron wrapper is crucial because browsers restrict screen capture to user-initiated actions and specific contexts. Electron's desktopCapturer API bypasses these limitations, letting the app enumerate displays and windows programmatically:
import { desktopCapturer } from 'electron';
const sources = await desktopCapturer.getSources({
types: ['screen', 'window'],
thumbnailSize: { width: 150, height: 150 }
});
const stream = await navigator.mediaDevices.getUserMedia({
audio: false,
video: {
mandatory: {
chromeMediaSource: 'desktop',
chromeMediaSourceId: sources[0].id,
minWidth: 1280,
maxWidth: 3840,
minHeight: 720,
maxHeight: 2160
}
}
});
This stream becomes a MediaStreamTrack attached to an RTCPeerConnection. But the interesting piece is the multi-cursor feature, which requires a separate RTCDataChannel running alongside the video stream. While the screen pixels flow through the standard media pipeline (encoded as VP8/VP9/H.264 depending on browser support), cursor positions need microsecond-latency updates that can tolerate packet loss—you want the freshest position, not a reliable queue of stale coordinates.
The data channel configuration likely looks like this:
const peerConnection = new RTCPeerConnection({
iceServers: [
{ urls: 'stun:stun.l.google.com:19302' },
{ urls: 'turn:relay.example.com', username: 'user', credential: 'pass' }
]
});
const cursorChannel = peerConnection.createDataChannel('cursor', {
ordered: false, // Don't wait for retransmits
maxRetransmits: 0, // Drop lost packets immediately
});
cursorChannel.onopen = () => {
document.addEventListener('mousemove', (e) => {
const position = {
x: e.clientX / window.innerWidth, // Normalize to viewport percentage
y: e.clientY / window.innerHeight,
timestamp: Date.now()
};
cursorChannel.send(JSON.stringify(position));
});
};
The normalization to viewport percentages is critical for handling resolution mismatches—a 4K sender and 1080p receiver need relative coordinates, not absolute pixels. The ordered: false and maxRetransmits: 0 settings trade reliability for latency: if a cursor update packet drops, the next frame (arriving 16ms later at 60Hz) makes it irrelevant anyway.
The signaling flow is where the 'no server' claim gets murky. WebRTC requires an out-of-band mechanism to exchange SDP (Session Description Protocol) offers and ICE (Interactive Connectivity Establishment) candidates. p2p.kiwi uses a centralized WebSocket server for this handshake. When you create a session, the app generates a room ID (likely a UUID or short hash), sends an SDP offer to the signaling server, and gives you a URL like p2p.kiwi/session/abc123. The receiver hits that URL, their browser fetches the SDP offer from the server, generates an answer, and sends it back through the same WebSocket channel. Only after this exchange does the actual P2P connection form.
The genius is in the ephemerality: once the WebRTC connection is established, the signaling server can theoretically vanish without disrupting the media flow. But NAT traversal is the catch—if both peers are behind symmetric NATs, direct connection fails and the TURN relay becomes mandatory. At that point, all your screen data flows through a third-party server, making 'peer-to-peer' a best-case scenario rather than a guarantee.
Gotcha
The peer-to-peer promise collapses in real-world network conditions more often than the marketing suggests. Studies show 8-15% of WebRTC connections require TURN relays due to restrictive NATs, particularly in corporate and university networks. When p2p.kiwi falls back to TURN, your screen content passes through the relay server unencrypted beyond standard DTLS-SRTP. While the encryption prevents casual interception, the relay operator could theoretically log media packets. The project doesn't document which TURN servers it uses or whether they're self-hostable, which matters enormously for trust.
Cross-platform cursor synchronization is fragile in ways that only surface with specific hardware combinations. Mac Retina displays report logical pixels (2x scaling), Windows machines vary from 100% to 300% DPI scaling, and Linux multi-monitor setups can mix resolutions. The viewport-percentage normalization helps, but cursor rendering on the receiver side can still desync when screen aspect ratios don't match—sharing a 21:9 ultrawide to someone on a 16:9 laptop will show the cursor in the wrong position whenever it's near the edges. The codebase would need to transmit screen dimensions alongside cursor positions and apply letterboxing calculations, which adds complexity that 'simple' tools often skip.
The Electron dependency means the sender needs to download and run a 150MB+ application (typical Electron bundle size with Chromium). This contradicts the 'no download' convenience pitch—only receivers get the browser-based experience. For developers who already have Slack or VS Code installed (both Electron apps), the overhead is familiar, but it's worth noting that this could've been a pure browser extension if not for the desktopCapturer requirement. The trade-off is necessary but limits adoption compared to truly zero-install solutions.
Verdict
Use if: You're doing frequent pair programming sessions across open-source projects where participants don't share corporate accounts, you need sub-second latency for real-time code reviews, or you're allergic to Zoom's 40-minute limits and want ephemeral sessions that disappear after use. The multi-cursor feature genuinely improves collaboration over pointing with your mouse and saying 'see that thing there?' Skip if: You need more than three participants simultaneously (WebRTC mesh topology melts CPUs beyond small groups), you require verified end-to-end encryption for compliance reasons (TURN fallback breaks trust guarantees), you want remote control rather than view-only sharing (use RustDesk or Tuple instead), or you're exclusively working in terminals (tmux + tmate is lighter and truly P2P via SSH). The tool delivers on 'simple' for the happy path but inherits every complexity of WebRTC for edge cases.