Day 3: The handshake and the wire protocol
Yesterday ended with EPMD handing out a port number. Today covers what happens on that port: the handshake that turns a TCP socket into a distribution channel, and the protocol that runs over it for the lifetime of the connection. This is the deepest layer the talk touches, and the layer the Wireshark demo lives in, so today’s goal is being able to narrate a capture of a full connection from SYN to tick without looking anything up.
One framing note before the mechanics. The handshake is where the “transparency layer” claim gets tested against reality. Everything Elixir gives you across nodes, send/2 to a remote pid, Process.monitor/1, Node.spawn/2, GenServer.call({name, node}, msg), compiles down to a dozen or so numbered control messages on one TCP stream. There is no RPC framework, no per-call connection, no serialization negotiation per message. Everything rides on a single socket with one framing rule and one term format, which is why the talk can show real packets and have them be legible.
Version context for everything below: OTP 29 is current (29.0 shipped May 2026, 29.1 since). The audience will be mostly on OTP 26 through 29. Everything here is the version 6 (“N”) world, which is the only world that exists since OTP 25.
1. The handshake, packet by packet
Section titled “1. The handshake, packet by packet”After PORT_PLEASE2_REQ to EPMD, the initiating node (call it A) opens a plain TCP connection to the port B registered. What follows is a five-message exchange, all framed with a 2-byte big-endian length prefix. The prefix is 2 bytes during the handshake and 4 bytes afterwards, which trips people up in Wireshark constantly.
Message 1, A to B: send_name. Tag byte 'N' (0x4E), then:
'N' | Flags (8 bytes) | Creation (4 bytes) | Nlen (2 bytes) | NameFlags is a 64-bit big-endian capability bitfield (section 3 below). Creation is a random 32-bit value chosen at boot that identifies this incarnation of the node. Name is the full node name, a@host, as UTF-8.
Message 2, B to A: send_status. Tag 's', then a status string:
ok: carry on.ok_simultaneous: carry on; B was also mid-handshake toward A (simultaneous connect) and is killing its own attempt. The tie is broken deterministically by comparing node names, so exactly one of the two crossing connections survives.nok: stop; B is in a simultaneous-connect handshake that it is keeping, and A’s attempt loses the tiebreak.alive: B believes it already has a live connection to A. This happens when A crashed and restarted fast enough that B has not yet noticed the old connection die. A answerstrue(kill the old one, continue) orfalse(abort).not_allowed: rejected, e.g.-kernel allow_hostsstyle restrictions.named:Name:Creation: only when A setDFLAG_NAME_ME, i.e. A started with a dynamic name (erl -sname undefinedterritory; this is what remsh-style tooling uses). B assigns A its name and creation.
The alive status is worth 15 seconds in the talk, since it is the handshake-level trace of the “stale incarnation” problem that Creation exists to solve.
Message 3, B to A: send_challenge. Also tagged 'N':
'N' | BFlags (8) | Challenge (4) | BCreation (4) | Nlen (2) | BNameThe challenge is a random 32-bit integer. This message also carries B’s flags, and at this point both sides compute the effective capability set (bitwise AND of optional flags) or abort if the other side is missing a mandatory flag.
Message 4, A to B: send_challenge_reply. Tag 'r':
'r' | AChallenge (4) | Digest (16)The digest is the entire authentication mechanism of Erlang distribution:
Digest = MD5(Cookie_as_text ++ Challenge_as_decimal_text)In Elixir, verifiably, against a real capture:
challenge = 0x2B3C4D5Ecookie = "monster"digest = :erlang.md5("#{cookie}#{challenge}")# the 16 bytes you see in packet 4Note the challenge is rendered as a decimal string rather than 4 raw bytes. MD5("monster" <> "725632350"), not MD5("monster" <> <<0x2B,0x3C,0x4D,0x5E>>). If you demo digest recomputation live, this is the detail that makes it work on the first try.
A also sends its own fresh challenge in the same message, because authentication is mutual.
Message 5, B to A: send_challenge_ack. Tag 'a', 16-byte digest of A’s challenge with B’s cookie. A verifies it, and the connection is up. From the next byte onward, the framing switches to 4-byte lengths and the control-message protocol takes over.
Three points about this design belong in the talk. The cookie never crosses the wire: challenge-response means a passive sniffer sees two random numbers and two MD5 hashes, never the secret. That is where the protection ends, though. After message 5, everything is plaintext. No session key is derived, and nothing is encrypted or integrity-protected. The handshake authenticates the peer and then trusts the TCP stream completely, so an active attacker who can MITM the stream after the handshake owns both nodes. This comes back on Oct 10 (cookie = RCE, TLS distribution). And MD5 is not a meaningful weakness in itself, since preimage resistance is not the problem here. The problem is a 1998 trust model. Put it that way and you preempt the “but MD5 is broken” question with something more precise.
Where this lives in OTP: the whole state machine is lib/kernel/src/dist_util.erl. handshake_we_started/1 and handshake_other_started/1 are the two entry points, gen_digest/2 is the two-liner that computes the MD5, and con_loop/2 at the bottom of the file is the post-handshake connection supervisor that handles ticks. It is readable in one sitting, and citing it on a slide backs up the claim that none of this is magic.
History checkpoint: the old 'n' (version 5) handshake, with 2-byte version field and 32-bit flags, died in OTP 25. OTP 23 and 24 spoke both and used DFLAG_HANDSHAKE_23 to upgrade. Since OTP 25 only version 6 is accepted, which also means an OTP 25+ node cannot connect to anything older than OTP 23. This cadence (introduce at N, both at N+1, drop at N+2) is OTP’s standard pattern for distribution changes, and it deserves one sentence in the talk because it explains how BEAM clusters do rolling upgrades across major versions at all.
2. Distribution flags: the capability negotiation
Section titled “2. Distribution flags: the capability negotiation”The 64-bit flags field is how two nodes agree on dialect. The effective set is the intersection, except for mandatory flags, where absence kills the connection during the handshake.
The mandatory set as of OTP 26+ (and so for anything the audience runs): EXTENDED_REFERENCES, FUN_TAGS, NEW_FUN_TAGS, EXTENDED_PIDS_PORTS, EXPORT_PTR_TAG, BIT_BINARIES, NEW_FLOATS, UTF8_ATOMS, MAP_TAG, BIG_CREATION, HANDSHAKE_23, UNLINK_ID, V4_NC. There is also DFLAG_MANDATORY_25_DIGEST, a meta-flag meaning “I support everything that was mandatory in OTP 25”, which exists so the ever-growing mandatory list can be compressed into one bit. It became mandatory itself in OTP 27.
None of this goes on a slide. What is worth knowing cold:
DFLAG_V4_NC(mandatory since 26): pids, ports and references with 64-bit data words, “V4 node container”. This is whyis_pidround-trips survive very long-lived nodes with huge pid counters.DFLAG_BIG_CREATION: the 32-bit creation. Pre-history: creation was 2 bits (values 1 to 3!) handed out by EPMD, which meant a node restarting four times could collide with its own ghost. Random 32-bit creation chosen by the node itself fixed that, and is why EPMD’s role shrank (ties back to yesterday).DFLAG_DIST_HDR_ATOM_CACHE: enables the atom cache (section 5).DFLAG_FRAGMENTS: fragmentation of large messages, OTP 22+ (section 6).DFLAG_SPAWN: remote spawn as a first-class control message, used byerpcandNode.spawn/2style calls on modern OTP.DFLAG_SEND_SENDER,DFLAG_EXIT_PAYLOAD: variants of SEND and EXIT that changed the control-message shape for optimization reasons (section 4).DFLAG_ALTACT_SIG: new in OTP 28, carries “alternate action” signals, which is the wire-level foundation for the priority-messages work in recent OTP (and the designated successor to the alias-send opcodes). The talk needs one sentence on this at most: the protocol is still evolving, 2026 included.
The negotiation story in one line for the stage: the handshake is also a feature negotiation, which is how a heterogeneous cluster (OTP 26 and OTP 29 nodes mixed) speaks a common dialect without any config.
3. Framing and the two header styles
Section titled “3. Framing and the two header styles”Post-handshake, every message on the socket is:
Length (4 bytes, big-endian) | DistributionHeader | ControlMessage [| Payload]Length counts everything after itself. A length of 0 is a tick (section 7), nothing follows.
Two header styles exist. The old “pass-through” style is a single byte 112 ($p) followed by two ETF terms, each with its own leading 131 version byte: the control message and, if present, the payload. You still see this style from libraries that implement distribution in userspace and in some alternative carriers. The modern style, used between real ERTS nodes, is the distribution header proper:
131 | 68 ('D') | NumberOfAtomCacheRefs | Flags... | AtomCacheRefs...followed by the control message and payload as ETF terms without their own 131 prefixes (the version byte is established once, by the header). Tags 69 ('E') and 70 ('F') are the fragmented variants (section 6).
For the Wireshark demo this is the practical decode key: handshake frames are 2-byte-length framed with ASCII tags (N, s, r, a), data frames are 4-byte-length framed and start with 131,68 (or 131,112 in pass-through captures), and empty frames are ticks. That one sentence makes a raw hex dump navigable, which is exactly the fallback documented for when the Erlang dissector is flaky (Decode As, then hex extraction plus :erlang.binary_to_term/2 with [:used] in IEx).
4. The control messages: the instruction set of distribution
Section titled “4. The control messages: the instruction set of distribution”The control message is a small ETF-encoded tuple whose first element is an integer opcode. This is the complete “instruction set” that all of distributed Elixir compiles to. The ones that matter, with their actual shapes:
| Op | Name | Shape |
|---|---|---|
| 1 | LINK | {1, FromPid, ToPid} |
| 2 | SEND | {2, Unused, ToPid} + payload |
| 3 | EXIT | {3, FromPid, ToPid, Reason} |
| 6 | REG_SEND | {6, FromPid, Unused, ToName} + payload |
| 7 | GROUP_LEADER | {7, FromPid, ToPid} |
| 8 | EXIT2 | {8, FromPid, ToPid, Reason} (this is Process.exit/2) |
| 19 | MONITOR_P | {19, FromPid, ToProc, Ref} |
| 20 | DEMONITOR_P | {20, FromPid, ToProc, Ref} |
| 21 | MONITOR_P_EXIT | {21, FromProc, ToPid, Ref, Reason} |
| 22 | SEND_SENDER | {22, FromPid, ToPid} + payload |
| 29 | SPAWN_REQUEST | {29, ReqId, From, GL, {M,F,A}, OptList} + args |
| 31 | SPAWN_REPLY | {31, ReqId, To, Flags, Result} |
| 33 | ALIAS_SEND | {33, FromPid, Alias} + payload |
| 35 | UNLINK_ID | {35, Id, FromPid, ToPid} |
| 36 | UNLINK_ID_ACK | {36, Id, FromPid, ToPid} |
| 37 | ALTACT_SIG_SEND | OTP 28+, alias/priority sends |
Several rows in this table carry stories, listed here in rough order of stage value.
SEND (2) has no sender. The shape is {2, Unused, ToPid}, so a plain send(remote_pid, msg) does not tell the receiving node who sent it. Message provenance is simply not part of the model. The Unused slot is a fossil: in ancient OTP it carried the cookie, on every single message. SEND_SENDER (22) was added so the VM could pass the sender for scheduling optimizations, and modern nodes use it when negotiated, but the semantic model still has no envelope sender. Good 20-second aside: people assume self() travels implicitly, when it only travels because OTP conventions (GenServer.call) put a reply address inside the payload.
The GenServer.call demo decodes differently than people expect. GenServer.call({MyServer, :"b@host"}, :ping) goes out as REG_SEND (6) carrying {:"$gen_call", {from_pid, alias_ref}, :ping}. The reply does not come back as SEND, though. Since OTP 24, gen uses process aliases for replies, so the reply arrives as ALIAS_SEND (33) targeting the alias ref, which is how late replies to timed-out calls get dropped at the VM level instead of polluting your mailbox. On OTP 28+ hardware you may see opcode 37 instead, as alias sends migrate to ALTACT_SIG. Verify in your own capture on your demo OTP version before stage; the pair 6-then-33 is what I would expect through OTP 27, and wire_demo.exs will settle it in 30 seconds.
UNLINK grew an ACK, and the reason is a concurrency story. Old UNLINK (4) was fire-and-forget, which left a race: an EXIT signal already in flight when you unlink could still kill you after you thought the link was gone. OTP 23 introduced UNLINK_ID / UNLINK_ID_ACK (35/36), a two-phase unlink where the unlinking side ignores incoming EXITs from that peer between ID and ACK. It became mandatory in OTP 26 and opcode 4 is dead. This example supports the talk’s thesis directly: even inside the “transparent” layer, signal ordering across a network needed a protocol fix that purely local semantics never needed.
Monitors are a round trip plus a standing entry. Process.monitor({name, node}) or of a remote pid emits MONITOR_P (19). The down notification arrives as MONITOR_P_EXIT (21) with the reason (or PAYLOAD_MONITOR_P_EXIT, 28, where the reason rides as a separate payload term). Every cross-node monitor is state in both nodes’ distribution tables, cleaned up on disconnect, which is why nodedown storms with hundreds of thousands of monitors are a real CPU event. Relevant Thursday for monitor_nodes and the Cluster GenServer.
SPAWN_REQUEST (29) is why erpc is better than rpc. Classic :rpc.call funnels through the rex server process on the target node (REG_SEND to a named GenServer, a bottleneck and a failure coupling point). erpc and modern spawn/4 across nodes use SPAWN_REQUEST, a first-class VM operation that spawns directly with no intermediary server. Worth a sentence on the slide: even RPC on the BEAM dissolved into a primitive send.
EXIT (3) vs EXIT2 (8): 3 is link-propagated death, 8 is explicit Process.exit(pid, reason). Because the distinction is wire-visible, a capture can show the difference between “a link did this” and “someone called exit”, which is useful forensically and worth a moment in the demo.
The _TT variants (12, 13, 16, 18, 23, etc.) are the same operations carrying a sequential-tracing token. Ignore them unless a capture surprises you with one.
5. ETF and the atom cache
Section titled “5. ETF and the atom cache”Payloads and control messages are External Term Format: tag bytes like 97 (SMALL_INTEGER), 98 (INTEGER), 104 (SMALL_TUPLE), 106 (NIL), 107 (STRING, the list-of-bytes optimization), 108 (LIST), 109 (BINARY), 118 (ATOM_UTF8), 119 (SMALL_ATOM_UTF8), 88 (NEW_PID), 90 (NEWER_REFERENCE). The hex-reading reflexes are already there from the September deep dive. The one piece worth consolidating today is the atom cache, because it is the most alien thing in a capture.
Atoms dominate BEAM messages ({:"$gen_call", ...}, module names, record tags), and re-sending MyApp.Some.Long.Module.Name as UTF-8 on every message would be absurd. So each connection maintains a shared atom cache: 8 segments of 256 slots, 2048 entries total, kept in sync by the sender. The distribution header carries up to 255 atom cache references per message. Each reference is a flag nibble (1 bit NewCacheEntryFlag, 3 bits segment index) plus an internal segment index byte. A new entry additionally carries the atom text once, and every subsequent message that uses that atom sends roughly one byte instead.
In Wireshark this is exactly why a naive decode fails: the atoms in the control message are not in the packet. They are back-references into connection state established by earlier packets. The reliable fallback (capture from connection start, or decode with the cache misses visible) follows directly. For the stage, the point is that the connection itself has memory, so a packet from the middle of a conversation is not self-describing. It also previews honestly why packet-level debugging of dist links is harder than HTTP, where every request is self-contained.
One security footnote for Oct 10: atoms received over distribution are created on the receiving node, and the atom table is finite (default 1,048,576) and not garbage collected. Inside the trust boundary that is fine, and it is one more reason the trust boundary is absolute.
6. Fragmentation: fixing head-of-line blocking
Section titled “6. Fragmentation: fixing head-of-line blocking”Before OTP 22, a 500 MB message on a dist link (say, someone sending a huge binary to a remote pid, or :global syncing something big) serialized as one frame, and every other message between those two nodes queued behind it. Worse, ticks queued behind it too, so a big transfer could make two healthy nodes declare each other dead. This failure shape happens in real production systems and is worth telling as a war story.
OTP 22 added fragmentation (DFLAG_FRAGMENTS). Large messages are split into fragments, in current ERTS source 64 kB each (ERTS_DIST_FRAGMENT_SIZE, 64 * 1024 in erts/emulator/beam/dist.h; the protocol spec deliberately does not fix a size). The first fragment uses header tag 69 ('E'): it carries a SequenceId (8 bytes, effectively identifying the sending process), a FragmentId that starts at the total fragment count and counts down to 1, plus the full atom cache refs and the complete control message. Continuations use tag 70 ('F') with just SequenceId and FragmentId. Fragments from different sending processes interleave freely on the socket. A single process must finish one fragmented message before starting its next message, which preserves the per-sender-pair ordering guarantee the rest of the week leans on.
The countdown FragmentId is a neat detail: the receiver knows the total from fragment one and knows it is done at 1, no trailer needed.
A caveat that belongs in the talk’s drawbacks section: fragmentation un-blocks the socket, but the receiver still reassembles the whole term in memory before delivery, and the default distribution buffer busy limit (+zdbbl, erts_dist_buf_busy_limit, default 1 MB) means senders get suspended when the outbound buffer backs up. The summary for the stage is that distribution is not a message queue. Big payloads belong in object storage or chunked protocols, with the dist link carrying references. This sets up both the netsplit material (Oct 9) and the “when not to cluster” close.
7. Ticks: the heartbeat you can see
Section titled “7. Ticks: the heartbeat you can see”With no traffic, nodes exchange ticks: a frame whose 4-byte length field is zero and which contains nothing, so the whole keepalive is four bytes of 0x00. con_loop in dist_util sends one if nothing else was sent during the last quarter of net_ticktime. The check runs every net_ticktime / 4 (default 60 s, so 15 s subintervals), and a peer from which nothing has arrived for four consecutive subintervals is declared down. Hence the canonical detection window of roughly 45 to 75 seconds, which is Friday’s (Oct 9) opening number. Today just anchor the mechanism: any traffic counts as liveness, ticks only fill silence, and detection latency is a function of net_ticktime, not of ticks per se.
For the capture demo: let two idle connected nodes sit for a minute and you get a clean rhythm of zero-length frames in both directions. It reads well on a projector and is the simplest packet in the protocol to explain.
8. Assembling today into the talk
Section titled “8. Assembling today into the talk”The section 1 arc now has its full shape. EPMD is the phone book (yesterday), the handshake is the border control (identity, incarnation, capabilities, shared-secret proof), and the connected protocol is a tiny instruction set (a dozen live opcodes) over one framed TCP stream with connection-local memory (atom cache) and a heartbeat (ticks). Everything in sections 2 and 3 of the talk is sugar over opcodes 2, 6, 19, 21, 29, and 33.
The honest-tradeoffs thread picks up three items today: authentication without encryption (one secret, all nodes, plaintext after the handshake), no per-message sender or provenance, and the buffer/fragmentation behavior that makes dist links the wrong transport for bulk data. All three pay off Oct 9 and 10.
References
Section titled “References”- Distribution Protocol, erl_dist_protocol (OTP 29 docs): handshake, flags, complete control-message table.
- External Term Format, erl_ext_dist (OTP 29 docs): distribution headers 68/69/70, atom cache encoding, fragmentation.
- dist_util.erl in OTP master: the handshake state machine,
gen_digest/2,con_looptick handling. - dist.h in OTP master:
ERTS_DIST_FRAGMENT_SIZE(64 kB),ERTS_DE_BUSY_LIMIT(1 MB). - How to Implement an Alternative Carrier for the Erlang Distribution: useful for understanding what is pluggable vs fixed.
- Erlang/OTP 29.0 release notes and OTP 29.1: current-version context.
Active block (45 min, repo section 00_epmd_wireshark)
Section titled “Active block (45 min, repo section 00_epmd_wireshark)”Capture and hand-verify one complete handshake. Start tcpdump -i lo0 -w handshake.pcap port not 4369 (exclude EPMD noise), then from wire_demo.exs connect a@localhost to b@localhost with a known cookie. In Wireshark: (1) identify the five handshake frames by their 2-byte framing and tag bytes N, s, N, r, a; (2) copy B’s 4-byte challenge from frame 3 and the 16-byte digest from frame 4 as hex; (3) in IEx, confirm :erlang.md5("#{cookie}#{challenge}") |> Base.encode16() matches the captured digest, remembering the challenge is decimal text; (4) then send one GenServer.call across and find the REG_SEND (6) frame and its reply frame, noting which opcode the reply actually uses on your demo OTP version (expect 33, record what you see); (5) let the nodes idle 90 seconds and screenshot the zero-length tick frames. Commit the pcap, the IEx transcript, and a NOTES.md with the observed reply opcode into the repo. That pcap is your fallback if live capture misbehaves on stage.
Exit questions
Section titled “Exit questions”- Walk the five handshake messages in order, with their tag bytes, and state exactly what is MD5-hashed to produce the digest (including the text-encoding detail of the challenge).
- A plain
send/2to a remote pid, aGenServer.call({name, node}, msg), and its reply: which control opcodes does each produce on the wire, and why does the reply not use opcode 2? - Why did UNLINK become a two-message protocol with an ID and an ACK, and what does that change tell you about the limits of location transparency?
Say it out loud
Section titled “Say it out loud”In 60 seconds, to the Haarlem room: “When two BEAM nodes connect, there is no TLS-style key exchange and no API negotiation like you know from HTTP. Here is what actually happens in five small packets, and here is the one MD5 hash that stands between the internet and remote code execution on your cluster.” Walk the five frames from memory, end on “and after that, every feature of distributed Elixir is one of about twelve numbered tuples on this socket”.