Day 2: EPMD, the departure board
Yesterday was the thesis: distribution gives you reach but not agreement. Today is the first hop of reach, the part that happens before two BEAM nodes have exchanged a single term: the Erlang Port Mapper Daemon. This is also the section of the talk you already shot a video for, so today’s job is making sure the mechanics under the Tulln departure board analogy hold up, because this is the material where someone in the Haarlem audience will have read the same docs you did.
The problem EPMD solves, stated precisely
Section titled “The problem EPMD solves, stated precisely”When a BEAM node starts distribution, its distribution listener binds an ephemeral port: it asks for port 0 and the kernel picks one. That is a deliberate choice. It means any number of nodes can run on one host without port coordination, and it means node identity lives in the name rather than the port. app@host is stable across restarts. The port is not.
That creates a rendezvous problem. a@berlin wants to connect to b@berlin. It can resolve berlin via DNS, but DNS gives it a host, not a port. Something on that host has to answer the question “what port is the node named b listening on right now?” That something is EPMD: a small standalone C daemon, one per host, listening on the one port everybody agrees on in advance, 4369.
Two properties matter and both are easy to get wrong on stage:
- EPMD is per host, not per cluster. There is no global registry. Each host’s EPMD knows only about the nodes on that host. Cluster-wide discovery, meaning which hosts exist at all, is a different problem, solved by DNS, libcluster, static config, whatever. This is also the cleanest way to position libcluster for the Kubernetes crowd: libcluster finds hosts, EPMD maps names to ports on each host. Two different layers solving two different problems, and they get conflated all the time.
- EPMD knows nothing about Erlang. It never sees a term, never checks a cookie, and holds no state beyond name, port, and a few bytes of metadata. It does no consensus, no authentication, and no replication. That minimalism is the design, and it previews the thesis: even the discovery layer only promises reach.
The canonical analogy in the Unix world is the Sun RPC portmapper (rpcbind, port 111): services on dynamic ports register with a fixed-port daemon, clients ask the daemon first and then talk to the service directly. EPMD is the same pattern with a four-decade pedigree. The departure board version is better for a mixed audience, and it maps precisely: the board (EPMD) tells you which platform (port) the train named “REX 3 to Krems” (node name) leaves from today, and tomorrow the same train may use a different platform. You never board the departure board itself. You walk to the platform, and the actual journey (the handshake, tomorrow’s topic) happens there.
A lifecycle detail worth knowing cold: erl -sname foo checks whether an EPMD is already running on the host and spawns one if not (-start_epmd true is the default). The daemon then outlives the node. Kill every node on a host and EPMD keeps running with an empty database. That surprises people who expect it to be part of the VM. It is a separate OS process, and epmd -names works from a plain shell with no BEAM anywhere in sight.
The protocol, byte by byte
Section titled “The protocol, byte by byte”Everything below is from the current erl_dist_protocol reference. The framing rule: every request to EPMD is prefixed with a 2-byte big-endian length, and responses are not length-prefixed. That asymmetry bites you the first time you hand-roll a client, so it is worth saying explicitly in the talk if you demo raw bytes.
Registration: ALIVE2_REQ (120)
Section titled “Registration: ALIVE2_REQ (120)”When a node starts distribution, erl_epmd (the kernel module acting as EPMD client) opens a TCP connection to localhost:4369 and sends:
| Field | Size | Meaning |
|---|---|---|
| 120 | 1 | ALIVE2_REQ opcode |
| PortNo | 2 | the distribution listener’s port |
| NodeType | 1 | 77 = normal node, 72 = hidden node |
| Protocol | 1 | 0 = TCP/IPv4 (also used for IPv6 in practice) |
| HighestVersion | 2 | highest dist protocol version, 6 since OTP 23 |
| LowestVersion | 2 | lowest accepted, 6 since OTP 25 (version 5 dropped) |
| Nlen | 2 | name length |
| NodeName | Nlen | the short part only: foo, not foo@host |
| Elen | 2 | extra length, 0 in practice |
| Extra | Elen | reserved |
Three things to internalize:
- The registered name has no host part. EPMD is host-local, so
@hostwould be redundant. When you demoepmd -namesyou seename foo at port 50123, bare names. - NodeType 72 is how hidden nodes stay findable. Hidden nodes (
-hidden, and anything with-dist_listen false, which implies hidden) still register with EPMD if they listen. Hidden-ness is about the mesh andNode.list/0visibility, not about discovery. Observer, remsh, and debug tooling depend on this. - The version fields preview the handshake. OTP 25 raised LowestVersion to 6, which is why an OTP 25+ node refuses a pre-OTP-23 peer outright. That negotiation happens here, before any cookie is involved.
The response is where creations enter. Modern EPMD replies with ALIVE2_X_RESP:
<<118, Result::8, Creation::32>> # ALIVE2_X_RESP, OTP 23+<<121, Result::8, Creation::16>> # legacy ALIVE2_RESPResult == 0 is success and anything else is failure. The symptom at the Elixir level is the familiar the name foo@host seems to be in use by another Erlang node, which is EPMD returning a nonzero Result because the name is already in its table with a live registration socket.
Deregistration: closing the socket
Section titled “Deregistration: closing the socket”A node unregisters by closing the TCP connection it registered on. No UNREGISTER opcode exists. The socket opened for ALIVE2_REQ stays open for the node’s entire lifetime, and the registration is exactly as alive as that socket. The name ALIVE2 is literal: the open connection is the liveness signal. Give this fact its own moment in the talk. Consequences:
kill -9a beam process and the kernel sends FIN/RST on its sockets, so EPMD drops the registration within milliseconds. No timeout, no garbage collection, no stale-entry reaper needed. The OS does the cleanup for free.- There is no heartbeat between node and EPMD, because TCP connection state serves that purpose.
- A
SIGSTOPped node stays registered (socket still open), which is consistent: the node is not dead, just not scheduled. The same logic shows up at cluster level on Day 6 with ticks. - This is why
epmd -stop Namebarely works. Forcibly unregistering a live node is such an unnatural operation that EPMD refuses it unless started with-relaxed_command_check, and always refuses it from remote hosts.
Worth one sentence in the talk: instead of adding an application-level keepalive protocol, EPMD borrows TCP’s connection semantics, the smallest mechanism that is still correct.
Lookup: PORT_PLEASE2_REQ (122)
Section titled “Lookup: PORT_PLEASE2_REQ (122)”When Node.connect(:"b@berlin") runs, :net_kernel resolves berlin (that is address_please, plain inet/DNS, not EPMD), then asks the EPMD on berlin:
<<Len::16, 122, "b">>Success response (119 = PORT2_RESP):
<<119, 0, PortNo::16, NodeType::8, Protocol::8, HighestVersion::16, LowestVersion::16, Nlen::16, NodeName::binary-size(Nlen), Elen::16, Extra::binary>>Failure is just <<119, Result>> with Result > 0. EPMD closes the connection after replying, so lookups are one-shot, unlike the long-lived registration socket. Your node then connects directly to PortNo and the handshake (Day 3) begins. EPMD is out of the picture after that. It is consulted once per connection attempt and never sees a message.
The query opcodes
Section titled “The query opcodes”NAMES_REQ (110) returns EPMD’s own port as 4 bytes followed by human-readable lines, one name X at port Y per registered node. This single opcode powers epmd -names from the shell, :net_adm.names/0,1 and :erl_epmd.names/1 from Elixir. DUMP_REQ (100) adds file descriptors and shows old/unused slots. KILL_REQ (107) returns the ASCII bytes "OK" and kills the daemon, but only if the node database is empty or -relaxed_command_check was given, and like -stop it is refused from remote hosts: remote connections get query commands answered and nothing else.
Note that NAMES is answered for anyone who can reach port 4369. Day 10’s security section comes back to this: an internet-exposed EPMD is a free enumeration service for node names and distribution ports. The RCE still requires the cookie, but EPMD hands over the map. CVE-2022-24706 (CouchDB) started exactly this way.
Raw lookup in forty lines
Section titled “Raw lookup in forty lines”This belongs in 00_epmd_wireshark as the “no magic” demo. No dependencies and no :net_kernel, just bytes on a socket:
defmodule EpmdRaw do @moduledoc "Speak the EPMD protocol by hand. Requests: <<len::16, body>>. Responses: raw."
def names(host \\ ~c"localhost") do {:ok, s} = :gen_tcp.connect(host, 4369, [:binary, active: false]) :ok = :gen_tcp.send(s, frame(<<110>>)) {:ok, <<_epmd_port::32, text::binary>>} = :gen_tcp.recv(s, 0) IO.puts(text) end
def port_please(name, host \\ ~c"localhost") do {:ok, s} = :gen_tcp.connect(host, 4369, [:binary, active: false]) :ok = :gen_tcp.send(s, frame(<<122, name::binary>>))
case :gen_tcp.recv(s, 0) do {:ok, <<119, 0, port::16, type, 0, hi::16, lo::16, nlen::16, ^name::binary-size(nlen), _::binary>>} -> {:ok, %{port: port, hidden?: type == 72, versions: lo..hi}}
{:ok, <<119, err>>} -> {:error, err} end end
defp frame(body), do: <<byte_size(body)::16, body::binary>>endRun iex --sname a, then from a plain iex (no name, no distribution) call EpmdRaw.port_please("a"). The demo shows the audience that the discovery layer of a BEAM cluster is simple enough to speak with :gen_tcp in a coffee break. Pair it with your Wireshark capture filter on tcp.port == 4369. Per your earlier findings, the EPMD side dissects reliably even on builds where the dist dissector is flaky, since 4369 is a fixed port and needs no Decode As.
Creation values: the incarnation stamp
Section titled “Creation values: the incarnation stamp”Creation is the field everyone skims past in ALIVE2_X_RESP, and it matters more for correctness than that suggests. The problem it solves: pids escape their node. They sit in remote ETS tables, in :global’s registry, in messages queued on other nodes. Now b@berlin crashes and restarts with the same name. The old pids still reference b@berlin, and pid slots get reused, so an old pid could structurally collide with a live process in the new incarnation. Without extra information, a message aimed at a process that died with the old node could be delivered to an unrelated process in the new one.
The fix: every pid, port, and reference carries a creation alongside its node name, an integer identifying the incarnation of that node. The external term format makes this concrete: NEW_PID_EXT is node atom, 4 bytes of id, 4 bytes of serial, and 4 bytes of creation. Same name plus different creation equals different node identity, and a message addressed to a dead incarnation’s pid is dropped rather than misdelivered.
Where does the number come from? In the classic setup, from EPMD, in the ALIVE2 response. History matters here because the talk’s audience may know the old behavior:
- Pre OTP 23: creation was 2 bits in the external format, so EPMD cycled through 1, 2, 3 as a name was re-registered. Three restarts and you wrap, so a sufficiently old pid could alias a new incarnation. A known, tolerated wart.
- OTP 23 (“big creation”): the external format grew
NEW_PID_EXT/NEWER_REFERENCE_EXTwith 32-bit creations, EPMD grew ALIVE2_X_RESP to hand back a 32-bit value, and wrap-around stopped being a practical concern. Values 0 through 3 are kept out of the new range for legacy reasons, and 0 historically served as a wildcard “unset” creation (verify the exact reserved-range wording against erl_dist_protocol before putting numbers on a slide).
You can see all of it from IEx:
:erlang.system_info(:creation)# => e.g. 1696401732 (32-bit, handed out at registration)
<<131, 88, _atom_prefix::binary>> = :erlang.term_to_binary(self())# 88 = NEW_PID_EXT; the last 4 bytes of the pid encoding are the creationRestart the node with the same --sname, print system_info(:creation) again, and you get a different number. That demo covers the whole concept in about ten seconds.
One subtlety for EPMD-less setups (next section): if nobody hands you a creation, you make your own. Custom epmd_module implementations return a creation from register_node, and nodes with dynamic names get theirs assigned by the peer during handshake. The invariant that matters is uniqueness across incarnations of the same name, not where the number comes from.
Life without EPMD
Section titled “Life without EPMD”This is the part of today’s material most relevant to your Kubernetes-adjacent audience, and the part where blog posts age worst. Here is the current (OTP 27/28) state of the escape hatches, in increasing order of radicalism.
-
-start_epmd false, keep EPMD. The node does not auto-spawn the daemon; you run it yourself (epmd -daemonin the container entrypoint, or a systemd unit). Nothing else changes. The node still registers and looks up as usual, and fails at startup if no daemon is reachable. Why bother: process supervision hygiene. The auto-spawned EPMD is owned by whichever node started it first, logs nowhere useful, and in containers you generally want PID 1 discipline rather than surprise daemons. -
Fix the distribution port, keep EPMD.
-kernel inet_dist_listen_min 9100 inet_dist_listen_max 9100pins the listener. EPMD still runs and still answers lookups, but your firewall rules become writable by humans. This is the pragmatic production default: one node per host or pod, one well-known port, EPMD reduced to a formality that costs one daemon. Most k8s Elixir deployments live here whether they know it or not. -
Replace the client:
epmd_module. The node-side EPMD client is just a module (erl_epmdby default) implementing a small behavior:start_link/0,register_node/2,3(return{ok, Creation}yourself),port_please/2,3(return{port, Port, Version}from wherever you like),address_please/3,listen_port_please/2,names/1. Swap it with the kernel parameterepmd_moduleand EPMD the daemon never enters the picture. The classic recipe (the Erlang Solutions post, theepmdlesslibrary) is a module that derives the port from the node name or a static map. The config ergonomics improved recently:-epmd_moduleas anerlflag is now deprecated in favor of the kernel application parameter, which meansconfig/runtime.exsorapplication:set_env(kernel, epmd_module, M)beforenet_kernel.start/2works, no VM args needed. That change landed via OTP PR #8671, aimed at OTP 27.2 (verify the exact release before stage if you cite the version number). -
The official one-liner:
erl_epmd_node_listen_port. Current kernel docs: configures the porterl_epmduses both to listen and to assume when connecting to other nodes. Set it (the old-erl_epmd_portflag is the deprecated spelling) together with-start_epmd falseand you get a fully EPMD-less node with stock OTP: registration becomes a no-op with a self-assigned creation, and lookups short-circuit to “the same port I use”. The built-in assumption is uniform ports across the cluster, which is exactly the k8s one-node-per-pod shape:
# rel/env.sh.eex or vm args# ELIXIR_ERL_OPTIONS="-start_epmd false"config :kernel, erl_epmd_node_listen_port: 9100(Flag the exact interplay of -start_epmd false plus the kernel param on your OTP patch level as “verify before stage”; run it in the companion repo rather than trusting my summary or anyone’s blog post.)
- Dynamic node names:
-name undefined. OTP 23+ lets a node start with no name at all. It requests one from the first node it connects to, via theDFLAG_NAME_MEhandshake flag, and gets its creation assigned by the peer too. This is whatiex --remshrelies on when you give it no name (it defaults to-sname undefined). Such a node is a pure client, and needs no EPMD, no listener, and no registration, because nothing ever dials in to it. Worth thirty seconds in the talk because remsh is the distribution feature everyone uses daily without knowing it is distribution.
What you give up without EPMD: epmd -names as a debugging tool, :net_adm.names/1, and the ability to run several nodes per host without port coordination, i.e. precisely the dev-machine conveniences. Hence the sane default posture: EPMD-less or fixed-port in production where ports are policy, stock EPMD on laptops where convenience wins. A good interop footnote: because the protocol is this simple, non-BEAM dist implementations (Go’s ergo, Python’s pyrlang) implement EPMD’s client side in a few dozen lines, which is decent evidence that there is nothing VM-specific or magical about it.
Stage notes
Section titled “Stage notes”For a 40-minute talk, EPMD is realistically 6 to 8 minutes. The arc that fits: the departure board analogy (you have the video), one Wireshark or EpmdRaw moment showing ALIVE2 and PORT_PLEASE2 are a handful of bytes, the fact that deregistration is just a TCP close, creation as the hidden incarnation stamp in every pid, and one slide on EPMD-less setups for the k8s crowd. The thesis thread to pull: EPMD is the simplest possible form of reach, a name, a port, and an open socket. It promises you can find a node. It says nothing about whether the node agrees with anyone about anything, and it does not promise the node will still be there when you connect. Every layer above it inherits the same limits.
Trap questions Haarlem might ask, worth having answers ready: “Why not just use DNS SRV records instead of EPMD?” (you can, that is essentially what a custom epmd_module does; EPMD exists so multiple nodes per host with dynamic ports work with zero config), and “Does EPMD see my messages / my cookie?” (never; it is out of the loop after PORT_PLEASE2, and the cookie belongs to tomorrow’s handshake).
References
Section titled “References”- Distribution Protocol reference (erl_dist_protocol), ERTS docs: the authoritative byte layouts for ALIVE2_REQ, ALIVE2_X_RESP, PORT_PLEASE2_REQ, PORT2_RESP, NAMES_REQ, KILL_REQ, and the EPMD section preceding the handshake spec.
- epmd man page (epmd_cmd), ERTS docs: daemon flags,
-names,-relaxed_command_check,ERL_EPMD_PORT,ERL_EPMD_ADDRESS, remote-host command restrictions. - erl_epmd module docs, kernel: the client-side callback surface you implement to replace EPMD.
- kernel app configuration reference:
epmd_module,erl_epmd_node_listen_port,start_distribution,net_ticktime. - erl command reference:
-start_epmd,-dist_listen, dynamic node names via-name undefined, deprecation notes for-epmd_moduleand-erl_epmd_port. - Erlang (and Elixir) distribution without epmd, Erlang Solutions: the classic custom-epmd_module walkthrough (predates the kernel params; read with that in mind).
- OTP PR #8671: move -epmd_module and -erl_epmd_port into kernel parameters: background and motivation for the config change.
- epmdless library: a packaged epmd_module replacement, useful as reference code.
Active block (30 to 45 min, repo section 00_epmd_wireshark)
Section titled “Active block (30 to 45 min, repo section 00_epmd_wireshark)”Build EpmdRaw from today’s snippet into 00_epmd_wireshark/epmd_raw.exs and script this sequence in the justfile:
- Start
tcpdump -i lo0 -w epmd.pcap port 4369(or Wireshark on loopback, filtertcp.port == 4369). iex --sname a, theniex --sname b, then fromb:Node.connect(:a@yourhost).- From a third, undistributed
iex:EpmdRaw.names()andEpmdRaw.port_please("a"). kill -9nodea’s beam process; immediately runEpmdRaw.names()again and confirm the registration vanished with no explicit deregistration message, then find the bare FIN in the capture.- In the pcap, annotate one ALIVE2_REQ (find the
120opcode after the length prefix), its ALIVE2_X_RESP with the 32-bit creation, and one PORT_PLEASE2/PORT2_RESP pair. Note the creation value, restart nodea, re-capture, and show the creation changed.
Deliverable: the annotated pcap plus epmd_raw.exs checked into the repo. That pcap is a candidate stage asset.
Exit questions
Section titled “Exit questions”- A node registers with EPMD exactly once but EPMD always knows instantly when it dies. What is the mechanism, and why does it need no timeout or heartbeat?
- Walk the full path of
Node.connect(:"b@berlin")up to (not including) the handshake: which component resolves the host, which resolves the port, which opcode is used, and who talks to whom over which port? - Your pod runs one release per container and the security team wants port 4369 closed everywhere. Name two stock-OTP configurations that get you a working cluster with no EPMD daemon, and the main constraint each imposes.
Say it out loud
Section titled “Say it out loud”In sixty seconds, to the Haarlem room: explain why there is no such thing as unregistering from EPMD, starting from the fact that the registration is the open TCP connection, and end on what this says about BEAM design taste: the smallest mechanism that is still correct, borrowed from the OS instead of reinvented. Try to work the departure board in without straining the metaphor.