Skip to content

Day 2: EPMD, the departure board

Yesterday was the thesis: distribution gives you reach but not agreement. Today is the first hop of reach, the part that happens before two BEAM nodes have exchanged a single term: the Erlang Port Mapper Daemon. This is also the section of the talk you already shot a video for, so today’s job is making sure the mechanics under the Tulln departure board analogy hold up, because this is the material where someone in the Haarlem audience will have read the same docs you did.

When a BEAM node starts distribution, its distribution listener binds an ephemeral port: it asks for port 0 and the kernel picks one. That is a deliberate choice. It means any number of nodes can run on one host without port coordination, and it means node identity lives in the name rather than the port. app@host is stable across restarts. The port is not.

That creates a rendezvous problem. a@berlin wants to connect to b@berlin. It can resolve berlin via DNS, but DNS gives it a host, not a port. Something on that host has to answer the question “what port is the node named b listening on right now?” That something is EPMD: a small standalone C daemon, one per host, listening on the one port everybody agrees on in advance, 4369.

Two properties matter and both are easy to get wrong on stage:

  1. EPMD is per host, not per cluster. There is no global registry. Each host’s EPMD knows only about the nodes on that host. Cluster-wide discovery, meaning which hosts exist at all, is a different problem, solved by DNS, libcluster, static config, whatever. This is also the cleanest way to position libcluster for the Kubernetes crowd: libcluster finds hosts, EPMD maps names to ports on each host. Two different layers solving two different problems, and they get conflated all the time.
  2. EPMD knows nothing about Erlang. It never sees a term, never checks a cookie, and holds no state beyond name, port, and a few bytes of metadata. It does no consensus, no authentication, and no replication. That minimalism is the design, and it previews the thesis: even the discovery layer only promises reach.

The canonical analogy in the Unix world is the Sun RPC portmapper (rpcbind, port 111): services on dynamic ports register with a fixed-port daemon, clients ask the daemon first and then talk to the service directly. EPMD is the same pattern with a four-decade pedigree. The departure board version is better for a mixed audience, and it maps precisely: the board (EPMD) tells you which platform (port) the train named “REX 3 to Krems” (node name) leaves from today, and tomorrow the same train may use a different platform. You never board the departure board itself. You walk to the platform, and the actual journey (the handshake, tomorrow’s topic) happens there.

A lifecycle detail worth knowing cold: erl -sname foo checks whether an EPMD is already running on the host and spawns one if not (-start_epmd true is the default). The daemon then outlives the node. Kill every node on a host and EPMD keeps running with an empty database. That surprises people who expect it to be part of the VM. It is a separate OS process, and epmd -names works from a plain shell with no BEAM anywhere in sight.

Everything below is from the current erl_dist_protocol reference. The framing rule: every request to EPMD is prefixed with a 2-byte big-endian length, and responses are not length-prefixed. That asymmetry bites you the first time you hand-roll a client, so it is worth saying explicitly in the talk if you demo raw bytes.

When a node starts distribution, erl_epmd (the kernel module acting as EPMD client) opens a TCP connection to localhost:4369 and sends:

Field Size Meaning
120 1 ALIVE2_REQ opcode
PortNo 2 the distribution listener’s port
NodeType 1 77 = normal node, 72 = hidden node
Protocol 1 0 = TCP/IPv4 (also used for IPv6 in practice)
HighestVersion 2 highest dist protocol version, 6 since OTP 23
LowestVersion 2 lowest accepted, 6 since OTP 25 (version 5 dropped)
Nlen 2 name length
NodeName Nlen the short part only: foo, not foo@host
Elen 2 extra length, 0 in practice
Extra Elen reserved

Three things to internalize:

  • The registered name has no host part. EPMD is host-local, so @host would be redundant. When you demo epmd -names you see name foo at port 50123, bare names.
  • NodeType 72 is how hidden nodes stay findable. Hidden nodes (-hidden, and anything with -dist_listen false, which implies hidden) still register with EPMD if they listen. Hidden-ness is about the mesh and Node.list/0 visibility, not about discovery. Observer, remsh, and debug tooling depend on this.
  • The version fields preview the handshake. OTP 25 raised LowestVersion to 6, which is why an OTP 25+ node refuses a pre-OTP-23 peer outright. That negotiation happens here, before any cookie is involved.

The response is where creations enter. Modern EPMD replies with ALIVE2_X_RESP:

<<118, Result::8, Creation::32>> # ALIVE2_X_RESP, OTP 23+
<<121, Result::8, Creation::16>> # legacy ALIVE2_RESP

Result == 0 is success and anything else is failure. The symptom at the Elixir level is the familiar the name foo@host seems to be in use by another Erlang node, which is EPMD returning a nonzero Result because the name is already in its table with a live registration socket.

A node unregisters by closing the TCP connection it registered on. No UNREGISTER opcode exists. The socket opened for ALIVE2_REQ stays open for the node’s entire lifetime, and the registration is exactly as alive as that socket. The name ALIVE2 is literal: the open connection is the liveness signal. Give this fact its own moment in the talk. Consequences:

  • kill -9 a beam process and the kernel sends FIN/RST on its sockets, so EPMD drops the registration within milliseconds. No timeout, no garbage collection, no stale-entry reaper needed. The OS does the cleanup for free.
  • There is no heartbeat between node and EPMD, because TCP connection state serves that purpose.
  • A SIGSTOPped node stays registered (socket still open), which is consistent: the node is not dead, just not scheduled. The same logic shows up at cluster level on Day 6 with ticks.
  • This is why epmd -stop Name barely works. Forcibly unregistering a live node is such an unnatural operation that EPMD refuses it unless started with -relaxed_command_check, and always refuses it from remote hosts.

Worth one sentence in the talk: instead of adding an application-level keepalive protocol, EPMD borrows TCP’s connection semantics, the smallest mechanism that is still correct.

When Node.connect(:"b@berlin") runs, :net_kernel resolves berlin (that is address_please, plain inet/DNS, not EPMD), then asks the EPMD on berlin:

<<Len::16, 122, "b">>

Success response (119 = PORT2_RESP):

<<119, 0, PortNo::16, NodeType::8, Protocol::8,
HighestVersion::16, LowestVersion::16,
Nlen::16, NodeName::binary-size(Nlen), Elen::16, Extra::binary>>

Failure is just <<119, Result>> with Result > 0. EPMD closes the connection after replying, so lookups are one-shot, unlike the long-lived registration socket. Your node then connects directly to PortNo and the handshake (Day 3) begins. EPMD is out of the picture after that. It is consulted once per connection attempt and never sees a message.

NAMES_REQ (110) returns EPMD’s own port as 4 bytes followed by human-readable lines, one name X at port Y per registered node. This single opcode powers epmd -names from the shell, :net_adm.names/0,1 and :erl_epmd.names/1 from Elixir. DUMP_REQ (100) adds file descriptors and shows old/unused slots. KILL_REQ (107) returns the ASCII bytes "OK" and kills the daemon, but only if the node database is empty or -relaxed_command_check was given, and like -stop it is refused from remote hosts: remote connections get query commands answered and nothing else.

Note that NAMES is answered for anyone who can reach port 4369. Day 10’s security section comes back to this: an internet-exposed EPMD is a free enumeration service for node names and distribution ports. The RCE still requires the cookie, but EPMD hands over the map. CVE-2022-24706 (CouchDB) started exactly this way.

This belongs in 00_epmd_wireshark as the “no magic” demo. No dependencies and no :net_kernel, just bytes on a socket:

defmodule EpmdRaw do
@moduledoc "Speak the EPMD protocol by hand. Requests: <<len::16, body>>. Responses: raw."
def names(host \\ ~c"localhost") do
{:ok, s} = :gen_tcp.connect(host, 4369, [:binary, active: false])
:ok = :gen_tcp.send(s, frame(<<110>>))
{:ok, <<_epmd_port::32, text::binary>>} = :gen_tcp.recv(s, 0)
IO.puts(text)
end
def port_please(name, host \\ ~c"localhost") do
{:ok, s} = :gen_tcp.connect(host, 4369, [:binary, active: false])
:ok = :gen_tcp.send(s, frame(<<122, name::binary>>))
case :gen_tcp.recv(s, 0) do
{:ok, <<119, 0, port::16, type, 0, hi::16, lo::16, nlen::16,
^name::binary-size(nlen), _::binary>>} ->
{:ok, %{port: port, hidden?: type == 72, versions: lo..hi}}
{:ok, <<119, err>>} ->
{:error, err}
end
end
defp frame(body), do: <<byte_size(body)::16, body::binary>>
end

Run iex --sname a, then from a plain iex (no name, no distribution) call EpmdRaw.port_please("a"). The demo shows the audience that the discovery layer of a BEAM cluster is simple enough to speak with :gen_tcp in a coffee break. Pair it with your Wireshark capture filter on tcp.port == 4369. Per your earlier findings, the EPMD side dissects reliably even on builds where the dist dissector is flaky, since 4369 is a fixed port and needs no Decode As.

Creation is the field everyone skims past in ALIVE2_X_RESP, and it matters more for correctness than that suggests. The problem it solves: pids escape their node. They sit in remote ETS tables, in :global’s registry, in messages queued on other nodes. Now b@berlin crashes and restarts with the same name. The old pids still reference b@berlin, and pid slots get reused, so an old pid could structurally collide with a live process in the new incarnation. Without extra information, a message aimed at a process that died with the old node could be delivered to an unrelated process in the new one.

The fix: every pid, port, and reference carries a creation alongside its node name, an integer identifying the incarnation of that node. The external term format makes this concrete: NEW_PID_EXT is node atom, 4 bytes of id, 4 bytes of serial, and 4 bytes of creation. Same name plus different creation equals different node identity, and a message addressed to a dead incarnation’s pid is dropped rather than misdelivered.

Where does the number come from? In the classic setup, from EPMD, in the ALIVE2 response. History matters here because the talk’s audience may know the old behavior:

  • Pre OTP 23: creation was 2 bits in the external format, so EPMD cycled through 1, 2, 3 as a name was re-registered. Three restarts and you wrap, so a sufficiently old pid could alias a new incarnation. A known, tolerated wart.
  • OTP 23 (“big creation”): the external format grew NEW_PID_EXT/NEWER_REFERENCE_EXT with 32-bit creations, EPMD grew ALIVE2_X_RESP to hand back a 32-bit value, and wrap-around stopped being a practical concern. Values 0 through 3 are kept out of the new range for legacy reasons, and 0 historically served as a wildcard “unset” creation (verify the exact reserved-range wording against erl_dist_protocol before putting numbers on a slide).

You can see all of it from IEx:

:erlang.system_info(:creation)
# => e.g. 1696401732 (32-bit, handed out at registration)
<<131, 88, _atom_prefix::binary>> = :erlang.term_to_binary(self())
# 88 = NEW_PID_EXT; the last 4 bytes of the pid encoding are the creation

Restart the node with the same --sname, print system_info(:creation) again, and you get a different number. That demo covers the whole concept in about ten seconds.

One subtlety for EPMD-less setups (next section): if nobody hands you a creation, you make your own. Custom epmd_module implementations return a creation from register_node, and nodes with dynamic names get theirs assigned by the peer during handshake. The invariant that matters is uniqueness across incarnations of the same name, not where the number comes from.

This is the part of today’s material most relevant to your Kubernetes-adjacent audience, and the part where blog posts age worst. Here is the current (OTP 27/28) state of the escape hatches, in increasing order of radicalism.

  1. -start_epmd false, keep EPMD. The node does not auto-spawn the daemon; you run it yourself (epmd -daemon in the container entrypoint, or a systemd unit). Nothing else changes. The node still registers and looks up as usual, and fails at startup if no daemon is reachable. Why bother: process supervision hygiene. The auto-spawned EPMD is owned by whichever node started it first, logs nowhere useful, and in containers you generally want PID 1 discipline rather than surprise daemons.

  2. Fix the distribution port, keep EPMD. -kernel inet_dist_listen_min 9100 inet_dist_listen_max 9100 pins the listener. EPMD still runs and still answers lookups, but your firewall rules become writable by humans. This is the pragmatic production default: one node per host or pod, one well-known port, EPMD reduced to a formality that costs one daemon. Most k8s Elixir deployments live here whether they know it or not.

  3. Replace the client: epmd_module. The node-side EPMD client is just a module (erl_epmd by default) implementing a small behavior: start_link/0, register_node/2,3 (return {ok, Creation} yourself), port_please/2,3 (return {port, Port, Version} from wherever you like), address_please/3, listen_port_please/2, names/1. Swap it with the kernel parameter epmd_module and EPMD the daemon never enters the picture. The classic recipe (the Erlang Solutions post, the epmdless library) is a module that derives the port from the node name or a static map. The config ergonomics improved recently: -epmd_module as an erl flag is now deprecated in favor of the kernel application parameter, which means config/runtime.exs or application:set_env(kernel, epmd_module, M) before net_kernel.start/2 works, no VM args needed. That change landed via OTP PR #8671, aimed at OTP 27.2 (verify the exact release before stage if you cite the version number).

  4. The official one-liner: erl_epmd_node_listen_port. Current kernel docs: configures the port erl_epmd uses both to listen and to assume when connecting to other nodes. Set it (the old -erl_epmd_port flag is the deprecated spelling) together with -start_epmd false and you get a fully EPMD-less node with stock OTP: registration becomes a no-op with a self-assigned creation, and lookups short-circuit to “the same port I use”. The built-in assumption is uniform ports across the cluster, which is exactly the k8s one-node-per-pod shape:

config/runtime.exs
# rel/env.sh.eex or vm args
# ELIXIR_ERL_OPTIONS="-start_epmd false"
config :kernel, erl_epmd_node_listen_port: 9100

(Flag the exact interplay of -start_epmd false plus the kernel param on your OTP patch level as “verify before stage”; run it in the companion repo rather than trusting my summary or anyone’s blog post.)

  1. Dynamic node names: -name undefined. OTP 23+ lets a node start with no name at all. It requests one from the first node it connects to, via the DFLAG_NAME_ME handshake flag, and gets its creation assigned by the peer too. This is what iex --remsh relies on when you give it no name (it defaults to -sname undefined). Such a node is a pure client, and needs no EPMD, no listener, and no registration, because nothing ever dials in to it. Worth thirty seconds in the talk because remsh is the distribution feature everyone uses daily without knowing it is distribution.

What you give up without EPMD: epmd -names as a debugging tool, :net_adm.names/1, and the ability to run several nodes per host without port coordination, i.e. precisely the dev-machine conveniences. Hence the sane default posture: EPMD-less or fixed-port in production where ports are policy, stock EPMD on laptops where convenience wins. A good interop footnote: because the protocol is this simple, non-BEAM dist implementations (Go’s ergo, Python’s pyrlang) implement EPMD’s client side in a few dozen lines, which is decent evidence that there is nothing VM-specific or magical about it.

For a 40-minute talk, EPMD is realistically 6 to 8 minutes. The arc that fits: the departure board analogy (you have the video), one Wireshark or EpmdRaw moment showing ALIVE2 and PORT_PLEASE2 are a handful of bytes, the fact that deregistration is just a TCP close, creation as the hidden incarnation stamp in every pid, and one slide on EPMD-less setups for the k8s crowd. The thesis thread to pull: EPMD is the simplest possible form of reach, a name, a port, and an open socket. It promises you can find a node. It says nothing about whether the node agrees with anyone about anything, and it does not promise the node will still be there when you connect. Every layer above it inherits the same limits.

Trap questions Haarlem might ask, worth having answers ready: “Why not just use DNS SRV records instead of EPMD?” (you can, that is essentially what a custom epmd_module does; EPMD exists so multiple nodes per host with dynamic ports work with zero config), and “Does EPMD see my messages / my cookie?” (never; it is out of the loop after PORT_PLEASE2, and the cookie belongs to tomorrow’s handshake).

Active block (30 to 45 min, repo section 00_epmd_wireshark)

Section titled “Active block (30 to 45 min, repo section 00_epmd_wireshark)”

Build EpmdRaw from today’s snippet into 00_epmd_wireshark/epmd_raw.exs and script this sequence in the justfile:

  1. Start tcpdump -i lo0 -w epmd.pcap port 4369 (or Wireshark on loopback, filter tcp.port == 4369).
  2. iex --sname a, then iex --sname b, then from b: Node.connect(:a@yourhost).
  3. From a third, undistributed iex: EpmdRaw.names() and EpmdRaw.port_please("a").
  4. kill -9 node a’s beam process; immediately run EpmdRaw.names() again and confirm the registration vanished with no explicit deregistration message, then find the bare FIN in the capture.
  5. In the pcap, annotate one ALIVE2_REQ (find the 120 opcode after the length prefix), its ALIVE2_X_RESP with the 32-bit creation, and one PORT_PLEASE2/PORT2_RESP pair. Note the creation value, restart node a, re-capture, and show the creation changed.

Deliverable: the annotated pcap plus epmd_raw.exs checked into the repo. That pcap is a candidate stage asset.

  1. A node registers with EPMD exactly once but EPMD always knows instantly when it dies. What is the mechanism, and why does it need no timeout or heartbeat?
  2. Walk the full path of Node.connect(:"b@berlin") up to (not including) the handshake: which component resolves the host, which resolves the port, which opcode is used, and who talks to whom over which port?
  3. Your pod runs one release per container and the security team wants port 4369 closed everywhere. Name two stock-OTP configurations that get you a working cluster with no EPMD daemon, and the main constraint each imposes.

In sixty seconds, to the Haarlem room: explain why there is no such thing as unregistering from EPMD, starting from the fact that the registration is the open TCP connection, and end on what this says about BEAM design taste: the smallest mechanism that is still correct, borrowed from the OS instead of reinvented. Try to work the departure board in without straining the metaphor.