Registry
Connected nodes can call any function on each other. The registry is what tells a node which node has the function it wants, in a compatible version.
NL.Cluster.Registry (netzlive.cluster/lib/nl/cluster/registry.ex) is a
Phoenix.Tracker. Each node
runs one, registered as NL.Cluster.V3.Registry. Trackers gossip over
NL.Cluster.V3.PubSub and converge on the same set of entries — a CRDT, no leader. Every
node answers lookups from its own copy; a lookup never leaves the node.
What an entry is
Section titled “What an entry is”Every entry sits under the topic "netzlive.services" (:18), keyed by service name, and
carries:
%{service: %NL.Cluster.Service{...}, node: node()}(:62-67). The struct is built at compile time from the use NL.Cluster.Service options
and the defapi declarations (netzlive.cluster/lib/nl/cluster/service.ex:223-240). A real one, read off a
compiled redispatch build:
%NL.Cluster.Service{ name: Redispatch.Service.MasterDataV2, version: "2.2.0", compat: :semver, node: nil, functions: [ {{Redispatch.Service.MasterDataV2, :all_sr_ids_by_source, 2}, [retry: [limit: 210000, after: 10000], timeout: 60000]}, {{Redispatch.Service.MasterDataV2, :get_sr, 2}, [retry: [limit: 210000, after: 10000], timeout: 60000]} ], __cluster_version__: %Version{major: 2, minor: 10, patch: 1}}The source declares retry: [after: 1000, limit: 10] on use; the struct shows the library
defaults instead. That is a library bug, described under
Dienste und Clients.
Worth noticing:
functionscarries the call options. Timeout and retry policy travel with the service description. Clients do not hard-code them; they read them from the registry entry on every call. A provider can change its retry policy and every client picks it up on the provider’s next deploy.- The MFA is the provider’s module. A client never needs the provider’s code. It sends
{Redispatch.Service.MasterDataV2, :get_sr, args}to the node, where that module exists. __cluster_version__is the library version the provider was compiled against (netzlive.cluster/lib/nl/cluster/service.ex:82-84). Nothing reads it today; it is there for future compatibility checks.nodeisnilin the compiled struct. The registry stores the node next to the struct, andlookup/3puts it in on the way out.
Registration
Section titled “Registration”A service module’s child_spec/1 starts an NL.Cluster.Monitor, not the module itself
(netzlive.cluster/lib/nl/cluster/service.ex:214-219). The service has no process of its own — defapi
functions run in the RPC handler process on the provider node. The monitor exists to own the
registration:
init/1setstrap_exitand callsNL.Cluster.register(service, self())(netzlive.cluster/lib/nl/cluster/monitor.ex:48-61). The tracker links the entry to the monitor’s pid.terminate/2untracks it (:97-101). Because oftrap_exit, this runs on a normal supervisor shutdown too.- If the monitor crashes, the tracker drops entries owned by the dead pid. When the supervisor restarts it, it registers again.
So “the service is available” means exactly “its monitor is alive”. Nothing checks that the code behind the functions works.
How entries spread and expire
Section titled “How entries spread and expire”This is Phoenix.Tracker behaviour, not netzlive_cluster code. Defaults from
phoenix_pubsub 2.3.0 (Phoenix.Tracker.Shard.init/1); netzlive_cluster passes no
overrides outside its tests:
| Event | When the other nodes see it |
|---|---|
| A service registers | Next delta broadcast — broadcast_period, 1.5 s |
| Monitor untracks (service supervisor restarts it, app shuts down) | Next delta broadcast — if the tracker is still running to send it |
| Node crashes, is killed, or the network splits | After down_period: 1.5 s × 10 silent periods × 2 = 30 s |
| Node stays away | Its entries are purged after permdown_period, 20 min |
On a pod’s graceful shutdown the monitors untrack first, but the delta only goes out on the
tracker’s next heartbeat. If the node exits before that, the others fall back to the 30 s
down_period. Which of the two happens in practice was not observed for these pages. The
middle rows decide what clients experience during a deploy — see
Ein Aufruf.
Every join and leave is also re-published on the node’s local PubSub
(netzlive.cluster/lib/nl/cluster/registry.ex:35-49), on topic "netzlive.service.up" as
{:service_up, service} and "netzlive.service.down" as {:service_down, service}, with
an info log line (“Service … became available in version …”). Local processes can
subscribe to react to a dependency appearing.
Lookup
Section titled “Lookup”lookup/3 (:104-119):
- Fetch every entry under the service name.
- None →
{:error, :noservice}. - Keep entries whose version matches the requirement.
- None left →
{:error, :noversion}. - Otherwise take the highest version and return its struct with
nodefilled in.
The requirement comes from the client’s version and compat
(:174-181):
| Client declares | Requirement | Matches |
|---|---|---|
version: "2.1.0" (default compat: :semver) |
"~> 2.0" |
any 2.x — including 2.0.0, lower than asked |
version: "2.1.0", compat: "== 2.1.0" |
"== 2.1.0" |
exactly that |
compat: ">= 2.1.0 and < 3.0.0" |
as given | any Version requirement string |
The :semver default only looks at the major. A client that needs a function added in
2.1 can still be routed to a 2.0 provider, and gets :no_service_function (see
Ein Aufruf). Set compat: if that matters.
Inspecting it
Section titled “Inspecting it”From a remote console on any node:
NL.Cluster.Registry.services() # every service the node knowsNL.Cluster.Registry.services(:"name@host") # only those on one nodeNL.Cluster.lookup(Redispatch.Service.MasterDataV2, "2.0.0")Node.list() # who we are connected toservices/0 and services/1: netzlive.cluster/lib/nl/cluster/registry.ex:137-168. NL.Cluster.lookup/3
delegates to Registry.lookup/3 (netzlive.cluster/lib/nl/cluster.ex:44-45). The legacy registry has the
same functions under Netzlive.Cluster.Registry.