Skip to content

Registry

Connected nodes can call any function on each other. The registry is what tells a node which node has the function it wants, in a compatible version.

NL.Cluster.Registry (netzlive.cluster/lib/nl/cluster/registry.ex) is a Phoenix.Tracker. Each node runs one, registered as NL.Cluster.V3.Registry. Trackers gossip over NL.Cluster.V3.PubSub and converge on the same set of entries — a CRDT, no leader. Every node answers lookups from its own copy; a lookup never leaves the node.

Every entry sits under the topic "netzlive.services" (:18), keyed by service name, and carries:

%{service: %NL.Cluster.Service{...}, node: node()}

(:62-67). The struct is built at compile time from the use NL.Cluster.Service options and the defapi declarations (netzlive.cluster/lib/nl/cluster/service.ex:223-240). A real one, read off a compiled redispatch build:

%NL.Cluster.Service{
name: Redispatch.Service.MasterDataV2,
version: "2.2.0",
compat: :semver,
node: nil,
functions: [
{{Redispatch.Service.MasterDataV2, :all_sr_ids_by_source, 2},
[retry: [limit: 210000, after: 10000], timeout: 60000]},
{{Redispatch.Service.MasterDataV2, :get_sr, 2},
[retry: [limit: 210000, after: 10000], timeout: 60000]}
],
__cluster_version__: %Version{major: 2, minor: 10, patch: 1}
}

The source declares retry: [after: 1000, limit: 10] on use; the struct shows the library defaults instead. That is a library bug, described under Dienste und Clients. Worth noticing:

  • functions carries the call options. Timeout and retry policy travel with the service description. Clients do not hard-code them; they read them from the registry entry on every call. A provider can change its retry policy and every client picks it up on the provider’s next deploy.
  • The MFA is the provider’s module. A client never needs the provider’s code. It sends {Redispatch.Service.MasterDataV2, :get_sr, args} to the node, where that module exists.
  • __cluster_version__ is the library version the provider was compiled against (netzlive.cluster/lib/nl/cluster/service.ex:82-84). Nothing reads it today; it is there for future compatibility checks.
  • node is nil in the compiled struct. The registry stores the node next to the struct, and lookup/3 puts it in on the way out.

A service module’s child_spec/1 starts an NL.Cluster.Monitor, not the module itself (netzlive.cluster/lib/nl/cluster/service.ex:214-219). The service has no process of its own — defapi functions run in the RPC handler process on the provider node. The monitor exists to own the registration:

  • init/1 sets trap_exit and calls NL.Cluster.register(service, self()) (netzlive.cluster/lib/nl/cluster/monitor.ex:48-61). The tracker links the entry to the monitor’s pid.
  • terminate/2 untracks it (:97-101). Because of trap_exit, this runs on a normal supervisor shutdown too.
  • If the monitor crashes, the tracker drops entries owned by the dead pid. When the supervisor restarts it, it registers again.

So “the service is available” means exactly “its monitor is alive”. Nothing checks that the code behind the functions works.

This is Phoenix.Tracker behaviour, not netzlive_cluster code. Defaults from phoenix_pubsub 2.3.0 (Phoenix.Tracker.Shard.init/1); netzlive_cluster passes no overrides outside its tests:

Event When the other nodes see it
A service registers Next delta broadcast — broadcast_period, 1.5 s
Monitor untracks (service supervisor restarts it, app shuts down) Next delta broadcast — if the tracker is still running to send it
Node crashes, is killed, or the network splits After down_period: 1.5 s × 10 silent periods × 2 = 30 s
Node stays away Its entries are purged after permdown_period, 20 min

On a pod’s graceful shutdown the monitors untrack first, but the delta only goes out on the tracker’s next heartbeat. If the node exits before that, the others fall back to the 30 s down_period. Which of the two happens in practice was not observed for these pages. The middle rows decide what clients experience during a deploy — see Ein Aufruf.

Every join and leave is also re-published on the node’s local PubSub (netzlive.cluster/lib/nl/cluster/registry.ex:35-49), on topic "netzlive.service.up" as {:service_up, service} and "netzlive.service.down" as {:service_down, service}, with an info log line (“Service … became available in version …”). Local processes can subscribe to react to a dependency appearing.

lookup/3 (:104-119):

  1. Fetch every entry under the service name.
  2. None → {:error, :noservice}.
  3. Keep entries whose version matches the requirement.
  4. None left → {:error, :noversion}.
  5. Otherwise take the highest version and return its struct with node filled in.

The requirement comes from the client’s version and compat (:174-181):

Client declares Requirement Matches
version: "2.1.0" (default compat: :semver) "~> 2.0" any 2.x — including 2.0.0, lower than asked
version: "2.1.0", compat: "== 2.1.0" "== 2.1.0" exactly that
compat: ">= 2.1.0 and < 3.0.0" as given any Version requirement string

The :semver default only looks at the major. A client that needs a function added in 2.1 can still be routed to a 2.0 provider, and gets :no_service_function (see Ein Aufruf). Set compat: if that matters.

From a remote console on any node:

NL.Cluster.Registry.services() # every service the node knows
NL.Cluster.Registry.services(:"name@host") # only those on one node
NL.Cluster.lookup(Redispatch.Service.MasterDataV2, "2.0.0")
Node.list() # who we are connected to

services/0 and services/1: netzlive.cluster/lib/nl/cluster/registry.ex:137-168. NL.Cluster.lookup/3 delegates to Registry.lookup/3 (netzlive.cluster/lib/nl/cluster.ex:44-45). The legacy registry has the same functions under Netzlive.Cluster.Registry.