Fallstricke
Everything on the other pages of this section that can cost someone an afternoon, in one
list, worst first. Each item links to where it is explained. Nothing here has been
reported upstream by these pages — check the netzlive.cluster issues before filing.
Defects
Section titled “Defects”1. Delegate-form defapi ignores use-level timeout/retry.
netzlive.cluster/lib/nl/cluster/service.ex:365 validates per-function options with defaults, which then
override the module options. Every NL service in NETZlive is affected and runs on
60 s / 10 s / 210 s regardless of what it declares. In 2.5.0 and 2.10.1.
→ Dienste und Clients,
Dienstkarte
2. Timeouts are retried, so slow calls run several times. With the defaults, a call that
keeps exceeding 60 s is started five times over about 5½ minutes. Non-idempotent functions
(NL.UserManagement.EmailService.send_email/1) can execute more than once.
→ Ein Aufruf
3. update_external_user arguments are swapped between netzkoordinator’s client and
user-management’s service. Not caught by tests because the double follows the provider.
→ Dienstkarte
4. A hard-coded Node.connect to a developer machine ships in the library.
netzlive.cluster/lib/nl/cluster.ex:37-38:
after Node.connect(:"right@xps-13-9310") endRuns on every application start, on every node, since commit 60e45a2 (2025-08-25,
“feature: Adds config option for phoenix tracker”). In production the host does not
resolve and the call returns false; the cost is a failed name lookup at boot. In an
environment where xps-13-9310 resolves, nodes would try to join it. Looks like leftover
debugging.
5. MS switch-state service delegates to a module that does not exist.
netzkoordinator/lib/netzkoordinator/cluster_services/switch_state_projection_ms_v2.ex:54.
→ Dienstkarte
6. The startup: probe on services is never run. Accepted and validated, never
scheduled (netzlive.cluster/lib/nl/cluster/monitor.ex:63-91).
→ Registry
Behaviour that surprises
Section titled “Behaviour that surprises”Retries do not survive a node-name change. The node is looked up once, before the retry loop. Deployments (netzkoordinator, user-management) get new pod names, hence new node names, on every rollout; their clients see errors instead of a delay. → Ein Aufruf
:noservice is never retried. If the registry no longer lists the provider, the client
returns an error at once.
No load balancing. All callers on a node go to the same replica of a service, and a crashed one can keep being chosen for up to 30 s. → Registry
:semver compat matches lower minors. A client declaring "2.1.0" accepts a 2.0.0
provider.
→ Registry
Legacy and NL services cannot see each other. Two registries. A :noservice that makes
no sense often means the provider registered under the other namespace.
→ Overview
Bare :error from a provider arrives as {:ok, :error}.
→ Dienste und Clients
A per-call retry: replaces the provider’s whole retry list and falls back to the
client’s compiled-in defaults for the keys it leaves out.
→ Ein Aufruf
Changing rpc_* or error_reporter config needs the dependency recompiled.
→ Cluster-Bildung
A Kubernetes 403 disconnects the node from everyone. Other API errors keep the last known peers. → Cluster-Bildung
Pods without any annotations are silently ignored by discovery. → Cluster-Bildung
Client stubs gain an arity. defdelegate_api f(x) defines f/1 and f/2.
→ Dienste und Clients