Skip to content

Fallstricke

Everything on the other pages of this section that can cost someone an afternoon, in one list, worst first. Each item links to where it is explained. Nothing here has been reported upstream by these pages — check the netzlive.cluster issues before filing.

1. Delegate-form defapi ignores use-level timeout/retry. netzlive.cluster/lib/nl/cluster/service.ex:365 validates per-function options with defaults, which then override the module options. Every NL service in NETZlive is affected and runs on 60 s / 10 s / 210 s regardless of what it declares. In 2.5.0 and 2.10.1. → Dienste und Clients, Dienstkarte

2. Timeouts are retried, so slow calls run several times. With the defaults, a call that keeps exceeding 60 s is started five times over about 5½ minutes. Non-idempotent functions (NL.UserManagement.EmailService.send_email/1) can execute more than once. → Ein Aufruf

3. update_external_user arguments are swapped between netzkoordinator’s client and user-management’s service. Not caught by tests because the double follows the provider. → Dienstkarte

4. A hard-coded Node.connect to a developer machine ships in the library. netzlive.cluster/lib/nl/cluster.ex:37-38:

after
Node.connect(:"right@xps-13-9310")
end

Runs on every application start, on every node, since commit 60e45a2 (2025-08-25, “feature: Adds config option for phoenix tracker”). In production the host does not resolve and the call returns false; the cost is a failed name lookup at boot. In an environment where xps-13-9310 resolves, nodes would try to join it. Looks like leftover debugging.

5. MS switch-state service delegates to a module that does not exist. netzkoordinator/lib/netzkoordinator/cluster_services/switch_state_projection_ms_v2.ex:54. → Dienstkarte

6. The startup: probe on services is never run. Accepted and validated, never scheduled (netzlive.cluster/lib/nl/cluster/monitor.ex:63-91). → Registry

Retries do not survive a node-name change. The node is looked up once, before the retry loop. Deployments (netzkoordinator, user-management) get new pod names, hence new node names, on every rollout; their clients see errors instead of a delay. → Ein Aufruf

:noservice is never retried. If the registry no longer lists the provider, the client returns an error at once.

No load balancing. All callers on a node go to the same replica of a service, and a crashed one can keep being chosen for up to 30 s. → Registry

:semver compat matches lower minors. A client declaring "2.1.0" accepts a 2.0.0 provider. → Registry

Legacy and NL services cannot see each other. Two registries. A :noservice that makes no sense often means the provider registered under the other namespace. → Overview

Bare :error from a provider arrives as {:ok, :error}. → Dienste und Clients

A per-call retry: replaces the provider’s whole retry list and falls back to the client’s compiled-in defaults for the keys it leaves out. → Ein Aufruf

Changing rpc_* or error_reporter config needs the dependency recompiled. → Cluster-Bildung

A Kubernetes 403 disconnects the node from everyone. Other API errors keep the last known peers. → Cluster-Bildung

Pods without any annotations are silently ignored by discovery. → Cluster-Bildung

Client stubs gain an arity. defdelegate_api f(x) defines f/1 and f/2. → Dienste und Clients