Skip to content

Ein Aufruf

What happens between Client.search(query) on one node and the return value, in order, with the place in the code for each step.

1. Client stub — generated by defdelegate_api. Calls NL.Cluster.Client.call(@service, :search, [query], call_opts), where @service is a struct holding only name, version and compat (netzlive.cluster/lib/nl/cluster/client.ex:217, :233-239).

2. Lookup — the struct has node: nil, so call/4 asks the local registry (netzlive.cluster/lib/nl/cluster/client.ex:285-299). On :noservice / :noversion it returns {:error, %ServiceError{}} immediately. On success it continues with the provider’s struct: node, function list, call defaults.

3. Find the function — Service.fetch_mfa_and_opts/3 looks for name and arity in the provider’s functions (netzlive.cluster/lib/nl/cluster/service.ex:143-162). Missing → {:error, %ServiceError{reason: :no_service_function}}, returned.

4. Merge options — Keyword.merge(service_call_opts, call_opts) (netzlive.cluster/lib/nl/cluster/client.ex:304): the provider’s defaults, overridden by whatever the caller passed. Unknown keys are dropped (:356-363).

5. RPC with retries — NL.Cluster.RPC.call/5 (netzlive.cluster/lib/nl/cluster/rpc.ex:133-146) wraps :erpc.call(node, module, fun, args, timeout) (:155) in a constant-backoff retry loop (retry library).

6. Provider side — :erpc spawns a process on the provider node that runs the defapi-generated function: user code in try, result classified and wrapped (netzlive.cluster/lib/nl/cluster/service.ex:418-452). See Dienste und Clients.

7. Classify on the way back — do_call/5 (netzlive.cluster/lib/nl/cluster/rpc.ex:151-197):

  • :erpc raised {:erpc, reason} → transport failure → {:error, %RPCError{reason: reason}}
  • the remote function raised past the defapi wrapper (should not happen) → {:ok, {:error, %ServiceError{}}}
  • otherwise → {:ok, result}

8. Unwrap — maybe_unwrap_rpc_result / maybe_unwrap_service_result (netzlive.cluster/lib/nl/cluster/client.ex:365-371) raise on transport errors and provider crashes, and flatten everything else to the shapes in the result table. Errors get service and service_function filled in for the message (:373-387).

9. unwrap: and do block, if declared.

Unless the provider or caller says otherwise:

Value Source
timeout per attempt 60 s rpc_timeout, netzlive.cluster/lib/nl/cluster/config.ex:24-28
pause between attempts 10 s rpc_retry_after, :29-36
total budget 210 s rpc_retry_limit, :37-44 — 3 * (1 min + 10 s)

The budget looks sized for three full-length attempts. It is not what happens. The retry library (0.19.0) starts the budget clock when it asks for the first delay — after attempt 1 has already failed — and always grants one last attempt when the budget runs out (Retry.DelayStreams.expiry/3). Measured with the library at 1/10 scale (6 s timeout, 1 s pause, 21 s budget), multiplied back:

Attempt Starts at
1 0 s times out at 60 s; budget clock starts now, ends at 270 s
2 70 s
3 140 s
4 210 s
5 271 s “one last try”
— 331 s RPCError{reason: :timeout} raised

So with defaults a call that keeps timing out runs five times and blocks the caller for about 5½ minutes. A call that fails instantly (:noconnection, node gone) is attempted every 10 s for about 210 s — about twenty attempts.

retry?/1 (netzlive.cluster/lib/nl/cluster/rpc.ex:148-149): any RPCError except reasons :badarg and :notsup. In practice that is :timeout, :noconnection and :system_limit.

Not retried:

  • lookup failures (:noservice, :noversion) and :no_service_function — they happen before the retry loop starts
  • anything the provider returned, including {:error, _}
  • provider crashes (ServiceError)

Provider (preferred — the provider knows how long its restarts and its work take):

use NL.Cluster.Service, name: …, version: …, timeout: 5_000, retry: [after: 2_000, limit: 30_000]
defapi send_mail(to, body), to: Mailer, retry: false

Caller, per call:

Client.search(q, timeout: 120_000)
Client.search(q, retry: false)

The retry loop is advertised as bridging “short downtimes of the remote node, such as a restart or redeploy” (README). Whether it does depends on how the provider goes away, because the node is resolved once, in step 2, before the loop — every retry goes to the same node name.

Provider… Registry on the client Client sees
is replaced by a pod with the same node name, within the budget still lists the old entry, or the new one retries hit :noconnection until the new pod is up, then succeed ✓
is replaced by a pod with a different node name (IP-based or pod-name-based node names, Deployment rollouts) old entry until it expires retries go to the dead name until the budget runs out → RPCError after ~210 s
has already been removed from the registry (untrack delivered, or 30 s down_period passed) no entry {:error, %ServiceError{reason: :noservice}} at once, no retry
is one of several replicas, one dies dead one may still be picked for up to 30 s as row 2 for calls routed to it

Only the first row is the case the retry was written for. Pod names in a Deployment change on every rollout, and both node-naming schemes in the README put the pod name before the @. So for Deployments, a node name usually does not survive a redeploy; for StatefulSets (app-0, app-1) it does. How each application is deployed is on the Dienstkarte.

This table is derived from the code, not observed in production. A client that must survive a provider rollout should treat :noservice as transient and handle it — the library will not.

Both exception types render a long, structured message (netzlive.cluster/lib/nl/cluster/errors/rpc_error.ex:37-60, netzlive.cluster/lib/nl/cluster/errors/service_error.ex:49-72):

The RPC call resulted in an error!
The following service call caused the error:
NMK.Service.Search.search/1
The reason was:
:noconnection
The connection to the node was lost or could not be established. The function
may or may not be applied.
The service:
Name: NMK.Service.Search
Node: :"nmk-7d9c…@10.1.2.3"
Functions:
- …

The Node: line tells you which node the registry picked — the first thing to check against Node.list() when a call fails.

Reason Kind Means
:noservice returned nobody registered that name — provider down, not yet started, or a different namespace (Netzlive.Cluster vs NL.Cluster)
:noversion returned registered, but not in a major the client accepts
:no_service_function returned provider version lacks that function/arity
:noconnection raised RPCError node not connected, after retries
:timeout raised RPCError every attempt ran into the timeout
:unhandled_exception / _exit / _throw raised ServiceError provider code crashed; it was reported on the provider’s error_reporter