Ein Aufruf
What happens between Client.search(query) on one node and the return value, in order,
with the place in the code for each step.
Step by step
Section titled “Step by step”1. Client stub — generated by defdelegate_api. Calls
NL.Cluster.Client.call(@service, :search, [query], call_opts), where @service is a
struct holding only name, version and compat (netzlive.cluster/lib/nl/cluster/client.ex:217,
:233-239).
2. Lookup — the struct has node: nil, so call/4 asks the local registry
(netzlive.cluster/lib/nl/cluster/client.ex:285-299). On :noservice / :noversion it returns
{:error, %ServiceError{}} immediately. On success it continues with the provider’s
struct: node, function list, call defaults.
3. Find the function — Service.fetch_mfa_and_opts/3 looks for name and arity in the
provider’s functions (netzlive.cluster/lib/nl/cluster/service.ex:143-162). Missing →
{:error, %ServiceError{reason: :no_service_function}}, returned.
4. Merge options — Keyword.merge(service_call_opts, call_opts)
(netzlive.cluster/lib/nl/cluster/client.ex:304): the provider’s defaults, overridden by whatever the caller
passed. Unknown keys are dropped (:356-363).
5. RPC with retries — NL.Cluster.RPC.call/5 (netzlive.cluster/lib/nl/cluster/rpc.ex:133-146) wraps
:erpc.call(node, module, fun, args, timeout) (:155) in a constant-backoff retry loop
(retry library).
6. Provider side — :erpc spawns a process on the provider node that runs the
defapi-generated function: user code in try, result classified and wrapped
(netzlive.cluster/lib/nl/cluster/service.ex:418-452). See
Dienste und Clients.
7. Classify on the way back — do_call/5 (netzlive.cluster/lib/nl/cluster/rpc.ex:151-197):
:erpcraised{:erpc, reason}→ transport failure →{:error, %RPCError{reason: reason}}- the remote function raised past the
defapiwrapper (should not happen) →{:ok, {:error, %ServiceError{}}} - otherwise →
{:ok, result}
8. Unwrap — maybe_unwrap_rpc_result / maybe_unwrap_service_result
(netzlive.cluster/lib/nl/cluster/client.ex:365-371) raise on transport errors and provider crashes, and
flatten everything else to the shapes in
the result table.
Errors get service and service_function filled in for the message (:373-387).
9. unwrap: and do block, if declared.
Timeouts und Retries
Section titled “Timeouts und Retries”Defaults
Section titled “Defaults”Unless the provider or caller says otherwise:
| Value | Source | |
|---|---|---|
| timeout per attempt | 60 s | rpc_timeout, netzlive.cluster/lib/nl/cluster/config.ex:24-28 |
| pause between attempts | 10 s | rpc_retry_after, :29-36 |
| total budget | 210 s | rpc_retry_limit, :37-44 — 3 * (1 min + 10 s) |
The budget looks sized for three full-length attempts. It is not what happens. The retry
library (0.19.0) starts the budget clock when it asks for the first delay — after
attempt 1 has already failed — and always grants one last attempt when the budget runs out
(Retry.DelayStreams.expiry/3). Measured with the library at 1/10 scale (6 s timeout, 1 s
pause, 21 s budget), multiplied back:
| Attempt | Starts at | |
|---|---|---|
| 1 | 0 s | times out at 60 s; budget clock starts now, ends at 270 s |
| 2 | 70 s | |
| 3 | 140 s | |
| 4 | 210 s | |
| 5 | 271 s | “one last try” |
| — | 331 s | RPCError{reason: :timeout} raised |
So with defaults a call that keeps timing out runs five times and blocks the caller for
about 5½ minutes. A call that fails instantly (:noconnection, node gone) is attempted
every 10 s for about 210 s — about twenty attempts.
What gets retried
Section titled “What gets retried”retry?/1 (netzlive.cluster/lib/nl/cluster/rpc.ex:148-149): any RPCError except reasons :badarg and
:notsup. In practice that is :timeout, :noconnection and :system_limit.
Not retried:
- lookup failures (
:noservice,:noversion) and:no_service_function— they happen before the retry loop starts - anything the provider returned, including
{:error, _} - provider crashes (
ServiceError)
Overriding
Section titled “Overriding”Provider (preferred — the provider knows how long its restarts and its work take):
use NL.Cluster.Service, name: …, version: …, timeout: 5_000, retry: [after: 2_000, limit: 30_000]defapi send_mail(to, body), to: Mailer, retry: falseCaller, per call:
Client.search(q, timeout: 120_000)Client.search(q, retry: false)Was bei einem Deploy passiert
Section titled “Was bei einem Deploy passiert”The retry loop is advertised as bridging “short downtimes of the remote node, such as a restart or redeploy” (README). Whether it does depends on how the provider goes away, because the node is resolved once, in step 2, before the loop — every retry goes to the same node name.
| Provider… | Registry on the client | Client sees |
|---|---|---|
| is replaced by a pod with the same node name, within the budget | still lists the old entry, or the new one | retries hit :noconnection until the new pod is up, then succeed ✓ |
| is replaced by a pod with a different node name (IP-based or pod-name-based node names, Deployment rollouts) | old entry until it expires | retries go to the dead name until the budget runs out → RPCError after ~210 s |
has already been removed from the registry (untrack delivered, or 30 s down_period passed) |
no entry | {:error, %ServiceError{reason: :noservice}} at once, no retry |
| is one of several replicas, one dies | dead one may still be picked for up to 30 s | as row 2 for calls routed to it |
Only the first row is the case the retry was written for. Pod names in a Deployment change
on every rollout, and both node-naming schemes in the
README put the pod name before the @. So
for Deployments, a node name usually does not survive a redeploy; for StatefulSets
(app-0, app-1) it does. How each application is deployed is on the
Dienstkarte.
This table is derived from the code, not observed in production. A client that must
survive a provider rollout should treat :noservice as transient and handle it — the
library will not.
Reading the errors
Section titled “Reading the errors”Both exception types render a long, structured message (netzlive.cluster/lib/nl/cluster/errors/rpc_error.ex:37-60,
netzlive.cluster/lib/nl/cluster/errors/service_error.ex:49-72):
The RPC call resulted in an error!
The following service call caused the error: NMK.Service.Search.search/1
The reason was: :noconnection
The connection to the node was lost or could not be established. The function may or may not be applied.
The service: Name: NMK.Service.Search Node: :"nmk-7d9c…@10.1.2.3" Functions: - …The Node: line tells you which node the registry picked — the first thing to check
against Node.list() when a call fails.
| Reason | Kind | Means |
|---|---|---|
:noservice |
returned | nobody registered that name — provider down, not yet started, or a different namespace (Netzlive.Cluster vs NL.Cluster) |
:noversion |
returned | registered, but not in a major the client accepts |
:no_service_function |
returned | provider version lacks that function/arity |
:noconnection |
raised RPCError |
node not connected, after retries |
:timeout |
raised RPCError |
every attempt ran into the timeout |
:unhandled_exception / _exit / _throw |
raised ServiceError |
provider code crashed; it was reported on the provider’s error_reporter |