Errors, retries, and idempotency

Ormuz separates three problems: automatically replaying a temporarily unavailable operation, creating an operator continuation after a terminal instance, and preventing duplication of the same trigger or external effect. These mechanisms have different guarantees and must not be conflated.

Automatic node retry

Node error
Replayable and safe?
Retry timer waiting
Terminal failure
New attempt of the same node

An error becomes a durable retry only when it is recognized as transient and the operation is safe to replay. Otherwise the node fails immediately.

When a node fails, Ormuz classifies the error before any new attempt. An explicitly replayable error may become a durable wait: the node remains waiting, a retry deadline is recorded, and the same logical invocation resumes later.

An unknown error is not assumed transient. It becomes terminal unless the capability that produced it supplies an explicit signal letting the runtime know replay is safe.

Retry budget

The effective budget starts from platform policy, can be configured at process level, and can then only be tightened by a node. A node cannot grant itself more time or attempts than the containing process.

The budget combines an attempt ceiling and a total deadline. Delays between attempts are bounded and may honor a Retry-After value supplied by the external system. When the deadline or ceiling is exhausted, the step becomes failed with the exhaustion indication and the instance may reach retry_exhausted.

A durable retry is a wait, not a busy loop

Between attempts, the instance is waiting. Backoff time is part of the observable resume contract; you do not need to add a Wait node around an operation already covered by this policy.

Replay safety

An idempotent HTTP read can be replayed without creating a new effect. A network write may be automatically replayed only when Ormuz can establish that it was not sent, when the call is protected by an idempotency key, or when the capability itself explicitly declares safe retry.

This rule avoids a classic trap: turning an ambiguous timeout into a double payment, duplicate email, or duplicate provider creation. A network error alone is therefore not authorization to replay.

Retryable ≠ side-effect free

Safety comes from the replay contract, not the word “transient”. For a write operation, check the idempotency strategy documented by the relevant node or extension.

Three different operator continuations

ActionAllowed sourceRevision usedContinuity
retryfailed / retry_exhausted / stoppedCurrent Working revisionSame input; fresh execution state; process-owned resource continuity transfers when still eligible.
rerunTerminal instanceCurrent Working revisionSame input, new environment snapshot, and independent execution; does not take over business ownership from the source.
forkInstance containing the requested checkpointSource-instance revisionRestarts around a selected node by reconstructing coherent state from the source audit.

All three operations create a new instance. They do not put the old instance back into running and do not rewrite its audit.

Manual retry

POST /v1/process-instances/:id/retry is available for an instance in failed, retry_exhausted, or stopped. The new instance starts with the same input but uses the definition's current Working revision, which may differ from the revision that failed.

HTTP
POST /v1/process-instances/pci_123/retry

Participants and the protection attached to the input are preserved. For a business resource whose handling was owned by the process, ownership may transfer to the new attempt; an explicitly stopped instance does not automatically regain ownership during retry.

Manual retry ≠ exact historical reproduction

If the Working definition changed since the source instance, the retry executes that new snapshot. Use a fork when you must remain on the source instance's historical revision.

Run again

POST /v1/process-instances/:id/rerun creates an independent execution from the same input on the current Working revision. It is available for terminal instances, including successful executions.

HTTP
POST /v1/process-instances/pci_123/rerun

The rerun starts with fresh node and variable state, rebuilds its environment snapshot, and does not become the business successor of the source's process-owned resource. Use it when the need is “run this process again today”, not “repair this attempt”.

Targeted fork

POST /v1/process-instances/:id/fork creates a continuation pinned to the source instance's revision. It reconstructs coherent state around a selected node and keeps fork_of to link the two instances.

HTTP
POST /v1/process-instances/pci_123/fork
{3 items
"action":"retry_step"
"node_id":"send_to_provider"
"reason":"Replay the call after external correction"
}
{
"action": "retry_step",
"node_id": "send_to_provider",
"reason": "Replay the call after external correction"
}

action may be retry_step, skip_step, or resume_after_step. The last mode requires a completed step that can serve as a checkpoint: Ormuz restores the audited state after that step before continuing.

A fork does not allow arbitrary history changes. It creates a new lineage branch with explicit provenance and leaves the source untouched.

Lineage, superseded state, and business ownership

The retry_of, fork_of, and start_source fields make it possible to reconstruct the relationship between instances. A corrective continuation may mark an earlier attempt as superseded when the new instance becomes the one continuing the handling.

For process-owned resources, only one instance should drive current handling. Continuations intended to repair that execution may transfer this role; an independent rerun does not acquire it. Use Observability to follow the lineage.

Trigger idempotency

Idempotency of a provider operation and deduplication of a process instance are two different guarantees. An event launcher automatically deduplicates the same event for the same launcher so that repeated delivery does not create several identical instances.

A direct POST /v1/processes/:id/runs call represents a new execution request each time. If your integration needs explicitly deduplicated instance creation, use the instance contract exposing dedupe_key rather than inventing a key in the /runs payload.