Skip to content

Commit 79e4c04

Browse files
authored
fix(core): re-route runs delivered to the wrong deployment (#2960)
## Summary & Motivation A queue callback that reaches a deployment other than the one its run is pinned to derives the per-run encryption key from the wrong master key, so the delivery fails before user code runs and the run dies as a blank "exceeded max retries". The delivery is re-enqueued explicitly addressed to the run's own deployment — strictly better-targeted than the send that misrouted — and the run is failed with the new `DEPLOYMENT_MISMATCH` error code only once `WORKFLOW_DEPLOYMENT_MISMATCH_MAX_RETRIES` (default 3) is spent. Gated on the new World capability `deploymentAffinity`, so worlds with synthetic or version-tagged deployment ids are unaffected. ## Test Plan Unit tests added for the guard and both runtime paths; local vitest and typechecks pass.
1 parent 8d47928 commit 79e4c04

21 files changed

Lines changed: 1243 additions & 24 deletions
Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,8 @@
1+
---
2+
"@workflow/core": patch
3+
"@workflow/errors": minor
4+
"@workflow/world": minor
5+
"@workflow/world-vercel": patch
6+
---
7+
8+
Re-route a queue delivery that reaches a deployment other than the one its run is pinned to, retry transient publishing failures through normal queue redelivery, and fail the run with the new `DEPLOYMENT_MISMATCH` error code once the recovery budget is spent.
Lines changed: 71 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,71 @@
1+
---
2+
title: deployment-mismatch
3+
description: A workflow run was delivered to a deployment other than the one it is pinned to.
4+
type: troubleshooting
5+
summary: Understand how Workflow recovers from a misrouted delivery, and why a run eventually fails with DEPLOYMENT_MISMATCH.
6+
prerequisites:
7+
- /docs/foundations/workflows-and-steps
8+
related:
9+
- /docs/foundations/versioning
10+
- /docs/errors/runtime-decryption-failed
11+
- /docs/foundations/errors-and-retries
12+
---
13+
14+
Every run is pinned to a single deployment when it starts. When a queued workflow or step callback is delivered to a **different** deployment, Workflow does not execute it there. Instead it re-routes the message to the deployment the run is pinned to, and only if the run keeps arriving elsewhere does it fail with the `DEPLOYMENT_MISMATCH` classification.
15+
16+
This is an SDK/runtime signal, not an error thrown by your workflow code, and it is not catchable inside a workflow function.
17+
18+
## Error Message
19+
20+
```
21+
Workflow run "wrun_..." is pinned to deployment "dpl_A", but was received by deployment "dpl_B". The runtime re-routed the message to "dpl_A" 3 times and it kept arriving elsewhere, so the run was stopped to protect against code-skew errors. Verify that the run's deployment is still available and that queue callbacks are routed to it.
22+
```
23+
24+
When the queue definitively reports that the run's deployment cannot be reached — it was deleted, or aged out of its retention window — no re-route is possible and the message omits the re-routing clause. Transient or unknown publishing failures leave the current delivery unacknowledged so the queue can redeliver it; they do not fail the run or consume this recovery budget.
25+
26+
## Why A Run Is Pinned
27+
28+
A run's deployment is chosen once, at [`start()`](/docs/api-reference/workflow-api/start):
29+
30+
- By default it is the deployment that called `start()` — see [Versioning](/docs/foundations/versioning) for why runs are pinned this way.
31+
- With `start(workflow, args, { deploymentId })` it is the id you pass, so a run can deliberately target a deployment other than the one that created it.
32+
- With `deploymentId: "latest"` it is the most recent deployment for the current environment, resolved at start time.
33+
34+
Whichever it is, that `deploymentId` is recorded on the run, and every subsequent workflow replay and step execution must happen on that deployment. Continuing on a different one is unsafe:
35+
36+
1. **Code skew.** The workflow and step bundles on the receiving deployment may not match the code that produced the run's recorded history, so replay could diverge or produce incorrect results.
37+
2. **Encryption.** Step inputs and other event-log payloads are encrypted with a per-run key derived from the pinned deployment's key material. A different deployment derives the wrong key and cannot decrypt them — previously the source of a confusing [runtime-decryption-failed](/docs/errors/runtime-decryption-failed) that exhausted retries with no clear cause.
38+
39+
So the runtime checks the pinned deployment before it executes anything, and `DEPLOYMENT_MISMATCH` names the result — instead of the mismatch surfacing later as an unrelated decryption failure.
40+
41+
## Automatic Recovery
42+
43+
A deployment that receives a run it does not own first tries to fix the delivery rather than fail the run:
44+
45+
1. It re-enqueues the message **explicitly addressed** to the run's own deployment. This is strictly better-addressed than the send that misrouted, which inherited the producing deployment's ambient id.
46+
2. Delivery is delayed with a short exponential backoff (1s, 2s, 4s).
47+
3. If the run keeps arriving at the wrong deployment, the run is failed with `DEPLOYMENT_MISMATCH` after `WORKFLOW_DEPLOYMENT_MISMATCH_MAX_RETRIES` attempts (default `3`). Set it to `0` to fail on the first misrouted delivery instead.
48+
49+
Nothing is executed on the wrong deployment during recovery: no workflow code, no step body, no `step_started`, and no hook resume. Whatever the delivery was carrying travels with it, so a pending step keeps its identity and a hook resume keeps its payload — they run on the deployment that can actually decrypt them.
50+
51+
Recovery attempts do not create events on the run, so a run that self-heals looks completely normal. They are reported on the invocation's trace span (`workflow.deployment.pinned_id`, `workflow.deployment_mismatch.retry_count`, `workflow.deployment_mismatch.recovered`) and as a runtime warning in your function logs.
52+
53+
## What To Do
54+
55+
- **Re-run from the current deployment.** Trigger the workflow again from your latest deployment (or use the **Re-run** button in the Workflow Dashboard). The new run is pinned to the current deployment.
56+
- **Keep a run's deployment available** for the lifetime of that run. A run whose deployment has been deleted or has aged out cannot be resumed and must be re-run — recovery cannot help, so these fail on the first misrouted delivery. This applies to runs started with an explicit `deploymentId` too: pinning a run to an older deployment keeps it dependent on that deployment for its whole lifetime.
57+
- **Report it** if the pinned deployment was still available. Include both deployment ids and the run id from the error message, plus the trace span attributes above — a run that failed this way despite a reachable target is a routing fault worth investigating rather than something to work around.
58+
59+
## This Error Cannot Be Caught
60+
61+
Like other runtime signals, `DEPLOYMENT_MISMATCH` is **not catchable** inside your workflow function — the run is failed before any workflow or step code executes on the receiving deployment. Check the run status from outside instead:
62+
63+
```typescript lineNumbers
64+
import { getRun } from "workflow/api";
65+
66+
const run = getRun("wrun_abc123");
67+
const status = await run.status;
68+
if (status === "failed") {
69+
console.error("Run failed");
70+
}
71+
```

‎docs/content/docs/v5/configuration/runtime-tuning.mdx‎

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -55,6 +55,14 @@ For example, a workflow can run a 10-minute inline step even with `WORKFLOW_REPL
5555
- The runtime falls back to the sequential path automatically when the consumer or backend does not attest dedup support (or the payload is too large to inline on the queue message). On the sequential path the event is written *before* dispatch and its failure fails the resume — the fallback trades that resilience away to stay safe when dedup is not enforced, it does not preserve it.
5656
- Set `1` to force the sequential path as a kill switch. The chosen strategy is reported on the resume span as `workflow.hook.resume_strategy`.
5757

58+
### `WORKFLOW_DEPLOYMENT_MISMATCH_MAX_RETRIES`
59+
60+
- Default: `3`
61+
- Times a delivery that reached a deployment other than the one its run is pinned to is re-routed to that deployment before the run is failed with [`DEPLOYMENT_MISMATCH`](/docs/errors/deployment-mismatch).
62+
- Re-routed deliveries back off exponentially (1s, 2s, 4s). Set to `0` to fail the run on the first misrouted delivery.
63+
- Only applies to Worlds with atomic, immutable deployments (the Vercel World). A run whose pinned deployment cannot be reached at all fails immediately regardless of this value.
64+
- Transient or unknown queue publishing failures use normal queue redelivery and do not consume this budget.
65+
5866
### `WORKFLOW_PRECONDITION_GUARD`
5967

6068
- Default: enabled
Lines changed: 71 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,71 @@
1+
---
2+
title: deployment-mismatch
3+
description: A workflow run was delivered to a deployment other than the one it is pinned to.
4+
type: troubleshooting
5+
summary: Understand how Workflow recovers from a misrouted delivery, and why a run eventually fails with DEPLOYMENT_MISMATCH.
6+
prerequisites:
7+
- /docs/foundations/workflows-and-steps
8+
related:
9+
- /docs/foundations/versioning
10+
- /docs/errors/runtime-decryption-failed
11+
- /docs/foundations/errors-and-retries
12+
---
13+
14+
Every run is pinned to a single deployment when it starts. When a queued workflow or step callback is delivered to a **different** deployment, Workflow does not execute it there. Instead it re-routes the message to the deployment the run is pinned to, and only if the run keeps arriving elsewhere does it fail with the `DEPLOYMENT_MISMATCH` classification.
15+
16+
This is an SDK/runtime signal, not an error thrown by your workflow code, and it is not catchable inside a workflow function.
17+
18+
## Error Message
19+
20+
```
21+
Workflow run "wrun_..." is pinned to deployment "dpl_A", but was received by deployment "dpl_B". The runtime re-routed the message to "dpl_A" 3 times and it kept arriving elsewhere, so the run was stopped to protect against code-skew errors. Verify that the run's deployment is still available and that queue callbacks are routed to it.
22+
```
23+
24+
When the queue definitively reports that the run's deployment cannot be reached — it was deleted, or aged out of its retention window — no re-route is possible and the message omits the re-routing clause. Transient or unknown publishing failures leave the current delivery unacknowledged so the queue can redeliver it; they do not fail the run or consume this recovery budget.
25+
26+
## Why A Run Is Pinned
27+
28+
A run's deployment is chosen once, at [`start()`](/docs/api-reference/workflow-api/start):
29+
30+
- By default it is the deployment that called `start()` — see [Versioning](/docs/foundations/versioning) for why runs are pinned this way.
31+
- With `start(workflow, args, { deploymentId })` it is the id you pass, so a run can deliberately target a deployment other than the one that created it.
32+
- With `deploymentId: "latest"` it is the most recent deployment for the current environment, resolved at start time.
33+
34+
Whichever it is, that `deploymentId` is recorded on the run, and every subsequent workflow replay and step execution must happen on that deployment. Continuing on a different one is unsafe:
35+
36+
1. **Code skew.** The workflow and step bundles on the receiving deployment may not match the code that produced the run's recorded history, so replay could diverge or produce incorrect results.
37+
2. **Encryption.** Step inputs and other event-log payloads are encrypted with a per-run key derived from the pinned deployment's key material. A different deployment derives the wrong key and cannot decrypt them — previously the source of a confusing [runtime-decryption-failed](/docs/errors/runtime-decryption-failed) that exhausted retries with no clear cause.
38+
39+
So the runtime checks the pinned deployment before it executes anything, and `DEPLOYMENT_MISMATCH` names the result — instead of the mismatch surfacing later as an unrelated decryption failure.
40+
41+
## Automatic Recovery
42+
43+
A deployment that receives a run it does not own first tries to fix the delivery rather than fail the run:
44+
45+
1. It re-enqueues the message **explicitly addressed** to the run's own deployment. This is strictly better-addressed than the send that misrouted, which inherited the producing deployment's ambient id.
46+
2. Delivery is delayed with a short exponential backoff (1s, 2s, 4s).
47+
3. If the run keeps arriving at the wrong deployment, the run is failed with `DEPLOYMENT_MISMATCH` after `WORKFLOW_DEPLOYMENT_MISMATCH_MAX_RETRIES` attempts (default `3`). Set it to `0` to fail on the first misrouted delivery instead.
48+
49+
Nothing is executed on the wrong deployment during recovery: no workflow code, no step body, no `step_started`, and no hook resume. Whatever the delivery was carrying travels with it, so a pending step keeps its identity and a hook resume keeps its payload — they run on the deployment that can actually decrypt them.
50+
51+
Recovery attempts do not create events on the run, so a run that self-heals looks completely normal. They are reported on the invocation's trace span (`workflow.deployment.pinned_id`, `workflow.deployment_mismatch.retry_count`, `workflow.deployment_mismatch.recovered`) and as a runtime warning in your function logs.
52+
53+
## What To Do
54+
55+
- **Re-run from the current deployment.** Trigger the workflow again from your latest deployment (or use the **Re-run** button in the Workflow Dashboard). The new run is pinned to the current deployment.
56+
- **Keep a run's deployment available** for the lifetime of that run. A run whose deployment has been deleted or has aged out cannot be resumed and must be re-run — recovery cannot help, so these fail on the first misrouted delivery. This applies to runs started with an explicit `deploymentId` too: pinning a run to an older deployment keeps it dependent on that deployment for its whole lifetime.
57+
- **Report it** if the pinned deployment was still available. Include both deployment ids and the run id from the error message, plus the trace span attributes above — a run that failed this way despite a reachable target is a routing fault worth investigating rather than something to work around.
58+
59+
## This Error Cannot Be Caught
60+
61+
Like other runtime signals, `DEPLOYMENT_MISMATCH` is **not catchable** inside your workflow function — the run is failed before any workflow or step code executes on the receiving deployment. Check the run status from outside instead:
62+
63+
```typescript lineNumbers
64+
import { getRun } from "workflow/api";
65+
66+
const run = getRun("wrun_abc123");
67+
const status = await run.status;
68+
if (status === "failed") {
69+
console.error("Run failed");
70+
}
71+
```

‎packages/core/src/classify-error.test.ts‎

Lines changed: 13 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -7,6 +7,7 @@ import {
77
RuntimeDecryptionError,
88
ThrottleError,
99
TooEarlyError,
10+
WorkflowDeploymentMismatchError,
1011
WorkflowNotRegisteredError,
1112
WorkflowRuntimeError,
1213
WorkflowWorldError,
@@ -49,6 +50,18 @@ describe('classifyRunError', () => {
4950
);
5051
});
5152

53+
it('classifies WorkflowDeploymentMismatchError as DEPLOYMENT_MISMATCH', () => {
54+
expect(
55+
classifyRunError(
56+
new WorkflowDeploymentMismatchError(
57+
'wrun_test',
58+
'dpl_expected',
59+
'dpl_actual'
60+
)
61+
)
62+
).toBe(RUN_ERROR_CODES.DEPLOYMENT_MISMATCH);
63+
});
64+
5265
it('classifies plain Error as USER_ERROR', () => {
5366
expect(classifyRunError(new Error('user code broke'))).toBe(
5467
RUN_ERROR_CODES.USER_ERROR

‎packages/core/src/classify-error.ts‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -7,6 +7,7 @@ import {
77
RuntimeDecryptionError,
88
StepNotRegisteredError,
99
ThrottleError,
10+
WorkflowDeploymentMismatchError,
1011
WorkflowNotRegisteredError,
1112
WorkflowRuntimeError,
1213
WorkflowWorldError,
@@ -121,6 +122,10 @@ export function classifyRunError(err: unknown): RunErrorCode {
121122
return RUN_ERROR_CODES.MAX_EVENTS_EXCEEDED;
122123
}
123124

125+
if (WorkflowDeploymentMismatchError.is(err)) {
126+
return RUN_ERROR_CODES.DEPLOYMENT_MISMATCH;
127+
}
128+
124129
// World-layer faults — both a malformed response (contract violation) and a
125130
// transient infrastructure failure (throttle / 5xx / transport / timeout,
126131
// e.g. a firewall challenge) — are the backend's fault, not the user's.

‎packages/core/src/describe-error.test.ts‎

Lines changed: 20 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -3,6 +3,7 @@ import {
33
RUN_ERROR_CODES,
44
SerializationError,
55
StepNotRegisteredError,
6+
WorkflowDeploymentMismatchError,
67
WorkflowRuntimeError,
78
} from '@workflow/errors';
89
import { describe, expect, test } from 'vitest';
@@ -118,6 +119,15 @@ describe('describeError', () => {
118119
expect(result.hint).toContain('maximum number of events');
119120
});
120121

122+
test('WorkflowDeploymentMismatchError is attributed to the SDK with the deployment hint', () => {
123+
const result = describeError(
124+
new WorkflowDeploymentMismatchError('wrun', 'dpl_a', 'dpl_b')
125+
);
126+
expect(result.attribution).toBe('sdk');
127+
expect(result.errorCode).toBe(RUN_ERROR_CODES.DEPLOYMENT_MISMATCH);
128+
expect(result.hint).toContain('deployment it is pinned to');
129+
});
130+
121131
test('precomputed errorCode wins over classifyRunError when both are provided', () => {
122132
// A plain Error would classify as USER_ERROR, but passing REPLAY_TIMEOUT
123133
// explicitly overrides that — useful for callers that know the failure
@@ -220,6 +230,16 @@ describe('describeRunError', () => {
220230
expect(result.hint).toContain('SDK contract');
221231
});
222232

233+
test('DEPLOYMENT_MISMATCH errorCode is attributed to the SDK', () => {
234+
const result = describeRunError({
235+
errorCode: RUN_ERROR_CODES.DEPLOYMENT_MISMATCH,
236+
errorName: 'WorkflowDeploymentMismatchError',
237+
});
238+
expect(result.attribution).toBe('sdk');
239+
expect(result.errorCode).toBe(RUN_ERROR_CODES.DEPLOYMENT_MISMATCH);
240+
expect(result.hint).toContain('deployment it is pinned to');
241+
});
242+
223243
test('RUNTIME_ERROR code without errorName still lands as SDK', () => {
224244
const result = describeRunError({
225245
errorCode: RUN_ERROR_CODES.RUNTIME_ERROR,

‎packages/core/src/describe-error.ts‎

Lines changed: 15 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -87,6 +87,8 @@ const MAX_EVENTS_HINT =
8787
'The workflow exceeded the maximum number of events per run. This usually means unbounded work in the workflow function — e.g. a loop that keeps creating steps without terminating. Break long-running workflows into child workflows to stay under the limit.';
8888
const WORLD_CONTRACT_HINT =
8989
'The workflow backend returned data that violated the SDK contract. This is not retryable; please report it with the stack trace and runId.';
90+
const DEPLOYMENT_MISMATCH_HINT =
91+
"The run was delivered to a deployment other than the deployment it is pinned to, and was stopped to protect against code-skew errors after the runtime failed to re-route it there. Verify that the run's deployment is still available and that queue callbacks route to it.";
9092

9193
function normalizeErrorCode(code: string | undefined): RunErrorCode {
9294
// Values read back from persisted events are `string | undefined` — we
@@ -149,6 +151,9 @@ export function describeRunError(
149151
hint: WORLD_CONTRACT_HINT,
150152
};
151153
}
154+
if (errorCode === RUN_ERROR_CODES.DEPLOYMENT_MISMATCH) {
155+
return { attribution: 'sdk', errorCode, hint: DEPLOYMENT_MISMATCH_HINT };
156+
}
152157
if (name === 'WorkflowRuntimeError' || name === 'StepNotRegisteredError') {
153158
return { attribution: 'sdk', errorCode, hint: RUNTIME_ERROR_HINT };
154159
}
@@ -209,6 +214,16 @@ export function describeError(
209214
};
210215
}
211216

217+
// Check DEPLOYMENT_MISMATCH before the generic WorkflowRuntimeError branch —
218+
// WorkflowDeploymentMismatchError subclasses it, but has its own code + hint.
219+
if (effectiveCode === RUN_ERROR_CODES.DEPLOYMENT_MISMATCH) {
220+
return {
221+
attribution: 'sdk',
222+
errorCode: effectiveCode,
223+
hint: DEPLOYMENT_MISMATCH_HINT,
224+
};
225+
}
226+
212227
if (err instanceof WorkflowRuntimeError) {
213228
return {
214229
attribution: 'sdk',

0 commit comments

Comments
 (0)