Skip to content

Commit 2668e33

Browse files
karthikscale3pranaygpclaude
authored
Durable hook resume: write, then wake (vercel#3841)
* test(core): reproduce lazy resume disposal race * Fix durable hook resume race * Fail closed on unknown hook wakes * Improve unsupported hook wake diagnostics * Address durable hook resume review feedback * Harden producer-committed wake handling * Serialize durable hook resume: write, then wake resumeHook() now dispatches strictly serially: the hook_received event is made durable first, and the workflow wake is published only after the write is acknowledged. The wake is a plain runId message (the shape the sequential path always published), so the producer-committed wake barrier, its queue-message field, and the HOOK_RESUME_INPUT_VERSION bump are all removed — no consumer or backend coordination is needed, and either side rolls back independently to today's behavior. The pre-write ops flush now partitions serialization ops: producer-push uploads are awaited before the event commits (the payload must not point at bytes still in flight), while consumer-settled reader ops — a dehydrated WritableStream, e.g. a manual webhook's responseWritable — are backgrounded. Awaiting those deadlocked the resume against its own wake (webhookWorkflow failing across the whole e2e matrix). Also: wake retries stop on definitive 4xx errors instead of burning the retry budget; WORKFLOW_DISABLE_LAZY_HOOK_RESUME no longer gates anything and is ignored; the internal resumeHookDurable alias is removed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Address review: retry classification, wake dedup, 409 passthrough - Wake retry classification now actually fires against @vercel/queue: its errors carry no status field, so classify by the World's deployment-unavailable hook, then numeric status, then the queue client's definitive-4xx error names. - The wake publish carries idempotencyKey `hook-<resumeId>` on the claim path, so a retried publish whose response was lost dedups instead of costing a duplicate full replay. - EntityConflictError (HTTP 409) from the durable write is no longer re-keyed to HookNotFoundError: every 409 the backend emits on this write today is transient (slot conflict past the server's retry budget, claim race) and committed nothing, so it surfaces retryable instead of presenting as a permanent 404. - Stamp workflow.hook.resume_committed / wake_published span attributes after each leg resolves, making stranded resumes (committed event, no wake) queryable from traces. - Document on the public resumeHook signature that passing the token (not a cached Hook) is what makes the write idempotent-on-retry. - Changeset/changelog: note the ended-run behavior change (late webhook deliveries to finished runs now 404 instead of 202) and the 409 passthrough. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Pranay Prakash <pranay.gp@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
1 parent e9d5c56 commit 2668e33

28 files changed

Lines changed: 1186 additions & 1127 deletions

‎.changeset/durable-hook-resume.md‎

Lines changed: 18 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,18 @@
1+
---
2+
'@workflow/core': patch
3+
'@workflow/world': patch
4+
---
5+
6+
Make `resumeHook()` durable before it resolves: the `hook_received` event is
7+
written durably first, and the workflow wake is published only after the write
8+
is acknowledged. A disposal racing the queue delivery can no longer lose a
9+
resume the caller was told succeeded. The wake message is unchanged, so no
10+
consumer or backend coordination is needed; `WORKFLOW_DISABLE_LAZY_HOOK_RESUME`
11+
is now a no-op and the internal `resumeHookDurable()` alias is removed.
12+
13+
Behavior changes: a resume against an ended run now throws `HookNotFoundError`
14+
instead of resolving (the lazy path never observed the server's rejection, so a
15+
late webhook delivery to a finished run answered 202 where it now answers 404).
16+
A transient write conflict (HTTP 409, e.g. an event-slot conflict under
17+
contention) is no longer re-keyed to `HookNotFoundError`: it surfaces as a
18+
retryable error, since its transaction committed nothing.

‎.changeset/lazy-hook-resume-vitest.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -2,4 +2,4 @@
22
'@workflow/vitest': patch
33
---
44

5-
`waitForHook()` accepts `notHookId` to skip a hook the caller already resumed, whose `hook_received` may not be written yet.
5+
`waitForHook()` accepts `notHookId` to exclude a previously observed hook when a workflow creates several hooks with the same token.

‎docs/content/docs/v5/api-reference/workflow-api/resume-hook.mdx‎

Lines changed: 5 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -12,9 +12,11 @@ related:
1212

1313
Resumes a workflow run by sending a payload to a hook identified by its token.
1414

15-
It publishes a workflow invocation carrying the payload; the runtime creates the `hook_received` event and continues execution from it.
15+
It durably writes the `hook_received` event and only then publishes a workflow wake. The call resolves only after both operations succeed, in that order.
1616

17-
`resumeHook()` throws `HookNotFoundError` when no hook holds the token. A run that has already ended cannot be resumed, including one whose Hook is kept by `experimental_minRetention`, but whether the call reports that depends on the path it takes: a resume dispatched without reading the run resolves and the ended state is only detected once the payload arrives, while one that reads the run, or that falls back to writing the event up front, throws `HookNotFoundError`. See [lazy hook resume](/docs/changelog/lazy-hook-resume).
17+
`resumeHook()` throws `HookNotFoundError` when no hook holds the token or when its `hook_received` write is refused because the hook was disposed or the run ended. See [durable hook resume](/docs/changelog/lazy-hook-resume).
18+
19+
If `resumeHook()` throws any other error, the outcome is ambiguous only in dispatch, never in durability: the event may already be durable even though the workflow wake failed, and any later wake of the run delivers it. Calling `resumeHook()` again creates a new `resumeId` and can append a second `hook_received`. Callers that need at-most-once behavior across separate invocations must retain and deduplicate their own request key.
1820

1921
<Callout type="warn">
2022
`resumeHook` is a runtime function that must be called from outside a workflow function.
@@ -50,7 +52,7 @@ showSections={["parameters"]}
5052

5153
### Returns
5254

53-
Returns a `Promise<ResumedHook>`, a `Hook` extended with an optional `resilientResume` flag. Resolving means the resume was accepted for delivery: the payload rides the workflow queue message and the runtime materializes the `hook_received` event from it before replaying (see the [lazy hook resume changelog](/docs/changelog/lazy-hook-resume)). `resilientResume` is retained for source compatibility and is no longer set by any path. The resolved hook:
55+
Returns a `Promise<ResumedHook>`, a `Hook` extended with an optional `resilientResume` flag. Resolving means the payload is durably recorded as `hook_received` and the workflow wake was accepted. `resilientResume` is retained for source compatibility and is no longer set by any path. The resolved hook:
5456

5557
<TSDoc
5658
definition={`

‎docs/content/docs/v5/changelog/index.mdx‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -12,7 +12,7 @@ Stay up to date with the latest changes to Workflow SDK.
1212

1313
## 2026
1414

15-
- [Lazy hook resume](/docs/changelog/lazy-hook-resume) (August 2026)
15+
- [Durable hook resume](/docs/changelog/lazy-hook-resume) (August 2026)
1616
- [Resilient hook resume](/docs/changelog/resilient-resume) (July 2026)
1717
- [Eager processing of steps and incremental event replay](/docs/changelog/eager-processing) (March 2026)
1818
- Serializable AbortController and AbortSignal (March 12, 2026)
Lines changed: 60 additions & 33 deletions
Original file line numberDiff line numberDiff line change
@@ -1,51 +1,78 @@
11
---
2-
title: Lazy hook resume
3-
description: resumeHook() no longer writes hook_received itself. The queue consumer materializes the event from the message, so a resume costs one round trip.
2+
title: Durable hook resume
3+
description: resumeHook() durably writes hook_received and only then publishes the workflow wake, so a resolved call can never be lost to a disposal race.
44
---
55

6-
# Lazy hook resume
6+
# Durable hook resume
77

88
## Motivation
99

10-
[Resilient hook resume](/docs/changelog/resilient-resume) made `resumeHook()` write the `hook_received` event and publish the workflow queue message concurrently, with the queue consumer re-ensuring the event from the message's `hookInput` before replay. Both sides then wrote the same event, and a `(runId, resumeId)` constraint collapsed them onto one.
10+
The previous lazy path published the serialized hook payload on the workflow
11+
queue and left the queue consumer to create `hook_received`. If the hook was
12+
disposed after `resumeHook()` returned but before the consumer write committed,
13+
that write was rejected and the acknowledged queue delivery could not resume the
14+
workflow: the caller was told the resume succeeded, and it was lost.
1115

12-
Running the two concurrently removed the second round trip from the critical path, but the write itself stayed: every resume still spent a request on an event the consumer was about to write anyway, and the producer still had to classify its outcome (conflict, throttle, terminal run) to decide whether the resume had survived.
13-
14-
This change drops the producer's write entirely. On the lazy path `resumeHook()` publishes the queue message and nothing else.
16+
`resumeHook()` now resolves only after both the durable event write and the
17+
workflow wake have succeeded, in that order.
1518

1619
## Design
1720

18-
- `resumeHook()` publishes one message carrying `hookInput`: the dehydrated payload, a client-minted `resumeId`, the hook token, and a payload digest. It writes no event.
19-
- The queue consumer materializes `hook_received` from `hookInput` before replay, keyed by `resumeId`. This is the same write it already performed; it is now the only one.
20-
- The `(runId, resumeId)` constraint still matters: a queue redelivery, or a delivery re-routed for deployment affinity, repeats the write with the same key, and the backend collapses those onto exactly one committed event.
21-
- **A failed publish fails the resume.** The message carries both the trigger and the only copy of the payload, so `resumeHook()` throws and nothing is persisted for a later delivery to pick up. This replaces the previous rule where a failed event write could still be recovered through the queue.
22-
- `ResumedHook.resilientResume` is retained on the type but is never set: with a single writer there is no partial outcome to report. The `workflow.hook.resilient_resume` span attribute is likewise no longer emitted.
23-
- The resume span reports `workflow.hook.resume_strategy: lazy` (previously `parallel`).
24-
25-
## Behavior change: the event is not visible when `resumeHook()` returns
26-
27-
`resumeHook()` used to await its own `hook_received` write, so by the time it resolved the event was in the log. It no longer writes, so **resolving means the message was published, not that the event exists**. The event appears when the run picks the resume up.
28-
29-
Code that reads the run back immediately after resuming now races. The pattern that breaks is a loop that resumes and then looks for the next thing to resume, keying off "this hook has no `hook_received` yet": it can be handed back the hook it just resumed and deliver a second payload to it. Wait for something that implies the run made progress instead. `waitForHook()` in `@workflow/vitest` takes a `notHookId` option for exactly this.
30-
31-
The runtime has one caller that needs the old guarantee. A step that aborts a shared `AbortController` resumes a hook to record the abort in the event log, and that write is an ordering barrier: it must land before the step completes, or the continuation `step_completed` enqueues can dispatch the next step with a stale, non-aborted signal. That path uses an internal durable resume which keeps the eager write and reports `resume_fallback_reason: durable_required`.
32-
33-
Nothing about delivery changes. The payload is on the queue message and reaches the workflow exactly once.
21+
The dispatch is strictly serial:
22+
23+
1. The hook is resolved by token. An unknown token throws `HookNotFoundError`.
24+
2. The producer writes `hook_received` durably into the run's event log. A
25+
client-minted `resumeId` and payload digest ride the write when the backend
26+
supports atomic resume claims, so transport-level retries of the same write
27+
converge on exactly one committed event. A write refused because the hook
28+
was disposed or the run ended throws `HookNotFoundError`.
29+
3. Only after the write is acknowledged does the producer publish the workflow
30+
wake. The wake carries no payload — the payload lives in the event log — so
31+
nothing rides on the queue message but the trigger. Publication is retried
32+
a bounded number of times.
33+
34+
Because the event is committed before the wake exists, a disposal or run
35+
completion racing the queue delivery cannot erase a resume the caller was told
36+
succeeded: the delivery replays the committed event from the log.
37+
38+
- `ResumedHook.resilientResume` remains on the type for source compatibility
39+
and is no longer set. The internal `resumeHookDurable()` entry point is
40+
removed; `resumeHook()` itself now provides the durable guarantee.
41+
42+
A resolved call proves that the event is durable and the wake was accepted.
43+
`HookNotFoundError` proves this invocation committed no event. Any other thrown
44+
error is ambiguous only in *dispatch*, never in durability: a wake failure
45+
after the write leaves the event committed, and any later wake of the run
46+
(from any source) delivers it. A fresh `resumeHook()` invocation mints a new
47+
`resumeId`, so blindly retrying a failed call can append a second
48+
`hook_received`; callers that need at-most-once behavior across separate
49+
invocations must deduplicate on their own request key.
3450

3551
## Behavior change: resumes against an ended run
3652

37-
The hook lookup is unchanged: `resumeHook(token, ...)` still resolves the token through `hooks.getByToken()`, which throws `HookNotFoundError` when no hook holds it. Hook existence, and the token's binding to a run, are still validated before anything is published.
38-
39-
What the lookup does not carry is the run's *mutable* status. `HookResumeContext` is deliberately an immutable slice of the run, so a resume that runs off it never learns whether the run is still live. That used to be caught by the `hook_received` write being rejected. With no write, **a resume against an ended run resolves instead of throwing `HookNotFoundError`**.
40-
41-
This is only reachable when the hook record outlives its run, since otherwise the lookup itself fails: a hook kept by `experimental_minRetention`, or one whose token has not been released yet. Resumes that fall back to reading the run keep their terminal pre-check, as does the sequential path, so a resume on either of those still fails loudly.
53+
The lazy path never observed the server's rejection — it published a message
54+
and resolved, so a resume against a run that had already ended reported
55+
success (reachable whenever the hook record outlives its run, e.g. token
56+
retention). The durable write restores the check: **a resume against an ended
57+
run now throws `HookNotFoundError`**, and a late webhook delivery to a
58+
finished run answers 404 where it previously answered 202. Senders that treat
59+
4xx as terminal will stop retrying such deliveries; that is the correct
60+
signal, since nothing can resume an ended run.
4261

43-
Nothing resumes either way. The consumer's write is rejected the same way and the delivery is consumed, so the ended run is untouched. Only the producer's report changes: an accepted publish means the resume was dispatched, not that the run was still live when it arrived. A [webhook](/docs/api-reference/workflow-api/resume-webhook) whose run has ended can answer `202` rather than surfacing an error.
44-
45-
Callers that need the distinction have to read the run.
62+
A transient write conflict (HTTP 409, e.g. an event-slot conflict that
63+
escaped the server's internal retry budget under contention) is no longer
64+
re-keyed to `HookNotFoundError`. It surfaces as a retryable error, and its
65+
rejected transaction committed nothing, so retrying the resume is safe.
4666

4767
## Compatibility
4868

49-
The gating is unchanged: the lazy path activates only when the target run's queue consumer and the live backend both attest support, re-checked on every resume. Oversized payloads, legacy runs, non-CBOR transports, and `WORKFLOW_DISABLE_LAZY_HOOK_RESUME=1` fall back to the sequential write-then-publish path.
69+
Nothing about the queue message changes: the wake has the same shape the
70+
sequential path always published, so no consumer, backend, or server
71+
coordination is needed and either side can roll back independently.
72+
73+
Consumers continue to accept legacy `hookInput` messages from older producers,
74+
materializing their payload before replay. This permits rolling upgrades
75+
without a coordinated producer and consumer deployment.
5076

51-
Consumers still accept a message from an older producer that wrote the event itself: such a message reports `strategy: parallel`, and the consumer's write converges on the producer's committed event exactly as before. No coordinated deploy is needed in either direction.
77+
`WORKFLOW_DISABLE_LAZY_HOOK_RESUME` no longer gates anything and is ignored:
78+
there is no lazy path left to disable.

‎docs/content/docs/v5/changelog/resilient-resume.mdx‎

Lines changed: 8 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -6,11 +6,13 @@ description: resumeHook() now tolerates transient event storage failures when th
66
# Resilient `resumeHook()`
77

88
<Callout type="info">
9-
Superseded by [lazy hook resume](/docs/changelog/lazy-hook-resume):
10-
`resumeHook()` no longer writes the `hook_received` event at all on the fast
11-
path, so the two-writer design and the `resilientResume` flag described below
12-
are historical. The `(runId, resumeId)` constraint and the queue-carried
13-
payload remain.
9+
Superseded by [durable hook resume](/docs/changelog/lazy-hook-resume):
10+
`resumeHook()` now writes the `hook_received` event durably and only then
11+
publishes the workflow wake, so the two-writer design, the queue-carried
12+
payload, and the `resilientResume` flag described below are historical. The
13+
`(runId, resumeId)` constraint remains, converging transport-level retries
14+
of the producer's own write (and legacy `hookInput` redeliveries from older
15+
producers).
1416
</Callout>
1517

1618
## Motivation
@@ -27,4 +29,4 @@ description: resumeHook() now tolerates transient event storage failures when th
2729

2830
## Compatibility
2931

30-
The parallel fast path is gated per resume: it activates only when both the target run's queue consumer and the live backend independently attest dedup support (re-checked on every resume, so rollout and rollback both degrade safely). Otherwise (for oversized payloads, legacy runs, or with `WORKFLOW_DISABLE_LAZY_HOOK_RESUME=1`), `resumeHook()` falls back to the original sequential write-then-dispatch path. Because runs keep executing on the deployment they were created on, a resume targeting a run from an older deployment uses the sequential path.
32+
The parallel fast path is gated per resume: it activates only when both the target run's queue consumer and the live backend independently attest dedup support (re-checked on every resume, so rollout and rollback both degrade safely). Otherwise (for oversized payloads or legacy runs), `resumeHook()` falls back to the original sequential write-then-dispatch path. Because runs keep executing on the deployment they were created on, a resume targeting a run from an older deployment uses the sequential path.

‎docs/content/docs/v5/configuration/runtime-tuning.mdx‎

Lines changed: 1 addition & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -81,14 +81,6 @@ For example, a workflow can run a 10-minute inline step even with `WORKFLOW_REPL
8181
- Default: `3`
8282
- Recovery replays before replay divergence is recorded as corruption.
8383

84-
### `WORKFLOW_DISABLE_LAZY_HOOK_RESUME`
85-
86-
- Default: enabled (lazy hook resume on)
87-
- Resuming a hook publishes the workflow invocation and writes no event, so the resume costs one round trip. The queue message carries the payload, and the consumer creates the `hook_received` event from it before replay. A backend `(runId, resumeId)` constraint keeps redeliveries of that message converging on exactly one event. Because the message is the only copy of the payload, a failed publish fails the resume.
88-
- The runtime falls back to the sequential path when the consumer or backend does not attest dedup support (or the payload is too large to inline on the queue message). On the sequential path, the event is written *before* dispatch, and its failure fails the resume.
89-
- The two paths differ on ended runs: the sequential write is rejected and surfaces as `HookNotFoundError`, while the lazy path resolves because it never writes. Neither resumes the run.
90-
- Set `1` to force the sequential path as a kill switch. The chosen strategy is reported on the resume span as `workflow.hook.resume_strategy`.
91-
9284
### `WORKFLOW_DEPLOYMENT_MISMATCH_MAX_RETRIES`
9385

9486
- Default: `3`
@@ -100,7 +92,7 @@ For example, a workflow can run a 10-minute inline step even with `WORKFLOW_REPL
10092
### `WORKFLOW_RESILIENT_STEP_DISPATCH`
10193

10294
- Default: disabled
103-
- When a suspension hands newly created steps to the queue, the runtime publishes each step's execution message in parallel with its `step_created` event write instead of sequencing them, cutting a round trip per dispatched step. The message also carries the serialized step input (`stepInput`), so a transient `step_created` write failure (429 / 5xx / transport) still executes the step. The queue consumer idempotently re-ensures the event before running it, converging with the producer's write on the step's correlation ID. This mirrors resilient start (`runInput`) and the lazy hook resume (`hookInput`).
95+
- When a suspension hands newly created steps to the queue, the runtime publishes each step's execution message in parallel with its `step_created` event write instead of sequencing them, cutting a round trip per dispatched step. The message also carries the serialized step input (`stepInput`), so a transient `step_created` write failure (429 / 5xx / transport) still executes the step. The queue consumer idempotently re-ensures the event before running it, converging with the producer's write on the step's correlation ID. This mirrors resilient start (`runInput`) and the legacy lazy hook resume's `hookInput` (which current producers no longer send; see [durable hook resume](/docs/changelog/lazy-hook-resume)).
10496
- It is off by default because the publish races the create's verdict, and a create can come back refused: as a duplicate this replay should stop pursuing, or as a [stale write](#stale-reads-and-why-nothing-has-to-be-rejected) on a World that refuses rather than reports. Either way the message carrying the payload is already out, so the consumer can materialize a step whose create was refused, and nothing orders the verdict before the consumer's redelivery re-ensure. The sequential path is the only one that gives the message a happens-after edge over it.
10597
- Even when enabled, the runtime falls back to the sequential create-then-publish dispatch when the step input is too large to inline on the queue message, or when the run's queue transport cannot carry binary payloads (pre-CBOR spec versions).
10698
- Producer-side recoveries are reported on the suspension span as `workflow.step.resilient_dispatch_recovered`; a consumer that materialized the event reports `workflow.step.resilient_dispatch_materialized`.

0 commit comments

Comments
 (0)