|
1 | 1 | --- |
2 | | -title: Lazy hook resume |
3 | | -description: resumeHook() no longer writes hook_received itself. The queue consumer materializes the event from the message, so a resume costs one round trip. |
| 2 | +title: Durable hook resume |
| 3 | +description: resumeHook() durably writes hook_received and only then publishes the workflow wake, so a resolved call can never be lost to a disposal race. |
4 | 4 | --- |
5 | 5 |
|
6 | | -# Lazy hook resume |
| 6 | +# Durable hook resume |
7 | 7 |
|
8 | 8 | ## Motivation |
9 | 9 |
|
10 | | -[Resilient hook resume](/docs/changelog/resilient-resume) made `resumeHook()` write the `hook_received` event and publish the workflow queue message concurrently, with the queue consumer re-ensuring the event from the message's `hookInput` before replay. Both sides then wrote the same event, and a `(runId, resumeId)` constraint collapsed them onto one. |
| 10 | +The previous lazy path published the serialized hook payload on the workflow |
| 11 | +queue and left the queue consumer to create `hook_received`. If the hook was |
| 12 | +disposed after `resumeHook()` returned but before the consumer write committed, |
| 13 | +that write was rejected and the acknowledged queue delivery could not resume the |
| 14 | +workflow: the caller was told the resume succeeded, and it was lost. |
11 | 15 |
|
12 | | -Running the two concurrently removed the second round trip from the critical path, but the write itself stayed: every resume still spent a request on an event the consumer was about to write anyway, and the producer still had to classify its outcome (conflict, throttle, terminal run) to decide whether the resume had survived. |
13 | | - |
14 | | -This change drops the producer's write entirely. On the lazy path `resumeHook()` publishes the queue message and nothing else. |
| 16 | +`resumeHook()` now resolves only after both the durable event write and the |
| 17 | +workflow wake have succeeded, in that order. |
15 | 18 |
|
16 | 19 | ## Design |
17 | 20 |
|
18 | | -- `resumeHook()` publishes one message carrying `hookInput`: the dehydrated payload, a client-minted `resumeId`, the hook token, and a payload digest. It writes no event. |
19 | | -- The queue consumer materializes `hook_received` from `hookInput` before replay, keyed by `resumeId`. This is the same write it already performed; it is now the only one. |
20 | | -- The `(runId, resumeId)` constraint still matters: a queue redelivery, or a delivery re-routed for deployment affinity, repeats the write with the same key, and the backend collapses those onto exactly one committed event. |
21 | | -- **A failed publish fails the resume.** The message carries both the trigger and the only copy of the payload, so `resumeHook()` throws and nothing is persisted for a later delivery to pick up. This replaces the previous rule where a failed event write could still be recovered through the queue. |
22 | | -- `ResumedHook.resilientResume` is retained on the type but is never set: with a single writer there is no partial outcome to report. The `workflow.hook.resilient_resume` span attribute is likewise no longer emitted. |
23 | | -- The resume span reports `workflow.hook.resume_strategy: lazy` (previously `parallel`). |
24 | | - |
25 | | -## Behavior change: the event is not visible when `resumeHook()` returns |
26 | | - |
27 | | -`resumeHook()` used to await its own `hook_received` write, so by the time it resolved the event was in the log. It no longer writes, so **resolving means the message was published, not that the event exists**. The event appears when the run picks the resume up. |
28 | | - |
29 | | -Code that reads the run back immediately after resuming now races. The pattern that breaks is a loop that resumes and then looks for the next thing to resume, keying off "this hook has no `hook_received` yet": it can be handed back the hook it just resumed and deliver a second payload to it. Wait for something that implies the run made progress instead. `waitForHook()` in `@workflow/vitest` takes a `notHookId` option for exactly this. |
30 | | - |
31 | | -The runtime has one caller that needs the old guarantee. A step that aborts a shared `AbortController` resumes a hook to record the abort in the event log, and that write is an ordering barrier: it must land before the step completes, or the continuation `step_completed` enqueues can dispatch the next step with a stale, non-aborted signal. That path uses an internal durable resume which keeps the eager write and reports `resume_fallback_reason: durable_required`. |
32 | | - |
33 | | -Nothing about delivery changes. The payload is on the queue message and reaches the workflow exactly once. |
| 21 | +The dispatch is strictly serial: |
| 22 | + |
| 23 | +1. The hook is resolved by token. An unknown token throws `HookNotFoundError`. |
| 24 | +2. The producer writes `hook_received` durably into the run's event log. A |
| 25 | + client-minted `resumeId` and payload digest ride the write when the backend |
| 26 | + supports atomic resume claims, so transport-level retries of the same write |
| 27 | + converge on exactly one committed event. A write refused because the hook |
| 28 | + was disposed or the run ended throws `HookNotFoundError`. |
| 29 | +3. Only after the write is acknowledged does the producer publish the workflow |
| 30 | + wake. The wake carries no payload — the payload lives in the event log — so |
| 31 | + nothing rides on the queue message but the trigger. Publication is retried |
| 32 | + a bounded number of times. |
| 33 | + |
| 34 | +Because the event is committed before the wake exists, a disposal or run |
| 35 | +completion racing the queue delivery cannot erase a resume the caller was told |
| 36 | +succeeded: the delivery replays the committed event from the log. |
| 37 | + |
| 38 | +- `ResumedHook.resilientResume` remains on the type for source compatibility |
| 39 | + and is no longer set. The internal `resumeHookDurable()` entry point is |
| 40 | + removed; `resumeHook()` itself now provides the durable guarantee. |
| 41 | + |
| 42 | +A resolved call proves that the event is durable and the wake was accepted. |
| 43 | +`HookNotFoundError` proves this invocation committed no event. Any other thrown |
| 44 | +error is ambiguous only in *dispatch*, never in durability: a wake failure |
| 45 | +after the write leaves the event committed, and any later wake of the run |
| 46 | +(from any source) delivers it. A fresh `resumeHook()` invocation mints a new |
| 47 | +`resumeId`, so blindly retrying a failed call can append a second |
| 48 | +`hook_received`; callers that need at-most-once behavior across separate |
| 49 | +invocations must deduplicate on their own request key. |
34 | 50 |
|
35 | 51 | ## Behavior change: resumes against an ended run |
36 | 52 |
|
37 | | -The hook lookup is unchanged: `resumeHook(token, ...)` still resolves the token through `hooks.getByToken()`, which throws `HookNotFoundError` when no hook holds it. Hook existence, and the token's binding to a run, are still validated before anything is published. |
38 | | - |
39 | | -What the lookup does not carry is the run's *mutable* status. `HookResumeContext` is deliberately an immutable slice of the run, so a resume that runs off it never learns whether the run is still live. That used to be caught by the `hook_received` write being rejected. With no write, **a resume against an ended run resolves instead of throwing `HookNotFoundError`**. |
40 | | - |
41 | | -This is only reachable when the hook record outlives its run, since otherwise the lookup itself fails: a hook kept by `experimental_minRetention`, or one whose token has not been released yet. Resumes that fall back to reading the run keep their terminal pre-check, as does the sequential path, so a resume on either of those still fails loudly. |
| 53 | +The lazy path never observed the server's rejection — it published a message |
| 54 | +and resolved, so a resume against a run that had already ended reported |
| 55 | +success (reachable whenever the hook record outlives its run, e.g. token |
| 56 | +retention). The durable write restores the check: **a resume against an ended |
| 57 | +run now throws `HookNotFoundError`**, and a late webhook delivery to a |
| 58 | +finished run answers 404 where it previously answered 202. Senders that treat |
| 59 | +4xx as terminal will stop retrying such deliveries; that is the correct |
| 60 | +signal, since nothing can resume an ended run. |
42 | 61 |
|
43 | | -Nothing resumes either way. The consumer's write is rejected the same way and the delivery is consumed, so the ended run is untouched. Only the producer's report changes: an accepted publish means the resume was dispatched, not that the run was still live when it arrived. A [webhook](/docs/api-reference/workflow-api/resume-webhook) whose run has ended can answer `202` rather than surfacing an error. |
44 | | - |
45 | | -Callers that need the distinction have to read the run. |
| 62 | +A transient write conflict (HTTP 409, e.g. an event-slot conflict that |
| 63 | +escaped the server's internal retry budget under contention) is no longer |
| 64 | +re-keyed to `HookNotFoundError`. It surfaces as a retryable error, and its |
| 65 | +rejected transaction committed nothing, so retrying the resume is safe. |
46 | 66 |
|
47 | 67 | ## Compatibility |
48 | 68 |
|
49 | | -The gating is unchanged: the lazy path activates only when the target run's queue consumer and the live backend both attest support, re-checked on every resume. Oversized payloads, legacy runs, non-CBOR transports, and `WORKFLOW_DISABLE_LAZY_HOOK_RESUME=1` fall back to the sequential write-then-publish path. |
| 69 | +Nothing about the queue message changes: the wake has the same shape the |
| 70 | +sequential path always published, so no consumer, backend, or server |
| 71 | +coordination is needed and either side can roll back independently. |
| 72 | + |
| 73 | +Consumers continue to accept legacy `hookInput` messages from older producers, |
| 74 | +materializing their payload before replay. This permits rolling upgrades |
| 75 | +without a coordinated producer and consumer deployment. |
50 | 76 |
|
51 | | -Consumers still accept a message from an older producer that wrote the event itself: such a message reports `strategy: parallel`, and the consumer's write converges on the producer's committed event exactly as before. No coordinated deploy is needed in either direction. |
| 77 | +`WORKFLOW_DISABLE_LAZY_HOOK_RESUME` no longer gates anything and is ignored: |
| 78 | +there is no lazy path left to disable. |
0 commit comments