Guide · Service business operations
What happens when a workflow runs twice or fails halfway?
Follow actual duplicate, repeated-approval and failed-connection outcomes, then see what a real recovery design still needs.
Written by Fireflax · 8 minute read
Reviewed by Codex — source and implementation review ·
A retry needs to know what already happened.
An enquiry may arrive twice. A person may approve twice. A connection may fail after part of a workflow has run. These are different situations: before repeating anything, identify the request, the intended action and the evidence of its outcome.
The example used here
This guide uses the fictional Northstar Plumbing enquiry sample, checked on 8 October 2026. Its rules and templates run locally. No email, CRM write or live connection exists. Its failed-connection switch is a controlled simulation, not a production outage test.
The sample keeps the record and asks again.
Received → Needs your approval
Approval succeeds → Demo completed, or Waiting for details if information is missing.
Simulated failure → Connection paused
Restore example connection → Needs your approval → read the draft → approve again.
In the browser demo, reset and run New job. Check Try a failed connection on this approval, then Approve demo reply. Inspect Connection paused and its history. Select Restore example connection, review the text again and approve. The restore button clears the failure checkbox so the next approval can finish the sample.
Reproduce the failure and restore in the demoWhat is retained, and what cannot repeat?
| Path | Expected | Observed |
|---|---|---|
| Same request again | Recognise normalized email, area and message in the current queue. | Existing ID returned; no extra task or history. Changed wording may become a new request. |
| Approval again | Only a record waiting for review accepts approval. | A completed, blocked or waiting-for-details record cannot be approved directly again. |
| Failed approval | Keep enough context to review recovery. | Same record, original input, draft, approved text and history retained in blocked state. |
| Restore | Require fresh review rather than replaying automatically. | Same record and draft retained; prior approval cleared; state returns to review. |
Same request again
Expected: Recognise normalized email, area and message in the current queue.
Observed: Existing ID returned; no extra task or history. Changed wording may become a new request.
Approval again
Expected: Only a record waiting for review accepts approval.
Observed: A completed, blocked or waiting-for-details record cannot be approved directly again.
Failed approval
Expected: Keep enough context to review recovery.
Observed: Same record, original input, draft, approved text and history retained in blocked state.
Restore
Expected: Require fresh review rather than replaying automatically.
Observed: Same record and draft retained; prior approval cleared; state returns to review.
Missing information remains missing after restore. If an incomplete enquiry is approved after recovery, it returns to Waiting for details. A retry does not invent an area, confirm a booking or receive a customer’s answer.
A real partial failure is harder.
Imagine a connected workflow updates a job record, then loses the confirmation before recording success locally. Its local status may be uncertain even though the destination changed. Repeating the whole workflow could repeat an external effect.
That is a hypothetical production case; the Fireflax sample does not simulate it. A real implementation needs durable identifiers for requests and intended actions, a retained record of outcomes, checks of the destination and a rule for uncertain results. Browser memory and matching message text alone are not enough.
The offline starter preserves a local file and offers an explicit pause command, but it is still a single-process rehearsal. It does not establish safe concurrent retries, exactly-once delivery, production uptime or rollback across systems.
Use this recovery checklist.
- Pause and preserve. Stop the relevant workflow through its agreed procedure. Keep the request, versions, error and outcome evidence.
- Inspect the destination. Check whether the message, record or other intended change happened. Classify it as done, not done or uncertain.
- Avoid a blind resend. Resolve uncertainty with the responsible person and the system’s evidence. Check the request/action identifiers and duplicate rule.
- Restore with the permission owner. Fix or reauthorize the actual connection as needed. Restoring access does not authorize replaying every action.
- Review again. Recheck the latest request and any changed conversation, draft or contact preference. Obtain fresh approval where required.
- Verify one intended result. Perform only the authorized next action, inspect its outcome and record what happened before resuming normal work.
Name the owner and backup for these steps. A message already sent cannot be unsent by resetting an internal record. Any correction or reversal needs its own permissions and confirmed limits.
Keep the evidence beside the handover.
The retained source and starter runs cover five defined cases each, including repeated approval and failed approval/recovery. They record synthetic inputs, exact states and history. The separate 21-test repository suite also checks missing-information recovery. These finite checks support the descriptions above; they are not a delivery guarantee.
Sources and review notes
Drafted with Codex and checked against executed fictional sample tests, source code and an independent Codex starter rehearsal. These results do not establish a live integration or customer outcome.
- Retained source test
Actual local duplicate and recovery outcomes.
- Starter README
Inspectable code, replay commands and boundaries.