A PRACTICAL BUSINESS GUIDE

Reliable automation knows what to do when the outcome is uncertain.

A timeout does not always mean an action failed. Before repeating an operation, the system should establish whether it might create a duplicate or leave connected tools in disagreement.

Distinguish rejection, failure and an unknown outcome

If an API rejects data because a field is missing, repeating the same request will not repair it. If the service responds slowly and the connection times out, the operation may have completed without confirmation reaching the caller. Treating both outcomes as “retry” can produce incorrect results.

For each integration, define observable states: received, validated, processing, completed, rejected and requiring verification. Preserve the external operation reference when available. The person responding should understand which step succeeded rather than see only a generic “workflow failed” notification. Record enough context to distinguish a new business request from another delivery of an existing one.

Repeat attempts without multiplying their effects

Idempotency allows attempts at the same operation to represent one intention. Amazon describes this principle for safer retries. Actual support depends on the API: some accept a key, while others need checks or an operation register. Do not assume every connector automatically prevents duplicates.

Where an API offers an idempotency key, reuse it only for the same logical operation under that API’s contract. A different request should not inherit an earlier request’s key. Stripe documentation illustrates provider-specific rules rather than a guarantee applicable to other services.

Scroll the table to compare all columns.

Behaviour to define for each outcome
OutcomeSuggested actionAvoid
Invalid dataCorrect it or request reviewRepeating unchanged input
Temporary limit or unavailabilityBounded retries with delay under the API contractInfinite or simultaneous retry loops
Timeout after submissionCheck status or use supported deduplicationBlindly create another operation
Partially completed stepsReconcile states and resume from a safe pointRerunning the entire flow without checking

Illustrative example: order created, response lost

  1. Initial attempt

    The workflow submits an order. The business system records it, but confirmation does not reach the workflow within the expected time.

  2. Verification

    The workflow preserves the same operation identity. If supported, it checks status using an external reference or retries with the required idempotency key.

  3. Intervention

    If the outcome remains unknown, the case enters a queue showing data and attempts. A person checks the business system before authorising another creation.

Safe retries are not a promise of exactly-once execution across the whole system. Check the boundaries and guarantees of every component.

Periodically check that systems agree

Reconciliation compares expected outcomes with recorded ones: a completed workflow case should have the correct reference in the business system, and vice versa. Define meaningful discrepancies, who examines them and how corrections preserve a trace of the change.

Not every step can be undone. A draft can be deleted; a sent email cannot be recovered as easily. Order actions and design corrective steps around this difference. Rolling back code does not undo effects already produced in external systems. Include those outstanding effects in any recovery plan.

Tests to request before release

Document these tests alongside the operating runbook. The benefit is practical: the team should know how to recover a case without asking the supplier to reconstruct every incident from scratch. Include the person authorised to restart processing once reconciliation is complete.

  • Simulate timeouts before and after the external record is created.
  • Deliver the same event twice and verify the result.
  • Stop processing between two systems and resume without duplicates.
  • Exhaust retries and check the queue, alert and case assignment.
  • Compare external and internal states after a manual correction.
FROM IDEAS TO A BRIEF

A template to work from.

System integration map

One connection record containing field mappings, permissions, duplicate handling and a reconciliation check.

Download the Markdown template

Automation exception register

An exception queue with status, priority, ownership and closure evidence; repeated issues become inputs to workflow improvements.

Download the Markdown template

AI operations runbook

Instructions available to the responsible team: signals, alerts, initial diagnosis, recovery and periodic review.

Download the Markdown template

Practical questions

Does a ready-made connector guarantee no duplicates?

No. Check the specific operation, including retries, timeouts, key lifetime and duplicate event handling. Test the complete integration path.

Should every error be retried?

No. Permanent errors and invalid data require correction. Unknown outcomes require verification; transient errors may permit bounded retries under the service rules.

Why keep a manual exception queue?

Some outcomes cannot be resolved automatically with sufficient evidence. An assigned queue with history and references supports controlled recovery instead of another blind submission.

References and method

Amazon Builders’ Library — Making retries safe with idempotent APIs

Idempotency and retry principles. The order scenario is illustrative.

Stripe — Idempotent requests

A concrete idempotency contract example; verify the rules for each API used.

Which process should improve first?

Start with a concrete process, the systems you use and the people who will operate it every day.

Let’s discuss your process