← Retour au journalJOURNAL / 05
Business & technologie3 min de lecture

Reliable AI automation: retries, approvals and recovery

Design business automation that can recover from interruptions, avoid duplicate actions and hand exceptions to the right person.

Cet article est présenté dans sa version originale anglaise.

The short answer

A reliable automation needs more than a successful first run. It must know what has already happened, which actions can be repeated and when a person needs to decide. Split the job into observable steps, preserve progress, and make recovery part of the design. Use AI only where interpretation or judgment adds value to the workflow.

Start with the work people actually do

Map one complete task with the people who perform it. Record the trigger, information needed, systems touched and definition of done. Pay attention to the exceptions they handle from memory: a missing invoice number, an unfamiliar supplier or a request arriving twice.

For an illustrative supplier-onboarding process, AI could help extract fields from documents. Deterministic checks can validate required fields. A named employee can approve the new supplier before any downstream action. This is a proposed pattern, not a claim about a completed client project.

Break the workflow at meaningful boundaries

A useful sequence is receive, extract, validate, request approval, write and notify. Store the result of each completed stage. If notification fails after a record is saved, recovery should resend the notification without recreating the supplier.

Cloudflare’s Workflows documentation describes steps as individually retryable components and stresses idempotency: repeating an operation should not create an additional effect. Read the Workflows rules.

The product decision is which boundaries matter to your team. A technically successful API call may still leave the business task incomplete. Model the state your operators need to see.

Design retries around the action

Give each business operation a stable identifier, and use downstream idempotency support where available. Persist the association between that identifier and the completed result. When a request times out, check whether the destination accepted it before issuing another action that could duplicate work.

Retry temporary failures with a bounded policy. Send permanent validation failures to an exception queue with a clear reason. Cloudflare documents per-step retry configuration and pauses; the appropriate values depend on the downstream system and the time the business can tolerate waiting. See retry configuration.

Make the human handover useful

An approval should show what will happen, why it was proposed and the relevant evidence. Include the exact fields that will be submitted. If those fields change after approval, require a new decision. Avoid a vague “approve automation” button that hides the action.

  • Assign an owner to each exception type.
  • Show the last completed step and the reason for stopping.
  • Provide a safe way to resume, cancel or correct the task.
  • Set an escalation time for work that remains unattended.

These are our suggested design checks. Validate them with the people operating the process, including colleagues who did not build it.

Measure time saved honestly

Compare equivalent work before and after automation. Separate elapsed time from staff time, record the number of items processed and include the time spent reviewing exceptions. A faster happy path can conceal extra manual work elsewhere.

Does every automation need an AI agent?

No. If the sequence and rules are known, a conventional workflow may be simpler to operate. Add a model where it helps interpret variable input or propose a decision.

What makes a good first project?

A repeated task with clear inputs, an observable outcome and an owner who can evaluate the result. Explore our automation capabilities or tell us which task takes too much of your team’s time.