Retries, fallback paths and alerts

You will be able to configure a workflow so failures retry where safe and alert a person when they do not.

After the week of missing enquiries in lesson 7.1, Mei Ling wanted two things. If a step failed for a passing reason, such as a brief outage, she wanted the flow to try again by itself. If it failed for a real reason, she wanted to know that day, not when a parent rang to complain. This lesson covers both, plus a third safeguard that catches the failures that never raise an error at all.

Retry, or take another path

Most automation tools give you two ways to respond when a step fails.

The first is a retry: run the failed step again, either automatically after a short wait or by hand from the run history. Some tools retry certain errors on their own. Others let you switch on automatic retries per step or per flow, or replay a failed run once you have fixed the cause. Retries are the right answer for failures that go away by themselves, such as a timeout, a rate limit or a short outage.

The second is an error path, sometimes called an error handler or fallback route. It is a separate branch that runs only when a step fails. Instead of the whole run stopping, the flow goes down the error path and does something useful: sends an alert, writes the failed record to a "needs attention" sheet, or tries an alternative action. Mei Ling's error path on the "add row" step writes the enquiry to a backup sheet and alerts her, so the parent's details are never lost even if the main sheet is unreachable.

Look in your tool's help pages for terms such as error handling, retries, replay, error routes or error workflows to see what yours offers.

Retry only what is safe to repeat

Retries carry a risk that is easy to overlook. Some steps change something every time they run. If the first attempt actually worked but the tool did not get a clear response, a retry does it a second time.

Think about which steps in your flow are safe to repeat. Looking up a row is safe, because doing it twice changes nothing. Updating a row to set a status is usually safe, because setting it to "sent" twice leaves it as "sent". Adding a row is not safe, because a retry can create a duplicate. Sending an email is not safe, because the parent gets two. Making a payment is very much not safe.

For steps that are not safe to repeat, either switch automatic retries off and send an alert instead, or put a check in front, like the lookup from lesson 4.3, Stop duplicates and missed rows, so a second attempt finds the first and stops. Never set automatic retries on a payment or a deletion. Those failures go to a person.

Alerts that reach someone

An alert only works if a person sees it in time. The automation tool's own error emails are a start, but they often go to whoever created the account, land among dozens of other notifications, or arrive in a format that says little more than "your flow had an error".

Build your own alert as part of the error path. Send it to a channel you actually read every day. That might be email, if you live in your inbox, or a team chat app such as Slack or Telegram, if that is where your team talks. Mei Ling sends hers to a small Telegram group with herself and the centre manager.

Make the alert useful on its own. Include the flow name, the step that failed, the error message the tool reported, the record ID and the time. With the record ID from lesson 2.2, Data mapping: passing fields from one step to the next, the person reading the alert can find the exact enquiry in seconds. An alert that just says "something failed" sends them hunting through the run history instead.

Keep alerts for things a person needs to act on. If every successful run also sends a message, people stop reading them, and the one that matters gets lost among the rest.

A weekly summary catches the quiet failures

Retries and alerts handle failures the tool can see. They do nothing for the trigger that stops firing, the filter that silently rejects good records, or the mapping that writes blanks and reports success. Lesson 7.1, How automations fail, and why you often do not notice, explained why those quiet failures go unnoticed.

The defence is a regular summary. Once a week, have a scheduled flow, or a calendar reminder for you, report a few numbers for each flow: how many runs happened, how many failed, and the count you expect from the source, such as form responses received. Lesson 4.3 introduced the form-versus-sheet comparison. A summary turns it into a habit.

If the enquiry flow normally runs about fifteen times a week and this week shows two runs, something is wrong even though no error appeared. Mei Ling's weekly summary would have caught her broken connection within days instead of after a parent's call.

Add your first alert

In the activity below you will add an error alert to your enquiry flow, sending you the step name, the error and the record ID. Once it is in place, trigger a failure on purpose to make sure the alert arrives where you expect it.

Add an error alert to your enquiry flow that sends you the step name, the error and the record ID.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).