• ISA provides technical resources and standards to help industrial automation professionals advance their careers and the field. We enable automation professionals worldwide to solve problems and enhance their skills by bringing people together to create new technologies and share best practices with future automation professionals.
    • Industry Insights

  • We attract over 140,000 unique automation professionals monthly, making us the premier online content provider and the only dedicated electronic magazine in the automation industry.

    Monthly Magazine

    Back
    Back
  • M logo for Automation.com Monthly. Link to current issue.

A Backup Is Not a Recovery Plan: Measuring What It Really Takes To Restore OT After an Incident

By: Scott Taylor
Source: Materion Corporation
02 October, 2026
5 min read
Hands typing on a keyboard with illustrated icons for disaster recovery, data protection, information security, and digital strategies.
In addition to a successful backup, nine conditions determine whether OT recovery takes hours, days or weeks.

Segmentation, monitoring, access control and vulnerability management reduce risk, but none eliminates every incident. When one gets through, the operational consequence is measured in the hours between the stop and the restart. If recovery has not been tested, that number is still an estimate.

Programmable logic controller (PLC), human-machine interface (HMI) and configuration backups are common. Measured restoration times are rare. NIST CSF 2.0 includes Recover among its six core functions, and ISA/IEC 62443 treats resource availability as a foundational security requirement. Recovery sits inside the security program.

A backup confirms that a file was captured. Putting that file back into production takes hardware, firmware, licensing, credentials, documentation, vendor support and people. After a security event, it also requires confidence that the file can be trusted.

On one line, an HMI watchdog alarm signaled a loss of communication with the PLC and the process stopped. The controller had failed. The site had a spare purchased with the equipment, a current backup and a troubleshooting laptop. But the spare carried different firmware, and the technician on the floor had never loaded a program. The team had to choose between converting the backup and matching the original firmware, and then reach a senior technician who was traveling to talk through the lower-risk path. A swap expected to take one hour took four.

Nothing in that sequence was a backup problem.

The file is the easy part

A backup is a PLC program, an HMI application or a configuration export sitting on a server or in a cloud archive. Recovery is the work that surrounds that file, everything required to get the system restored, validated and returned to production safely. The nine conditions below are where that work stalls.

Advertisement

Nine conditions that determine recovery time

  1. Firmware compatibility. The backup was taken on a controller running one firmware version. Replacement hardware, pulled from spares or a distributor, may run another. Restoring the program may then require a firmware change, a project file migration or vendor confirmation that the versions are compatible.
  2. Replacement hardware availability. The failed controller may have been discontinued years ago. Either there is a spare on site, or procurement is sourcing an obsolete part with a lead time measured in weeks. Asset lifecycle planning should catch this, but spares strategy and backup strategy are often managed separately.
  3. Software availability and licensing. Supervisory control and data acquisition (SCADA) platforms, HMI runtimes and engineering workstations may depend on specific installation media, versions, license keys or online activation. If the license lived on the machine that failed, reactivation becomes the delay.
  4. Credentials. The recovery team needs access to the replacement device, switches, servers, historian and remote access gateway. If those credentials live in one engineer’s password manager or an unmaintained spreadsheet, recovery stalls at the login screen.
  5. Current documentation. Current documentation is not the commissioning drawing, but the one reflecting the changes made since. Outdated documentation sends people down paths that no longer match the plant.
  6. Active vendor support. The contract has to be current, someone has to know the account number or which distributor holds the relationship, and the person opening the ticket has to be an authorized contact. Administrative questions become urgent in the middle of an outage.
  7. Network configuration. IP addressing, VLAN assignments, firewall rules and switch configurations have to be right and backed up, or the recovered controller comes back online and still cannot talk to anything. Restoring control logic and restoring the network are different skill sets, and recovery can stall at that handoff.
  8. People who know the restoration sequence. Recovery assumes someone knows the order of operations: which systems come up first, what gets verified before the next step and what working correctly looks like when everything is back online. That knowledge is rarely documented in enough detail, and it is exactly what is lost when a senior engineer retires, changes employers or is traveling when something fails.
  9. Usable, trusted backup copy. After a cyber incident, the team must know which backup predates the compromise and whether it has been validated. A copy stored where an attacker can alter, encrypt or delete it may not be recoverable.

 

  What a backup covers versus what it takes to recover.

Why this gap persists

Backup success is easy to confirm. Where it is automated, a job completes and a dashboard turns green; where it is manual, someone confirms the file was pulled. Either way, the confirmation is about capture. Recovery capability depends on all of those conditions holding at once, and most of them are invisible until someone attempts a restoration.

During a cyber incident, several of those conditions degrade together. Backups stored on the production network may be encrypted or suspect. The credentials may have been rotated during containment or may be what the adversary used. Restoration cannot begin until someone can vouch for a validated copy.

Part of the gap is organizational. Business continuity and disaster recovery programs may cover enterprise information technology (IT) while leaving plant control systems outside the recovery scope. When operational technology (OT) recovery has no clear owner, testing and maintenance are easily deferred.

Demonstrated recovery time belongs among the measured security outcomes for systems whose failure materially affects production.

What tested recovery looks like

  • Treat recovery as an operational capability. Define ownership, set testing frequency by system criticality, document results and assign responsibility for closing findings. For systems that stop production, repeat testing on a defined schedule and after material changes.
  • Rehearse the restoration. Perform it on a bench or during a planned outage with the hardware, firmware and licensing the site would actually have on hand and measure the result.
  • Document the restoration sequence. A network diagram shows how the system looks when it is running. It says little about the order in which systems come back or what gets verified along the way. That sequence deserves a short, plain-language document written for whoever may have to perform the restoration on an off-shift.
Advertisement

 

  A four-question self-check for the systems where downtime matters most.

Recovery should follow criticality

Not every system warrants equal rigor. A low-volume asset does not carry the urgency of a line responsible for a significant share of plant output. Where equivalent cells can share the load, a longer recovery target is defensible. A single machine that is the only path to a finished product earns the shortest target.

This is the logic behind recovery time objectives (RTOs) in formal disaster recovery planning. An RTO is a business target. A controlled recovery exercise shows whether the organization can meet the target, and the two numbers are often different because the assumptions behind the target were never validated.

Recovery point objectives deserve the same attention. Backup frequency determines how much program and configuration history may be lost, while backup architecture determines whether a usable copy survives the incident. For critical systems, at least one validated copy should be offline, immutable or otherwise isolated from the credentials and network paths an attacker could compromise.

The real question

Whether backups exist is the easiest question to answer, and it is no longer the most important one.

In systems where downtime significantly affects production, customers or safety, the key questions are these: How many hours will pass before the system is restored, validated and returned to production, and has anyone demonstrated that number under controlled conditions?

In my experience, it has not been demonstrated. The site with the spare on the shelf, the current backup and the laptop on the bench had every reason to believe the number was one hour. Take the system that matters most, attempt the recovery under controlled conditions and time it honestly.

Advertisement

Trending Articles

Advertisement

Related Articles

View all Articles and News
Advertisement
Advertisement