Water Utilities Face Growing Cybersecurity Pressure as OT Risks Rise Reading Automation Can Create a Hidden Single-Point-of-Failure Risk

Automation Can Create a Hidden Single-Point-of-Failure Risk

Automation Can Create a Hidden Single-Point-of-Failure Risk

As factories become more automated, the biggest operational risk may not always be a failed controller, damaged drive or network outage. In some facilities, the more difficult problem is the person who is not there when something goes wrong.

A production line can have operators on duty, maintenance technicians on site and engineers available by phone, yet still lack the expertise required to recover a critical automated process. One controls engineer may be the only employee who understands a legacy PLC program, a customized robot sequence, an undocumented network architecture or the history behind several years of software changes.

When that employee is unavailable, the plant can suddenly discover that its apparent staffing level does not reflect its actual technical coverage.

This form of concentration risk is becoming increasingly relevant as manufacturers deploy more sophisticated industrial automation systems, connected control platforms, robotics, historians, supervisory systems and industrial networks. Automation reduces many forms of manual work, but it can also make specialized knowledge more important. A stable production environment may therefore conceal a fragile dependency on a small number of people.

The problem is not necessarily that a plant has too few employees. The issue is whether the employees available during a failure have the right capabilities to respond.

A technician may be able to replace an I/O module but have no authorization to modify a PLC program. An operator may recognize an alarm but not know how to retrieve the correct configuration backup. An IT specialist may understand the network but have little experience with a machine-level control system. Meanwhile, the engineer who knows the complete system may be working at another site or unavailable until the next shift.

This creates what manufacturers should treat as a human single point of failure.

Traditional workforce planning often focuses on job titles, headcount and training completion. Those measurements remain useful, but they do not necessarily show whether a critical automated asset has genuine backup coverage.

A plant might have five maintenance employees and three controls engineers, for example, but still depend almost entirely on one individual for a particular PLC platform. Another facility may have several automation specialists but no one qualified to support a legacy interface connecting a production machine with an older database.

The difference is capability concentration.

For manufacturers operating PLC systems, DCS platforms, safety systems, robotic cells and industrial communication networks, the first step is to identify which assets require specialist support. The list should be based on operational consequences rather than technology popularity.

A system may be considered critical because its failure affects worker safety, product quality, environmental controls, production throughput or recovery time. In one facility, that could mean a safety controller or robot cell. In another, it could be a process-control system, historian, supervisory platform or custom integration between production equipment and enterprise software.

Once those systems are identified, companies need to define what a successful response actually looks like.

Simply stating that someone needs “PLC knowledge” is not enough. The required capability may involve several different tasks, including recognizing the fault, collecting diagnostic information, checking alarms, locating an approved backup, restoring a known-good configuration, verifying system status and escalating the issue when the problem exceeds the employee's authority.

Software modification is another matter.

A person capable of restoring an approved configuration may not be qualified to make a program change. Similarly, someone who can diagnose a communication failure may not have the authorization to alter network settings. Separating these capabilities creates a more realistic picture of workforce readiness and reduces the temptation to treat every technical skill as interchangeable.

The same principle applies to robotics and other advanced automation platforms. Knowing that a robot exists in the plant does not mean a technician is ready to recover it after a controller fault or software update. Practical competence can depend on the exact robot model, controller generation, programming environment, safety configuration and maintenance procedures used at the facility.

The situation becomes even more complicated across multiple shifts.

A company may have two or three employees who can independently troubleshoot a critical control system, but if all of them work during the day, the night shift still has a coverage problem. An on-call arrangement can reduce that exposure, but only if the person can respond within the required time and has remote access, documentation and the necessary authority.

This is why workforce assessments should be performed at the system, shift and site level, rather than simply at the organizational level.

A useful capability assessment can identify the primary technical owner, the available backup, the backup's demonstrated skill level and the escalation path. It should also show whether the backup has access to current documentation, configuration files, software tools, passwords or other authorized resources needed for the assigned response.

The distinction between nominal backup and functional backup is important.

Adding another employee's name to a spreadsheet does not create redundancy. A backup becomes meaningful only when that person can perform the defined task under realistic operating conditions.

Manufacturers do not need to deliberately disrupt production to test this. Controlled troubleshooting exercises, planned maintenance windows, tabletop scenarios and approved test environments can provide evidence without creating unnecessary operational risk.

For example, a plant could ask a backup technician to locate the current PLC documentation, identify the latest approved configuration, explain the recovery sequence and determine when escalation is required. Another exercise could test whether the technician knows which vendor or system integrator should be contacted and what diagnostic information needs to be collected before making that call.

These exercises often uncover surprisingly basic weaknesses.

The documentation may reference a server that no longer exists. The backup file may be stored in a location the technician cannot access. A vendor support agreement may only cover normal business hours even though the company's recovery plan assumes overnight assistance. A software license may be tied to one engineer's workstation. A password or configuration archive may exist but have no documented ownership.

None of these problems necessarily represents a major technology failure. Together, however, they can significantly extend downtime.

This is particularly important for plants with customized or aging automation infrastructure. A modern controller supported by a large ecosystem may have several available service resources. A legacy application with undocumented modifications can be much harder to transfer between employees.

The value of redundancy therefore depends on the system's characteristics.

Manufacturers should prioritize coverage according to three practical factors: the consequence of losing support, the required response time and the difficulty of transferring the knowledge.

A safety-critical control system with a response requirement measured in minutes deserves a different staffing model from a production reporting application that can remain offline overnight. Likewise, a standard automation platform may be easier to support through an external service provider than a highly customized legacy control environment.

This approach also changes how companies should think about training.

Training completion is not the same as operational readiness.

An employee may have completed a vendor course on a particular PLC system but never performed an independent recovery. Another technician may have worked with the equipment for years but have no experience with the latest controller revision. A third employee may understand the technical process but lack the access rights required to execute it.

A meaningful readiness assessment therefore needs evidence of capability, not just evidence of attendance.

The most effective organizations are likely to combine formal training with supervised troubleshooting, equipment-specific practice, documented recovery procedures and periodic competency checks. This also makes it easier to identify where additional cross-training will produce the greatest reduction in operational risk.

Not every specialist needs a full replacement.

In many plants, the objective should not be to make every technician capable of performing every engineering task. That would be expensive and, in some cases, unnecessary. The more practical goal is to ensure that the first person responding to a problem can safely perform the tasks appropriate to their role and reach the right specialist without avoidable delays.

A tiered support model can help.

An operator may be responsible for identifying the abnormal condition and securing the process. A maintenance technician may perform approved hardware replacement or restoration procedures. A controls engineer may handle program-level diagnostics and authorized changes. An external system integrator or OEM may be brought in when the problem requires deeper expertise.

Such a structure makes responsibilities clearer and prevents companies from assuming that one employee must be capable of handling every layer of the automation stack.

Succession planning should also extend beyond job positions.

A company may have a succession plan for its automation manager while having no replacement for the engineer who understands a particular legacy integration. The job may be covered, but the system knowledge is not.

For this reason, organizations should consider building capability maps around critical automation assets rather than relying solely on conventional organizational charts. The question is not only who can replace an employee. It is who can support the systems that employee currently supports.

These capability maps also need regular maintenance.

Automation environments rarely remain static. A controller upgrade can alter troubleshooting procedures. A new production cell can introduce additional communication dependencies. A cybersecurity project can change access permissions. A software migration can invalidate an old backup process. Even a change in vendor support arrangements can affect the reliability of an escalation plan.

Consequently, workforce coverage should be reviewed after major technology changes, incidents, near misses and significant organizational changes rather than only during an annual training cycle.

The review should have a clear owner. Someone needs to verify that the documentation is current, backups are accessible, permissions remain valid, escalation contacts are correct and designated backups can still perform their assigned tasks.

This matters particularly as manufacturers increase the use of connected automation, remote support and data-driven production systems. The more systems become interconnected, the more difficult it can be to define where technical responsibility begins and ends.

A controls engineer may own the PLC. IT may manage the network. An OEM may support the machine. A cybersecurity team may control remote access. Operations may own the production process. During normal operation, these boundaries can work well. During an incident, however, unclear responsibility can create additional delays.

The answer is not to eliminate specialization. Deep expertise remains essential to modern manufacturing.

Instead, companies need to understand where that expertise is concentrated and decide which dependencies are acceptable.

The most resilient automated plants will not necessarily be those with the largest engineering teams. They will be the ones that understand the capabilities required to keep critical systems operating and have credible alternatives when a key specialist is unavailable.

Automation is designed to reduce dependence on manual intervention. That does not mean manufacturers can ignore dependence on human expertise.

If a production system can only be recovered by one employee, the plant has a risk that may not appear on its automation architecture, maintenance schedule or organizational chart. The equipment may be redundant. The network may be resilient. The backups may exist.

But if the knowledge required to use them belongs to only one person, the operation still has a single point of failure.

For industrial operators, the practical test is straightforward: when the primary specialist is unavailable, can another appropriately trained and authorized employee perform the required response safely, using current documentation and within the expected recovery window?

If the answer has never been demonstrated, the organization should treat that capability as a risk rather than assuming it is covered.

Written by: Daniel Mercer — An industrial automation and controls technology writer with more than a decade of experience covering PLC systems, robotics, process control, industrial networking and manufacturing workforce strategy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Please note, comments need to be approved before they are published.