Alarm Management.
Alarm management is the discipline of designing, deploying, and maintaining the alarm system in a process plant so that operators receive only the alarms they need, when they need them. Bad alarm management causes alarm floods that overwhelm the operator and contributed to several major industrial incidents, Texaco Milford Haven 1994, BP Texas City 2005. ANSI/ISA 18.2 codifies the lifecycle.
Three drawings included on a new account.
What Alarm Management means.
Alarm management is the discipline of making sure an operator receives only the alarms that demand a response, and receives them in time to act. It exists because the marginal cost of adding an alarm in a modern control system is effectively zero, so without discipline the alarm list grows until a process upset produces a flood the operator cannot read, which is precisely the failure mode that contributed to incidents such as Texaco Milford Haven in 1994 and BP Texas City in 2005. ANSI/ISA 18.2 codifies the answer as a lifecycle rather than a one-time configuration. A philosophy that defines what an alarm is, identification of the conditions that warrant one, rationalization that validates each alarm carries a defined operator response and priority, detailed design, implementation, operation, monitoring against rate-based metrics, management of change, and audit. The performance targets the lifecycle works toward, drawn from ISA 18.2 and the EEMUA 191 guidance, are demanding, an average well below one alarm per operator every ten minutes, peaks held in single digits during an upset, almost no alarms left standing for days and most plants miss them without active intervention. The part of this that touches the drawing is narrow. A P&ID may show that an alarm exists and its priority, but the rationalization data, the response, the time to act, the consequence of inaction lives in a separate alarm register. The I/O list anchors which tags can alarm. The alarm register defines how each one should behave.
Lifecycle stages per ISA 18.2
Philosophy, define what an alarm is and isn't, identification, which conditions need alarms, rationalization, validate each alarm is meaningful, prioritize, document the operator response, detailed design, engineering parameters, implementation, deploy in BPCS, operation, monitor performance, maintenance, manage changes, monitoring & assessment, rate-based metrics, MOC, track changes, audit. Most plants live somewhere on the spectrum from fully implemented to philosophy-on-paper-only-and-no-rationalization-ever.
What good alarm-system performance looks like
Average alarm rate per operator. Under 1 per 10 minutes during normal operation. Peak alarm rate during upset. Under 10 per 10 minutes. No more than 1% of alarms standing for over 24 hours, chattering or ignored. No more than 5 alarms in any 10-minute peak. These targets come from EEMUA 191 and ISA 18.2 industry data and are drastically harder to hit than they look. The 10-minute-peak target alone fails on most plants without active intervention.
The ten ISA 18.2 alarm management lifecycle stages
The stages form a loop, not a line. Monitoring feeds management of change, which feeds rationalization again, which is what stops an alarm system decaying after commissioning.
| Stage | What happens | What it produces |
|---|---|---|
| Philosophy | The site defines priorities, classes, performance targets and the rules everything else is measured against | The alarm philosophy document |
| Identification | Candidate alarms are collected from hazard studies, P&IDs, incidents and operating procedures | The candidate alarm list |
| Rationalization | Each candidate is tested against the philosophy: is there a consequence, an operator action, and time to take it | A master alarm database entry with priority, class, setpoint and the documented response |
| Detailed design | The rationalized entry becomes a configuration, including how it appears on the display | The alarm design specification |
| Implementation | The alarm is configured, tested, commissioned, and the operator is trained on it | Commissioned alarms and training records |
| Operation | The alarm is in service and the operator responds to it | Operator response, and the response time achieved |
| Maintenance | Alarms are repaired, tested and periodically verified, and out of service alarms are controlled | Test records and the out of service register |
| Monitoring and assessment | Rate, standing, chattering and the top contributors are measured against the targets | Performance reports against the philosophy |
| Management of change | Any change to an alarm goes back through an authorised route into rationalization | Change records against the master alarm database |
| Audit | Periodic review of the whole system, including the philosophy itself | The audit report and the improvement plan |
The bad alarm patterns, and what each one is
Every one of these is measurable from the alarm history alone, which is why monitoring comes before any redesign.
| Pattern | What it is | Usual cause |
|---|---|---|
| Alarm flood | More alarms in ten minutes than the operator can read, let alone act on | One upset propagating through dependent measurements with no state based suppression |
| Chattering alarm | An alarm that repeatedly annunciates and clears within a short period | A setpoint sitting inside the normal noise band, with no deadband or on delay |
| Fleeting alarm | An alarm that clears before the operator can even see it | A transient that needs a short delay rather than an alarm |
| Stale or standing alarm | An alarm that has been active for hours or days | A known condition nobody has an action for, or a failed instrument left in service |
| Nuisance alarm | An alarm the operator has learned to acknowledge without looking | An alarm that survived commissioning without passing rationalization |
| Duplicate alarm | Several alarms for one condition | The same condition alarmed in the control system, the package and the safety system |
| Priority inflation | Most alarms marked high priority | Priority assigned by the requesting discipline rather than by consequence and time to act |
Common questions
What is an alarm flood.
Where does alarm-management data come from in a P&ID extraction.
When is alarm rationalization required by a standard or regulation.
Get the I/O list template.
Fourteen columns with the signal class as a dropdown, auto-filter on every column, frozen header. Plain .xlsx.