Welcome to Unit 3
OTHM Level 6 International Diploma in Occupational Health and Safety. Choose the section you want to study. Each section opens its complete explanations, workplace examples, calculations, activities and interactive tools.
Systems Reliability and Failure Tracing
- 1.6.3Common reliability calculations
- 1.6.4Analyse system performance
- 1.6.5Use evidence to trace failures
- ToolsSystem explainer, nine calculators, analysis and tracing activities
Strategies and Techniques of Risk Control
- 2.1Evaluate common risk-management strategies
- 2.2Justify when each strategy should be used
- 2.3Develop SSOWs and SOPs
- ToolsAssessment formats, risk matrix, strategy planner, SSOW and SOP builders
Loss Causation and Loss-Data Analysis
- 3.1Outline loss-causation theories and techniques
- 3.2Justify quantitative analysis of loss data
- ModelsBird, multi-causality, Swiss Cheese, FTA, ETA, Bowtie and behavioural RCA
- ToolsCause classifier, barrier simulator, logic calculators, rate and trend analysis
Systems Reliability and Failure Tracing
Welcome to Section 01. Begin with system reliability, complete all nine calculations in 1.6.3, analyse what the results mean in 1.6.4, and use the evidence to trace failures in 1.6.5.
CEO & Founder, DB HSE International
Meet Your Tutor
Debjyoti Biswas is an Approved Tutor of the OTHM Level 6 International Diploma in Occupational Health and Safety Qualification and a CertIOSH Member. His teaching approach connects Level 6 knowledge with practical workplace application, evidence-based judgement, leadership and professional development.
What Is Reliability?
Let us begin with one easy question:
If you ask a machine, alarm, pump or complete safety system to do its job, how confident are you that it will perform correctly without failing?
That confidence—expressed as a probability—is reliability.
The exact job the item must perform. A fire pump must deliver water; a gas detector must detect gas.
The real environment and duty: temperature, pressure, dust, workload, operating mode and maintenance condition.
The period for which successful operation is required—for example, a 100-hour mission or an emergency demand.
The item completes the required function without stopping, leaking, giving a false signal or becoming unavailable.
The answer is normally written between 0 and 1, or as a percentage from 0% to 100%.
Now apply that idea to a complete system.
A system is a group of connected parts—equipment, power, controls, information, people and procedures—working together. System reliability asks whether the complete required function succeeds.
Explore One Safety Function—One Step at a Time
Imagine a gas-release protection system. Select each part to understand its job.
Component reliability
Asks whether one item—such as a sensor, bearing, seal or pump—will perform successfully.
System reliability
Asks whether the complete required function succeeds. Reliable individual parts do not automatically guarantee a reliable system because connections, power, controls, procedures and common causes also matter.
Not only quality
A well-made item may still be unreliable if it is used outside its design conditions or maintained poorly.
Not the same as availability
An item can fail frequently yet remain highly available when every repair is very fast.
Not proof of safety
Reliability supports risk decisions, but safety also depends on consequence, design, inspection, people, procedures and independent protection.
The DB HSE Emergency-Protection System
Yes—we are going to use one connected workplace scene to explain all nine calculation techniques.
Imagine that you and I have entered this chemical warehouse. The protection system contains gas detection, a control panel, an alarm, an isolation valve, replaceable sensor cartridges and two fire-water pumps. We will not keep changing the story. We will simply ask nine different reliability questions about the relevant parts of this one system.
| Part of our scene | Information collected | Why we need it |
|---|---|---|
| Repairable protection pump | 5,000 operating hours; 10 failures; 40 total repair hours | MTBF, MTTR, λ, availability, R(t) and F(t) |
| Required mission | The pump must operate for the next 100 hours | Reliability and failure probability over time |
| Five replaceable sensor cartridges | Lives: 1,800; 2,100; 1,900; 2,200; 2,000 hours | MTTF for non-repairable items |
| Series alarm chain | Detector 98%; control panel 97%; alarm 99% | Reliability when every required link must work |
| Two backup fire-water pumps | Each pump has 90% reliability for the stated demand | Parallel reliability when at least one independent pump must work |
Watch how the same scene gives us nine different questions
How long does the repairable pump operate, on average, between failures?
How long does one pump repair take, on average?
How frequently does the pump fail per operating hour?
What percentage of time is the pump ready for use?
What is the chance that it completes 100 hours without failure?
What is the chance that it fails during those 100 hours?
What is the average life of a replaceable sensor cartridge?
Will the detector, control panel and alarm all work together?
Will at least one of the two independent fire pumps work?
Why Do HSE Professionals Need These Calculations?
Statements such as “the machine fails frequently” or “the alarm is usually reliable” are opinions. Reliability calculations convert operating, failure and repair records into evidence that can be compared, investigated and improved.
Measure
Determine how often equipment fails, how long it runs and how long repairs take.
Trace
Use trends and component records to locate recurring failures and investigate their causes.
Improve
Verify whether maintenance, redesign, training, spare parts or redundancy improved performance.
Questions the calculations answer
- How frequently does the equipment fail?
- How long does it operate before failing?
- How long is normally needed to repair it?
- What proportion of time is it available?
- What is the chance of completing a task without failure?
- Which component repeats?
- Did maintenance improve performance?
Safety-critical examples
Fire pumps, emergency alarms, pressure-relief systems, local exhaust ventilation, gas detectors, lifting equipment, emergency generators, rescue equipment and emergency shutdown systems.
The required reliability should reflect the consequence of failure and the availability of independent protection.
Understand Every Symbol Before Calculating
Select a symbol to see what it means and where it is used.
Total operating time — T
T represents the total time for which equipment actually operated. It must use the same time unit throughout the calculation, such as hours, days or cycles.
Explore Each Reliability Calculation
Choose a calculation, change the figures and examine how the result and interpretation respond.
Mean Time Between Failures — MTBF
Mean means average. MTBF is the average operating time between one failure and the next failure of repairable equipment.
T = total operating time
N = number of failures
Why it is relevant
- Compares the reliability of machines or operating periods.
- Supports preventive-maintenance planning.
- Shows whether failures are becoming more frequent.
- Provides evidence for investigating deterioration or repeated component failure.
Important interpretation
A higher MTBF is generally better. A falling MTBF—such as 1,000 → 600 → 300 hours—means failures are occurring more frequently and should be investigated.
Live MTBF calculator
See the complete worked example
- Record total operating time: T = 5,000 hours.
- Record failures: N = 10.
- Divide 5,000 by 10.
- MTBF = 500 hours.
- Interpretation: the machine operated for an average of 500 hours between failures.
MTBF trend inspector
What can make MTBF fall?
Equipment deterioration, ageing parts, ineffective maintenance, overloading, poor lubrication, incorrect operation, harsh conditions or repeated failure of the same component. Use maintenance records, failure reports and operating evidence to test these possibilities.
Mean Time to Repair — MTTR
MTTR is the average time required to detect, diagnose, repair, test and safely return failed equipment to service.
Why it is relevant
- Evaluates maintenance response and recovery performance.
- Supports staffing, competence and spare-parts planning.
- Identifies access, diagnostic or permit-to-work delays.
- Helps reduce downtime without compromising safe isolation and testing.
Failure-tracing clue
If MTBF remains stable but MTTR increases, the machine is not necessarily failing more often; the organisation is taking longer to restore it. Investigate the repair process.
Live MTTR calculator
What can increase MTTR?
Slow diagnosis, unavailable spare parts, poor equipment access, insufficient competence, delayed authorisation, inadequate tools, weak procedures or complex testing and restart requirements.
Failure Rate — Lambda (λ)
Failure rate measures how frequently a system or component fails during its operating time. The Greek letter λ is pronounced “lambda”.
λ = failure rate
N = number of failures
T = total operating time
Relationship with MTBF
When a constant failure rate is assumed, higher MTBF normally means a lower failure rate.
Important limitation
The simple relationship λ = 1 ÷ MTBF assumes a reasonably constant failure rate. It may not represent early-life defects, rapid deterioration or age-related wear.
Live failure-rate calculator
Why express it per 1,000 hours?
A rate of 0.002 failures per hour may feel abstract. Multiplying it by 1,000 gives two failures per 1,000 operating hours, which is easier to communicate to managers and workers.
Where is failure rate used?
To compare equipment, monitor deterioration, establish inspection intervals, evaluate maintenance, estimate future failure probability and prioritise improvement. If a seal records 8 of 12 pump failures, trace seal selection, alignment, pressure, temperature, installation and maintenance.
Availability — A
Availability is the proportion of time that equipment is operational and ready to perform its required function.
A = availability
U = uptime
D = downtime
Why it is relevant
Availability is critical for fire pumps, emergency generators, gas detectors, alarms, ventilation and emergency communication systems that must be ready when required.
Availability is not reliability
A machine may fail regularly yet still show high availability if it is repaired quickly. Reliability asks whether it will operate without failure; availability asks whether it is ready for use.
Live availability calculator
How can availability be improved?
Increase MTBF; safely reduce MTTR; provide competent maintainers, fault detection and essential spares; improve preventive maintenance; remove repeated causes; and provide reliable, independently tested backup where justified.
Reliability Probability — R(t)
R(t) is the probability that a system performs its required function without failure for a specified time, t.
R(t) = reliability during time t
e = mathematical constant, approximately 2.718
λ = failure rate
t = required operating time
Why it is relevant
It estimates whether equipment can complete a continuous production period, lifting operation, emergency duty, shutdown task or inspection interval without failure.
Assumption
This exponential formula assumes a constant failure rate. It is an estimate based on recorded data, not a guarantee of future performance.
Live reliability calculator
Probability of Failure — F(t)
F(t), also called unreliability, is the probability that equipment will fail during a specified operating period.
F(t) = probability of failure
1 = total probability, equal to 100%
R(t) = probability of successful operation
Why it is relevant
Management can compare the probability of failure with risk tolerance and decide whether maintenance, inspection, standby equipment, replacement or additional controls are required.
Always check the total
Reliability and failure probability must add to 100%: R(t) + F(t) = 1.
Live failure-probability calculator
Connect this with the previous example
If R(100) = 81.87%, then F(100) = 100% − 81.87% = 18.13%. The machine therefore has an estimated 18.13% probability of failure during the 100-hour period.
What decisions can F(t) support?
Whether risk is acceptable; whether maintenance or inspection is needed before operation; whether independent standby equipment is required; whether replacement is justified; and whether emergency arrangements must be strengthened.
Mean Time to Failure — MTTF
MTTF is the average operating life of a non-repairable item—an item that is normally replaced after failure.
Why it is relevant
MTTF supports component selection, life-cycle planning, stock control, purchasing and the preventive replacement of fuses, disposable sensors, certain batteries and other non-repairable parts.
MTBF versus MTTF
MTBF is mainly for repairable equipment and measures time between failures. MTTF is mainly for non-repairable items and measures time until failure.
Live MTTF calculator
Examples of non-repairable items
Disposable sensors, fuses, certain electronic parts, light bulbs, single-use protective devices and some batteries. They are normally replaced, so MTTF—not MTBF—is the appropriate average life measure.
Reliability of a Series System — Rs
In a series system, every component must work for the complete system to succeed. If one component fails, the complete function fails.
Rs = complete series-system reliability
R1, R2, R3 = reliability of individual components
Workplace example
A safety alarm may require a sensor, control unit and audible alarm. All three must operate successfully. The system’s reliability will therefore be lower than the reliability of its best component.
Live series-system builder
Why does the overall value fall?
Each additional required component creates another point at which the system can fail. Improving the weakest component can have an important effect on the complete system.
Reliability of a Parallel System — Rp
Rp is pronounced “R sub p”. The subscript p means parallel. A parallel system contains backup or redundant components and continues to operate if at least one independent component remains functional.
Rp = parallel-system reliability
1 − R = probability that a component fails
Why it is relevant
Redundancy may improve the reliability of fire pumps, emergency generators, ventilation, alarms, communication systems and emergency shutdown arrangements.
Common-cause failure
Two backups are not fully independent if they share the same electrical supply, fuel source, control panel, cooling system, room or maintenance error.
Live parallel-system builder
Combined Interpretation Dashboard
These figures describe different aspects of the same machine. One figure alone does not provide a complete reliability assessment.
Level 6 interpretation
The machine has high availability because failures can be repaired quickly, but its probability of completing 100 continuous hours without failure is only 81.87%. High availability must not be incorrectly presented as proof of high reliability or safety.
Use Calculations for Failure Tracing
Collect operating and failure data
Gather operating hours, downtime, repair duration, failure type, failed component, operating condition and maintenance history. Calculations are only as reliable as the data used.
Move From a Number to a Defensible Judgement
Is it acceptable?
Compare the result with the performance target, manufacturer information, legal or organisational requirements and the safety function.
Is it changing?
Compare time periods. Decide whether MTBF is falling, failure rate or MTTR is rising, and whether availability or mission reliability is deteriorating.
What should happen?
Identify weak equipment, consider consequences, investigate causes, prioritise action and allocate competent people, time, spares and budget.
Questions an HSE professional should ask
- Is the result acceptable for the required safety function?
- Is performance improving, stable or deteriorating?
- How does this period compare with earlier periods?
- Which equipment or component is weakest?
- What are the safety, health, environmental and business consequences?
- Is further inspection or investigation required?
- What management action and resources are justified?
- Is the data accurate, complete and collected on a consistent basis?
Four Ways to Analyse Performance
1. Establish a baseline
A baseline is the starting performance against which future results are compared. Example: MTBF 1,000 h; MTTR 4 h; λ 0.001/h; availability 99.60%; R(100) 90.48%.
The same definition of failure, operating-time boundary, repair-time boundary and units must be used in every later comparison.
2. Analyse the trend
One result is a snapshot. A sequence reveals direction. Falling MTBF together with rising λ and MTTR is stronger evidence of deterioration than one isolated value.
3. Compare equipment
Compare like with like, then explain differences in age, condition, duty, workload, environment, design, operator competence, maintenance, materials and modifications.
4. Compare with a target
A fire pump may have a target availability of at least 99.5%, MTTR below 3 h and λ below 0.001/h. Actual values of 98.7%, 5 h and 0.002/h are adverse gaps requiring action.
| Indicator | Q1 | Q2 | Q3 | Q4 | Interpretation |
|---|---|---|---|---|---|
| MTBF (h) | 1,000 | 800 | 600 | 400 | Failures are becoming more frequent. |
| MTTR (h) | 4.0 | 4.5 | 5.0 | 7.0 | Recovery is becoming slower. |
| λ (failures/h) | 0.00100 | 0.00125 | 0.00167 | 0.00250 | The deterioration is consistent across indicators. |
When should performance be treated as abnormal?
Do not judge a figure in isolation. Compare it with all relevant reference points:
- The asset’s previous performance
- Similar equipment performing the same duty
- Manufacturer specifications and design limits
- Internal standards and maintenance targets
- Industry benchmarks and recognised good practice
- The reliability required for the safety function
A statistically unusual result, a continuing adverse trend or any failure that threatens a critical protection function requires investigation—even when the percentage appears numerically high.
Compare Two Performance Periods
Change the data. The tool calculates MTBF, MTTR, λ, availability, R(t) and F(t), then explains the direction of change.
Period 1 — baseline
Period 2 — current
| Indicator | Period 1 | Period 2 | Direction |
|---|
One Indicator Can Mislead
High availability can hide frequent failure
If MTBF = 500 h and MTTR = 2 h, availability is 99.60%. Yet, with λ = 0.002/h, R(100) is only 81.87%. Rapid repair creates high availability but does not make the system failure-free.
Consequence changes the priority
Risk combines likelihood and consequence. The same failure probability may be tolerable for a non-critical printer but unacceptable for a gas detector, fire pump or emergency shutdown system.
| Indicator | Pump A | Pump B | Meaning |
|---|---|---|---|
| MTBF | 1,200 h | 450 h | Pump B fails more frequently. |
| MTTR | 3 h | 8 h | Pump B is slower to restore. |
| λ | 0.00083/h | 0.00222/h | Pump B has the higher failure rate. |
| Availability | 99.75% | 98.25% | Pump B is less ready for use. |
Prioritise for investigation when evidence shows:
- High or rising failure rate
- Falling MTBF
- High or rising MTTR
- Availability below target
- Unacceptable F(t)
- Repeated failure of one component
- Safety-critical function
- No effective backup
Start With a Clearly Defined Failure Event
Failure identification
States what failed, where and when: for example, “Fire Pump B failed to start during the weekly proof test at 09:20.”
Failure tracing
Examines how and why the failure developed, how it moved through the system and which immediate, underlying and root causes must be controlled.
Common triggers
- Sudden or continuous reduction in MTBF
- Rising λ or MTTR
- Availability below target
- Unacceptable failure probability
- Repeated failure of one component
- Large differences between similar assets
- Failure soon after maintenance
- Failure of a safety-critical item
- Primary and backup equipment failing together
- Deterioration in previously reliable equipment
Evidence to collect before concluding
- Operating hours, cycles and duty
- Failure date, time, type and alarms
- Repair time, tests and replaced parts
- Maintenance, inspection and proof-test history
- Operators, contractors and shift conditions
- Manufacturer information and design limits
- Temperature, pressure, vibration and process data
- Dust, moisture, corrosion and other environmental conditions
- Photos and preserved damaged components
- Permit-to-work and isolation records
- Training, competence and handover records
- Previous investigations and management-of-change records
Select the Result That Changed
Twelve Steps From Function to Verification
Select each step. A competent investigation may move back and forth as new evidence appears.
Do Not Stop at the First Technical Explanation
Immediate cause
The direct event closest to the failure: a bearing overheated and seized.
Underlying cause
The local condition that allowed it: insufficient lubrication and no condition warning.
Root / organisational cause
The management-system weakness: the maintenance system did not specify, schedule or verify lubrication and monitoring.
Five Whys — pump example
Fault Tree Analysis — “fire pump fails to start”
- Electrical: loss of supply, open protection, starter defect.
- Mechanical: motor seizure, pump obstruction, coupling failure.
- Control: sensor, logic, signal, set-point or interlock failure.
FTA helps test multiple credible paths instead of accepting the first visible defect.
Failure Modes and Effects Analysis — FMEA
FMEA is a forward-looking method. For each component or process step, identify the possible failure mode, its effect, its likely cause, existing controls and any further action. It helps prioritise what could fail before an incident occurs.
| Item | Failure mode | Possible effect | Possible cause | Existing / further control |
|---|---|---|---|---|
| Pump seal | Leakage or face damage | Loss of containment and pump shutdown | Misalignment, vibration or unsuitable material | Alignment verification, material review and vibration monitoring |
RPN is a prioritisation aid, not a universal risk-acceptance rule. Follow the organisation’s FMEA rating definitions; a catastrophic severity may require action even when the total RPN is not the highest.
Find the Component Dominating the Failure Record
Example total = 12 failures. The tool sorts the components and calculates percentage and cumulative percentage.
| Component | Failures | Share | Cumulative |
|---|
Human, Organisational, Common-Cause and Hidden Failures
Human & organisational factors
Examine workload, fatigue, supervision, shift handover, competence, interface design, communication, production pressure, resources, contractor control, unclear roles and ignored warnings.
Common-cause failure
Parallel equipment may share electricity, control panel, fuel, cooling, room, procedure, maintenance error or defective component batch. Redundancy is valuable only when independence is credible.
Hidden failure
An alarm, emergency generator, relief valve, gas detector, interlock or shutdown may fail without being noticed until demanded. Inspection and proof testing must reveal dormant failure.
Recurring Pump-Seal Failure: Before and After
Initial evidence — 6,000 operating hours
12 failures, 60 repair hours; 8 of 12 failures involved the seal.
- MTBF = 6,000 ÷ 12 = 500 h
- MTTR = 60 ÷ 12 = 5 h
- λ = 12 ÷ 6,000 = 0.002/h
- A = 500 ÷ 505 × 100 = 99.01%
- R(100) = e−0.2 = 81.87%
Tracing findings
Evidence showed excessive vibration and shaft misalignment. Seals had been replaced repeatedly without checking alignment; no vibration monitoring existed, and the maintenance procedure addressed replacement but not the failure cause.
Failure
Pump stopped because leakage exceeded the safe limit.
Mode & immediate cause
Seal faces damaged by excessive vibration.
Underlying & root cause
Misalignment plus a procedure that omitted alignment verification and vibration monitoring.
Corrective actions
Realign the pump, replace damaged parts, introduce vibration monitoring, revise the maintenance procedure, train maintainers, verify alignment after intervention and review similar pumps for the same weakness.
| Indicator | Before | After next 6,000 h | Change |
|---|---|---|---|
| Failures | 12 | 3 | 75% reduction |
| MTBF | 500 h | 2,000 h | Longer between failures |
| MTTR | 5 h | 3 h | Faster safe recovery |
| λ | 0.002/h | 0.0005/h | 75% reduction |
| Availability | 99.01% | 99.85% | Improved readiness |
| R(100) | 81.87% | 95.12% | Improved mission reliability |
Verification: the recalculated results support the conclusion that the actions improved performance. Continue monitoring to confirm that the improvement is sustained and has not introduced another risk.
What the Calculations Can and Cannot Prove
They can show
- Frequency, average duration and direction of change
- Comparison with another asset, period or target
- Which component dominates the recorded failures
- Whether performance improved after action
They cannot prove alone
- The physical, human or organisational root cause
- That the data definition and records are accurate
- That a high availability figure means the system is safe
- That a short-term improvement will continue
Failure-tracing record
Record the asset and required function; defined failure event; operating hours and duty; failed component and failure mode; immediate, underlying and root causes; evidence reviewed; risk and consequence; corrective actions, owner and due date; post-action calculations; proof of effectiveness; residual risk; and lessons shared with similar systems.
Reliability Case Studies
Compressor comparison
Compressor A operates for 12 months without failure. Compressor B requires repair every six weeks.
LEV performance deterioration
Capture velocity falls from 0.52 m/s to 0.28 m/s while fume concentration rises from 1.2 mg/m³ to 4.8 mg/m³.
Test Your Understanding
Understand the Strategies and Techniques of Risk Control
Welcome to the next part of our learning journey. Section 01 helped us identify, calculate and trace risk-related evidence. Section 02 asks the next management question: What should we do about the risk, and how can we justify that decision?
Learning outcomes: 2.1 Evaluate the use of common risk-management strategies. 2.2 Justify when to use risk avoidance, risk reduction, risk transfer, risk analysis, risk evaluation and risk review strategies. 2.3 Explain the development and characteristics of safe systems of work and safe operating procedures.
What Is Risk Control?
Imagine that we find an unguarded moving part on a machine.
The moving part is the hazard. The possibility and seriousness of someone being injured is the risk. The guard, isolation system or redesign used to prevent contact is the control.
Risk control means selecting, implementing and checking measures that eliminate the hazard or reduce the risk to an acceptable or required level.
Risk assessment is not the final product
A completed form does not protect anyone. Protection appears only when suitable controls are implemented, communicated, resourced, used and verified.
Easy memory: Find it → Understand it → Control it → Check it.
Hazard
Something with the potential to cause injury, ill health, damage or another unwanted outcome.
Risk
The combination of how likely harm is and how serious the consequence may be, considering exposure and existing controls.
Control
A measure that removes the hazard, prevents exposure, reduces likelihood, limits consequence or supports safe recovery.
The Risk-Assessment Process—Seven Clear Steps
Different organisations may group the stages differently. Here, we use the seven stages indicated for this learning outcome, followed by a continuous review loop. Select each step to hear the portal explain it.
Control Standards, Action Plans and Priority
What is a risk-control standard?
It is the benchmark that tells us what level or type of control is required. Sources may include legislation, approved guidance, exposure limits, engineering codes, manufacturer instructions, industry good practice, internal rules and the hierarchy of controls.
Why it matters: A coloured risk score alone cannot decide whether a mandatory guard, ventilation system, permit or exposure limit is required.
What makes an action plan useful?
- Specific control action—not “be careful”.
- Named responsible owner and adequate resources.
- Realistic completion date and interim protection.
- Priority based on risk, standards, people and uncertainty.
- Method to verify completion and effectiveness.
Priority is not “highest score only”
Give prompt attention to imminent danger, serious consequences, legal or control-standard gaps, many people exposed, vulnerable persons, ineffective controls and high uncertainty. A lower matrix score must not be used to postpone a non-negotiable requirement.
Generic, Specific and Dynamic Risk Assessments
These are not three levels of quality. They are three ways of matching the assessment to the work situation.
| Assessment type | Easy meaning | Use it when… | Do not rely on it when… | Main strength and limitation |
|---|---|---|---|---|
| Generic | A baseline assessment for similar activities, hazards or locations. | Work is routine, repeated and genuinely comparable; common controls can be standardised. | People, equipment, substances, environment or task conditions differ materially. | Strength: efficient and consistent. Limitation: may overlook local or individual differences. |
| Specific | An assessment for one task, site, machine, substance, project or person. | Work is unusual, complex, high risk, legally specific, non-routine or affected by vulnerability. | A suitable generic assessment already covers truly identical low-risk work—although local checks are still needed. | Strength: detailed and relevant. Limitation: requires time, information and competence. |
| Dynamic | A continuous, in-the-moment judgement as conditions change. | Emergency response, rapidly changing work or an unexpected condition requires immediate reassessment. | Work is planned and foreseeable. It must not replace a suitable formal assessment, method statement or permit. | Strength: responds to reality. Limitation: time pressure and incomplete information may weaken judgement. |
Interactive Assessment-Type Selector
Describe the situation. The tool will recommend a starting approach and explain why.
What Do Generic, Specific and Dynamic Assessments Actually Look Like?
Think of these as three different lenses. A generic assessment establishes a reusable baseline. A specific assessment focuses that baseline on one real task, place, item, substance or person. A dynamic assessment keeps checking the live situation while work or an incident develops.
Interactive Assessment-Format Explorer
Select a type. The portal will show its purpose, its header, the columns a suitable form normally needs, a completed example and the final decision that must be recorded.
Core fields every formal assessment needs
- Clear scope, boundaries, location, activity and version.
- Assessor, competent contributors, workers consulted and approval.
- Hazards and credible harm—not only a list of objects.
- Who may be harmed, including contractors, visitors and vulnerable people.
- Existing controls and evidence that they are really in place.
- Risk judgement before further action, with the reasoning shown.
- Further controls, owner, due date, interim protection and priority.
- Residual risk, communication, verification and review triggers.
A blank box is not automatically a bad form
Good forms create space for sound thinking; they do not replace it. The assessor must walk the task, consult the people who understand the work, use relevant evidence and test assumptions.
Quality test: Could a competent supervisor read this record and understand what may happen, who is exposed, which controls must exist, what remains to be done and when work must stop?
Do Not Miss Long-Term Hazards to Health
An injury hazard may produce an immediate event. Many health hazards are quieter: exposure can accumulate and illness may appear months or years later. “Nothing happened today” is not evidence that the risk is controlled.
Exposure pattern
Frequency, duration, intensity, route, peaks, recovery time and combined exposure.
Evidence
Monitoring, sampling, health surveillance, absence records, worker reports and historical data.
Control
Prevent exposure at source. PPE and surveillance support control; they do not replace elimination or engineering where these are required.
Qualitative, Semi-Quantitative and Quantitative Assessment
The difference is how the risk is described and analysed—not how seriously the assessor takes it.
Qualitative
Uses reasoned descriptions such as low, medium, high, unlikely or severe.
Useful for: straightforward work, screening, discussion and situations where numerical data would add little.
Limitation: categories may be subjective and different risks may appear equal.
Semi-quantitative
Assigns ranked numbers to likelihood and severity, often calculating a matrix score such as L × S.
Useful for: consistent comparison and action planning across many hazards.
Limitation: numbers can create false precision; score boundaries and multiplication rules are organisational conventions.
Quantitative
Uses numerical estimates of exposure, probability, frequency or consequence based on data and models.
Useful for: complex, high-consequence or technical decisions and comparison with numerical standards.
Limitation: needs valid data, competence and transparent assumptions; an exact-looking number may still be uncertain.
From Professional Description to Ranked Scores and Probability Models
There is no honest single-name answer to “who invented risk assessment?” People have judged danger for as long as organised work has existed. Modern methods developed gradually as industries needed decisions that were more systematic, comparable and technically defensible.
The methods are a progression of detail—not a competition
Qualitative thinking helps us name what may happen and judge its importance. Semi-quantitative scoring helps a group rank and compare many concerns using defined categories. Quantitative analysis estimates frequency, probability, exposure or consequence when the decision needs numerical evidence.
A major quantitative study still begins with qualitative questions: Which scenarios matter? What assumptions are credible? Who may be affected? A risk matrix still needs professional judgement. The methods often work together.
What the historical landmarks do—and do not—prove
System-safety programmes helped formalise ranked categories and risk matrices, but no single standard invented every semi-quantitative method. By 1971, NASA and aircraft manufacturers were using fault-tree tools. The 1975 US Reactor Safety Study, WASH-1400, directed by MIT professor Norman Rasmussen and AEC staff member Saul Levine, became a landmark probabilistic risk assessment. Later criticism of some numerical claims also taught an essential lesson: model structure, data gaps and uncertainty must be made visible.
Historical teaching sources: US Nuclear Regulatory Commission histories of the Reactor Safety Study and risk-informed regulation; NASA System Safety Handbook risk-matrix material; MIL-STD-882 system-safety practice.
The Qualitative, Semi-Quantitative and Quantitative Assessment Workshop
We will use one scene throughout: pedestrians and forklifts interact in a busy warehouse loading area. This makes the difference easy to see—the hazard does not change, but the depth and form of the analysis changes.
Open qualitative form ↓ 2. Semi-quantitative formDefined rankings combined in the interactive 5 × 5 matrix.
Open matrix ↓ 3. Quantitative formModelled event frequency, uncertainty range and numerical criterion.
Open quantitative form ↓
Reasoned words
Judgement: “Collision risk is high because pedestrians frequently cross an active vehicle route and a collision could cause fatal injury.”
Defined ranks
Judgement: Likelihood 4 (likely) × severity 5 (catastrophic) = 20, “very high” under the example organisation’s approved matrix.
Measured or modelled values
Evidence: vehicle movements/hour, crossing frequency, near-miss rate, speed, exposure time, barrier reliability and predicted collision consequence.
Interactive Method-Format Explorer
Select a method to see the complete form structure. Notice that each higher-data method retains the basic hazard, people, controls, action and review fields.
Which Method Is Proportionate?
Answer five questions. The recommendation is a starting point for competent judgement—not an automatic approval.
Build a Defensible Qualitative Judgement
“High risk” alone is weak. Link the descriptor to exposure, consequence, controls, evidence and uncertainty.
Create One Complete Risk-Assessment Record
This tool mirrors the essential columns of a practical assessment register. It helps learners see that a risk rating sits inside a much larger management record.
Interactive Qualitative Risk-Assessment Form
A qualitative assessment uses defined words and reasoned professional judgement. It does not multiply scores. The record must explain why the likelihood, consequence and overall priority descriptions fit the evidence.
A. Assessment identity and scope
B. Hazard, people and evidence
C. Initial qualitative judgement
D. Treatment, responsibility and residual judgement
Example likelihood meanings
- Rare: exceptional under the defined conditions.
- Unlikely: foreseeable but not expected during normal activity.
- Possible: could occur during the activity or assessment period.
- Likely: expected to occur repeatedly unless controls improve.
- Almost certain: occurs frequently or conditions make occurrence imminent.
Example consequence meanings
- Minor: limited, short-term harm.
- Moderate: treatment or restricted work may be required.
- Serious: major injury or significant occupational ill health.
- Major: life-changing harm or single fatality potential.
- Catastrophic: multiple fatalities or major widespread impact.
Interactive 5 × 5 Semi-Quantitative Risk Matrix
Compare the initial risk with the residual risk after proposed controls. The example bands are for learning only; an organisation must define and approve its own criteria.
Choose the ratings
Solid outline: initial rating. Dashed white outline: residual rating.
Why might severity stay at 5?
A guard or interlock may make contact much less likely, but if contact still occurs the possible injury may remain catastrophic. Do not automatically reduce both numbers simply because controls were proposed. Rate the real effect of the selected controls and verify them.
A Simple Quantitative Comparison Tool
Measured value compared with an applicable limit
This demonstration calculates an exposure ratio. Both figures must use the same unit and come from a valid assessment.
How to interpret carefully
A ratio of 0.70 means the measured value is 70% of the selected limit. It does not automatically mean there is no risk or that controls can be relaxed.
Check sampling quality, uncertainty, peak exposure, routes of exposure, combined substances, vulnerable people, the legal meaning of the limit and whether further reduction is required.
Quantitative Scenario-Frequency Calculator
This teaching model follows one event path. It estimates how often the defined harmful outcome may occur by combining an initiating-event frequency with conditional probabilities. It is useful for learning the logic of event trees; it is not a substitute for a validated QRA.
How to Read Every Symbol—and Why It Is Used
It means the estimated frequency of the defined harmful outcome, normally stated per year. It appears on the left because this is the answer the model is calculating.
It means how many times the event that starts the harmful scenario occurs during a stated period, usually one year. The small word below f is a label—it is not multiplication.
P shows that the value inside the brackets is a chance from 0 to 1. For example, 0.25 means a 25% chance.
It means the chance that a person is present or exposed when the initiating event occurs. It is used because an initiating event cannot harm a person who is not in the exposure path.
It means the chance that the intended barrier, safeguard or protective control fails, is unavailable or does not stop the event path.
The vertical bar | means “given that”. It asks: once the event has reached the exposed person, what is the chance of the defined harm?
Multiplication is used because the harmful outcome follows a sequence: the initiating event occurs, exposure exists, the control fails and harm follows. Each stage narrows the original frequency.
It separates the answer on the left from the factors used to calculate it on the right. Both sides describe the same estimated harmful-outcome frequency.
This is the unit, not a percentage. A result of 0.12 per year is a modelled event frequency; it does not mean a 12% annual risk unless a valid model specifically supports that interpretation.
Four checks before trusting the output
- Scenario: Is the initiating event and harmful outcome defined without ambiguity?
- Dependence: Are the probabilities really independent, or can one common cause defeat several controls?
- Data: Are frequencies based on comparable equipment, tasks, people and operating conditions?
- Uncertainty: Would reasonable lower and upper assumptions materially change the decision?
Level 6 point: A calculated frequency is an estimate conditional on a model. Report units, source, time period, assumptions, uncertainty and sensitivity—not only the final number.
Interactive Quantitative Risk-Assessment Form
A quantitative assessment records more than a calculation. It defines the decision, scenario and model; gives every input a source and unit; shows uncertainty; compares the result with an approved criterion; and records treatment and verification.
A. Decision, scope and model boundary
B. Numerical model inputs
The pronunciation and meaning of every symbol are explained in the symbol guide immediately above.
C. Assumptions, treatment and assurance
What the uncertainty factor means here
An uncertainty factor of 2 displays a teaching range from the central estimate ÷ 2 to the central estimate × 2. A real QRA may require probability distributions, confidence intervals, alternative models or structured expert judgement. The factor is a learning device—not a universal scientific rule.
What must be reviewed independently?
Scenario completeness, units, input provenance, relevance of data, dependencies and common causes, human-reliability assumptions, consequence model, uncertainty treatment, sensitivity, numerical criterion and whether the model is valid for the decision.
Build a Complete Risk-Control Action
An action without an owner, date, interim measure and verification method is only an intention. Complete the fields and create a practical action-plan entry.
Meet the Six Strategies
Select a card. The portal will explain when to use it, when not to depend on it, and what a Level 6 justification should recognise.
When Each Strategy Is Most Relevant
| Strategy | Use when… | Evidence needed | Key limitation or warning |
|---|---|---|---|
| Avoidance | Exposure can be removed by stopping, substituting, redesigning or choosing another objective. | Severity, feasibility of alternatives, standards, lifecycle and risk-transfer effects. | May create a different risk or sacrifice an essential activity; assess the replacement. |
| Reduction | The activity is necessary and the risk can be controlled using the hierarchy of controls. | Control performance, human factors, maintenance, residual risk and verification. | Administrative controls and PPE are vulnerable to failure; reduce at source where possible. |
| Transfer | Specialist competence or financial sharing is appropriate—for example, a competent contractor or insurance. | Competence, contract scope, interfaces, supervision, insurance and monitoring. | Legal and ethical responsibility for protecting people is not simply transferred away. |
| Analysis | The risk, causes, exposure, failure paths or options are uncertain or complex. | Measurements, incidents, task information, models, assumptions and uncertainty. | Do not delay obvious immediate controls while waiting for perfect data. |
| Evaluation | Analysed risk must be compared with legal, technical or organisational criteria to set priority and treatment. | Approved criteria, control standards, consequence, affected groups and tolerance. | A matrix colour cannot override a mandatory requirement or conceal uncertainty. |
| Review | Time has passed or change, incident, failure, new evidence, worker concern or new requirements may affect validity. | Inspection, monitoring, incidents, health data, assurance findings and change information. | A scheduled annual review is insufficient when a trigger requires immediate reassessment. |
Interactive Strategy Decision Lab
Choose a workplace situation and the strategy you think should lead the response. The tool will explain the strongest answer and supporting strategies.
Risk Reduction Must Follow the Hierarchy of Controls
When a risk cannot be avoided completely, begin with controls that act on the hazard and exposure pathway. Measures lower in the hierarchy usually depend more heavily on consistent human behaviour.
Eliminate
Remove the hazard from the work—for example, design out the need to enter a vessel.
Substitute
Replace it with a safer material, method, machine or energy source, then assess the substitute’s hazards.
Engineering
Isolate people through guarding, enclosure, segregation, automation, extraction or fail-safe design.
Administrative
Use planning, permits, procedures, competence, supervision, scheduling, signage and restricted access.
PPE
Protect the individual when exposure remains. Select, fit, maintain and supervise its use; do not make it the automatic first answer.
Risk Transfer—The Point Learners Must Not Miss
What can be transferred or shared?
Some financial loss may be insured. Specialist work may be contracted to an organisation with suitable equipment and competence. Contract terms may allocate defined responsibilities.
What does not disappear?
The hazard remains until controlled. The client or employer must still select competent parties, provide information, coordinate interfaces, monitor work and meet applicable legal duties. A signature on a contract is not a physical control.
Build a Level 6 Justification
Use the structure Decision → Because → Evidence → Limitation → Review. This produces a learning scaffold that you should explain in your own professional words.
Analysis, Evaluation and Review—Do Not Mix Them Up
One simple example
Analysis: solvent-vapour measurements, duration and ventilation performance show the nature and level of exposure. Evaluation: the evidence is compared with applicable exposure criteria and good practice to decide whether treatment is adequate. Review: monitoring and reassessment confirm whether the new local exhaust ventilation continues to control exposure after process or maintenance changes.
Common Errors—and the Better Approach
| Weak approach | Why it is weak | Better Level 6 approach |
|---|---|---|
| Copy a generic assessment without checking the site. | Local hazards, people and conditions may be different. | Use it as a baseline, then verify and adapt it before work. |
| Use a dynamic assessment for planned high-risk work. | It avoids proper planning, consultation and control design. | Complete a formal specific assessment, then use dynamic checks for real-time change. |
| Reduce both likelihood and severity scores automatically. | The selected control may affect only one dimension. | Explain how each control changes exposure, failure path or consequence. |
| Call a risk “low” because no one has yet been harmed. | Absence of recorded harm is weak evidence, especially for latent health risks. | Consider exposure data, potential severity, under-reporting and control reliability. |
| Transfer work and stop managing it. | Contracting does not eliminate the hazard or all duties. | Assess competence, coordinate, monitor and verify contractor controls. |
| Review only once a year. | A change or failure can make the assessment invalid immediately. | Use scheduled and event-triggered review. |
Can You Make and Justify the Decision?
Learn One Control Layer at a Time
The portal starts with the wider safe system of work, develops it, then moves to the more focused safe operating procedure. Only after both are clear do we connect the document family and explore the complete permit-to-work lifecycle.
DB HSE learning resource. Prepared solely by Debjyoti Biswas for teaching Unit 3, Section 2.3. This independent learning portal is not produced by OTHM and does not issue workplace authority.
A Risk Assessment Decides What Must Be Controlled; the System of Work Decides How
Imagine a chemical-transfer pump that needs maintenance.
The assessment identifies hazardous chemical residue, stored pressure, electricity, moving parts, restricted access, contractors and possible conflict with nearby operations. That information is essential—but it does not yet tell the maintenance team exactly how the job will be prepared, authorised, completed, checked and handed back.
The safe system of work connects the people, equipment, controls, communication and sequence. The safe operating procedure gives the approved steps for a defined operation inside that wider system.
The paperwork test
A document does not make work safe merely because it has been signed. The controls must exist at the workplace, the people must understand them, and a responsible person must verify that they remain effective.
Easy memory: Assess → Design → Explain → Do → Check → Improve.
One Scenario, Four Stages of Control
We will use the same chemical-transfer pump throughout Section 2.3. This allows you to see how one risk picture is converted into a complete working arrangement.
What Is a Safe System of Work—SSOW?
Safe System of Work — SSOW
Pronounced: “S-S-O-W,” or simply “safe system of work.”
An SSOW is a deliberately organised method for completing work so that foreseeable hazards are controlled throughout preparation, execution, completion and foreseeable abnormal conditions.
It is a system because it joins people, plant, materials, environment, controls, responsibilities, communication, competence, supervision and review. It is not only a list of steps.
What an SSOW is not
It is not a risk-assessment form, a signature, a list of PPE, a copied method statement or an instruction to “be careful.” It must convert risk decisions into a realistic arrangement that people can understand and use.
Simple test: Could a competent team use this system to know who does what, in what order, with which controls, when to stop, and how the work returns safely to normal?
| SSOW format field | What to record | Why the field matters |
|---|---|---|
| Identity and control | Title, number, owner, version, approval and review date. | Prevents obsolete or unapproved instructions being used. |
| Purpose, scope and limits | Activity, plant, location, persons, conditions, interfaces and exclusions. | Stops the system being applied outside the conditions it was designed for. |
| Roles and competence | Who plans, authorises, performs, supervises, verifies, hands back and reviews. | Prevents gaps, duplication and unverified assumptions. |
| Hazards and control basis | Assessment reference, standards, energy sources, substances, exposure and credible failures. | Shows that the method is risk-based and technically supported. |
| Preparation and resources | Access, barriers, isolations, tools, staffing, communication, permits and PPE. | Creates the conditions needed before work starts. |
| Safe sequence and hold points | Ordered actions, responsible role, required result and checks before progression. | Some controls only work when applied in the correct order. |
| Stop, abnormal and emergency rules | Conditions that suspend work, safe state, escalation, rescue and recovery. | Prevents unsafe improvisation when assumptions change. |
| Completion, handback and review | Inspection, reinstatement, records, acceptance, monitoring and review triggers. | Controls the return to normal operation and captures learning. |
Interactive Jargon Translator
Select any term. The portal will pronounce it, define it and explain why it matters.
Do Not Mix Up the Document Family
The documents are connected, but they do different jobs. The level of formality should be proportionate to the risk, complexity and need for coordination.
| Document or arrangement | Easy purpose | Typical use | Important limitation |
|---|---|---|---|
| Risk assessment | Identifies hazards, people, existing controls, risk and further action. | Before deciding the safe method and whenever relevant change occurs. | A completed form does not implement the controls. |
| SSOW | Coordinates the whole method, people, controls and interfaces. | Where risks require a defined safe way of working, particularly complex or significant activities. | It fails if impractical, unknown, unsupervised or not followed. |
| SOP | Standardises the safe steps for a defined operation. | Routine or repeated operation, inspection, start-up, shutdown, cleaning or maintenance. | It cannot predict every abnormal condition; stop and escalation rules are needed. |
| Method statement | Describes how a particular job or project stage will be carried out. | Construction, installation, maintenance and contractor work. | A generic copied statement may not match the real site or sequence. |
| JSA/JHA | Breaks a job into steps, hazards and controls. | Task planning and workforce discussion. | Step-by-step analysis must still consider interactions and emergencies. |
| Permit-to-work — PTW | Formally authorises specified work, location, time and precautions. | Defined high-risk or tightly coordinated work under site rules. | A permit is not a guarantee of safety and does not replace risk assessment. |
| Checklist | Confirms that required checks were completed. | Pre-start, inspection, handover and verification. | Ticking boxes without observation gives false assurance. |
| Emergency procedure | Explains response when control is lost or conditions become unsafe. | Credible abnormal and emergency situations. | It must be resourced, communicated and tested—not only filed. |
Tool: Which Arrangement Does This Job Need?
Fifteen Stages for Developing an Effective Safe System of Work
The stages are grouped into five phases. Select a phase to explore what must happen and why.
What Is a Safe Operating Procedure—SOP?
Safe Operating Procedure — SOP
Pronounced: “S-O-P,” or “safe operating procedure.”
An SOP is an approved, controlled and repeatable set of instructions for performing a defined operation safely and consistently. It tells the authorised user what conditions must exist, what to do in sequence, what result to confirm and when to stop.
Typical SOPs cover start-up, normal operation, sampling, cleaning, inspection, safe shutdown, isolation preparation, testing and return to service.
How it fits inside the SSOW
The SSOW coordinates the complete job—including teams, interfaces, permits, isolations, emergency arrangements and handback. The SOP standardises one defined operation within that system.
A pump-maintenance SSOW may refer to separate SOPs for shutdown, electrical isolation, line draining, gas testing and controlled recommissioning.
| SOP format field | What it should contain | Quality question |
|---|---|---|
| Document identity | Title, equipment or process, number, version, owner, approver and review date. | Can the user confirm this is the current approved procedure? |
| Purpose and scope | Intended result, authorised users, operating range, location and exclusions. | Is it clear when the SOP applies—and when it does not? |
| Responsibilities and competence | Operator, supervisor, verifier, specialist and required authorisation. | Does each person understand their role and limit of authority? |
| Prerequisites | Plant state, permits, isolations, tools, inspections, guards, ventilation and PPE. | What must be true before Step 1? |
| Ordered steps | One clear action per step, responsible role, location, setting, safe limit and expected result. | Can the action and its successful outcome be observed? |
| Warnings and hold points | Critical hazards, prohibited actions and mandatory verification before continuing. | Are the most safety-critical instructions easy to find? |
| Operating limits | Pressure, temperature, concentration, speed, time or other acceptance criteria. | Does the user know when a result is outside the safe range? |
| Stop and escalation | Unexpected state, failed check, alarm, leak, defect, safe shutdown and person to contact. | Does the SOP prevent improvisation? |
| Completion and records | Final checks, housekeeping, status communication, log entries and retained evidence. | Can another person confirm the operation ended safely? |
| Review and change | Scheduled date plus triggers such as incident, modification, feedback or new evidence. | Will the SOP remain aligned with the real process? |
How Is an SOP Developed?
Select each phase. The five phases contain ten connected stages: define → observe → assess → sequence → write → validate → approve → train → use → review.
Characteristics of a Strong SSOW and SOP
A good document must be technically correct and usable by the people who depend on it. These characteristics are evidence of quality—not decorative features.
| Characteristic | Easy meaning | Why it matters | Warning sign |
|---|---|---|---|
| Risk-based | Controls come from a suitable assessment and required standards. | The system addresses credible harm rather than copying another job. | The procedure mentions PPE but not the main energy or exposure source. |
| Task-specific | It matches the actual plant, people, place and conditions. | Local differences can change the failure path. | Wrong equipment number, location, substance or isolation point. |
| Proportionate | Detail and formality reflect risk and complexity. | Too little detail leaves gaps; excessive paperwork hides critical controls. | A simple task has 50 pages, while a major intervention has one vague paragraph. |
| Clear and sequential | Actions are unambiguous and in the correct order. | Sequence can determine whether energy or exposure is controlled. | “Make safe” without stating who, how or how safety is verified. |
| Practical | Controls can be applied with available time, access, tools and resources. | Impossible instructions encourage deviation and workarounds. | The required test point cannot be reached safely. |
| Participative | People who understand the work contribute to development and review. | Worker knowledge reveals practical hazards and foreseeable shortcuts. | Written remotely without observing or discussing the job. |
| Role-defined | Authority, responsibility and handover are explicit. | Prevents gaps, overlaps and assumptions. | Everyone believes someone else verified the isolation. |
| Competence-based | Required knowledge, skill, experience and supervision are stated. | The same instruction may not be safe for an inexperienced person. | “Trained person” is stated but competence is never checked. |
| Human-centred | It considers workload, fatigue, usability, communication and predictable error. | Controls must work in real human conditions. | Critical information is buried, contradictory or unreadable. |
| Inclusive and accessible | Users can find, read and understand it. | Language, literacy, disability or unfamiliar terminology can affect safe use. | Only one complex-language copy exists away from the workplace. |
| Abnormal-condition ready | It states when to stop, withdraw, isolate, escalate or use emergency arrangements. | People must not improvise when normal conditions disappear. | No instruction for a leak, failed test or unexpected pressure. |
| Controlled and current | Only the approved version is available and changes are traceable. | Obsolete instructions may conflict with modified plant or controls. | Different versions are posted at the same workplace. |
| Verified | Critical controls are checked before reliance. | An assumed control may be absent, failed or incorrectly applied. | The permit is signed without a field check. |
| Monitored and reviewed | Use and effectiveness are checked over time and after triggers. | Work, people, equipment and evidence change. | Repeated deviations are normalised without investigation. |
When a detailed written system is normally needed
- Significant or high-risk work
- Complex or non-routine tasks
- Several people, teams or contractors
- Critical sequence, isolation or verification
- Permit-controlled work
- Serious foreseeable abnormal conditions
When simplicity may be appropriate
Straightforward low-risk work may be controlled through concise instruction, training and normal supervision. Simplicity must come from low complexity—not from ignoring significant hazards. The arrangement still needs to be understood and effective.
How the Six Risk Strategies Shape the System of Work
These strategies are not six competing documents. They influence different decisions during development, authorisation and assurance.
| Strategy | When it is used | Pump-maintenance application | What the SSOW or SOP must show |
|---|---|---|---|
| Avoidance | The exposure or activity can be removed or redesigned. | Use remote condition monitoring to avoid unnecessary intrusive inspection. | Why the hazardous step is no longer required and whether the alternative creates new risk. |
| Reduction | Necessary work continues under stronger controls. | Isolate, depressurise, drain, purge, verify, segregate and supervise. | Control hierarchy, sequence, responsibilities, verification and residual risk. |
| Transfer | Specialist competence, equipment or financial sharing is appropriate. | Use a competent specialist for seal replacement or hazardous cleaning. | Selection, information, coordination, interfaces, monitoring and retained duties. |
| Analysis | Causes, exposure, failure paths or uncertainty require deeper understanding. | Analyse chemical residue, pressure, isolation effectiveness and previous failures. | Evidence, assumptions and how findings affected the method. |
| Evaluation | Evidence must be compared with criteria to decide adequacy and authorisation. | Compare proposed precautions with legal, technical, manufacturer and site requirements. | Acceptance criteria, decision authority and unresolved gaps. |
| Review | Time, change, incident, feedback or failed control may affect validity. | Revise after a leak, near miss, plant modification, contractor concern or recurring deviation. | Review triggers, owner, evidence, revised version and communication. |
Tool: Build a Combined Risk-Strategy Route
Build an Educational Safe System of Work Draft
Adjust every field to your own workplace example. The output helps you understand the structure; it must be reviewed by competent people against the real workplace and applicable requirements before use.
Build an Educational Safe Operating Procedure Format
An SOP should tell the right person what to do, in what order, under which conditions, and when to stop. It should not ask the user to make undefined safety decisions during a critical step.
Tool: Put the Pump-Maintenance Controls in a Defensible Order
Use the arrow buttons to move each step. The “correct” route is the approved teaching sequence for this example; a real installation may require a different, technically validated sequence.
Permit-to-Work—PTW: Meaning, Purpose, Issue, Use, Handback and Closure
What is PTW?
Pronounced: “P-T-W,” meaning permit-to-work.
A PTW is a formal, time-limited system for authorising specified work at a defined place, on identified plant or equipment, by named or competent parties, subject to stated precautions and conditions.
It communicates an agreement: this exact work may proceed, within this exact boundary, while these verified conditions remain true.
What PTW does not mean
A permit is not a risk assessment, an instruction to begin automatically, a substitute for isolation, a certificate that danger has disappeared, or a transfer of all responsibility to the worker.
Issuing the paper alone does not make a job safe. The assessment, SSOW, competence, communication, physical controls, field checks, supervision and stop-work response must all function.
Why is a permit-to-work used?
A PTW creates disciplined communication where mistakes in plant identity, isolation, timing, coordination or handover could cause serious harm. It defines ownership, prevents incompatible simultaneous activities, records critical precautions, controls the period of work and manages the return to normal operation.
| Work category | Why formal control may be needed | Typical linked controls or certificates |
|---|---|---|
| Hot work | Flame, arc, spark or heat may ignite flammable material or damage adjacent systems. | Gas testing, area preparation, fire protection, fire watch and post-work monitoring. |
| Confined-space or vessel entry | Atmosphere, engulfment, restricted access, energy and rescue hazards can change rapidly. | Isolation certificate, atmospheric test, ventilation, entry log, attendant and rescue plan. |
| Line breaking / hazardous containment | Opening pipework or equipment can release pressure, temperature, toxic, corrosive or flammable material. | Process isolation, drain/vent/purge, decontamination, test and line-break controls. |
| Electrical or mechanical work | Contact, arc, unexpected start, gravity, pressure or stored energy may be fatal. | Isolation plan, lockout/tagout, prove-dead or zero-energy verification and controlled reinstatement. |
| Excavation / ground disturbance | Underground services, collapse, water, contaminated ground and vehicle interaction may be present. | Service drawings, detection and marking, trial holes, shoring, access and inspection. |
| Other site-defined work | Work at height, lifting, radiography, roof access, energised testing or unusual simultaneous work may need coordination. | Site-specific permits, certificates, exclusion zones, specialist plans and interfaces. |
Who is involved?
| Typical role | Easy meaning | Main responsibility |
|---|---|---|
| Area / operating authority | The person controlling the plant or area. | Confirms operating status, interfaces and whether the area can be released and later accepted back. |
| Permit issuer / issuing authority | The competent authorised person who issues the permit. | Checks scope, assessment, precautions, isolations, conflicts, validity and field conditions before authorising. |
| Performing authority / permit receiver | The person accepting the permit for the work party. | Understands the permit, briefs the team, keeps within boundaries, monitors conditions and stops when conditions change. |
| Isolating authority | The person controlling required energy or process isolations. | Applies, records, proves and later removes isolation under the approved process. |
| Authorised gas tester | A competent person approved to test atmosphere. | Uses suitable equipment, records results and understands limits, frequency and conditions of testing. |
| Work party | The persons carrying out the authorised task. | Attend the briefing, follow the SSOW/SOP and permit, protect controls and report change or uncertainty. |
| Permit / SIMOPS coordinator | The person who sees the whole work picture. | Prevents conflicts between permits, operations, contractors, isolations and emergency arrangements. |
Terminology varies: a site may use different role names. The essential point is that authority, competence, accountability, communication and handover cannot be vague.
The Full PTW Lifecycle—14 Phases
Select a phase to see what must happen, why it matters and what evidence should exist.
What should a complete permit form contain?
| Permit field | Information required | Control purpose |
|---|---|---|
| Identification | Permit number/type, work order, exact plant/equipment tag, location and description. | Prevents work on the wrong item or outside the authorised task. |
| Validity | Issue date/time, start, expiry, shift and any rules for extension or revalidation. | Stops an old permit being treated as continuing permission. |
| Supporting documents | Risk assessment, SSOW, SOP/method, drawings, certificates and rescue/emergency plans. | Connects authorisation to the technical control basis. |
| Hazards and interfaces | Energy, substances, atmosphere, access, environment, nearby work and SIMOPS. | Makes foreseeable interactions visible to both parties. |
| Isolations and tests | Isolation points/certificate, lock and tag references, drain/vent/purge, test type, result, time and tester. | Provides traceable evidence of critical plant preparation. |
| Precautions | Barriers, ventilation, fire controls, access, tools, PPE, monitoring and prohibited actions. | Defines conditions that must remain in place. |
| Emergency and communication | Alarm, withdrawal, rescue, contact, stop-work rule, briefing and handover method. | Supports response when normal assumptions fail. |
| Authorisation and acceptance | Issuer and receiver names/signatures, date/time and declarations of understanding. | Confirms that authority and shared understanding are explicit. |
| Suspension / extension / handover | Reason, safe state, new conditions, outgoing/incoming parties and revalidation. | Prevents work continuing across a change without control. |
| Completion and handback | Work complete/incomplete, people/tools cleared, guards restored, plant status, inspection and acceptance. | Controls transfer back to operations and reinstatement. |
| Cancellation and records | Closure time, permit cancellation, linked documents closed, defects/actions and retained record. | Ends authority clearly and preserves evidence for audit and learning. |
Interactive: Is the Permit Ready for Issue?
Tick only items verified by evidence. A high total cannot compensate for a missing critical condition.
Interactive: Shift, Alarm and Change Decision
Interactive: Build a Complete PTW Learning Form
Complete every field. This produces a teaching draft—not a workplace permit or authorisation.
Interactive: Does the Work Need Formal Permit Control?
A permit-to-work is a formal authorisation and communication system for defined work. Select the features that apply. Site procedures and applicable law make the final decision.
Interactive: Stress-Test the Working Arrangement
Rate each condition from 1 (weak) to 5 (strong). This is a learning diagnostic, not a risk-acceptance formula.
PTW Knowledge Check
Tools: SSOW Quality Diagnostic and Review Trigger
Tool 1: Check the Quality of an Existing SSOW
Tool 2: What Should Trigger Review?
Weak Wording Versus Defensible Wording
| Weak wording | Why it fails | Stronger wording principle |
|---|---|---|
| “Make the pump safe.” | No person, isolation method, condition or verification is defined. | Name the equipment, energy sources, authorised role, approved isolation method and verification requirement. |
| “Wear proper PPE.” | “Proper” is undefined and PPE may not control the main hazard. | Select PPE from the assessment after applying higher-order controls; state type, limitation and checks. |
| “Be careful when opening.” | It transfers responsibility to behaviour without controlling stored pressure or residue. | Prevent opening until depressurisation, drainage, safe condition and authorisation are verified. |
| “Experienced workers only.” | Experience is not defined or verified. | State required authorisation, task knowledge, skill, experience and supervision. |
| “In an emergency, act accordingly.” | No stop, alarm, withdrawal or escalation route is given. | Define credible abnormal conditions and the immediate response expected from each role. |
| “Review annually.” | Waits for a date even after a change, failure or near miss. | Use both scheduled and event-triggered review. |
How to Explain 2.3 at Level 6
A strong explanation normally contains
- Clear definitions of SSOW and SOP.
- The relationship with risk assessment and supporting documents.
- A connected development lifecycle.
- Characteristics explained with reasons—not only listed.
- Application of the six risk strategies.
- How risk assessment, SSOW, SOP, PTW, isolation and handback connect.
- The PTW lifecycle from request and field verification to suspension, handback and closure.
- A suitable workplace example.
- Limitations, implementation and review arrangements.
Use this paragraph structure
Point → Meaning → How developed/applied → Why it matters → Workplace example → Consequence if missing → Review.
Write in your own professional words and relate the explanation to a genuine or realistic workplace. A list of headings alone does not fully satisfy “explain.”
Explain how a safe system of work and supporting safe operating procedures should be developed for intrusive maintenance of a chemical-transfer pump. Your response should address consultation, risk strategies, control sequence, competence, human factors, permit interfaces, abnormal conditions, implementation and review.
Can You Convert a Risk Decision Into Safe Work?
Understand the Models of Loss Causation, Analysis of Loss Data and the Importance of Incident Investigation
Section 03 asks us to move from what happened, to how the event developed, why the controls were vulnerable, and what the loss data can—and cannot—prove.
Learning outcomes covered now: 3.1 Outline a range of loss-causation theories and techniques. 3.2 Justify the use of quantitative methods in analysing loss data.
What Do “Loss” and “Causation” Mean?
Loss means an unwanted outcome that removes or damages something of value. It may include injury, ill health, death, environmental harm, property damage, production interruption, legal exposure, financial cost, lost information or damaged trust.
Causation means the way conditions, decisions, actions, failures and interactions combine to produce an event or outcome.
An incident normally has more than one relevant cause. The final action or failed component may be easy to see, but a Level 6 analysis asks what shaped that action, why the control was absent or ineffective, and what management-system conditions allowed the vulnerability to remain.
The central learning rule
A model organises thinking; it does not manufacture evidence.
Investigators must still preserve the scene, gather reliable evidence, test competing explanations, consult involved people, identify controls and verify corrective action.
Immediate cause
The unsafe act, condition, energy transfer or equipment state directly connected to the unwanted event.
Ask: What directly triggered or enabled the contact?
Underlying cause
The job, workplace or organisational factor that allowed the immediate condition or action to arise.
Ask: What influenced the work and weakened control?
Root cause
A deeper management-system or organisational failing whose correction can prevent a wider class of recurrence.
Ask: Why did the system create, accept or fail to detect the vulnerability?
Interactive Section 03 Jargon Translator
Select a term to hear how it is pronounced and understand why it matters.
Our Master Case: Forklift Contact With a Process Line
A forklift enters a congested transfer-area route, passes a damaged low barrier and contacts a valve manifold. A small chemical release occurs. The alarm operates, workers withdraw and the response team isolates the area. No model should be used to blame the driver or to assume a cause before evidence is gathered.
| Evidence source | Initial finding | What must still be tested? |
|---|---|---|
| CCTV and scene measurements | Pallets narrowed the route; forklift contacted the low pipe barrier. | Actual speed, visibility, pedestrian interaction and why storage entered the route. |
| Inspection records | Barrier damage had been recorded twice but not permanently repaired. | Risk classification, escalation, ownership, resources and closure verification. |
| Driver and worker interviews | Peak dispatch created queuing and radio interruptions. | Work-as-done, production pressure, route rules, competence and normal adaptations. |
| Alarm and response log | Detection and emergency isolation limited the release. | Alarm timing, response reliability, exposure and opportunities to strengthen recovery. |
| Six-month loss data | Vehicle–route near-miss reports increased, especially during peak dispatch. | Reporting quality, exposure hours, location clustering and whether risk really increased. |
Accident and Incident Ratio Studies
Ratio studies arrange recorded events by consequence level. They helped organisations recognise that low-consequence events and near misses can reveal control weaknesses before a major loss occurs.
Who, when, why—and what changed?
Why it was created: to show that the serious injury at the top is only part of the recorded experience. Lower-consequence events may provide more frequent opportunities to discover exposure and weak controls.
How thinking changed: Bird widened the categories to include property damage and near misses. Later research showed that the shape changes with definitions, industry, severity threshold and reporting practice, and that fatal and non-fatal events can follow different causal pathways.
Historical context: DNV tribute to Frank Bird. Critical evidence: Salminen, Saari, Saarela and Räsänen (1992) and Marshall, Hirmas and Singer (2018).
Heinrich’s historical triangle
Often presented as 1 major injury : 29 minor injuries : 300 no-injury accidents. Read the colon “:” as “to”: one major injury to 29 minor injuries to 300 no-injury accidents. It was derived from historical insurance and accident records and promoted attention to the larger body of less-serious events.
Bird’s historical ratio
Often presented as 1 serious or major injury : 10 minor injuries : 30 property-damage events : 600 near-miss incidents. It broadened attention to damage and no-loss events. The figures describe Bird’s historical dataset; they do not calculate the probability of the next accident.
| Potential benefit | Why it helps | Limitation to explain |
|---|---|---|
| Encourages near-miss reporting | Weak signals can reveal exposure and failing controls before serious harm. | More reports can mean better trust and reporting—not necessarily worsening safety. |
| Supports prevention | Recurring lower-level events can direct inspection and improvement. | Preventing minor slips does not automatically control a low-frequency catastrophic process event. |
| Communicates scale simply | The triangle is memorable and helps introduce proactive learning. | Simplicity may hide different causal pathways and consequence mechanisms. |
| Provides trend categories | Organisations can compare reporting levels and event types over time. | Changed definitions, workforce hours or reporting systems can create a false trend. |
| Challenges injury-only thinking | Damage and near misses can contain valuable control information. | A ratio does not replace investigation, risk assessment or barrier assurance. |
Interactive Tool: See How Reporting Culture Changes the Triangle
The “true opportunities for learning” remain constant in this teaching example. Adjust the percentage that gets reported. Notice how the visible triangle changes even when the underlying events do not.
Bird’s Loss-Causation Model and Multi-Causality
Bird’s expanded domino approach connects management control with basic causes, immediate causes, the incident and the final loss. Removing or strengthening an earlier “domino” can interrupt the sequence.
Where it came from
Bird developed accident-prevention thinking beyond the earlier person-centred domino sequence. The model placed lack of management control at the beginning and widened “injury” into loss, including harm to people, property, environment and production. Bird’s later work with George Germain was published in Practical Loss Control Leadership in 1985.
How it developed
The sequence was adapted using energy-exchange concepts associated with William Haddon. Read it left to right to explain how loss developed; work from right to left during investigation to ask what control should have interrupted each step.
Multi-causality: More Than One Path Can Matter
Multi-causality rejects the idea that one unsafe act is normally a complete explanation. Several conditions may combine, interact or increase one another’s effect. Causes can exist at task, equipment, environmental, individual, supervisory and organisational levels.
Immediate examples
- Forklift contacts low barrier.
- Route is narrowed by pallet storage.
- Valve manifold remains exposed to vehicle energy.
Underlying examples
- Peak traffic and pedestrian interaction were not reassessed.
- Temporary storage became normal.
- Damaged barrier repair was delayed.
Root examples
- Defect priority criteria ignored major-consequence potential.
- No effective owner verified corrective-action closure.
- Layout-change governance excluded operations and workforce evidence.
Interactive Tool: Classify the Cause—Then Look Deeper
The Swiss Cheese Model
James Reason described safety as several layers of defence, barrier and safeguard. Each layer can contain weaknesses—shown as “holes.” An adverse trajectory can pass through when weaknesses in different layers align.
James Reason: from human error to organisational defences
Who and when: British psychologist James Reason developed the model across the 1990s, beginning with ideas in Human Error (1990) and refining the familiar defence-layer image through the decade. Safety practitioner John Wreathall also influenced its development.
Why: Reason wanted to move investigation beyond “the operator made an error.” The model asks how front-line actions combine with weaknesses created earlier by design, staffing, maintenance, supervision and organisational decisions.
How it changed: diagrams and terminology evolved between 1990 and 2000. This matters: Swiss Cheese is a family of evolving explanations, not one frozen diagram. Later safety thinking also stresses that defences interact dynamically and that a simple line through holes must not replace evidence.
Read the system approach in Reason (2000), Human error: models and management, and the historical critique in Larouzee and Le Coze (2020).
Active failure
An action or omission close in time and place to the event, such as an incorrect control input or missed check. It may trigger the event, but it is rarely the whole explanation.
Latent condition
A deeper weakness created by design, staffing, maintenance, procurement, priorities, procedures, supervision or management decisions. It may remain hidden until combined with local conditions.
Interactive Barrier Alignment Simulator
Mark a layer as failed or ineffective. The event pathway opens only when every selected defence in this simplified example has a relevant weakness.
Limitation: The cheese image is a communication model, not a detailed dynamic simulation. It can oversimplify interactions unless each layer, threat, dependency, owner and performance requirement is defined with evidence.
Fault Tree Analysis and Event Tree Analysis
How structured logic entered safety analysis
Fault Tree Analysis: commonly traced to H. A. Watson at Bell Laboratories in 1962 during the Minuteman missile programme. It was created because complex systems could fail through combinations that a simple checklist might miss.
Event Tree Analysis: developed as a forward, consequence-oriented partner to fault-tree reasoning. Probabilistic risk work such as the US Reactor Safety Study, WASH-1400 (1975), helped establish combined fault-tree and event-tree methods for complex high-hazard systems.
What changed: the techniques expanded from defence and nuclear applications into aviation, process safety and other industries. Software can now evaluate very large trees, but the logic, data, dependencies and uncertainty still need competent human review.
Authoritative background: NRC Fault Tree Handbook and NRC history of WASH-1400.
FTA and ETA Symbol Decoder + Logic-Gate Starter
Symbols are a form of shorthand. They make a large analysis easier to read, but only after every symbol has been defined. Start by reading the symbol aloud, identify what it represents, check its unit or range, and then ask why it is present in the equation.
First: know what every common mark means
Letters name defined events. A might mean “detector fails”; B might mean “isolation valve fails.” The letter has no meaning until the analyst defines it.
The total number of relevant demands, tests, opportunities or observations. It is commonly the denominator. Example: N = 100 valid detector tests.
A count from the defined set. Example: n = 3 recorded failures. A subscript makes the count more specific, such as nF for number of failures.
The subscript is a label, not multiplication. nA counts event A; nF counts failures. Thus nF ÷ N means failures divided by all valid opportunities.
The probability that defined event A occurs on the stated demand, mission or period. Probability has no unit and must lie from 0 to 1 inclusive.
p commonly represents success probability and q the complementary failure probability. When success and failure cover all possibilities, q = 1 − p and p + q = 1.
i is an index meaning “the particular item or branch being considered.” p1, p2 and p3 can represent three different input probabilities.
f is an event frequency with a unit such as per year. fI is the initiating-event frequency used at the start of an event tree.
T often represents total observed exposure time; t often represents the mission time being evaluated. The analyst must state the unit—hours, days or years.
A failure rate, normally stated per unit time. A value of λ = 0.002 per hour means the assumed rate basis is 0.002 failures per operating hour—not a 0.2% certainty for every hour.
A mathematical constant approximately equal to 2.71828. In e−λt, it supports the exponential reliability model; it does not mean an event count.
ETA branches are often labelled S and F. These labels say whether the defined barrier performs its required function at that branch point.
∩ means AND—events occur together. ∪ means OR—at least one of the defined events occurs, including the possibility that both occur.
The vertical bar means “given that.” It is used when B’s probability is evaluated under the condition that A has already occurred.
Add the listed values. In ETA, Σfbranch means add the frequencies of the mutually exclusive terminal branches.
Multiply the listed values. ∏pi means p1 × p2 × p3 and so on; it is not the circle constant π.
Equals states that both sides represent the same value. Multiplication combines required path factors. Division creates a proportion or rate from a count and denominator.
The complement of p. If p is the probability of success, 1 − p is the probability of failure only when the two states are mutually exclusive and cover all defined possibilities.
Calculate the expression inside them first. They group terms and prevent the calculation from being performed in the wrong order.
Zero means impossible within the defined model; one means certain within that model. Multiply a decimal by 100 to express it as a percentage.
Second: what is a logic gate and why do we need it?
A logic gate is a rule that explains how input events combine to create an output event. It is not a physical gate and it does not prove causation by itself. It converts a full English sentence into an exact visual instruction, helping an FTA remain consistent when many failure paths are connected.
If detector failure by itself can cause the output, or valve failure by itself can cause it, connect them through OR. A normal inclusive OR also allows both events to occur.
If the output occurs only when the detector fails and the automatic isolation also fails, connect them through AND. AND describes a required combination, not automatically a time sequence.
The output occurs when at least k of n similar channels meet the stated condition. A 2-out-of-3 trip system needs any two of its three channels. Here n means total channels inside this gate.
NOT, exclusive-OR, inhibit and priority-AND can express special logic or sequence conditions. Use them only when their exact meaning is required and defined; most introductory FTAs begin with AND and OR.
| Event A | Event B | A AND B | A OR B | Plain meaning |
|---|---|---|---|---|
| 0 — does not occur | 0 — does not occur | 0 | 0 | No input occurred. |
| 0 | 1 — occurs | 0 | 1 | Only B occurred; OR opens, AND does not. |
| 1 — occurs | 0 | 0 | 1 | Only A occurred; OR opens, AND does not. |
| 1 | 1 | 1 | 1 | Both occurred; both rules are satisfied. |
Interactive Logic-Rule Coach
Choose the English statement you need to represent. The coach will identify the logic and explain how the symbols are used.
Where Does the Probability Come From?
FTA and ETA do not create probability simply because a box is drawn. Every input needs a defined event, population or equipment item, time or demand basis, data source and uncertainty. The first question is therefore not “Which formula shall I use?” It is “What exactly does this number describe?”
Use this empirical estimate when each opportunity is clearly defined. If a detector failed 3 of 100 valid proof-test demands, the observed failure-on-demand estimate is 3 ÷ 100 = 0.03, or 3%.
If p is success probability, q is failure probability. A barrier that succeeds with probability 0.90 has a complementary failure probability of 1 − 0.90 = 0.10.
The vertical bar means given that. An ETA branch asks for the chance of the next success or failure after the initiating event and earlier branch conditions have occurred.
Under a justified constant-rate exponential model, this estimates the probability of at least one failure by mission time t. It is not suitable automatically for ageing, repair, dependence or changing conditions.
Frequency carries a unit such as events per year. Probability has no unit and stays between 0 and 1. A frequency of 0.5 per year must not automatically be called a 50% annual probability.
This means A and B occur together. The general rule is P(A ∩ B) = P(A) × P(B | A). It becomes P(A) × P(B) only when independence is defensible.
| Possible source | How the value may be obtained | Essential quality question |
|---|---|---|
| Operating or test data | Defined failures ÷ valid demands, or events ÷ exposure time. | Are definitions, equipment, conditions and reporting sufficiently comparable? |
| Reliability model | Use a justified distribution, failure rate and mission time—for example 1 − e−λt. | Do the model assumptions fit ageing, repair, maintenance and operating conditions? |
| Fault-tree calculation | A barrier-failure probability used in an ETA may itself come from a detailed FTA. | Were dependencies, shared utilities, human actions and common causes represented? |
| Expert judgement or analogous data | Elicit and document a defensible estimate or range when direct data are sparse. | Are the experts, evidence, assumptions, bias controls and uncertainty traceable? |
Interactive Probability Source Calculator
Choose one method. The portal will use only the fields required for that method and explain the units and assumptions.
Authoritative learning references: US NRC glossary definitions for fault trees and event trees, US NRC explanation of probabilistic risk assessment and NASA system-safety learning on logic models and probability.
Fault Tree Analysis — FTA
Pronounced: “F-T-A.” A Fault Tree Analysis is a structured, top-down method. Start with one precisely defined unwanted top event, then reason backwards to identify the equipment failures, human failures, external events and combinations capable of producing it.
Open the FTA equation and symbol guide
Real trees require validated logic, common-cause and dependency checks, suitable data, minimal cut sets and sensitivity analysis.
Event Tree Analysis — ETA
Pronounced: “E-T-A.” An Event Tree Analysis is a structured, forward or inductive method. Start with a defined initiating event, then move forwards through the conditional success or failure of safeguards, operator actions and recovery measures until every modelled path reaches an end state.
Open the ETA equation and symbol guide
ETA outputs are only as sound as the initiating frequency, branch definitions, conditional probabilities and dependency assumptions.
| Feature | Fault tree | Event tree |
|---|---|---|
| Direction | Backward from a defined unwanted top event. | Forward from a defined initiating event. |
| Main question | What combinations can cause this event? | What outcomes can follow as barriers succeed or fail? |
| Logic | AND, OR and other gates combine causal events. | Branches represent conditional success/failure pathways. |
| Input basis | Basic-event probabilities or frequencies on a consistent mission, demand or time basis. | Initiating-event frequency plus conditional branch probabilities. |
| Typical result | Qualitative cut sets and, when quantified, top-event probability or frequency. | Conditional path probabilities and frequencies for defined end states. |
| Useful for | Complex failure logic, critical combinations and design weaknesses. | Escalation, mitigation, consequence pathways and outcome frequency. |
| Limitation | A poor top-event definition or missed dependency creates false confidence. | Too many branches become difficult; dynamic interactions may be simplified. |
The Bowtie Model
Bowtie combines fault-tree thinking on the left and event-tree thinking on the right. It places the top event—the moment control over the hazard is lost—in the centre.
A collective industrial technique—not one person’s invention
There is no single uncontested Bowtie inventor or creation date. The visual method evolved collectively from fault-tree and event-tree thinking and was progressively adopted in high-hazard industries. It became popular because a multidisciplinary team could see the complete threat–control–loss pathway on one page.
Why it is used: to connect each threat to a preventive barrier, define the loss-of-control top event, connect consequences to mitigating barriers and make critical-control ownership visible.
How it changed: modern practice adds escalation factors, escalation controls, barrier owners, performance standards and assurance evidence. A decorative “bowtie picture” without these elements is not enough for control management.
See the UK Government Bowtie overview and the Office of Rail and Road’s health-risk application.
Vehicle contacts manifold and containment is lost
Escalation factor
A condition that can defeat or weaken a barrier—for example poor lighting reduces the reliability of a visual route check.
Escalation-factor control
A control that protects the main barrier—for example lighting inspection and emergency lighting support route visibility.
Interactive Bowtie Draft Builder
Behavioural Root-Cause Analysis
Behavioural RCA examines what a person did and the conditions that made the behaviour understandable or likely. It should not become a search for someone to blame.
Where the ABC idea came from
Who and when: there is no single creator of “behavioural root-cause analysis.” Its ABC structure grew from twentieth-century behavioural science. Psychologist B. F. Skinner’s work on operant conditioning helped explain how consequences can strengthen or weaken behaviour.
Why safety practitioners use it: repeated behaviour often makes sense when the antecedents are clear and the immediate consequence is easier, faster or socially accepted. ABC analysis makes these influences visible.
How it changed: modern human-factors practice rejects a behaviour-only explanation. ABC evidence should be combined with task design, competence, equipment, workload, supervision, leadership, culture and organisational controls. This prevents “the worker chose badly” from becoming the false root cause.
Interactive ABC and System-Factor Builder
Boundary: A fair system approach does not mean that every action is acceptable. Deliberate reckless conduct may require a just and proportionate response, but the investigation must still examine supervision, controls and organisational context.
Tool: Which Causation Technique Should Lead?
Complex investigations often combine techniques. Select the main question to see a defensible starting point.
From Raw Loss Records to Decision-Useful Evidence
Why a count can mislead
Site A records 8 cases and Site B records 5. Site A may appear worse, but if it worked four times as many hours, its exposure-normalised rate may be lower.
Why a rate can also mislead
A single event can move a small workforce’s rate sharply. A low rate may also reflect under-reporting, changed classification, outsourced exposure or chance—not strong control.
Symbols Used in the Calculations
Calculate and Interpret Loss Rates
Flexible Accident / Incident Frequency-Rate Calculator
Why calculate it? A count alone ignores how long people were exposed. Frequency rate converts defined events into a common hours-worked base, supporting more defensible comparison.
How to say and use every sign
The 200,000-hour convention is used in US OSHA/BLS incidence rates. A one-million-hour base is common for some international frequency measures. Confirm the required definition before comparison.
Severity and Average-Consequence Calculator
Why calculate it? Frequency tells how often events occur; severity shows the amount of defined consequence relative to exposure. Average days per case describes the typical recorded consequence per relevant case.
How to say and use every sign
“Lost day” and severity conventions differ. Record calendar/workday rules, caps, fatalities and restricted work consistently.
Accident Incidence-Rate Calculator
Why calculate it? When reliable hours are unavailable or the required convention uses workers, incidence expresses new defined accidents or cases per standard number of workers.
Do not confuse incidence with frequency: incidence uses workers or people; frequency uses hours worked. “New” means the case started within the reference period.
Ill-Health Prevalence-Rate Calculator
Why calculate it? Prevalence estimates the burden of a condition now—both new and continuing cases—so an organisation can plan health controls, surveillance, support and resources.
Prevalence is not incidence: prevalence counts all qualifying existing cases; ill-health incidence counts only new cases. Long-latency occupational disease may reflect exposures from many years earlier.
Interactive Tool: Which Rate Should I Use?
Formula conventions vary. Examples follow official explanations from the US Bureau of Labor Statistics, International Labour Organization, and UK HSE ill-health statistics guidance.
Tool: Compare Two Sites Fairly
Counts answer “how many?” Rates help answer “how many for the amount of exposure?” Use the same event definition, period and multiplier.
Site A
Site B
Trend, Moving Average and Pareto Analysis
Six-Period Trend Explorer
Enter six comparable monthly rates separated by commas.
Why calculate it? Δ (“delta”) shows relative endpoint change. MA₃ smooths short-term fluctuation by averaging the current and previous two periods. Neither proves the reason for change.
Pareto Priority Explorer
Pareto analysis orders categories from largest to smallest so the team can see where recorded loss is concentrated.
What every symbol means
Histogram, Pie Chart and Line Graph Laboratory
Different charts answer different questions. A chart must match the data structure; attractive graphics do not repair weak definitions or incomplete records.
Interactive Chart Explorer
Choose by question—not appearance
Statistical Variability, Distributions, Sampling and Data Validity
Statistical variability means observations and sample results naturally differ. Validity asks whether the measure actually represents the decision question. A precise calculation from biased or incorrectly classified data is still misleading.
Descriptive Statistics and Approximate CI Explorer
Enter numerical observations such as lost days per case. This tool describes the sample; it does not certify a population model.
Open every equation and symbol
What a distribution can reveal
Representative sample: a sample should reflect the population relevant to the question. Every important subgroup needs a fair chance of inclusion. A large convenience sample can still be biased.
Sampling a population: define the target population, build a suitable sampling frame, select participants or records using a defensible method, record non-response and compare the achieved sample with the population.
UK HSE explains sampling error and 95% confidence intervals in its Labour Force Survey guidance. See also the ONS guide to uncertainty.
Interactive Sample and Validity Check
This is a structured warning tool, not a formal sample-size calculation or statistical certification.
Common data errors to test
| Error | Easy meaning | Example |
|---|---|---|
| Coverage | Some of the population cannot appear in the data. | Night shift and contractors are missing. |
| Selection | The inclusion method favours certain people or records. | Only volunteers answer a wellbeing survey. |
| Non-response | Selected participants do not respond, and may differ from responders. | Workers with symptoms do not trust confidentiality. |
| Measurement | The question, instrument or observer produces inaccurate values. | Different clinics use different symptom questions. |
| Classification | The same case is placed in different categories. | Restricted work is recorded as first aid at one site. |
| Duplicate / missing | A case is counted twice or not counted. | One injury exists in two systems without a unique ID. |
| Denominator | The exposure base excludes relevant work. | Contractor cases included, contractor hours excluded. |
| Processing | Entry, coding, formula or transfer is wrong. | Hours are entered as 40,000 instead of 400,000. |
| Time lag | The measured harm appears long after exposure. | Current respiratory disease reflects earlier dust exposure. |
| Reporting culture | Trust and rules change what becomes visible. | A reporting campaign increases near-miss counts. |
Why Use Quantitative Methods—and Why Not Use Them Alone?
| Reason for using numbers | Decision value | Condition or limitation |
|---|---|---|
| Normalise exposure | Rates allow more defensible comparison across differently sized populations or periods. | Definitions, hours, workforce scope and base multiplier must match. |
| Detect change | Time-series analysis can show sustained deterioration, improvement, seasonality or unusual variation. | Short runs, rare events and changed reporting can produce unstable signals. |
| Measure disease burden | Ill-health prevalence estimates all qualifying existing cases and supports surveillance, control and resource planning. | Long latency, diagnostic access, worker turnover and healthy-worker effects can disconnect current prevalence from current exposure. |
| Prioritise investigation | Pareto and distribution analysis identify categories, locations or activities contributing most recorded loss. | Frequency must be considered with credible severity and major-hazard potential. |
| Express uncertainty | Spread, standard error and confidence intervals show that a sample estimate is not an exact population truth. | The calculation depends on a defensible sample, measurement quality and an appropriate statistical model. |
| Evaluate intervention | Before/after measures can test whether performance changed following control. | Control for exposure, operational change, reporting, regression to the mean and other influences. |
| Communicate performance | Defined indicators support dashboards, accountability and resource decisions. | Targets can encourage under-reporting or classification manipulation if poorly designed. |
| Model event pathways | FTA, ETA and quantitative risk methods estimate the contribution of failure combinations and barriers. | Models contain assumptions, dependencies and uncertainty; precision is not certainty. |
Small numbers
One event can double a rate in a small workforce. Use longer periods, confidence intervals or pooled evidence where appropriate.
Under-reporting
A “good” rate may reflect low trust or restricted definitions. Triangulate with audits, surveys, health data and workforce evidence.
Lagging-only bias
Injury data describes realised outcomes. Add leading evidence about exposure, critical controls, defects and corrective-action quality.
Severity randomness
Similar events can produce very different harm. Do not assume low historical injury means low potential consequence.
Changing denominator
Overtime, contractors, shutdowns and outsourcing change exposure. Record the population and hours consistently.
Metric fixation
Managing the number rather than the risk can distort behaviour. Indicators must serve learning and control—not replace them.
Interactive Level 6 Justification Builder
Why Incident Investigation Remains Essential
Protect people, prevent escalation and avoid disturbing evidence unnecessarily.
Define scope and competence; obtain physical evidence, photographs, measurements, documents, records, data and fair accounts from involved people.
Distinguish fact, interpretation and uncertainty. Check inconsistencies and alternative explanations.
Use suitable models to identify immediate, underlying and root causes, successful controls and failed or missing defences.
Prioritise higher-order and system-level action; avoid relying only on reminders or retraining.
Name owners and dates, manage interim risk, confirm effectiveness and share relevant learning.
Check similar equipment, tasks, sites and procedures so the organisation learns beyond the single event.
How to Answer 3.1 and 3.2 at Level 6
3.1 — Outline
For each theory or technique: state its name and origin/context, describe its principal structure, explain its direction or logic, apply it briefly to an incident, and state one useful feature and limitation.
Suggested structure: Name → main idea → components → application → usefulness → limitation.
3.2 — Justify
Identify the decision, explain why the selected quantitative method fits, show the data and formula, interpret the result, discuss reliability and limitations, combine it with qualitative evidence, and state the resulting action and review.
Suggested structure: Decision → method → evidence → calculation → interpretation → limitation → complementary evidence → action.
Using the forklift–process-line case, outline Bird’s model, multi-causality, Reason’s Swiss Cheese model, FTA, ETA, Bowtie and behavioural RCA. Then justify how incident rates, severity measures, trend and Pareto analysis could support—but not replace—the investigation.
Can You Connect Models, Data and Investigation?
Use Authoritative Information
This independent DB HSE explanation supports learning. Workplace decisions must use current applicable legislation, exposure limits, approved organisational criteria and competent specialist advice where required.