Prepared solely by Debjyoti Biswas · DB HSE International Independent teaching resource · Not produced by OTHM Level 6 · Unit 3 · Sections 01–03
DB HSE Interactive Learning Portal

Welcome to Unit 3

OTHM Level 6 International Diploma in Occupational Health and Safety. Choose the section you want to study. Each section opens its complete explanations, workplace examples, calculations, activities and interactive tools.

othmqualifications
OTHM Qualifications logo
DB HSE International logo
Unit 3 · Section 01

Systems Reliability and Failure Tracing

Welcome to Section 01. Begin with system reliability, complete all nine calculations in 1.6.3, analyse what the results mean in 1.6.4, and use the evidence to trace failures in 1.6.5.

Resource status: This portal was independently prepared solely by Debjyoti Biswas for DB HSE International. It is not an OTHM-produced learning resource. It supports DB HSE teaching for OTHM Level 6, Unit 3, Section 01.
Section 01 1.6.3–1.6.5 9 Live Calculators Analysis & Failure Tracing
Debjyoti Biswas, approved tutor and CertIOSH member
Debjyoti Biswas
CEO & Founder, DB HSE International
Approved OTHM Tutor CertIOSH Member

Meet Your Tutor

Debjyoti Biswas is an Approved Tutor of the OTHM Level 6 International Diploma in Occupational Health and Safety Qualification and a CertIOSH Member. His teaching approach connects Level 6 knowledge with practical workplace application, evidence-based judgement, leadership and professional development.

Approved OTHM Level 6 Tutor CertIOSH Member International HSE Trainer DB HSE International
Learner-support response: “Thank you for raising your concern. Please allow me a moment; I will get back to you.” Questions requiring tutor judgement should be referred to Debjyoti Biswas rather than treated as an official OTHM response.
Start here · Let us talk about reliability first

What Is Reliability?

Let us begin with one easy question:

If you ask a machine, alarm, pump or complete safety system to do its job, how confident are you that it will perform correctly without failing?

That confidence—expressed as a probability—is reliability.

Required function
The exact job the item must perform. A fire pump must deliver water; a gas detector must detect gas.
Stated conditions
The real environment and duty: temperature, pressure, dust, workload, operating mode and maintenance condition.
Specified time
The period for which successful operation is required—for example, a 100-hour mission or an emergency demand.
Without failure
The item completes the required function without stopping, leaking, giving a false signal or becoming unavailable.
Probability
The answer is normally written between 0 and 1, or as a percentage from 0% to 100%.
Reliability = Probability that an item performs its required function, under stated conditions, for a specified time, without failure

Now apply that idea to a complete system.

A system is a group of connected parts—equipment, power, controls, information, people and procedures—working together. System reliability asks whether the complete required function succeeds.

Industrial engineers inspecting a gas detector, control panel, alarm beacon and shutdown valve as one safety system
One safety function may depend on detection, decision, warning and action. Failure of a required link can prevent the complete system from protecting people.

Explore One Safety Function—One Step at a Time

Imagine a gas-release protection system. Select each part to understand its job.

1. Detect the hazardThe gas detector must sense the hazardous concentration accurately and send a valid signal. A blocked sensor, loss of power or overdue calibration can break this first link.

Component reliability

Asks whether one item—such as a sensor, bearing, seal or pump—will perform successfully.

System reliability

Asks whether the complete required function succeeds. Reliable individual parts do not automatically guarantee a reliable system because connections, power, controls, procedures and common causes also matter.

Not only quality

A well-made item may still be unreliable if it is used outside its design conditions or maintained poorly.

Not the same as availability

An item can fail frequently yet remain highly available when every repair is very fast.

Not proof of safety

Reliability supports risk decisions, but safety also depends on consequence, design, inspection, people, procedures and independent protection.

DB HSE learning note: Independently prepared solely by Debjyoti Biswas for Level 6, Unit 3, Section 01. This explanation supports learning and is not an official OTHM interpretation.
Section 1.6.3

Common Reliability Calculations Used in Systems Reliability and Failure Tracing

Use operating, failure and repair data to create evidence. This section explains every symbol, formula, step, answer, limitation and workplace use for all nine calculations.

Independent DB HSE teaching content: prepared solely by Debjyoti Biswas for Level 6 · Unit 3 · Section 01.

1. CalculateUse valid data and correct units. 2. InterpretExplain what the answer means. 3. ApplyUse the evidence to improve control.

One scene · Nine techniques

The same master workplace scene will stay with you throughout Section 1.6.3. We will use the repairable pump, repair records, replaceable sensors, alarm chain and backup pumps to explain all nine techniques without changing the story.

Our master example · One workplace scene

The DB HSE Emergency-Protection System

Yes—we are going to use one connected workplace scene to explain all nine calculation techniques.

Imagine that you and I have entered this chemical warehouse. The protection system contains gas detection, a control panel, an alarm, an isolation valve, replaceable sensor cartridges and two fire-water pumps. We will not keep changing the story. We will simply ask nine different reliability questions about the relevant parts of this one system.

Master reliability example showing gas detection, alarm and control equipment, replaceable sensor cartridges, engineers and two parallel fire-water pumps in one industrial protection system
Our single master scene: one facility, one connected protection system and nine different reliability questions.
Part of our sceneInformation collectedWhy we need it
Repairable protection pump5,000 operating hours; 10 failures; 40 total repair hoursMTBF, MTTR, λ, availability, R(t) and F(t)
Required missionThe pump must operate for the next 100 hoursReliability and failure probability over time
Five replaceable sensor cartridgesLives: 1,800; 2,100; 1,900; 2,200; 2,000 hoursMTTF for non-repairable items
Series alarm chainDetector 98%; control panel 97%; alarm 99%Reliability when every required link must work
Two backup fire-water pumpsEach pump has 90% reliability for the stated demandParallel reliability when at least one independent pump must work

Watch how the same scene gives us nine different questions

1. MTBF

How long does the repairable pump operate, on average, between failures?

2. MTTR

How long does one pump repair take, on average?

3. λ

How frequently does the pump fail per operating hour?

4. Availability

What percentage of time is the pump ready for use?

5. R(t)

What is the chance that it completes 100 hours without failure?

6. F(t)

What is the chance that it fails during those 100 hours?

7. MTTF

What is the average life of a replaceable sensor cartridge?

8. Rs

Will the detector, control panel and alarm all work together?

9. Rp

Will at least one of the two independent fire pumps work?

Important: One scene does not mean one formula fits everything. We select the formula that matches the question, the equipment type and the data available.
Begin with the workplace problem

Why Do HSE Professionals Need These Calculations?

Statements such as “the machine fails frequently” or “the alarm is usually reliable” are opinions. Reliability calculations convert operating, failure and repair records into evidence that can be compared, investigated and improved.

Measure

Determine how often equipment fails, how long it runs and how long repairs take.

Trace

Use trends and component records to locate recurring failures and investigate their causes.

Improve

Verify whether maintenance, redesign, training, spare parts or redundancy improved performance.

Questions the calculations answer

  • How frequently does the equipment fail?
  • How long does it operate before failing?
  • How long is normally needed to repair it?
  • What proportion of time is it available?
  • What is the chance of completing a task without failure?
  • Which component repeats?
  • Did maintenance improve performance?

Safety-critical examples

Fire pumps, emergency alarms, pressure-relief systems, local exhaust ventilation, gas detectors, lifting equipment, emergency generators, rescue equipment and emergency shutdown systems.

The required reliability should reflect the consequence of failure and the availability of independent protection.

Interactive symbol guide

Understand Every Symbol Before Calculating

Select a symbol to see what it means and where it is used.

Total operating time — T

T represents the total time for which equipment actually operated. It must use the same time unit throughout the calculation, such as hours, days or cycles.

Live calculation laboratory

Explore Each Reliability Calculation

Choose a calculation, change the figures and examine how the result and interpretation respond.

Use the same easy method every time: ① Understand what is being measured → ② identify every symbol → ③ place the figures in the formula → ④ calculate → ⑤ explain what the answer means for the workplace.
1 of 9 1 of 9 calculations explored

Mean Time Between Failures — MTBF

Mean means average. MTBF is the average operating time between one failure and the next failure of repairable equipment.

MTBF = T ÷ N

T = total operating time
N = number of failures

Why it is relevant

  • Compares the reliability of machines or operating periods.
  • Supports preventive-maintenance planning.
  • Shows whether failures are becoming more frequent.
  • Provides evidence for investigating deterioration or repeated component failure.

Important interpretation

A higher MTBF is generally better. A falling MTBF—such as 1,000 → 600 → 300 hours—means failures are occurring more frequently and should be investigated.

Maintenance technicians checking vibration, bearing condition, lubrication and alignment on an industrial compressor
Example: A falling compressor MTBF is the clue. Technicians then inspect vibration, bearings, seals, lubrication, alignment, workload and maintenance history to find the cause.

Live MTBF calculator

See the complete worked example
  1. Record total operating time: T = 5,000 hours.
  2. Record failures: N = 10.
  3. Divide 5,000 by 10.
  4. MTBF = 500 hours.
  5. Interpretation: the machine operated for an average of 500 hours between failures.

MTBF trend inspector

What can make MTBF fall?

Equipment deterioration, ageing parts, ineffective maintenance, overloading, poor lubrication, incorrect operation, harsh conditions or repeated failure of the same component. Use maintenance records, failure reports and operating evidence to test these possibilities.

Read the complete picture

Combined Interpretation Dashboard

These figures describe different aspects of the same machine. One figure alone does not provide a complete reliability assessment.

MTBF500 hAverage operating time between failures
MTTR4 hAverage time needed to repair
Availability99.21%Time ready and operational
R(100)81.87%Chance of completing 100 hours without failure

Level 6 interpretation

The machine has high availability because failures can be repaired quickly, but its probability of completing 100 continuous hours without failure is only 81.87%. High availability must not be incorrectly presented as proof of high reliability or safety.

Interactive investigation sequence

Use Calculations for Failure Tracing

Step 1 of 6

Collect operating and failure data

Gather operating hours, downtime, repair duration, failure type, failed component, operating condition and maintenance history. Calculations are only as reliable as the data used.

Section 1.6.4

Using Reliability Calculations to Analyse System Performance

Calculation is only the beginning. Level 6 analysis explains what the figures show, why the change matters, what evidence is still needed and which management action is proportionate.

Independent DB HSE teaching content: prepared solely by Debjyoti Biswas; not produced by OTHM.

CalculateProduce valid indicators. InterpretCompare trends, targets and consequences. Decide & actPrioritise control and verify improvement.
The Level 6 questions

Move From a Number to a Defensible Judgement

Is it acceptable?

Compare the result with the performance target, manufacturer information, legal or organisational requirements and the safety function.

Is it changing?

Compare time periods. Decide whether MTBF is falling, failure rate or MTTR is rising, and whether availability or mission reliability is deteriorating.

What should happen?

Identify weak equipment, consider consequences, investigate causes, prioritise action and allocate competent people, time, spares and budget.

Questions an HSE professional should ask

  • Is the result acceptable for the required safety function?
  • Is performance improving, stable or deteriorating?
  • How does this period compare with earlier periods?
  • Which equipment or component is weakest?
  • What are the safety, health, environmental and business consequences?
  • Is further inspection or investigation required?
  • What management action and resources are justified?
  • Is the data accurate, complete and collected on a consistent basis?
Baseline, trend, comparison and target

Four Ways to Analyse Performance

1. Establish a baseline

A baseline is the starting performance against which future results are compared. Example: MTBF 1,000 h; MTTR 4 h; λ 0.001/h; availability 99.60%; R(100) 90.48%.

The same definition of failure, operating-time boundary, repair-time boundary and units must be used in every later comparison.

2. Analyse the trend

One result is a snapshot. A sequence reveals direction. Falling MTBF together with rising λ and MTTR is stronger evidence of deterioration than one isolated value.

MTBF: 1,000 → 800 → 600 → 400 h

3. Compare equipment

Compare like with like, then explain differences in age, condition, duty, workload, environment, design, operator competence, maintenance, materials and modifications.

4. Compare with a target

A fire pump may have a target availability of at least 99.5%, MTTR below 3 h and λ below 0.001/h. Actual values of 98.7%, 5 h and 0.002/h are adverse gaps requiring action.

IndicatorQ1Q2Q3Q4Interpretation
MTBF (h)1,000800600400Failures are becoming more frequent.
MTTR (h)4.04.55.07.0Recovery is becoming slower.
λ (failures/h)0.001000.001250.001670.00250The deterioration is consistent across indicators.

When should performance be treated as abnormal?

Do not judge a figure in isolation. Compare it with all relevant reference points:

  • The asset’s previous performance
  • Similar equipment performing the same duty
  • Manufacturer specifications and design limits
  • Internal standards and maintenance targets
  • Industry benchmarks and recognised good practice
  • The reliability required for the safety function

A statistically unusual result, a continuing adverse trend or any failure that threatens a critical protection function requires investigation—even when the percentage appears numerically high.

Live system-performance analyser

Compare Two Performance Periods

Change the data. The tool calculates MTBF, MTTR, λ, availability, R(t) and F(t), then explains the direction of change.

Period 1 — baseline

Period 2 — current

IndicatorPeriod 1Period 2Direction
Before and after reliability comparison A grouped bar chart comparing normalised reliability indicators.
Combine indicators and consequences

One Indicator Can Mislead

High availability can hide frequent failure

If MTBF = 500 h and MTTR = 2 h, availability is 99.60%. Yet, with λ = 0.002/h, R(100) is only 81.87%. Rapid repair creates high availability but does not make the system failure-free.

Availability ≠ reliability ≠ safety

Consequence changes the priority

Risk combines likelihood and consequence. The same failure probability may be tolerable for a non-critical printer but unacceptable for a gas detector, fire pump or emergency shutdown system.

IndicatorPump APump BMeaning
MTBF1,200 h450 hPump B fails more frequently.
MTTR3 h8 hPump B is slower to restore.
λ0.00083/h0.00222/hPump B has the higher failure rate.
Availability99.75%98.25%Pump B is less ready for use.

Prioritise for investigation when evidence shows:

  • High or rising failure rate
  • Falling MTBF
  • High or rising MTTR
  • Availability below target
  • Unacceptable F(t)
  • Repeated failure of one component
  • Safety-critical function
  • No effective backup
Section 1.6.5

Using Reliability Calculations in Failure Tracing

Calculations show that performance changed; they do not, by themselves, explain why. Failure tracing follows the evidence through the equipment, task, conditions, human factors and management system until controllable causes are identified.

Independent DB HSE teaching content: prepared solely by Debjyoti Biswas for Level 6 · Unit 3 · Section 01.

When to trace and what to collect

Start With a Clearly Defined Failure Event

Failure identification

States what failed, where and when: for example, “Fire Pump B failed to start during the weekly proof test at 09:20.”

Failure tracing

Examines how and why the failure developed, how it moved through the system and which immediate, underlying and root causes must be controlled.

Common triggers

  • Sudden or continuous reduction in MTBF
  • Rising λ or MTTR
  • Availability below target
  • Unacceptable failure probability
  • Repeated failure of one component
  • Large differences between similar assets
  • Failure soon after maintenance
  • Failure of a safety-critical item
  • Primary and backup equipment failing together
  • Deterioration in previously reliable equipment

Evidence to collect before concluding

  • Operating hours, cycles and duty
  • Failure date, time, type and alarms
  • Repair time, tests and replaced parts
  • Maintenance, inspection and proof-test history
  • Operators, contractors and shift conditions
  • Manufacturer information and design limits
  • Temperature, pressure, vibration and process data
  • Dust, moisture, corrosion and other environmental conditions
  • Photos and preserved damaged components
  • Permit-to-work and isolation records
  • Training, competence and handover records
  • Previous investigations and management-of-change records
Interactive reliability clues

Select the Result That Changed

Choose a reliability clueThe tool will suggest a focused investigation direction. It does not declare a root cause.
Interactive investigation route

Twelve Steps From Function to Verification

Select each step. A competent investigation may move back and forth as new evidence appears.

Cause structure

Do Not Stop at the First Technical Explanation

Immediate cause

The direct event closest to the failure: a bearing overheated and seized.

Underlying cause

The local condition that allowed it: insufficient lubrication and no condition warning.

Root / organisational cause

The management-system weakness: the maintenance system did not specify, schedule or verify lubrication and monitoring.

Five Whys — pump example

Problem: Pump stopped.
Why 1: Bearing seized.
Why 2: Bearing overheated.
Why 3: Lubrication was insufficient.
Why 4: The lubrication task was not completed.
Why 5: The maintenance system did not generate and verify the task.

Fault Tree Analysis — “fire pump fails to start”

Top event = Electrical OR Mechanical OR Control failure
  • Electrical: loss of supply, open protection, starter defect.
  • Mechanical: motor seizure, pump obstruction, coupling failure.
  • Control: sensor, logic, signal, set-point or interlock failure.

FTA helps test multiple credible paths instead of accepting the first visible defect.

Failure Modes and Effects Analysis — FMEA

FMEA is a forward-looking method. For each component or process step, identify the possible failure mode, its effect, its likely cause, existing controls and any further action. It helps prioritise what could fail before an incident occurs.

ItemFailure modePossible effectPossible causeExisting / further control
Pump sealLeakage or face damageLoss of containment and pump shutdownMisalignment, vibration or unsuitable materialAlignment verification, material review and vibration monitoring
Risk Priority Number: RPN = S × O × D

RPN is a prioritisation aid, not a universal risk-acceptance rule. Follow the organisation’s FMEA rating definitions; a catastrophic severity may require action even when the total RPN is not the highest.

Live Pareto failure-frequency tool

Find the Component Dominating the Failure Record

Example total = 12 failures. The tool sorts the components and calculates percentage and cumulative percentage.

ComponentFailuresShareCumulative
Broaden the investigation

Human, Organisational, Common-Cause and Hidden Failures

Human & organisational factors

Examine workload, fatigue, supervision, shift handover, competence, interface design, communication, production pressure, resources, contractor control, unclear roles and ignored warnings.

Common-cause failure

Parallel equipment may share electricity, control panel, fuel, cooling, room, procedure, maintenance error or defective component batch. Redundancy is valuable only when independence is credible.

Hidden failure

An alarm, emergency generator, relief valve, gas detector, interlock or shutdown may fail without being noticed until demanded. Inspection and proof testing must reveal dormant failure.

Full worked failure-tracing case

Recurring Pump-Seal Failure: Before and After

Initial evidence — 6,000 operating hours

12 failures, 60 repair hours; 8 of 12 failures involved the seal.

  • MTBF = 6,000 ÷ 12 = 500 h
  • MTTR = 60 ÷ 12 = 5 h
  • λ = 12 ÷ 6,000 = 0.002/h
  • A = 500 ÷ 505 × 100 = 99.01%
  • R(100) = e−0.2 = 81.87%

Tracing findings

Evidence showed excessive vibration and shaft misalignment. Seals had been replaced repeatedly without checking alignment; no vibration monitoring existed, and the maintenance procedure addressed replacement but not the failure cause.

Failure event → Seal damage → Vibration → Misalignment → Maintenance-system omission

Failure

Pump stopped because leakage exceeded the safe limit.

Mode & immediate cause

Seal faces damaged by excessive vibration.

Underlying & root cause

Misalignment plus a procedure that omitted alignment verification and vibration monitoring.

Corrective actions

Realign the pump, replace damaged parts, introduce vibration monitoring, revise the maintenance procedure, train maintainers, verify alignment after intervention and review similar pumps for the same weakness.

IndicatorBeforeAfter next 6,000 hChange
Failures12375% reduction
MTBF500 h2,000 hLonger between failures
MTTR5 h3 hFaster safe recovery
λ0.002/h0.0005/h75% reduction
Availability99.01%99.85%Improved readiness
R(100)81.87%95.12%Improved mission reliability

Verification: the recalculated results support the conclusion that the actions improved performance. Continue monitoring to confirm that the improvement is sustained and has not introduced another risk.

Evidence boundaries and records

What the Calculations Can and Cannot Prove

They can show

  • Frequency, average duration and direction of change
  • Comparison with another asset, period or target
  • Which component dominates the recorded failures
  • Whether performance improved after action

They cannot prove alone

  • The physical, human or organisational root cause
  • That the data definition and records are accurate
  • That a high availability figure means the system is safe
  • That a short-term improvement will continue

Failure-tracing record

Record the asset and required function; defined failure event; operating hours and duty; failed component and failure mode; immediate, underlying and root causes; evidence reviewed; risk and consequence; corrective actions, owner and due date; post-action calculations; proof of effectiveness; residual risk; and lessons shared with similar systems.

Practical workplace application

Reliability Case Studies

Compressor comparison

Compressor A operates for 12 months without failure. Compressor B requires repair every six weeks.

Annotated welding local exhaust ventilation case showing airflow and fume measurements before, during and after failure correction

LEV performance deterioration

Capture velocity falls from 0.52 m/s to 0.28 m/s while fume concentration rises from 1.2 mg/m³ to 4.8 mg/m³.

Knowledge check

Test Your Understanding

1. What does a falling MTBF normally indicate?
2. Which calculation measures average repair time?
3. Why can availability be high while reliability is lower?
4. What can defeat a parallel backup arrangement?
5. A pump operates 6,000 hours and fails 12 times. What is MTBF?
6. MTBF is stable but MTTR is increasing. Where should tracing focus first?
7. If R(100) = 81.87%, what is F(100)?
8. What is the defining feature of a series system?
9. How is Rp pronounced?
10. Can reliability calculations alone prove the root cause?
11. Seal 8, bearing 2, electrical 1, control 1: which component should receive focused tracing?
12. Why proof-test an emergency shutdown or alarm?
Answer all twelve questions, then check your score.
Unit 3 · Section 02

Understand the Strategies and Techniques of Risk Control

Welcome to the next part of our learning journey. Section 01 helped us identify, calculate and trace risk-related evidence. Section 02 asks the next management question: What should we do about the risk, and how can we justify that decision?

Learning outcomes: 2.1 Evaluate the use of common risk-management strategies. 2.2 Justify when to use risk avoidance, risk reduction, risk transfer, risk analysis, risk evaluation and risk review strategies. 2.3 Explain the development and characteristics of safe systems of work and safe operating procedures.

Let us begin with the simplest meaning

What Is Risk Control?

Imagine that we find an unguarded moving part on a machine.

The moving part is the hazard. The possibility and seriousness of someone being injured is the risk. The guard, isolation system or redesign used to prevent contact is the control.

Risk control means selecting, implementing and checking measures that eliminate the hazard or reduce the risk to an acceptable or required level.

Risk assessment is not the final product

A completed form does not protect anyone. Protection appears only when suitable controls are implemented, communicated, resourced, used and verified.

Easy memory: Find it → Understand it → Control it → Check it.

Hazard

Something with the potential to cause injury, ill health, damage or another unwanted outcome.

Risk

The combination of how likely harm is and how serious the consequence may be, considering exposure and existing controls.

Control

A measure that removes the hazard, prevents exposure, reduces likelihood, limits consequence or supports safe recovery.

Level 6 lens: Do not only name a strategy. Evaluate its strengths and limitations, then justify why it is suitable for the people, hazard, evidence, standards and workplace conditions involved.
Section 2.1

Evaluate the Use of Common Risk-Management Strategies

Evaluation means more than describing what a strategy is. Ask whether it is suitable, what evidence supports it, what benefit it offers, where it may fail and how its effectiveness will be checked.

1. SuitabilityDoes it fit the risk and context? 2. EvidenceWhat information supports the choice? 3. LimitationsWhat uncertainty or weakness remains?

The Risk-Assessment Process—Seven Clear Steps

Different organisations may group the stages differently. Here, we use the seven stages indicated for this learning outcome, followed by a continuous review loop. Select each step to hear the portal explain it.

Step 1 — Identify risks and hazardsExamine routine and non-routine work, equipment, substances, people, environment, foreseeable misuse, maintenance, change and emergencies. Use observation, worker consultation, incident data, instructions and technical information.
AssessBuild a suitable and sufficient picture. ActImplement the planned controls. ReviewCheck effectiveness and reassess change.

Control Standards, Action Plans and Priority

What is a risk-control standard?

It is the benchmark that tells us what level or type of control is required. Sources may include legislation, approved guidance, exposure limits, engineering codes, manufacturer instructions, industry good practice, internal rules and the hierarchy of controls.

Why it matters: A coloured risk score alone cannot decide whether a mandatory guard, ventilation system, permit or exposure limit is required.

What makes an action plan useful?

  • Specific control action—not “be careful”.
  • Named responsible owner and adequate resources.
  • Realistic completion date and interim protection.
  • Priority based on risk, standards, people and uncertainty.
  • Method to verify completion and effectiveness.

Priority is not “highest score only”

Give prompt attention to imminent danger, serious consequences, legal or control-standard gaps, many people exposed, vulnerable persons, ineffective controls and high uncertainty. A lower matrix score must not be used to postpone a non-negotiable requirement.

Generic, Specific and Dynamic Risk Assessments

These are not three levels of quality. They are three ways of matching the assessment to the work situation.

Assessment typeEasy meaningUse it when…Do not rely on it when…Main strength and limitation
GenericA baseline assessment for similar activities, hazards or locations.Work is routine, repeated and genuinely comparable; common controls can be standardised.People, equipment, substances, environment or task conditions differ materially.Strength: efficient and consistent. Limitation: may overlook local or individual differences.
SpecificAn assessment for one task, site, machine, substance, project or person.Work is unusual, complex, high risk, legally specific, non-routine or affected by vulnerability.A suitable generic assessment already covers truly identical low-risk work—although local checks are still needed.Strength: detailed and relevant. Limitation: requires time, information and competence.
DynamicA continuous, in-the-moment judgement as conditions change.Emergency response, rapidly changing work or an unexpected condition requires immediate reassessment.Work is planned and foreseeable. It must not replace a suitable formal assessment, method statement or permit.Strength: responds to reality. Limitation: time pressure and incomplete information may weaken judgement.

Interactive Assessment-Type Selector

Describe the situation. The tool will recommend a starting approach and explain why.

Now let us see the paperwork behind the words

What Do Generic, Specific and Dynamic Assessments Actually Look Like?

Think of these as three different lenses. A generic assessment establishes a reusable baseline. A specific assessment focuses that baseline on one real task, place, item, substance or person. A dynamic assessment keeps checking the live situation while work or an incident develops.

Safety adviser and warehouse supervisor reviewing a baseline assessment beside segregated loading-bay traffic
Generic — the repeatable baselineUseful when the activities, hazards and standard controls are genuinely similar. The local supervisor must still check that the real people, place, equipment and conditions match.
Competent team carrying out a pre-entry assessment at an industrial confined space with gas testing, ventilation and rescue arrangements
Specific — one real job in its real contextThis vessel entry needs named isolations, atmospheric tests, entrants, rescue arrangements, permit interfaces and a time-limited decision. A generic confined-space form is only background.
Industrial response team behind an exclusion barrier while a supervisor checks changing leak conditions with a detector and radio
Dynamic — the situation is changing nowThe leader continually observes, reassesses, controls, communicates and decides whether to proceed, withdraw or escalate. The later debrief must feed lessons back into formal assessments.
History note: These three formats were not invented together by one person. Generic and task-specific forms grew from systematic safety-management practice. Dynamic risk assessment was developed particularly through fire-service incident command for dangerous, unpredictable environments and was later applied more widely. It supplements planned assessment; it does not excuse foreseeable work from proper planning.

Interactive Assessment-Format Explorer

Select a type. The portal will show its purpose, its header, the columns a suitable form normally needs, a completed example and the final decision that must be recorded.

Core fields every formal assessment needs

  • Clear scope, boundaries, location, activity and version.
  • Assessor, competent contributors, workers consulted and approval.
  • Hazards and credible harm—not only a list of objects.
  • Who may be harmed, including contractors, visitors and vulnerable people.
  • Existing controls and evidence that they are really in place.
  • Risk judgement before further action, with the reasoning shown.
  • Further controls, owner, due date, interim protection and priority.
  • Residual risk, communication, verification and review triggers.

A blank box is not automatically a bad form

Good forms create space for sound thinking; they do not replace it. The assessor must walk the task, consult the people who understand the work, use relevant evidence and test assumptions.

Quality test: Could a competent supervisor read this record and understand what may happen, who is exposed, which controls must exist, what remains to be done and when work must stop?

Do Not Miss Long-Term Hazards to Health

An injury hazard may produce an immediate event. Many health hazards are quieter: exposure can accumulate and illness may appear months or years later. “Nothing happened today” is not evidence that the risk is controlled.

Noise—hearing damage can accumulateConsider sound level, exposure duration, work pattern, combined sources and susceptible workers. Use competent measurement where needed, compare with the applicable standard, control at source and review hearing-protection and health-surveillance evidence.

Exposure pattern

Frequency, duration, intensity, route, peaks, recovery time and combined exposure.

Evidence

Monitoring, sampling, health surveillance, absence records, worker reports and historical data.

Control

Prevent exposure at source. PPE and surveillance support control; they do not replace elimination or engineering where these are required.

Qualitative, Semi-Quantitative and Quantitative Assessment

The difference is how the risk is described and analysed—not how seriously the assessor takes it.

Words

Qualitative

Uses reasoned descriptions such as low, medium, high, unlikely or severe.

Useful for: straightforward work, screening, discussion and situations where numerical data would add little.

Limitation: categories may be subjective and different risks may appear equal.

Ordered scores

Semi-quantitative

Assigns ranked numbers to likelihood and severity, often calculating a matrix score such as L × S.

Useful for: consistent comparison and action planning across many hazards.

Limitation: numbers can create false precision; score boundaries and multiplication rules are organisational conventions.

Measured or modelled values

Quantitative

Uses numerical estimates of exposure, probability, frequency or consequence based on data and models.

Useful for: complex, high-consequence or technical decisions and comparison with numerical standards.

Limitation: needs valid data, competence and transparent assumptions; an exact-looking number may still be uncertain.

Important: Quantitative does not automatically mean “better”. Use the simplest method that is sufficiently reliable for the decision, consequence, uncertainty and applicable standard.
How did these methods develop?

From Professional Description to Ranked Scores and Probability Models

There is no honest single-name answer to “who invented risk assessment?” People have judged danger for as long as organised work has existed. Modern methods developed gradually as industries needed decisions that were more systematic, comparable and technically defensible.

Industrial safety team considering descriptive hazard evidence, a risk matrix and modelled technical evidence together
One decision may need three kinds of evidence

The methods are a progression of detail—not a competition

Qualitative thinking helps us name what may happen and judge its importance. Semi-quantitative scoring helps a group rank and compare many concerns using defined categories. Quantitative analysis estimates frequency, probability, exposure or consequence when the decision needs numerical evidence.

A major quantitative study still begins with qualitative questions: Which scenarios matter? What assumptions are credible? Who may be affected? A risk matrix still needs professional judgement. The methods often work together.

Qualitative assessment—description came firstThere is no single inventor or creation date. Early safety decisions were commonly deterministic and descriptive: experience, rules, testing and expert judgement were used to ask what could go wrong and what consequences could follow. Qualitative assessment remains useful for screening and straightforward work, especially when numerical data would not improve the decision. Its terms must be defined so “unlikely” or “major” means something consistent.

What the historical landmarks do—and do not—prove

System-safety programmes helped formalise ranked categories and risk matrices, but no single standard invented every semi-quantitative method. By 1971, NASA and aircraft manufacturers were using fault-tree tools. The 1975 US Reactor Safety Study, WASH-1400, directed by MIT professor Norman Rasmussen and AEC staff member Saul Levine, became a landmark probabilistic risk assessment. Later criticism of some numerical claims also taught an essential lesson: model structure, data gaps and uncertainty must be made visible.

Historical teaching sources: US Nuclear Regulatory Commission histories of the Reactor Safety Study and risk-informed regulation; NASA System Safety Handbook risk-matrix material; MIL-STD-882 system-safety practice.

Full formats + worked comparison + decision tools

The Qualitative, Semi-Quantitative and Quantitative Assessment Workshop

We will use one scene throughout: pedestrians and forklifts interact in a busy warehouse loading area. This makes the difference easy to see—the hazard does not change, but the depth and form of the analysis changes.

1. Qualitative formReasoned descriptors, evidence and a narrative decision.
Open qualitative form ↓
2. Semi-quantitative formDefined rankings combined in the interactive 5 × 5 matrix.
Open matrix ↓
3. Quantitative formModelled event frequency, uncertainty range and numerical criterion.
Open quantitative form ↓
Qualitative

Reasoned words

Judgement: “Collision risk is high because pedestrians frequently cross an active vehicle route and a collision could cause fatal injury.”

Decision value: Fast, understandable screening that identifies the urgent need for physical segregation.
Semi-quantitative

Defined ranks

Judgement: Likelihood 4 (likely) × severity 5 (catastrophic) = 20, “very high” under the example organisation’s approved matrix.

Decision value: Helps compare this issue with other recorded hazards and apply action rules consistently.
Quantitative

Measured or modelled values

Evidence: vehicle movements/hour, crossing frequency, near-miss rate, speed, exposure time, barrier reliability and predicted collision consequence.

Decision value: Tests route designs or investment options where reliable data and a suitable model exist.

Interactive Method-Format Explorer

Select a method to see the complete form structure. Notice that each higher-data method retains the basic hazard, people, controls, action and review fields.

Which Method Is Proportionate?

Answer five questions. The recommendation is a starting point for competent judgement—not an automatic approval.

Build a Defensible Qualitative Judgement

“High risk” alone is weak. Link the descriptor to exposure, consequence, controls, evidence and uncertainty.

Complete the evidence and let the portal build the reasoning chain.

Create One Complete Risk-Assessment Record

This tool mirrors the essential columns of a practical assessment register. It helps learners see that a risk rating sits inside a much larger management record.

Complete the owner and due date, then build the assessment row.
Complete interactive format 1 of 3

Interactive Qualitative Risk-Assessment Form

A qualitative assessment uses defined words and reasoned professional judgement. It does not multiply scores. The record must explain why the likelihood, consequence and overall priority descriptions fit the evidence.

A. Assessment identity and scope

B. Hazard, people and evidence

C. Initial qualitative judgement

D. Treatment, responsibility and residual judgement

Enter the assessment and action dates, then generate the complete qualitative record.

Example likelihood meanings

  • Rare: exceptional under the defined conditions.
  • Unlikely: foreseeable but not expected during normal activity.
  • Possible: could occur during the activity or assessment period.
  • Likely: expected to occur repeatedly unless controls improve.
  • Almost certain: occurs frequently or conditions make occurrence imminent.

Example consequence meanings

  • Minor: limited, short-term harm.
  • Moderate: treatment or restricted work may be required.
  • Serious: major injury or significant occupational ill health.
  • Major: life-changing harm or single fatality potential.
  • Catastrophic: multiple fatalities or major widespread impact.
Important: These descriptor definitions are teaching examples. An organisation must approve definitions appropriate to its activities and use them consistently. Legal requirements and recognised control standards override a convenient “low” description.

Interactive 5 × 5 Semi-Quantitative Risk Matrix

Compare the initial risk with the residual risk after proposed controls. The example bands are for learning only; an organisation must define and approve its own criteria.

Choose the ratings

Initial risk
20
Residual risk
10
1–4 Low5–9 Moderate10–16 High17–25 Very high

Solid outline: initial rating. Dashed white outline: residual rating.

Why might severity stay at 5?

A guard or interlock may make contact much less likely, but if contact still occurs the possible injury may remain catastrophic. Do not automatically reduce both numbers simply because controls were proposed. Rate the real effect of the selected controls and verify them.

A Simple Quantitative Comparison Tool

Measured value compared with an applicable limit

This demonstration calculates an exposure ratio. Both figures must use the same unit and come from a valid assessment.

Exposure ratio = Measured value ÷ Applicable limit

How to interpret carefully

A ratio of 0.70 means the measured value is 70% of the selected limit. It does not automatically mean there is no risk or that controls can be relaxed.

Check sampling quality, uncertainty, peak exposure, routes of exposure, combined substances, vulnerable people, the legal meaning of the limit and whether further reduction is required.

Quantitative Scenario-Frequency Calculator

This teaching model follows one event path. It estimates how often the defined harmful outcome may occur by combining an initiating-event frequency with conditional probabilities. It is useful for learning the logic of event trees; it is not a substitute for a validated QRA.

fharm = finitiator × P(exposure) × P(control failure) × P(harm | event)
Read the full formula aloud: “F sub harm equals F sub initiator, multiplied by the probability of exposure, multiplied by the probability of control failure, multiplied by the probability of harm given the event.”

How to Read Every Symbol—and Why It Is Used

fharmPronounced: “f sub harm.”

It means the estimated frequency of the defined harmful outcome, normally stated per year. It appears on the left because this is the answer the model is calculating.

finitiatorPronounced: “f sub initiator” or “initiating-event frequency.”

It means how many times the event that starts the harmful scenario occurs during a stated period, usually one year. The small word below f is a label—it is not multiplication.

PPronounced: “probability.”

P shows that the value inside the brackets is a chance from 0 to 1. For example, 0.25 means a 25% chance.

P(exposure)Pronounced: “probability of exposure.”

It means the chance that a person is present or exposed when the initiating event occurs. It is used because an initiating event cannot harm a person who is not in the exposure path.

P(control failure)Pronounced: “probability of control failure.”

It means the chance that the intended barrier, safeguard or protective control fails, is unavailable or does not stop the event path.

P(harm | event)Pronounced: “probability of harm given the event.”

The vertical bar | means “given that”. It asks: once the event has reached the exposed person, what is the chance of the defined harm?

×Pronounced: “multiplied by.”

Multiplication is used because the harmful outcome follows a sequence: the initiating event occurs, exposure exists, the control fails and harm follows. Each stage narrows the original frequency.

=Pronounced: “equals.”

It separates the answer on the left from the factors used to calculate it on the right. Both sides describe the same estimated harmful-outcome frequency.

per year · /yearPronounced: “per year.”

This is the unit, not a percentage. A result of 0.12 per year is a modelled event frequency; it does not mean a 12% annual risk unless a valid model specifically supports that interpretation.

Easy memory: Start events per year × chance of exposure × chance the control fails × chance harm follows = estimated harmful outcomes per year.

Four checks before trusting the output

  1. Scenario: Is the initiating event and harmful outcome defined without ambiguity?
  2. Dependence: Are the probabilities really independent, or can one common cause defeat several controls?
  3. Data: Are frequencies based on comparable equipment, tasks, people and operating conditions?
  4. Uncertainty: Would reasonable lower and upper assumptions materially change the decision?

Level 6 point: A calculated frequency is an estimate conditional on a model. Report units, source, time period, assumptions, uncertainty and sensitivity—not only the final number.

Complete interactive format 3 of 3

Interactive Quantitative Risk-Assessment Form

A quantitative assessment records more than a calculation. It defines the decision, scenario and model; gives every input a source and unit; shows uncertainty; compares the result with an approved criterion; and records treatment and verification.

A. Decision, scope and model boundary

B. Numerical model inputs

fharm = finitiator × P(person exposed) × P(protection fails) × P(harm | exposure)

The pronunciation and meaning of every symbol are explained in the symbol guide immediately above.

C. Assumptions, treatment and assurance

Enter the assessment date, review the numerical inputs and generate the complete quantitative record.

What the uncertainty factor means here

An uncertainty factor of 2 displays a teaching range from the central estimate ÷ 2 to the central estimate × 2. A real QRA may require probability distributions, confidence intervals, alternative models or structured expert judgement. The factor is a learning device—not a universal scientific rule.

What must be reviewed independently?

Scenario completeness, units, input provenance, relevance of data, dependencies and common causes, human-reliability assumptions, consequence model, uncertainty treatment, sensitivity, numerical criterion and whether the model is valid for the decision.

Decision rule: Do not compare a result with an invented criterion. The criterion must come from an applicable legal, regulatory, technical or approved organisational framework. Even a result below a criterion does not remove the duty to apply required good practice and reasonably practicable controls.

Build a Complete Risk-Control Action

An action without an owner, date, interim measure and verification method is only an intention. Complete the fields and create a practical action-plan entry.

Complete the fields, then create the action-plan entry.
Section 2.2

Justify When to Use Six Risk-Management Strategies

Avoidance, reduction and transfer are treatment responses. Analysis, evaluation and review are decision and assurance activities that help us choose, prioritise and verify treatment. In practice, a strong risk-management decision may combine several of them.

To justify means: state the chosen strategy, connect it to evidence and criteria, explain why it suits the context, recognise its limitations and state how it will be reviewed.

Meet the Six Strategies

Select a card. The portal will explain when to use it, when not to depend on it, and what a Level 6 justification should recognise.

Risk avoidance—remove the decision to be exposedUse when: the risk is intolerable, the activity is unnecessary, reliable control is not reasonably achievable, or a safer design or method can remove the hazard. Do not misuse it: cancelling one activity may shift risk elsewhere. Check the alternative for new hazards and operational consequences. Example: redesign a roof-level valve so routine operation can be completed at ground level.

When Each Strategy Is Most Relevant

StrategyUse when…Evidence neededKey limitation or warning
AvoidanceExposure can be removed by stopping, substituting, redesigning or choosing another objective.Severity, feasibility of alternatives, standards, lifecycle and risk-transfer effects.May create a different risk or sacrifice an essential activity; assess the replacement.
ReductionThe activity is necessary and the risk can be controlled using the hierarchy of controls.Control performance, human factors, maintenance, residual risk and verification.Administrative controls and PPE are vulnerable to failure; reduce at source where possible.
TransferSpecialist competence or financial sharing is appropriate—for example, a competent contractor or insurance.Competence, contract scope, interfaces, supervision, insurance and monitoring.Legal and ethical responsibility for protecting people is not simply transferred away.
AnalysisThe risk, causes, exposure, failure paths or options are uncertain or complex.Measurements, incidents, task information, models, assumptions and uncertainty.Do not delay obvious immediate controls while waiting for perfect data.
EvaluationAnalysed risk must be compared with legal, technical or organisational criteria to set priority and treatment.Approved criteria, control standards, consequence, affected groups and tolerance.A matrix colour cannot override a mandatory requirement or conceal uncertainty.
ReviewTime has passed or change, incident, failure, new evidence, worker concern or new requirements may affect validity.Inspection, monitoring, incidents, health data, assurance findings and change information.A scheduled annual review is insufficient when a trigger requires immediate reassessment.

Interactive Strategy Decision Lab

Choose a workplace situation and the strategy you think should lead the response. The tool will explain the strongest answer and supporting strategies.

Choose the scenario and your leading strategy, then check the decision.

Risk Reduction Must Follow the Hierarchy of Controls

When a risk cannot be avoided completely, begin with controls that act on the hazard and exposure pathway. Measures lower in the hierarchy usually depend more heavily on consistent human behaviour.

1 · Most effective

Eliminate

Remove the hazard from the work—for example, design out the need to enter a vessel.

2

Substitute

Replace it with a safer material, method, machine or energy source, then assess the substitute’s hazards.

3

Engineering

Isolate people through guarding, enclosure, segregation, automation, extraction or fail-safe design.

4

Administrative

Use planning, permits, procedures, competence, supervision, scheduling, signage and restricted access.

5 · Last line

PPE

Protect the individual when exposure remains. Select, fit, maintain and supervise its use; do not make it the automatic first answer.

Avoidance and elimination can overlap, but the emphasis differs: avoidance changes the decision or objective so the exposure is not undertaken; elimination removes the hazard or hazardous step from work that continues. In both cases, check that the alternative does not introduce a new serious risk.

Risk Transfer—The Point Learners Must Not Miss

What can be transferred or shared?

Some financial loss may be insured. Specialist work may be contracted to an organisation with suitable equipment and competence. Contract terms may allocate defined responsibilities.

What does not disappear?

The hazard remains until controlled. The client or employer must still select competent parties, provide information, coordinate interfaces, monitor work and meet applicable legal duties. A signature on a contract is not a physical control.

Easy example: Hiring a specialist crane contractor may transfer performance of the lift to specialist hands, but the site still has to coordinate exclusion zones, ground conditions, permits, communication and emergency arrangements.

Build a Level 6 Justification

Use the structure Decision → Because → Evidence → Limitation → Review. This produces a learning scaffold that you should explain in your own professional words.

Complete the evidence, limitation and review trigger, then build the explanation.

Analysis, Evaluation and Review—Do Not Mix Them Up

Risk analysisHow can harm occur? What are the causes, likelihood, consequences, controls and uncertainty? Risk evaluationHow does that analysed risk compare with criteria? Is more treatment required and how urgent is it? Risk reviewIs the assessment still valid and are the controls present, used and effective?

One simple example

Analysis: solvent-vapour measurements, duration and ventilation performance show the nature and level of exposure. Evaluation: the evidence is compared with applicable exposure criteria and good practice to decide whether treatment is adequate. Review: monitoring and reassessment confirm whether the new local exhaust ventilation continues to control exposure after process or maintenance changes.

Common Errors—and the Better Approach

Weak approachWhy it is weakBetter Level 6 approach
Copy a generic assessment without checking the site.Local hazards, people and conditions may be different.Use it as a baseline, then verify and adapt it before work.
Use a dynamic assessment for planned high-risk work.It avoids proper planning, consultation and control design.Complete a formal specific assessment, then use dynamic checks for real-time change.
Reduce both likelihood and severity scores automatically.The selected control may affect only one dimension.Explain how each control changes exposure, failure path or consequence.
Call a risk “low” because no one has yet been harmed.Absence of recorded harm is weak evidence, especially for latent health risks.Consider exposure data, potential severity, under-reporting and control reliability.
Transfer work and stop managing it.Contracting does not eliminate the hazard or all duties.Assess competence, coordinate, monitor and verify contractor controls.
Review only once a year.A change or failure can make the assessment invalid immediately.Use scheduled and event-triggered review.
Section 02 knowledge check

Can You Make and Justify the Decision?

1. When is a generic assessment most suitable?
2. What is the main limitation of a dynamic assessment?
3. A 5 × 5 likelihood–severity matrix is normally…
4. Why may long-term health hazards be underestimated?
5. Relocating a roof valve to ground level is mainly…
6. What does risk transfer NOT mean?
7. What is risk evaluation?
8. When should a risk assessment be reviewed?
9. Why might residual severity remain unchanged?
10. Which is the strongest Level 6 justification?
Answer all ten questions, then check your score.
Section 2.3

Explain the Development and Characteristics of Safe Systems of Work and Safe Operating Procedures

You have identified the hazard, assessed the risk and selected a strategy. The next question is practical: How will people complete the work safely, consistently and under control?

UnderstandDistinguish the documents and their purposes. DevelopTurn risk decisions into a workable system. AssureTrain, supervise, verify and review.
Your guided learning route

Learn One Control Layer at a Time

The portal starts with the wider safe system of work, develops it, then moves to the more focused safe operating procedure. Only after both are clear do we connect the document family and explore the complete permit-to-work lifecycle.

DB HSE International logo

DB HSE learning resource. Prepared solely by Debjyoti Biswas for teaching Unit 3, Section 2.3. This independent learning portal is not produced by OTHM and does not issue workplace authority.

The bridge from assessment to action

A Risk Assessment Decides What Must Be Controlled; the System of Work Decides How

Imagine a chemical-transfer pump that needs maintenance.

The assessment identifies hazardous chemical residue, stored pressure, electricity, moving parts, restricted access, contractors and possible conflict with nearby operations. That information is essential—but it does not yet tell the maintenance team exactly how the job will be prepared, authorised, completed, checked and handed back.

The safe system of work connects the people, equipment, controls, communication and sequence. The safe operating procedure gives the approved steps for a defined operation inside that wider system.

The paperwork test

A document does not make work safe merely because it has been signed. The controls must exist at the workplace, the people must understand them, and a responsible person must verify that they remain effective.

Easy memory: Assess → Design → Explain → Do → Check → Improve.

Level 6 lens: “Explain” requires a connected account of what an SSOW and SOP are, how they are developed, why each characteristic matters, when different strategies are used, and what may happen if the arrangements are weak.

One Scenario, Four Stages of Control

We will use the same chemical-transfer pump throughout Section 2.3. This allows you to see how one risk picture is converted into a complete working arrangement.

Technician beginning unsafe chemical-transfer pump maintenance without verified isolation, barriers, supervision or coordinated controls
1. The uncontrolled starting pointThe task has begun before the energy, chemical, access and coordination risks have been converted into verified controls.
Four industrial professionals consult beside a chemical-transfer pump, reviewing isolation points and a pre-job plan
2. Consultation and developmentThe operator, maintainer, supervisor and HSE professional combine task knowledge, hazard information and practical experience.
Technician and supervisor verify isolated, locked-out chemical-transfer pump maintenance within a controlled exclusion zone
3. Controlled executionIsolation, safe condition, barriers, tools, competence, PPE and supervision are present and verified before intrusive work begins.
Responsible people jointly verify a permit, isolation points and safe handover before chemical-transfer pump maintenance
4. Authorisation and handoverThe issuing and performing parties share the same understanding of the job, limits, precautions, status and return-to-service requirements.
Start here · the wider control system

What Is a Safe System of Work—SSOW?

Safe System of Work — SSOW

Pronounced: “S-S-O-W,” or simply “safe system of work.”

An SSOW is a deliberately organised method for completing work so that foreseeable hazards are controlled throughout preparation, execution, completion and foreseeable abnormal conditions.

It is a system because it joins people, plant, materials, environment, controls, responsibilities, communication, competence, supervision and review. It is not only a list of steps.

What an SSOW is not

It is not a risk-assessment form, a signature, a list of PPE, a copied method statement or an instruction to “be careful.” It must convert risk decisions into a realistic arrangement that people can understand and use.

Simple test: Could a competent team use this system to know who does what, in what order, with which controls, when to stop, and how the work returns safely to normal?

1 · Scope and boundaryTask, equipment, location, start and finish points, exclusions and conditions of use.
2 · People and authorityRoles, responsibilities, competence, supervision and stop-work authority.
3 · Hazards and evidenceRisk assessment, legal and technical standards, manuals, monitoring and incident learning.
4 · Controls and sequenceElimination or reduction measures, isolation, access, PPE, hold points and safe order.
5 · CommunicationBriefing, language and literacy needs, interfaces, handovers and changes in plant status.
6 · Abnormal conditionsStop rules, failed tests, loss of control, alarm, withdrawal, rescue and emergency escalation.
7 · Verification and handbackPhysical checks, completion, accounting for people/tools, reinstatement and acceptance.
8 · Assurance and reviewMonitoring, supervision, document control, feedback, investigation and review triggers.
SSOW format fieldWhat to recordWhy the field matters
Identity and controlTitle, number, owner, version, approval and review date.Prevents obsolete or unapproved instructions being used.
Purpose, scope and limitsActivity, plant, location, persons, conditions, interfaces and exclusions.Stops the system being applied outside the conditions it was designed for.
Roles and competenceWho plans, authorises, performs, supervises, verifies, hands back and reviews.Prevents gaps, duplication and unverified assumptions.
Hazards and control basisAssessment reference, standards, energy sources, substances, exposure and credible failures.Shows that the method is risk-based and technically supported.
Preparation and resourcesAccess, barriers, isolations, tools, staffing, communication, permits and PPE.Creates the conditions needed before work starts.
Safe sequence and hold pointsOrdered actions, responsible role, required result and checks before progression.Some controls only work when applied in the correct order.
Stop, abnormal and emergency rulesConditions that suspend work, safe state, escalation, rescue and recovery.Prevents unsafe improvisation when assumptions change.
Completion, handback and reviewInspection, reinstatement, records, acceptance, monitoring and review triggers.Controls the return to normal operation and captures learning.

Interactive Jargon Translator

Select any term. The portal will pronounce it, define it and explain why it matters.

Competent personPronounced: “kom-puh-tent person.” A person with suitable knowledge, training, skill and experience—and the ability to recognise their own limits—for the assigned work. A certificate alone does not prove competence for every situation.

Do Not Mix Up the Document Family

The documents are connected, but they do different jobs. The level of formality should be proportionate to the risk, complexity and need for coordination.

1 · AssessIdentify hazards, people, risk, standards and further controls.
2 · DesignConvert the assessment into an SSOW for the complete activity.
3 · InstructUse SOPs or method steps for defined operations inside the SSOW.
4 · AuthoriseUse PTW where specified high-risk work needs formal time-and-place control.
5 · PerformBrief, supervise, follow controls and monitor changing conditions.
6 · Hand backInspect, reinstate, communicate status and return plant to its owner.
7 · LearnKeep records, investigate deviations and review every affected layer.
Risk assessmentWhat could cause harm? SSOWHow will the complete activity be controlled? SOPWhat approved operating steps must be followed? Permit-to-workWho authorises this defined high-risk job, where and when?
Document or arrangementEasy purposeTypical useImportant limitation
Risk assessmentIdentifies hazards, people, existing controls, risk and further action.Before deciding the safe method and whenever relevant change occurs.A completed form does not implement the controls.
SSOWCoordinates the whole method, people, controls and interfaces.Where risks require a defined safe way of working, particularly complex or significant activities.It fails if impractical, unknown, unsupervised or not followed.
SOPStandardises the safe steps for a defined operation.Routine or repeated operation, inspection, start-up, shutdown, cleaning or maintenance.It cannot predict every abnormal condition; stop and escalation rules are needed.
Method statementDescribes how a particular job or project stage will be carried out.Construction, installation, maintenance and contractor work.A generic copied statement may not match the real site or sequence.
JSA/JHABreaks a job into steps, hazards and controls.Task planning and workforce discussion.Step-by-step analysis must still consider interactions and emergencies.
Permit-to-work — PTWFormally authorises specified work, location, time and precautions.Defined high-risk or tightly coordinated work under site rules.A permit is not a guarantee of safety and does not replace risk assessment.
ChecklistConfirms that required checks were completed.Pre-start, inspection, handover and verification.Ticking boxes without observation gives false assurance.
Emergency procedureExplains response when control is lost or conditions become unsafe.Credible abnormal and emergency situations.It must be resourced, communicated and tested—not only filed.

Tool: Which Arrangement Does This Job Need?

Development is a lifecycle, not a typing exercise

Fifteen Stages for Developing an Effective Safe System of Work

The stages are grouped into five phases. Select a phase to explore what must happen and why.

Phase 1 — Understand the real work1. Define the task, purpose, location, boundaries and normal/abnormal conditions. 2. Gather legislation, technical information, manufacturer instructions, incident history and existing procedures. 3. Observe the job and consult the people who perform, supervise and maintain it. The aim is to understand work as done—not only work as imagined.
Define the task and boundaries. State what is included, excluded, where it happens, what must be achieved and when the system applies.
Gather reliable information. Use applicable law, guidance, designs, manuals, safety data, previous assessments, monitoring and incident learning.
Observe and consult. Involve workers, supervisors, contractors, engineers and specialists who understand the real task and foreseeable shortcuts.
Identify hazards and credible failure paths. Include people, plant, substances, energy, environment, human factors, simultaneous operations and emergencies.
Analyse and evaluate risk. Understand causes, exposure, likelihood, consequence, uncertainty and applicable criteria.
Select the risk strategy. Avoid where possible; otherwise reduce using the hierarchy. Use specialist transfer, analysis, evaluation and review appropriately.
Design the control sequence. Put controls in the order they must exist, identify hold points and define stop-work conditions.
Assign roles and competence. State who prepares, authorises, performs, supervises, verifies, hands over and reviews.
Plan abnormal and emergency conditions. Explain safe shutdown, alarm, withdrawal, containment, rescue and escalation where relevant.
Write the SSOW and supporting SOPs. Use clear language, diagrams and workplace terminology. Remove ambiguity and unnecessary complexity.
Walk through and test. A competent team checks the sequence against the real workplace before full implementation.
Approve and control the document. Record owner, approver, version, date, review date and controlled availability.
Communicate, train and confirm understanding. Adapt for language, literacy, learning needs and role-specific competence.
Implement and supervise. Provide time, equipment, staffing and authority to stop when conditions differ.
Monitor, learn and improve. Observe work, inspect controls, investigate deviation, consult users and revise after relevant triggers.
Now focus on a defined operation

What Is a Safe Operating Procedure—SOP?

Safe Operating Procedure — SOP

Pronounced: “S-O-P,” or “safe operating procedure.”

An SOP is an approved, controlled and repeatable set of instructions for performing a defined operation safely and consistently. It tells the authorised user what conditions must exist, what to do in sequence, what result to confirm and when to stop.

Typical SOPs cover start-up, normal operation, sampling, cleaning, inspection, safe shutdown, isolation preparation, testing and return to service.

How it fits inside the SSOW

The SSOW coordinates the complete job—including teams, interfaces, permits, isolations, emergency arrangements and handback. The SOP standardises one defined operation within that system.

A pump-maintenance SSOW may refer to separate SOPs for shutdown, electrical isolation, line draining, gas testing and controlled recommissioning.

Important boundary: An SOP supports consistent work; it does not make an unsuitable task safe, replace risk assessment, authorise permit-controlled work or remove the need to stop when actual conditions differ.
SOP format fieldWhat it should containQuality question
Document identityTitle, equipment or process, number, version, owner, approver and review date.Can the user confirm this is the current approved procedure?
Purpose and scopeIntended result, authorised users, operating range, location and exclusions.Is it clear when the SOP applies—and when it does not?
Responsibilities and competenceOperator, supervisor, verifier, specialist and required authorisation.Does each person understand their role and limit of authority?
PrerequisitesPlant state, permits, isolations, tools, inspections, guards, ventilation and PPE.What must be true before Step 1?
Ordered stepsOne clear action per step, responsible role, location, setting, safe limit and expected result.Can the action and its successful outcome be observed?
Warnings and hold pointsCritical hazards, prohibited actions and mandatory verification before continuing.Are the most safety-critical instructions easy to find?
Operating limitsPressure, temperature, concentration, speed, time or other acceptance criteria.Does the user know when a result is outside the safe range?
Stop and escalationUnexpected state, failed check, alarm, leak, defect, safe shutdown and person to contact.Does the SOP prevent improvisation?
Completion and recordsFinal checks, housekeeping, status communication, log entries and retained evidence.Can another person confirm the operation ended safely?
Review and changeScheduled date plus triggers such as incident, modification, feedback or new evidence.Will the SOP remain aligned with the real process?
Develop the instruction from real work

How Is an SOP Developed?

Select each phase. The five phases contain ten connected stages: define → observe → assess → sequence → write → validate → approve → train → use → review.

Stages 1–2 — Define and observeDefine the operation, purpose, users, plant, boundaries and safe result. Then observe competent people performing the real task and ask where variation, delay, confusion, error or abnormal conditions occur.
1. Define. State the operation, intended result, equipment, users, range and exclusions.
2. Observe. Watch the job as actually performed and consult operators, maintainers and supervisors.
3. Assess. Link hazards, exposure, failure modes, limits and controls to the wider assessment and SSOW.
4. Sequence. Put preparation, operation, verification, shutdown and completion into a safe order.
5. Write. Use direct actions, familiar terms, diagrams where useful, measurable limits and clear stop rules.
6. Validate. Competent users walk through or trial the draft in controlled conditions and report ambiguity.
7. Approve. The authorised owner accepts the technical basis and controls the version.
8. Train. Explain the purpose, critical steps, limits, abnormal response and evidence of competence.
9. Use and supervise. Make the current SOP accessible, provide resources and verify real application.
10. Review. Learn from change, deviation, incident, user feedback, audit, monitoring and scheduled review.

Characteristics of a Strong SSOW and SOP

A good document must be technically correct and usable by the people who depend on it. These characteristics are evidence of quality—not decorative features.

CharacteristicEasy meaningWhy it mattersWarning sign
Risk-basedControls come from a suitable assessment and required standards.The system addresses credible harm rather than copying another job.The procedure mentions PPE but not the main energy or exposure source.
Task-specificIt matches the actual plant, people, place and conditions.Local differences can change the failure path.Wrong equipment number, location, substance or isolation point.
ProportionateDetail and formality reflect risk and complexity.Too little detail leaves gaps; excessive paperwork hides critical controls.A simple task has 50 pages, while a major intervention has one vague paragraph.
Clear and sequentialActions are unambiguous and in the correct order.Sequence can determine whether energy or exposure is controlled.“Make safe” without stating who, how or how safety is verified.
PracticalControls can be applied with available time, access, tools and resources.Impossible instructions encourage deviation and workarounds.The required test point cannot be reached safely.
ParticipativePeople who understand the work contribute to development and review.Worker knowledge reveals practical hazards and foreseeable shortcuts.Written remotely without observing or discussing the job.
Role-definedAuthority, responsibility and handover are explicit.Prevents gaps, overlaps and assumptions.Everyone believes someone else verified the isolation.
Competence-basedRequired knowledge, skill, experience and supervision are stated.The same instruction may not be safe for an inexperienced person.“Trained person” is stated but competence is never checked.
Human-centredIt considers workload, fatigue, usability, communication and predictable error.Controls must work in real human conditions.Critical information is buried, contradictory or unreadable.
Inclusive and accessibleUsers can find, read and understand it.Language, literacy, disability or unfamiliar terminology can affect safe use.Only one complex-language copy exists away from the workplace.
Abnormal-condition readyIt states when to stop, withdraw, isolate, escalate or use emergency arrangements.People must not improvise when normal conditions disappear.No instruction for a leak, failed test or unexpected pressure.
Controlled and currentOnly the approved version is available and changes are traceable.Obsolete instructions may conflict with modified plant or controls.Different versions are posted at the same workplace.
VerifiedCritical controls are checked before reliance.An assumed control may be absent, failed or incorrectly applied.The permit is signed without a field check.
Monitored and reviewedUse and effectiveness are checked over time and after triggers.Work, people, equipment and evidence change.Repeated deviations are normalised without investigation.

When a detailed written system is normally needed

  • Significant or high-risk work
  • Complex or non-routine tasks
  • Several people, teams or contractors
  • Critical sequence, isolation or verification
  • Permit-controlled work
  • Serious foreseeable abnormal conditions

When simplicity may be appropriate

Straightforward low-risk work may be controlled through concise instruction, training and normal supervision. Simplicity must come from low complexity—not from ignoring significant hazards. The arrangement still needs to be understood and effective.

How the Six Risk Strategies Shape the System of Work

These strategies are not six competing documents. They influence different decisions during development, authorisation and assurance.

StrategyWhen it is usedPump-maintenance applicationWhat the SSOW or SOP must show
AvoidanceThe exposure or activity can be removed or redesigned.Use remote condition monitoring to avoid unnecessary intrusive inspection.Why the hazardous step is no longer required and whether the alternative creates new risk.
ReductionNecessary work continues under stronger controls.Isolate, depressurise, drain, purge, verify, segregate and supervise.Control hierarchy, sequence, responsibilities, verification and residual risk.
TransferSpecialist competence, equipment or financial sharing is appropriate.Use a competent specialist for seal replacement or hazardous cleaning.Selection, information, coordination, interfaces, monitoring and retained duties.
AnalysisCauses, exposure, failure paths or uncertainty require deeper understanding.Analyse chemical residue, pressure, isolation effectiveness and previous failures.Evidence, assumptions and how findings affected the method.
EvaluationEvidence must be compared with criteria to decide adequacy and authorisation.Compare proposed precautions with legal, technical, manufacturer and site requirements.Acceptance criteria, decision authority and unresolved gaps.
ReviewTime, change, incident, feedback or failed control may affect validity.Revise after a leak, near miss, plant modification, contractor concern or recurring deviation.Review triggers, owner, evidence, revised version and communication.
Essential distinction: analysis, evaluation and review support the decision; they do not physically isolate energy or prevent exposure. Transfer may allocate some work or financial impact, but it does not make the hazard or all responsibilities disappear.

Tool: Build a Combined Risk-Strategy Route

Select one or more strategies and let the portal test whether the combination fits the scenario.
Interactive document workshop

Build an Educational Safe System of Work Draft

Adjust every field to your own workplace example. The output helps you understand the structure; it must be reviewed by competent people against the real workplace and applicable requirements before use.

Adjust the fields, then generate a structured learning draft.

Build an Educational Safe Operating Procedure Format

An SOP should tell the right person what to do, in what order, under which conditions, and when to stop. It should not ask the user to make undefined safety decisions during a critical step.

Complete the procedure fields, then generate the format.

Tool: Put the Pump-Maintenance Controls in a Defensible Order

Use the arrow buttons to move each step. The “correct” route is the approved teaching sequence for this example; a real installation may require a different, technically validated sequence.

    Move the steps, then check whether critical preparation, verification and recommissioning controls appear in the right order.
    Complete control-of-work concept

    Permit-to-Work—PTW: Meaning, Purpose, Issue, Use, Handback and Closure

    What is PTW?

    Pronounced: “P-T-W,” meaning permit-to-work.

    A PTW is a formal, time-limited system for authorising specified work at a defined place, on identified plant or equipment, by named or competent parties, subject to stated precautions and conditions.

    It communicates an agreement: this exact work may proceed, within this exact boundary, while these verified conditions remain true.

    What PTW does not mean

    A permit is not a risk assessment, an instruction to begin automatically, a substitute for isolation, a certificate that danger has disappeared, or a transfer of all responsibility to the worker.

    Issuing the paper alone does not make a job safe. The assessment, SSOW, competence, communication, physical controls, field checks, supervision and stop-work response must all function.

    Requested / draftThe job is being scoped and assessed. It is not permission to start.
    Issued / accepted / activeThe authorised issuer and performing party have verified, communicated and accepted the permit conditions.
    SuspendedWork has stopped and the permit is not active—for example after alarm, shift change, changed condition or control loss.
    Completed / handed back / closedThe work party declares completion; the area is checked, plant status is handed back, and the permit is formally cancelled or closed.

    Why is a permit-to-work used?

    A PTW creates disciplined communication where mistakes in plant identity, isolation, timing, coordination or handover could cause serious harm. It defines ownership, prevents incompatible simultaneous activities, records critical precautions, controls the period of work and manages the return to normal operation.

    Work categoryWhy formal control may be neededTypical linked controls or certificates
    Hot workFlame, arc, spark or heat may ignite flammable material or damage adjacent systems.Gas testing, area preparation, fire protection, fire watch and post-work monitoring.
    Confined-space or vessel entryAtmosphere, engulfment, restricted access, energy and rescue hazards can change rapidly.Isolation certificate, atmospheric test, ventilation, entry log, attendant and rescue plan.
    Line breaking / hazardous containmentOpening pipework or equipment can release pressure, temperature, toxic, corrosive or flammable material.Process isolation, drain/vent/purge, decontamination, test and line-break controls.
    Electrical or mechanical workContact, arc, unexpected start, gravity, pressure or stored energy may be fatal.Isolation plan, lockout/tagout, prove-dead or zero-energy verification and controlled reinstatement.
    Excavation / ground disturbanceUnderground services, collapse, water, contaminated ground and vehicle interaction may be present.Service drawings, detection and marking, trial holes, shoring, access and inspection.
    Other site-defined workWork at height, lifting, radiography, roof access, energised testing or unusual simultaneous work may need coordination.Site-specific permits, certificates, exclusion zones, specialist plans and interfaces.
    Do not assume permit categories are identical everywhere. The organisation’s control-of-work procedure and applicable legal/technical requirements determine which work needs a permit, who may issue or accept it, and which supporting certificates are required.

    Who is involved?

    Typical roleEasy meaningMain responsibility
    Area / operating authorityThe person controlling the plant or area.Confirms operating status, interfaces and whether the area can be released and later accepted back.
    Permit issuer / issuing authorityThe competent authorised person who issues the permit.Checks scope, assessment, precautions, isolations, conflicts, validity and field conditions before authorising.
    Performing authority / permit receiverThe person accepting the permit for the work party.Understands the permit, briefs the team, keeps within boundaries, monitors conditions and stops when conditions change.
    Isolating authorityThe person controlling required energy or process isolations.Applies, records, proves and later removes isolation under the approved process.
    Authorised gas testerA competent person approved to test atmosphere.Uses suitable equipment, records results and understands limits, frequency and conditions of testing.
    Work partyThe persons carrying out the authorised task.Attend the briefing, follow the SSOW/SOP and permit, protect controls and report change or uncertainty.
    Permit / SIMOPS coordinatorThe person who sees the whole work picture.Prevents conflicts between permits, operations, contractors, isolations and emergency arrangements.

    Terminology varies: a site may use different role names. The essential point is that authority, competence, accountability, communication and handover cannot be vague.

    The Full PTW Lifecycle—14 Phases

    Select a phase to see what must happen, why it matters and what evidence should exist.

    Phase 1 — Request and planDescribe the proposed job, reason, location, equipment, work order, timing, people and likely interactions. The request starts planning; it is not permission to begin.

    What should a complete permit form contain?

    Permit fieldInformation requiredControl purpose
    IdentificationPermit number/type, work order, exact plant/equipment tag, location and description.Prevents work on the wrong item or outside the authorised task.
    ValidityIssue date/time, start, expiry, shift and any rules for extension or revalidation.Stops an old permit being treated as continuing permission.
    Supporting documentsRisk assessment, SSOW, SOP/method, drawings, certificates and rescue/emergency plans.Connects authorisation to the technical control basis.
    Hazards and interfacesEnergy, substances, atmosphere, access, environment, nearby work and SIMOPS.Makes foreseeable interactions visible to both parties.
    Isolations and testsIsolation points/certificate, lock and tag references, drain/vent/purge, test type, result, time and tester.Provides traceable evidence of critical plant preparation.
    PrecautionsBarriers, ventilation, fire controls, access, tools, PPE, monitoring and prohibited actions.Defines conditions that must remain in place.
    Emergency and communicationAlarm, withdrawal, rescue, contact, stop-work rule, briefing and handover method.Supports response when normal assumptions fail.
    Authorisation and acceptanceIssuer and receiver names/signatures, date/time and declarations of understanding.Confirms that authority and shared understanding are explicit.
    Suspension / extension / handoverReason, safe state, new conditions, outgoing/incoming parties and revalidation.Prevents work continuing across a change without control.
    Completion and handbackWork complete/incomplete, people/tools cleared, guards restored, plant status, inspection and acceptance.Controls transfer back to operations and reinstatement.
    Cancellation and recordsClosure time, permit cancellation, linked documents closed, defects/actions and retained record.Ends authority clearly and preserves evidence for audit and learning.

    Interactive: Is the Permit Ready for Issue?

    Tick only items verified by evidence. A high total cannot compensate for a missing critical condition.

    Do not issue yet. Confirm each condition using the real worksite and approved control-of-work procedure.

    Interactive: Shift, Alarm and Change Decision

    Choose the event, then decide whether the permit remains controlled, is suspended, needs revalidation or can proceed to handback.

    Interactive: Build a Complete PTW Learning Form

    Complete every field. This produces a teaching draft—not a workplace permit or authorisation.

    Complete and check the form, then generate the structured learning draft.

    Interactive: Does the Work Need Formal Permit Control?

    A permit-to-work is a formal authorisation and communication system for defined work. Select the features that apply. Site procedures and applicable law make the final decision.

    Interactive: Stress-Test the Working Arrangement

    Rate each condition from 1 (weak) to 5 (strong). This is a learning diagnostic, not a risk-acceptance formula.

    4
    4
    3
    4
    5
    3

    PTW Knowledge Check

    1. What is the main purpose of PTW?
    2. What does issuing a permit automatically do?
    3. What must happen before issue?
    4. The job continues into a new shift. What is required?
    5. What if plant conditions or work scope change?
    6. What completes the PTW lifecycle?
    Answer all six questions, then check your understanding of the full permit lifecycle.

    Tools: SSOW Quality Diagnostic and Review Trigger

    Tool 1: Check the Quality of an Existing SSOW

    Select only what is genuinely present and effective in the system you are reviewing.

    Tool 2: What Should Trigger Review?

    Choose the trigger and current status, then decide whether work should stop, be restricted or continue under the approved system.

    Weak Wording Versus Defensible Wording

    Weak wordingWhy it failsStronger wording principle
    “Make the pump safe.”No person, isolation method, condition or verification is defined.Name the equipment, energy sources, authorised role, approved isolation method and verification requirement.
    “Wear proper PPE.”“Proper” is undefined and PPE may not control the main hazard.Select PPE from the assessment after applying higher-order controls; state type, limitation and checks.
    “Be careful when opening.”It transfers responsibility to behaviour without controlling stored pressure or residue.Prevent opening until depressurisation, drainage, safe condition and authorisation are verified.
    “Experienced workers only.”Experience is not defined or verified.State required authorisation, task knowledge, skill, experience and supervision.
    “In an emergency, act accordingly.”No stop, alarm, withdrawal or escalation route is given.Define credible abnormal conditions and the immediate response expected from each role.
    “Review annually.”Waits for a date even after a change, failure or near miss.Use both scheduled and event-triggered review.
    Assessment-ready learning support

    How to Explain 2.3 at Level 6

    A strong explanation normally contains

    1. Clear definitions of SSOW and SOP.
    2. The relationship with risk assessment and supporting documents.
    3. A connected development lifecycle.
    4. Characteristics explained with reasons—not only listed.
    5. Application of the six risk strategies.
    6. How risk assessment, SSOW, SOP, PTW, isolation and handback connect.
    7. The PTW lifecycle from request and field verification to suspension, handback and closure.
    8. A suitable workplace example.
    9. Limitations, implementation and review arrangements.

    Use this paragraph structure

    Point → Meaning → How developed/applied → Why it matters → Workplace example → Consequence if missing → Review.

    Write in your own professional words and relate the explanation to a genuine or realistic workplace. A list of headings alone does not fully satisfy “explain.”

    Practice question

    Explain how a safe system of work and supporting safe operating procedures should be developed for intrusive maintenance of a chemical-transfer pump. Your response should address consultation, risk strategies, control sequence, competence, human factors, permit interfaces, abnormal conditions, implementation and review.

    Section 2.3 knowledge check

    Can You Convert a Risk Decision Into Safe Work?

    1. What best describes an SSOW?
    2. What is the main purpose of an SOP?
    3. What should development begin with?
    4. Why is worker consultation important?
    5. What does a permit-to-work NOT do?
    6. What is a hold point?
    7. Intrusive pump maintenance must continue. What is normally the leading strategy?
    8. Which is a human-factor characteristic?
    9. Which is an event-triggered review?
    10. Which best demonstrates the command word “explain”?
    Answer all ten questions, then check your score and explanations.
    Unit 3 · Section 03

    Understand the Models of Loss Causation, Analysis of Loss Data and the Importance of Incident Investigation

    Section 03 asks us to move from what happened, to how the event developed, why the controls were vulnerable, and what the loss data can—and cannot—prove.

    Learning outcomes covered now: 3.1 Outline a range of loss-causation theories and techniques. 3.2 Justify the use of quantitative methods in analysing loss data.

    Start with the language

    What Do “Loss” and “Causation” Mean?

    Loss means an unwanted outcome that removes or damages something of value. It may include injury, ill health, death, environmental harm, property damage, production interruption, legal exposure, financial cost, lost information or damaged trust.

    Causation means the way conditions, decisions, actions, failures and interactions combine to produce an event or outcome.

    An incident normally has more than one relevant cause. The final action or failed component may be easy to see, but a Level 6 analysis asks what shaped that action, why the control was absent or ineffective, and what management-system conditions allowed the vulnerability to remain.

    The central learning rule

    A model organises thinking; it does not manufacture evidence.

    Investigators must still preserve the scene, gather reliable evidence, test competing explanations, consult involved people, identify controls and verify corrective action.

    Immediate cause

    The unsafe act, condition, energy transfer or equipment state directly connected to the unwanted event.

    Ask: What directly triggered or enabled the contact?

    Underlying cause

    The job, workplace or organisational factor that allowed the immediate condition or action to arise.

    Ask: What influenced the work and weakened control?

    Root cause

    A deeper management-system or organisational failing whose correction can prevent a wider class of recurrence.

    Ask: Why did the system create, accept or fail to detect the vulnerability?

    Interactive Section 03 Jargon Translator

    Select a term to hear how it is pronounced and understand why it matters.

    LossPronounced: “loss.” The harmful or unwanted consequence of an event. Loss can affect people, environment, assets, operations, finance, compliance or reputation.
    One event · several ways to understand it

    Our Master Case: Forklift Contact With a Process Line

    A forklift enters a congested transfer-area route, passes a damaged low barrier and contacts a valve manifold. A small chemical release occurs. The alarm operates, workers withdraw and the response team isolates the area. No model should be used to blame the driver or to assume a cause before evidence is gathered.

    Industrial loading area before a forklift incident, showing interacting conditions around the route and process pipe
    1. Before the eventLook for conditions: layout, visibility, barrier condition, workload, traffic interaction, supervision and competing demands.
    Controlled aftermath of a forklift contact with a process line while workers withdraw and responders secure the area
    2. The loss eventSeparate the hazardous contact from its consequences. Notice which prevention, detection, mitigation and emergency barriers worked or failed.
    Multidisciplinary investigation team reviewing incident evidence, controls and loss data
    3. Investigation and learningCombine physical evidence, people’s accounts, documents, data and models. Test explanations before deciding actions.
    Evidence sourceInitial findingWhat must still be tested?
    CCTV and scene measurementsPallets narrowed the route; forklift contacted the low pipe barrier.Actual speed, visibility, pedestrian interaction and why storage entered the route.
    Inspection recordsBarrier damage had been recorded twice but not permanently repaired.Risk classification, escalation, ownership, resources and closure verification.
    Driver and worker interviewsPeak dispatch created queuing and radio interruptions.Work-as-done, production pressure, route rules, competence and normal adaptations.
    Alarm and response logDetection and emergency isolation limited the release.Alarm timing, response reliability, exposure and opportunities to strengthen recovery.
    Six-month loss dataVehicle–route near-miss reports increased, especially during peak dispatch.Reporting quality, exposure hours, location clustering and whether risk really increased.
    Learning outcome 3.1

    Outline a Range of Loss-Causation Theories and Techniques

    Outline means present the principal features and show the basic structure, purpose and application of each theory or technique. A Level 6 response should also distinguish models, use a relevant example and recognise important limitations.

    3.1
    Historical pattern studies

    Accident and Incident Ratio Studies

    Ratio studies arrange recorded events by consequence level. They helped organisations recognise that low-consequence events and near misses can reveal control weaknesses before a major loss occurs.

    Teaching illustration of H. W. Heinrich and Frank E. Bird Jr. beside historical safety records
    AI-generated teaching illustration—not an archival photograph. Heinrich is shown on the left; Bird on the right.

    Who, when, why—and what changed?

    1931 — HeinrichH. W. Heinrich published Industrial Accident Prevention: A Scientific Approach and presented the historical 1:29:300 relationship.
    1959–1965 — BirdFrank E. Bird Jr.’s Lukens Steel study examined about 90,000 incidents and reported a 1:100:500 pattern.
    1969 studyBird and colleagues analysed 1,753,498 accident reports from 297 companies and produced the widely taught 1:10:30:600 ratio.
    Modern useThe triangle is now best treated as a prompt for reporting and investigation—not a law, prediction or promise that minor-event reduction will automatically prevent catastrophe.

    Why it was created: to show that the serious injury at the top is only part of the recorded experience. Lower-consequence events may provide more frequent opportunities to discover exposure and weak controls.

    How thinking changed: Bird widened the categories to include property damage and near misses. Later research showed that the shape changes with definitions, industry, severity threshold and reporting practice, and that fatal and non-fatal events can follow different causal pathways.

    Historical context: DNV tribute to Frank Bird. Critical evidence: Salminen, Saari, Saarela and Räsänen (1992) and Marshall, Hirmas and Singer (2018).

    Heinrich’s historical triangle

    Often presented as 1 major injury : 29 minor injuries : 300 no-injury accidents. Read the colon “:” as “to”: one major injury to 29 minor injuries to 300 no-injury accidents. It was derived from historical insurance and accident records and promoted attention to the larger body of less-serious events.

    Bird’s historical ratio

    Often presented as 1 serious or major injury : 10 minor injuries : 30 property-damage events : 600 near-miss incidents. It broadened attention to damage and no-loss events. The figures describe Bird’s historical dataset; they do not calculate the probability of the next accident.

    Why use a ratio study? Use it to test whether people report near misses, identify recurring event types and decide where investigation may prevent loss. Do not multiply 600 near misses and claim that one major injury must follow. A ratio summarises a dataset; it is not a countdown.
    Essential caution: These historical ratios are not universal laws, probability predictions or targets. Event definitions, industry, hazard type, reporting culture and data quality change the observed pattern. Major-accident hazards may develop without a large visible base of minor personal injuries.
    Potential benefitWhy it helpsLimitation to explain
    Encourages near-miss reportingWeak signals can reveal exposure and failing controls before serious harm.More reports can mean better trust and reporting—not necessarily worsening safety.
    Supports preventionRecurring lower-level events can direct inspection and improvement.Preventing minor slips does not automatically control a low-frequency catastrophic process event.
    Communicates scale simplyThe triangle is memorable and helps introduce proactive learning.Simplicity may hide different causal pathways and consequence mechanisms.
    Provides trend categoriesOrganisations can compare reporting levels and event types over time.Changed definitions, workforce hours or reporting systems can create a false trend.
    Challenges injury-only thinkingDamage and near misses can contain valuable control information.A ratio does not replace investigation, risk assessment or barrier assurance.

    Interactive Tool: See How Reporting Culture Changes the Triangle

    The “true opportunities for learning” remain constant in this teaching example. Adjust the percentage that gets reported. Notice how the visible triangle changes even when the underlying events do not.

    25%
    70%
    90%
    Adjust a reporting slider to reveal the observed ratio.
    From a chain to a system

    Bird’s Loss-Causation Model and Multi-Causality

    Bird’s expanded domino approach connects management control with basic causes, immediate causes, the incident and the final loss. Removing or strengthening an earlier “domino” can interrupt the sequence.

    Where it came from

    Bird developed accident-prevention thinking beyond the earlier person-centred domino sequence. The model placed lack of management control at the beginning and widened “injury” into loss, including harm to people, property, environment and production. Bird’s later work with George Germain was published in Practical Loss Control Leadership in 1985.

    How it developed

    The sequence was adapted using energy-exchange concepts associated with William Haddon. Read it left to right to explain how loss developed; work from right to left during investigation to ask what control should have interrupted each step.

    Lack of controlWeak standards, responsibilities, planning, monitoring or correction within the management system.
    Basic causesJob factors and personal factors that shape exposure and performance.
    Immediate causesObservable substandard acts and conditions directly connected with the event.
    Incident or contactThe transfer of energy, substance or force—or loss of control—that creates the event.
    LossInjury, illness, environmental harm, damage, interruption or another unwanted consequence.
    Applied to our case: overdue corrective maintenance and weak closure assurance → damaged barrier and congested route → forklift enters a vulnerable path → vehicle contacts the manifold → small release and operational loss.
    Level 6 caution: A domino picture can appear too linear. Real events often involve feedback, several simultaneous pathways, successful controls and changing conditions. Use the five stages as organising headings, then use multi-causality to avoid forcing the evidence into one neat chain.

    Multi-causality: More Than One Path Can Matter

    Multi-causality rejects the idea that one unsafe act is normally a complete explanation. Several conditions may combine, interact or increase one another’s effect. Causes can exist at task, equipment, environmental, individual, supervisory and organisational levels.

    Immediate examples

    • Forklift contacts low barrier.
    • Route is narrowed by pallet storage.
    • Valve manifold remains exposed to vehicle energy.

    Underlying examples

    • Peak traffic and pedestrian interaction were not reassessed.
    • Temporary storage became normal.
    • Damaged barrier repair was delayed.

    Root examples

    • Defect priority criteria ignored major-consequence potential.
    • No effective owner verified corrective-action closure.
    • Layout-change governance excluded operations and workforce evidence.

    Interactive Tool: Classify the Cause—Then Look Deeper

    Choose the best classification, then read the explanation.
    Reason’s model of organisational accident causation

    The Swiss Cheese Model

    James Reason described safety as several layers of defence, barrier and safeguard. Each layer can contain weaknesses—shown as “holes.” An adverse trajectory can pass through when weaknesses in different layers align.

    James Reason: from human error to organisational defences

    Who and when: British psychologist James Reason developed the model across the 1990s, beginning with ideas in Human Error (1990) and refining the familiar defence-layer image through the decade. Safety practitioner John Wreathall also influenced its development.

    Why: Reason wanted to move investigation beyond “the operator made an error.” The model asks how front-line actions combine with weaknesses created earlier by design, staffing, maintenance, supervision and organisational decisions.

    How it changed: diagrams and terminology evolved between 1990 and 2000. This matters: Swiss Cheese is a family of evolving explanations, not one frozen diagram. Later safety thinking also stresses that defences interact dynamically and that a simple line through holes must not replace evidence.

    Read the system approach in Reason (2000), Human error: models and management, and the historical critique in Larouzee and Le Coze (2020).

    Teaching illustration of James Reason and H. A. Watson with defence layers and logic diagrams
    AI-generated teaching illustration—not an archival photograph. Reason is shown on the left; Watson on the right.

    Active failure

    An action or omission close in time and place to the event, such as an incorrect control input or missed check. It may trigger the event, but it is rarely the whole explanation.

    Latent condition

    A deeper weakness created by design, staffing, maintenance, procurement, priorities, procedures, supervision or management decisions. It may remain hidden until combined with local conditions.

    Defence layerA safeguard intended to prevent the event, detect it or reduce its consequence.
    HoleA relevant weakness in that layer. A hole can move or change as conditions, resources and work practices change.
    TrajectoryThe path by which a hazard passes through aligned weaknesses and reaches a harmful outcome.
    Person approach versus system approach: “The driver made an error” closes the question too early. A system approach asks what task, equipment, environment and organisational conditions shaped the action, and which defences should have prevented or contained its effect.

    Interactive Barrier Alignment Simulator

    Mark a layer as failed or ineffective. The event pathway opens only when every selected defence in this simplified example has a relevant weakness.

    Five defence layers availableA single weakness need not produce loss if another independent and effective layer interrupts the pathway.

    Limitation: The cheese image is a communication model, not a detailed dynamic simulation. It can oversimplify interactions unless each layer, threat, dependency, owner and performance requirement is defined with evidence.

    Structured logic techniques

    Fault Tree Analysis and Event Tree Analysis

    Teaching illustration of James Reason and Bell Laboratories engineer H. A. Watson with early fault-tree logic
    AI-generated teaching illustration—not an archival photograph. H. A. Watson is represented on the right.

    How structured logic entered safety analysis

    Fault Tree Analysis: commonly traced to H. A. Watson at Bell Laboratories in 1962 during the Minuteman missile programme. It was created because complex systems could fail through combinations that a simple checklist might miss.

    Event Tree Analysis: developed as a forward, consequence-oriented partner to fault-tree reasoning. Probabilistic risk work such as the US Reactor Safety Study, WASH-1400 (1975), helped establish combined fault-tree and event-tree methods for complex high-hazard systems.

    What changed: the techniques expanded from defence and nuclear applications into aviation, process safety and other industries. Software can now evaluate very large trees, but the logic, data, dependencies and uncertainty still need competent human review.

    Authoritative background: NRC Fault Tree Handbook and NRC history of WASH-1400.

    Read this before any equation

    FTA and ETA Symbol Decoder + Logic-Gate Starter

    Symbols are a form of shorthand. They make a large analysis easier to read, but only after every symbol has been defined. Start by reading the symbol aloud, identify what it represents, check its unit or range, and then ask why it is present in the equation.

    The important difference between capital N and small n: capital N normally means the total number of valid opportunities, demands, tests or items in the denominator. Lowercase n normally means the number actually counted. Therefore, in P(A) = nA ÷ N, nA is the number of times event A occurred and N is the total number of relevant opportunities. Symbols are local conventions, however: in a “k-out-of-n” gate, lowercase n means the total number of components in that gate. Always read the symbol key written for the equation being used.

    First: know what every common mark means

    A, B, CPronounced: “event A, event B, event C.”

    Letters name defined events. A might mean “detector fails”; B might mean “isolation valve fails.” The letter has no meaning until the analyst defines it.

    NPronounced: “capital N.”

    The total number of relevant demands, tests, opportunities or observations. It is commonly the denominator. Example: N = 100 valid detector tests.

    nPronounced: “lowercase n.”

    A count from the defined set. Example: n = 3 recorded failures. A subscript makes the count more specific, such as nF for number of failures.

    nA, nFPronounced: “n sub A” and “n sub F.”

    The subscript is a label, not multiplication. nA counts event A; nF counts failures. Thus nF ÷ N means failures divided by all valid opportunities.

    P(A)Pronounced: “probability of A.”

    The probability that defined event A occurs on the stated demand, mission or period. Probability has no unit and must lie from 0 to 1 inclusive.

    p, qPronounced: “p” and “q.”

    p commonly represents success probability and q the complementary failure probability. When success and failure cover all possibilities, q = 1 − p and p + q = 1.

    piPronounced: “p sub i.”

    i is an index meaning “the particular item or branch being considered.” p1, p2 and p3 can represent three different input probabilities.

    f, fIPronounced: “frequency” and “f sub initiator.”

    f is an event frequency with a unit such as per year. fI is the initiating-event frequency used at the start of an event tree.

    T, tPronounced: “capital T” and “lowercase t.”

    T often represents total observed exposure time; t often represents the mission time being evaluated. The analyst must state the unit—hours, days or years.

    λPronounced: “lambda.”

    A failure rate, normally stated per unit time. A value of λ = 0.002 per hour means the assumed rate basis is 0.002 failures per operating hour—not a 0.2% certainty for every hour.

    ePronounced: “Euler’s number” or simply “e.”

    A mathematical constant approximately equal to 2.71828. In e−λt, it supports the exponential reliability model; it does not mean an event count.

    S / FPronounced: “success” and “failure.”

    ETA branches are often labelled S and F. These labels say whether the defined barrier performs its required function at that branch point.

    ∩ / ∪Pronounced: “intersection” and “union.”

    ∩ means AND—events occur together. ∪ means OR—at least one of the defined events occurs, including the possibility that both occur.

    P(B | A)Pronounced: “probability of B given A.”

    The vertical bar means “given that.” It is used when B’s probability is evaluated under the condition that A has already occurred.

    ΣPronounced: “capital sigma” or “sum.”

    Add the listed values. In ETA, Σfbranch means add the frequencies of the mutually exclusive terminal branches.

    Pronounced: “capital pi” or “product.”

    Multiply the listed values. ∏pi means p1 × p2 × p3 and so on; it is not the circle constant π.

    =, ×, ÷Pronounced: “equals,” “multiplied by,” and “divided by.”

    Equals states that both sides represent the same value. Multiplication combines required path factors. Division creates a proportion or rate from a count and denominator.

    1 − pPronounced: “one minus p.”

    The complement of p. If p is the probability of success, 1 − p is the probability of failure only when the two states are mutually exclusive and cover all defined possibilities.

    [ ] and ( )Pronounced: “brackets” and “parentheses.”

    Calculate the expression inside them first. They group terms and prevent the calculation from being performed in the wrong order.

    0 ≤ P ≤ 1Pronounced: “probability is greater than or equal to zero and less than or equal to one.”

    Zero means impossible within the defined model; one means certain within that model. Multiply a decimal by 100 to express it as a percentage.

    Second: what is a logic gate and why do we need it?

    A logic gate is a rule that explains how input events combine to create an output event. It is not a physical gate and it does not prove causation by itself. It converts a full English sentence into an exact visual instruction, helping an FTA remain consistent when many failure paths are connected.

    OR gate — “at least one is enough.”
    If detector failure by itself can cause the output, or valve failure by itself can cause it, connect them through OR. A normal inclusive OR also allows both events to occur.
    AND gate — “all stated inputs are required.”
    If the output occurs only when the detector fails and the automatic isolation also fails, connect them through AND. AND describes a required combination, not automatically a time sequence.
    k-out-of-n voting gate.
    The output occurs when at least k of n similar channels meet the stated condition. A 2-out-of-3 trip system needs any two of its three channels. Here n means total channels inside this gate.
    Advanced gates.
    NOT, exclusive-OR, inhibit and priority-AND can express special logic or sequence conditions. Use them only when their exact meaning is required and defined; most introductory FTAs begin with AND and OR.
    Event AEvent BA AND BA OR BPlain meaning
    0 — does not occur0 — does not occur00No input occurred.
    01 — occurs01Only B occurred; OR opens, AND does not.
    1 — occurs001Only A occurred; OR opens, AND does not.
    1111Both occurred; both rules are satisfied.
    FTA versus ETA: FTA mainly uses logic gates to reason backwards from a top event. ETA normally does not join causes through AND/OR gates; it moves forwards from one initiating event and divides into conditional success/failure branches. We multiply along one ETA path and add mutually exclusive end-state paths when they belong to the same outcome group.

    Interactive Logic-Rule Coach

    Choose the English statement you need to represent. The coach will identify the logic and explain how the symbols are used.

    Choose a statement to decode its logic.
    Five-pass formula-reading habit: (1) name the outcome on the left of “=”; (2) pronounce every symbol; (3) replace each symbol with its defined value and unit; (4) calculate brackets first, then multiplication/division, then addition/subtraction; and (5) explain what the answer means, its assumptions and what decision it supports.
    Begin before the tree

    Where Does the Probability Come From?

    FTA and ETA do not create probability simply because a box is drawn. Every input needs a defined event, population or equipment item, time or demand basis, data source and uncertainty. The first question is therefore not “Which formula shall I use?” It is “What exactly does this number describe?”

    P(A) = nA ÷ NPronounced: “probability of A equals number of A events divided by total relevant opportunities.”

    Use this empirical estimate when each opportunity is clearly defined. If a detector failed 3 of 100 valid proof-test demands, the observed failure-on-demand estimate is 3 ÷ 100 = 0.03, or 3%.

    q = 1 − pPronounced: “q equals one minus p.”

    If p is success probability, q is failure probability. A barrier that succeeds with probability 0.90 has a complementary failure probability of 1 − 0.90 = 0.10.

    P(B | A)Pronounced: “probability of B given A.”

    The vertical bar means given that. An ETA branch asks for the chance of the next success or failure after the initiating event and earlier branch conditions have occurred.

    F(t) = 1 − e−λtPronounced: “F of t equals one minus e to the power minus lambda t.”

    Under a justified constant-rate exponential model, this estimates the probability of at least one failure by mission time t. It is not suitable automatically for ageing, repair, dependence or changing conditions.

    f = n ÷ TPronounced: “frequency equals number of events divided by exposure time.”

    Frequency carries a unit such as events per year. Probability has no unit and stays between 0 and 1. A frequency of 0.5 per year must not automatically be called a 50% annual probability.

    P(A ∩ B)Pronounced: “probability of A intersection B.”

    This means A and B occur together. The general rule is P(A ∩ B) = P(A) × P(B | A). It becomes P(A) × P(B) only when independence is defensible.

    Possible sourceHow the value may be obtainedEssential quality question
    Operating or test dataDefined failures ÷ valid demands, or events ÷ exposure time.Are definitions, equipment, conditions and reporting sufficiently comparable?
    Reliability modelUse a justified distribution, failure rate and mission time—for example 1 − e−λt.Do the model assumptions fit ageing, repair, maintenance and operating conditions?
    Fault-tree calculationA barrier-failure probability used in an ETA may itself come from a detailed FTA.Were dependencies, shared utilities, human actions and common causes represented?
    Expert judgement or analogous dataElicit and document a defensible estimate or range when direct data are sparse.Are the experts, evidence, assumptions, bias controls and uncertainty traceable?

    Interactive Probability Source Calculator

    Choose one method. The portal will use only the fields required for that method and explain the units and assumptions.

    Choose a probability-source method.
    Level 6 interpretation: A calculated value is not automatically valid evidence. State the event definition, denominator, period, units, data provenance, assumptions, missing information and uncertainty. Use sensitivity analysis to see whether a reasonable change in an uncertain input changes the decision.

    Authoritative learning references: US NRC glossary definitions for fault trees and event trees, US NRC explanation of probabilistic risk assessment and NASA system-safety learning on logic models and probability.

    Fault Tree Analysis — FTA

    Pronounced: “F-T-A.” A Fault Tree Analysis is a structured, top-down method. Start with one precisely defined unwanted top event, then reason backwards to identify the equipment failures, human failures, external events and combinations capable of producing it.

    1. Define top event2. Set boundary3. Ask how4. Connect gates5. Check + quantify
    Think of FTA as asking: “The unwanted event is here. What must have happened beneath it?” A basic event is the lowest event taken forward for data or action. An intermediate event is created by lower events. A minimal cut set is a smallest combination of basic events sufficient to produce the top event.
    Top event: vehicle reaches hazardous manifold
    OR
    Failure path A: vehicle enters prohibited zone
    Failure path B: impact barrier fails when demanded
    Enter independent teaching probabilities.
    How the two gates differ: With an AND gate, every stated input is needed, so independent probabilities are multiplied. With an OR gate, any input can produce the output. Exact addition must remove the overlap where both occur; the complement method does this automatically. Choose the gate from the real causal logic—never choose it merely to obtain a preferred answer.
    Open the FTA equation and symbol guide
    P(A)Pronounced “P of A.” Probability that input event A occurs during the defined mission or period.
    P(B)Pronounced “P of B.” Probability that input event B occurs during the same basis.
    ×Pronounced “multiplied by.” For independent events, AND probability is P(A) × P(B).
    1 − P(A)Pronounced “one minus P of A.” The complement: probability that A does not occur.
    AND gateP(A ∩ B) = P(A) × P(B) when A and B are independent. The symbol ∩ is pronounced “intersection.”
    OR gateP(A ∪ B) = 1 − [1 − P(A)][1 − P(B)]. The symbol ∪ is pronounced “union.”
    General AND ruleP(A ∩ B) = P(A) × P(B | A). Use the conditional value when B’s chance changes because A occurred.
    General OR ruleP(A ∪ B) = P(A) + P(B) − P(A ∩ B). Subtract the intersection because simple addition counts “both” twice.
    n-input ANDP(top) = ∏pi for independent required inputs. ∏ is capital pi and means multiply the listed probabilities.
    n-input ORP(top) = 1 − ∏(1 − pi) for independent alternative inputs.
    Why calculate it? To identify combinations that contribute most to the top event, test design alternatives and prioritise reliability improvement. If inputs share a power supply, environment or maintenance error, independence is false and simple multiplication may understate risk.

    Real trees require validated logic, common-cause and dependency checks, suitable data, minimal cut sets and sensitivity analysis.

    Event Tree Analysis — ETA

    Pronounced: “E-T-A.” An Event Tree Analysis is a structured, forward or inductive method. Start with a defined initiating event, then move forwards through the conditional success or failure of safeguards, operator actions and recovery measures until every modelled path reaches an end state.

    1. Define initiator2. Order barriers3. Split branches4. Name end states5. Calculate + sum
    Think of ETA as asking: “The initiating event has occurred. What may happen next?” Every branch probability is conditional on arriving at that branch. The terminal paths should be mutually exclusive, and together they should cover the defined possibilities.
    Initiating event: contact causes leak
    Detection succeeds / fails
    Isolation succeeds / fails
    Enter an initiating frequency and conditional barrier probabilities.
    Two calculations occur: first multiply probabilities along a path to obtain the conditional branch probability. Then multiply that branch probability by the initiating-event frequency to obtain a branch frequency with units such as per year.
    Open the ETA equation and symbol guide
    fIPronounced “f sub initiator” or “initiating-event frequency.” Expected initiating events per stated period, normally one year.
    pDPronounced “p sub D.” Conditional probability that detection succeeds after the initiating event.
    qD = 1 − pDPronounced “q sub D equals one minus p sub D.” Probability that detection fails.
    pIPronounced “p sub I.” Conditional probability that isolation succeeds on the branch where it is demanded.
    qI = 1 − pIPronounced “q sub I equals one minus p sub I.” Conditional probability that isolation fails when demanded.
    PbranchPronounced “P sub branch.” Multiply the appropriate success or failure probability at every branch point on the selected path.
    fbranchPronounced “f sub branch.” fI multiplied by every conditional success or failure probability along that branch.
    ΣfbranchPronounced “sum of the branch frequencies.” Σ is capital sigma and means add all mutually exclusive terminal branches.
    Uncontrolled branch: funcontrolled = fI × (1 − pD) × (1 − pI). We calculate it to estimate how often a defined consequence pathway may occur and to see which barrier improvement changes the outcome most.

    ETA outputs are only as sound as the initiating frequency, branch definitions, conditional probabilities and dependency assumptions.

    FeatureFault treeEvent tree
    DirectionBackward from a defined unwanted top event.Forward from a defined initiating event.
    Main questionWhat combinations can cause this event?What outcomes can follow as barriers succeed or fail?
    LogicAND, OR and other gates combine causal events.Branches represent conditional success/failure pathways.
    Input basisBasic-event probabilities or frequencies on a consistent mission, demand or time basis.Initiating-event frequency plus conditional branch probabilities.
    Typical resultQualitative cut sets and, when quantified, top-event probability or frequency.Conditional path probabilities and frequencies for defined end states.
    Useful forComplex failure logic, critical combinations and design weaknesses.Escalation, mitigation, consequence pathways and outcome frequency.
    LimitationA poor top-event definition or missed dependency creates false confidence.Too many branches become difficult; dynamic interactions may be simplified.
    Connect causes, controls and consequences

    The Bowtie Model

    Bowtie combines fault-tree thinking on the left and event-tree thinking on the right. It places the top event—the moment control over the hazard is lost—in the centre.

    A collective industrial technique—not one person’s invention

    There is no single uncontested Bowtie inventor or creation date. The visual method evolved collectively from fault-tree and event-tree thinking and was progressively adopted in high-hazard industries. It became popular because a multidisciplinary team could see the complete threat–control–loss pathway on one page.

    Why it is used: to connect each threat to a preventive barrier, define the loss-of-control top event, connect consequences to mitigating barriers and make critical-control ownership visible.

    How it changed: modern practice adds escalation factors, escalation controls, barrier owners, performance standards and assurance evidence. A decorative “bowtie picture” without these elements is not enough for control management.

    See the UK Government Bowtie overview and the Office of Rail and Road’s health-risk application.

    Teaching illustration of a multidisciplinary industrial safety team collaboratively developing a Bowtie barrier map
    AI-generated teaching illustration—not an archival photograph. The team represents collective Bowtie development; no single inventor is implied.
    HazardThreats + preventionTop eventMitigation + recoveryConsequences
    ThreatCongested vehicle route
    Preventive barrierStorage exclusion and route inspection
    ThreatDamaged impact protection
    Preventive barrierEngineered barrier and verified repair
    TOP EVENT
    Vehicle contacts manifold and containment is lost
    Mitigating barrierLeak detection and alarm
    ConsequenceWorker chemical exposure
    Mitigating barrierEmergency isolation, exclusion and response
    ConsequenceEnvironmental release and interruption
    HazardA source with the potential to cause harm, such as hazardous chemical inventory or moving vehicles.
    ThreatA credible cause that could release control of the hazard. It belongs on the left of the top event.
    Top eventThe first moment control is lost—not the threat and not the final injury.
    Preventive barrierActs before the top event to stop a threat from causing loss of control.
    Mitigating barrierActs after the top event to reduce escalation or consequence.
    ConsequenceA credible harmful outcome to people, health, environment, assets or operations.

    Escalation factor

    A condition that can defeat or weaken a barrier—for example poor lighting reduces the reliability of a visual route check.

    Escalation-factor control

    A control that protects the main barrier—for example lighting inspection and emergency lighting support route visibility.

    Interactive Bowtie Draft Builder

    Complete the fields and build a barrier-focused teaching record.
    Understand behaviour in context

    Behavioural Root-Cause Analysis

    Behavioural RCA examines what a person did and the conditions that made the behaviour understandable or likely. It should not become a search for someone to blame.

    Teaching illustration of B. F. Skinner observing behaviour in a historical laboratory and a safety team analysing barriers
    AI-generated teaching illustration—not an archival photograph. B. F. Skinner is represented on the left; no single inventor of behavioural RCA is claimed.

    Where the ABC idea came from

    Who and when: there is no single creator of “behavioural root-cause analysis.” Its ABC structure grew from twentieth-century behavioural science. Psychologist B. F. Skinner’s work on operant conditioning helped explain how consequences can strengthen or weaken behaviour.

    Why safety practitioners use it: repeated behaviour often makes sense when the antecedents are clear and the immediate consequence is easier, faster or socially accepted. ABC analysis makes these influences visible.

    How it changed: modern human-factors practice rejects a behaviour-only explanation. ABC evidence should be combined with task design, competence, equipment, workload, supervision, leadership, culture and organisational controls. This prevents “the worker chose badly” from becoming the false root cause.

    Describe factsDefine behaviourFind antecedentsTest consequencesCorrect the system
    A — AntecedentWhat existed before the behaviour? Instructions, signals, layout, targets, training, workload, tools, norms and supervision.
    B — BehaviourWhat observable action occurred? Describe it neutrally and specifically—do not label a person “careless.”
    C — ConsequenceWhat followed the behaviour immediately or repeatedly? Time saved, praise, delay avoided, discomfort, correction—or no response at all.
    Easy rule: Behaviour is something observable that a person did or did not do. “Careless,” “complacent” and “poor attitude” are interpretations, not observable behaviours. Ask what evidence supports them—and what system conditions made the action likely.
    Example: “The driver used the narrowed route” is the behaviour. Antecedents may include pallets in the approved lane, dispatch pressure and an accepted local workaround. The immediate consequence may have been faster completion on many previous occasions. The corrective action must address the system that reinforced the behaviour—not only tell the driver to be careful.

    Interactive ABC and System-Factor Builder

    Describe behaviour neutrally, then connect it to antecedents, consequences and system conditions.

    Boundary: A fair system approach does not mean that every action is acceptable. Deliberate reckless conduct may require a just and proportionate response, but the investigation must still examine supervision, controls and organisational context.

    Tool: Which Causation Technique Should Lead?

    Complex investigations often combine techniques. Select the main question to see a defensible starting point.

    Choose the investigation question and system complexity.
    Learning outcome 3.2

    Justify the Use of Quantitative Methods in Analysing Loss Data

    Quantitative analysis converts valid counts, exposure and consequences into comparable measures. The Level 6 requirement is not only to calculate; it is to justify why a numerical method is suitable, explain assumptions and limitations, interpret the result, and connect it to investigation and control.

    3.2
    Count correctly before calculating

    From Raw Loss Records to Decision-Useful Evidence

    CountHow many defined events occurred?
    RateHow many events occurred for a stated amount of exposure?
    SeverityHow much consequence, such as lost time, followed?
    DistributionWhere, when, to whom and through which mechanism did loss occur?

    Why a count can mislead

    Site A records 8 cases and Site B records 5. Site A may appear worse, but if it worked four times as many hours, its exposure-normalised rate may be lower.

    Why a rate can also mislead

    A single event can move a small workforce’s rate sharply. A low rate may also reflect under-reporting, changed classification, outsourced exposure or chance—not strong control.

    Symbols Used in the Calculations

    NPronounced: “capital N.” Number of cases matching the chosen definition.
    HPronounced: “capital H.” Total exposure hours actually worked for the population and period.
    WPronounced: “capital W.” Average workers or full-time-equivalent workers in the defined population.
    CPronounced: “capital C.” Existing cases: new and continuing cases present during the prevalence reference period.
    PPronounced: “capital P.” Population at risk of the defined ill-health condition.
    KPronounced: “capital K,” or “base multiplier.” The agreed standardising base, such as 1,000 workers, 100,000 people or 1,000,000 hours.
    FRPronounced: “F-R,” or “frequency rate.” Defined new events per stated number of exposure hours.
    IRPronounced: “I-R,” or “incidence rate.” Defined new cases per stated worker or population base.
    PrevPronounced: “prevalence.” Existing new and continuing ill-health cases per stated population base.
    DPronounced: “capital D.” Number of lost, restricted or otherwise defined consequence days.
    SRPronounced: “S-R,” or “severity rate.” SR = (D × K) ÷ H under the selected convention.
    Pronounced: “D bar.” Average defined days per case: D̄ = D ÷ N.
    pPronounced: “lower-case p.” Probability of a defined event or branch success; its value lies from 0 to 1.
    q = 1 − pPronounced: “q equals one minus p.” Complementary probability that the defined event does not occur.
    ΣPronounced: “capital sigma,” or “sum.” Add all listed values that follow the symbol.
    Pronounced: “x bar.” Arithmetic mean: add the observations and divide by their number.
    sPronounced: “lower-case s.” Sample standard deviation: typical spread of observations around the sample mean.
    SEPronounced: “S-E,” or “standard error.” Estimated sampling variability of a statistic; in the teaching tool, SE = s ÷ √n.
    CI₉₅Pronounced: “95 percent confidence interval.” An interval generated by a method expected to cover the population value in about 95% of repeated samples.
    Δ%Pronounced: “delta percent.” Percentage change between comparable periods.
    MA₃Pronounced: “M-A sub three.” Three-period moving average used to reduce short-term fluctuation.
    Definitions before arithmetic: Decide what counts as a case, which workers and hours are included, how contractors are treated, which period applies, and which base multiplier is required. Never compare rates that use different definitions without adjustment.
    Live calculation laboratory

    Calculate and Interpret Loss Rates

    Flexible Accident / Incident Frequency-Rate Calculator

    Why calculate it? A count alone ignores how long people were exposed. Frequency rate converts defined events into a common hours-worked base, supporting more defensible comparison.

    FR = (N × Kh) ÷ H
    The calculated rate will appear here.
    How to say and use every sign
    FR“F-R” or “frequency rate”: the result.
    =“equals”: the left side has the same calculated value as the right side.
    ( )“brackets”: complete the multiplication inside first.
    N“capital N”: defined new events in the period.
    דmultiplied by”: multiply N by the hours base Kh.
    ÷“divided by”: divide by exposure hours H.

    The 200,000-hour convention is used in US OSHA/BLS incidence rates. A one-million-hour base is common for some international frequency measures. Confirm the required definition before comparison.

    Severity and Average-Consequence Calculator

    Why calculate it? Frequency tells how often events occur; severity shows the amount of defined consequence relative to exposure. Average days per case describes the typical recorded consequence per relevant case.

    SR = (D × Kh) ÷ H   |   D̄ = D ÷ N
    Severity results will appear here.
    How to say and use every sign
    SR“S-R” or “severity rate”: defined lost days per chosen hours base.
    “D bar”: average defined days per relevant case.
    |“and separately”: the vertical line separates two related calculations; it is not division.

    “Lost day” and severity conventions differ. Record calendar/workday rules, caps, fatalities and restricted work consistently.

    Accident Incidence-Rate Calculator

    Why calculate it? When reliable hours are unavailable or the required convention uses workers, incidence expresses new defined accidents or cases per standard number of workers.

    IR = (N × Kw) ÷ W
    The worker-based incidence rate will appear here.

    Do not confuse incidence with frequency: incidence uses workers or people; frequency uses hours worked. “New” means the case started within the reference period.

    Ill-Health Prevalence-Rate Calculator

    Why calculate it? Prevalence estimates the burden of a condition now—both new and continuing cases—so an organisation can plan health controls, surveillance, support and resources.

    Prev = (C × Kp) ÷ P
    The prevalence rate will appear here.

    Prevalence is not incidence: prevalence counts all qualifying existing cases; ill-health incidence counts only new cases. Long-latency occupational disease may reflect exposures from many years earlier.

    Interactive Tool: Which Rate Should I Use?

    Choose the decision question and available denominator.

    Formula conventions vary. Examples follow official explanations from the US Bureau of Labor Statistics, International Labour Organization, and UK HSE ill-health statistics guidance.

    Tool: Compare Two Sites Fairly

    Counts answer “how many?” Rates help answer “how many for the amount of exposure?” Use the same event definition, period and multiplier.

    Site A

    Site B

    The exposure-normalised comparison will appear here.
    Find direction and concentration

    Trend, Moving Average and Pareto Analysis

    Six-Period Trend Explorer

    Enter six comparable monthly rates separated by commas.

    Δ% = [(new − old) ÷ old] × 100   |   MA₃ = (xt + xt−1 + xt−2) ÷ 3
    Enter six rates and analyse the direction.

    Why calculate it? Δ (“delta”) shows relative endpoint change. MA₃ smooths short-term fluctuation by averaging the current and previous two periods. Neither proves the reason for change.

    Pareto Priority Explorer

    Pareto analysis orders categories from largest to smallest so the team can see where recorded loss is concentrated.

    Category % = (nc ÷ N) × 100   |   Cumulative % = Σ category %
    Adjust a count to update priority and cumulative contribution.
    What every symbol means
    nc“n sub c”: count in one category.
    N“capital N”: total count across all displayed categories.
    Σ“capital sigma”: add percentages as you move down the ranked categories.
    Interpretation rule: A trend or Pareto chart tells you where to look. It does not prove causation. Investigate exposure, reporting practice, severity potential, control performance and the narratives behind the categories.
    Interpret loss data visually

    Histogram, Pie Chart and Line Graph Laboratory

    Different charts answer different questions. A chart must match the data structure; attractive graphics do not repair weak definitions or incomplete records.

    Interactive Chart Explorer

    Choose a chart type and update the values.

    Choose by question—not appearance

    HistogramShows the distribution of numerical observations grouped into ordered, usually equal-width intervals called bins. Bars touch because the scale is continuous. Example: days lost per case.
    Pie chartShows how mutually exclusive categories make up one whole at one point or period. Use few categories and show the denominator. Example: recorded event mechanisms this year.
    Line graphShows an ordered series, normally time. It helps reveal direction, cycles and unusual change. Keep definitions, exposure and intervals comparable.
    Important distinction: a category chart with separate bars is a bar chart, not a histogram. A histogram needs a numerical scale split into intervals. The tool begins with lost-day intervals so the histogram use is valid.
    Interpretation sentence: “The chart shows ___, across ___, using ___ as the denominator. The main pattern is ___. However, ___ may affect validity; therefore we should investigate ___ before deciding ___.”
    Do not confuse precision with truth

    Statistical Variability, Distributions, Sampling and Data Validity

    Statistical variability means observations and sample results naturally differ. Validity asks whether the measure actually represents the decision question. A precise calculation from biased or incorrectly classified data is still misleading.

    Descriptive Statistics and Approximate CI Explorer

    Enter numerical observations such as lost days per case. This tool describes the sample; it does not certify a population model.

    Enter at least two observations.
    Open every equation and symbol
    x̄ = Σx ÷ n“x bar equals sum of x divided by n.” The arithmetic mean.
    MedianThe middle ordered value; for an even n, average the two middle values. It is less affected by extreme values than the mean.
    Range = max − minLargest observation minus smallest observation.
    s = √[Σ(x − x̄)² ÷ (n − 1)]Sample standard deviation. √ is “square root”; ² is “squared.”
    SE = s ÷ √nEstimated standard error of the sample mean.
    Approx. CI₉₅ = x̄ ± 1.96 × SE“x bar plus or minus 1.96 times standard error.” A teaching normal approximation, not suitable for every dataset.

    What a distribution can reveal

    SymmetricalValues spread similarly around the centre; mean and median may be close.
    Right-skewedMany small values with a few large losses. The mean may exceed the median—common with days lost or costs.
    Clusters or multiple peaksMay indicate different workgroups, tasks, mechanisms or reporting systems that should not be pooled without examination.

    Representative sample: a sample should reflect the population relevant to the question. Every important subgroup needs a fair chance of inclusion. A large convenience sample can still be biased.

    Sampling a population: define the target population, build a suitable sampling frame, select participants or records using a defensible method, record non-response and compare the achieved sample with the population.

    UK HSE explains sampling error and 95% confidence intervals in its Labour Force Survey guidance. See also the ONS guide to uncertainty.

    Interactive Sample and Validity Check

    The teaching validity check will appear here.

    This is a structured warning tool, not a formal sample-size calculation or statistical certification.

    Common data errors to test

    ErrorEasy meaningExample
    CoverageSome of the population cannot appear in the data.Night shift and contractors are missing.
    SelectionThe inclusion method favours certain people or records.Only volunteers answer a wellbeing survey.
    Non-responseSelected participants do not respond, and may differ from responders.Workers with symptoms do not trust confidentiality.
    MeasurementThe question, instrument or observer produces inaccurate values.Different clinics use different symptom questions.
    ClassificationThe same case is placed in different categories.Restricted work is recorded as first aid at one site.
    Duplicate / missingA case is counted twice or not counted.One injury exists in two systems without a unique ID.
    DenominatorThe exposure base excludes relevant work.Contractor cases included, contractor hours excluded.
    ProcessingEntry, coding, formula or transfer is wrong.Hours are entered as 40,000 instead of 400,000.
    Time lagThe measured harm appears long after exposure.Current respiratory disease reflects earlier dust exposure.
    Reporting cultureTrust and rules change what becomes visible.A reporting campaign increases near-miss counts.
    The command word is “justify”

    Why Use Quantitative Methods—and Why Not Use Them Alone?

    Reason for using numbersDecision valueCondition or limitation
    Normalise exposureRates allow more defensible comparison across differently sized populations or periods.Definitions, hours, workforce scope and base multiplier must match.
    Detect changeTime-series analysis can show sustained deterioration, improvement, seasonality or unusual variation.Short runs, rare events and changed reporting can produce unstable signals.
    Measure disease burdenIll-health prevalence estimates all qualifying existing cases and supports surveillance, control and resource planning.Long latency, diagnostic access, worker turnover and healthy-worker effects can disconnect current prevalence from current exposure.
    Prioritise investigationPareto and distribution analysis identify categories, locations or activities contributing most recorded loss.Frequency must be considered with credible severity and major-hazard potential.
    Express uncertaintySpread, standard error and confidence intervals show that a sample estimate is not an exact population truth.The calculation depends on a defensible sample, measurement quality and an appropriate statistical model.
    Evaluate interventionBefore/after measures can test whether performance changed following control.Control for exposure, operational change, reporting, regression to the mean and other influences.
    Communicate performanceDefined indicators support dashboards, accountability and resource decisions.Targets can encourage under-reporting or classification manipulation if poorly designed.
    Model event pathwaysFTA, ETA and quantitative risk methods estimate the contribution of failure combinations and barriers.Models contain assumptions, dependencies and uncertainty; precision is not certainty.

    Small numbers

    One event can double a rate in a small workforce. Use longer periods, confidence intervals or pooled evidence where appropriate.

    Under-reporting

    A “good” rate may reflect low trust or restricted definitions. Triangulate with audits, surveys, health data and workforce evidence.

    Lagging-only bias

    Injury data describes realised outcomes. Add leading evidence about exposure, critical controls, defects and corrective-action quality.

    Severity randomness

    Similar events can produce very different harm. Do not assume low historical injury means low potential consequence.

    Changing denominator

    Overtime, contractors, shutdowns and outsourcing change exposure. Record the population and hours consistently.

    Metric fixation

    Managing the number rather than the risk can distort behaviour. Indicators must serve learning and control—not replace them.

    Interactive Level 6 Justification Builder

    Build a connected justification: decision → method → evidence → value → limitation → complementary evidence → action.
    Assessment-ready learning support

    How to Answer 3.1 and 3.2 at Level 6

    3.1 — Outline

    For each theory or technique: state its name and origin/context, describe its principal structure, explain its direction or logic, apply it briefly to an incident, and state one useful feature and limitation.

    Suggested structure: Name → main idea → components → application → usefulness → limitation.

    3.2 — Justify

    Identify the decision, explain why the selected quantitative method fits, show the data and formula, interpret the result, discuss reliability and limitations, combine it with qualitative evidence, and state the resulting action and review.

    Suggested structure: Decision → method → evidence → calculation → interpretation → limitation → complementary evidence → action.

    Integrated practice task

    Using the forklift–process-line case, outline Bird’s model, multi-causality, Reason’s Swiss Cheese model, FTA, ETA, Bowtie and behavioural RCA. Then justify how incident rates, severity measures, trend and Pareto analysis could support—but not replace—the investigation.

    Section 03 knowledge check

    Can You Connect Models, Data and Investigation?

    1. What is the best use of an incident triangle?
    2. What is the first domino in Bird’s expanded loss-causation sequence?
    3. What does multi-causality emphasise?
    4. What do the holes in Swiss Cheese represent?
    5. Which technique starts with a top event and reasons backwards?
    6. What does ETA mainly explore?
    7. What sits at the centre of a Bowtie?
    8. Which is a sound behavioural RCA approach?
    9. Why divide cases by hours worked?
    10. What is the strongest Level 6 conclusion?
    Answer all ten questions, then check your score and explanations.

    Use Authoritative Information

    This independent DB HSE explanation supports learning. Workplace decisions must use current applicable legislation, exposure limits, approved organisational criteria and competent specialist advice where required.

    Independent resource status: Sections 1.6.3–1.6.5, 2.1–2.3 and 3.1–3.2 are prepared solely by Debjyoti Biswas for DB HSE International to support OTHM Level 6, Unit 3 teaching. They are not produced by OTHM or Ofqual and do not replace the official specification, assessment guidance or applicable workplace requirements.