Prepared solely by Debjyoti Biswas, FIIRSM, CertIOSH · DB HSE International Independent teaching resource · Not produced by OTHM Level 6 · Unit 3 · Sections 01–04 + Assignment Brief
DB HSE Interactive Learning Portal

Welcome to Unit 3 · Sections 01–04 + Assignment Brief

Welcome to Unit 3 of the OTHM Level 6 International Diploma in Occupational Health and Safety. This unit develops complete risk-and-incident management judgement—from understanding system reliability and selecting effective controls to learning from loss events, maintaining defensible records and evaluating whether the organisation truly improves. The final assignment stage then shows learners how to convert that knowledge into all 17 required pieces of assessment evidence.

othmqualifications
The essence of Unit 3

See Risk Clearly · Control It Systematically · Learn from Every Event · Prove Improvement

The four sections form one connected professional journey. Learners move from understanding how systems fail, through choosing and designing controls, to investigating what happened and assuring that the complete incident-management process is lawful, reliable and effective.

Identify Assess Control Investigate Record Evaluate Improve
Section 01 · Understand and measureUse hazard information, reliability calculations and failure tracing to turn uncertainty into defensible evidence.
Section 02 · Select and design controlsEvaluate risk-control strategies and convert sound decisions into practical safe systems of work and SOPs.
Section 03 · Learn from lossApply causation models, loss data, reporting judgement and evidence-led investigation to understand why events occur.
Section 04 · Manage and assureCoordinate the incident lifecycle, govern responsibilities, maintain lawful records and evaluate process effectiveness.

Why it matters: Unit 3 connects prevention before an event with disciplined learning after it. A Level 6 practitioner must not only identify what went wrong, but decide what the organisation should change, demonstrate that the change was implemented and verify that it reduced risk.

01 AVAILABLE

Systems Reliability and Failure Tracing

  • 1.6.3Common reliability calculations
  • 1.6.4Analyse system performance
  • 1.6.5Use evidence to trace failures
  • ToolsSystem explainer, nine calculators, analysis and tracing activities
Open Section 01
02 2.1–2.3 READY

Strategies and Techniques of Risk Control

  • 2.1Evaluate common risk-management strategies
  • 2.2Justify when each strategy should be used
  • 2.3Develop SSOWs and SOPs
  • ToolsAssessment formats, risk matrix, strategy planner, SSOW and SOP builders
Open Section 02
03 3.1–3.4 READY

Loss Causation, Reporting and Investigation

Guest practitioner overview · Dr. Sarel Du Plessis
  • 3.1Completed foundation: loss-causation theories and techniques
  • 3.2Completed foundation: quantitative loss-data analysis
  • 3.3Who needs to know, when and why? Assess loss-event reporting
  • 3.4How does evidence become prevention? Explain incident investigation
  • ModelsBird, multi-causality, Swiss Cheese, FTA, ETA, Bowtie and behavioural RCA
  • ToolsCause, data, reporting, evidence, investigation and action-quality laboratories
Open Section 03
04 4.1–4.4 READY

Managing Incidents Through Organisational Systems

  • 4.1What must happen from the first alarm to verified closure?
  • 4.2Which policies make identification, investigation, reporting and recording consistent?
  • 4.3How are incident records kept lawful, trustworthy, secure and retrievable?
  • 4.4How do we evaluate whether the complete incident-management process really works?
  • ToolsResponse, policy, record-integrity, legal-profile, assurance, metrics and evaluation laboratories
Open Section 04
A FINAL STAGE READY

Unit 3 Assignment Preparation Masterclass

Official brief mapped · 4 tasks · 17 assessment criteria
  • Task 1Essay · AC 1.1–1.6 · Risk identification, assessment and reliability
  • Task 2Report · AC 2.1–2.3 · Risk control, SSOW and SOP
  • Task 3Essay · AC 3.1–3.4 · Loss causation, data, reporting and investigation
  • Task 4Report · AC 4.1–4.4 · Incident systems, records and evaluation
  • ToolsExact AC map, live word budgets, organisation planner and command-verb clinic
  • Assurance17-criterion coverage audit, academic-integrity rules and submission controls
Open Assignment Masterclass
OTHM Qualifications logo
DB HSE International logo
Unit 3 · Section 01

Systems Reliability and Failure Tracing

Welcome to Section 01. Begin with system reliability, complete all nine calculations in 1.6.3, analyse what the results mean in 1.6.4, and use the evidence to trace failures in 1.6.5.

Resource status: This portal was independently prepared solely by Debjyoti Biswas for DB HSE International. It is not an OTHM-produced learning resource. It supports DB HSE teaching for OTHM Level 6, Unit 3, Section 01.
Section 01 1.6.3–1.6.5 9 Live Calculators Analysis & Failure Tracing
Debjyoti Biswas, FIIRSM and CertIOSH, approved OTHM tutor and Fellow Member of IIRSM
Debjyoti Biswas, FIIRSM, CertIOSH
FIIRSM · CertIOSH · CEO & Founder, DB HSE International
Approved OTHM Tutor Fellow Member of IIRSM CertIOSH Member

Meet Your Tutor

Debjyoti Biswas, FIIRSM, CertIOSH is an Approved Tutor of the OTHM Level 6 International Diploma in Occupational Health and Safety Qualification, a Fellow Member of IIRSM and a CertIOSH Member. His teaching approach connects Level 6 knowledge with practical workplace application, evidence-based judgement, leadership and professional development.

Approved OTHM Level 6 Tutor Fellow Member of IIRSM CertIOSH Member International HSE Trainer DB HSE International
Learner-support response: “Thank you for raising your concern. Please allow me a moment; I will get back to you.” Questions requiring tutor judgement should be referred to Debjyoti Biswas rather than treated as an official OTHM response.
Start here · Let us talk about reliability first

What Is Reliability?

Let us begin with one easy question:

If you ask a machine, alarm, pump or complete safety system to do its job, how confident are you that it will perform correctly without failing?

That confidence—expressed as a probability—is reliability.

Required function
The exact job the item must perform. A fire pump must deliver water; a gas detector must detect gas.
Stated conditions
The real environment and duty: temperature, pressure, dust, workload, operating mode and maintenance condition.
Specified time
The period for which successful operation is required—for example, a 100-hour mission or an emergency demand.
Without failure
The item completes the required function without stopping, leaking, giving a false signal or becoming unavailable.
Probability
The answer is normally written between 0 and 1, or as a percentage from 0% to 100%.
Reliability = Probability that an item performs its required function, under stated conditions, for a specified time, without failure

Now apply that idea to a complete system.

A system is a group of connected parts—equipment, power, controls, information, people and procedures—working together. System reliability asks whether the complete required function succeeds.

Industrial engineers inspecting a gas detector, control panel, alarm beacon and shutdown valve as one safety system
One safety function may depend on detection, decision, warning and action. Failure of a required link can prevent the complete system from protecting people.

Explore One Safety Function—One Step at a Time

Imagine a gas-release protection system. Select each part to understand its job.

1. Detect the hazardThe gas detector must sense the hazardous concentration accurately and send a valid signal. A blocked sensor, loss of power or overdue calibration can break this first link.

Component reliability

Asks whether one item—such as a sensor, bearing, seal or pump—will perform successfully.

System reliability

Asks whether the complete required function succeeds. Reliable individual parts do not automatically guarantee a reliable system because connections, power, controls, procedures and common causes also matter.

Not only quality

A well-made item may still be unreliable if it is used outside its design conditions or maintained poorly.

Not the same as availability

An item can fail frequently yet remain highly available when every repair is very fast.

Not proof of safety

Reliability supports risk decisions, but safety also depends on consequence, design, inspection, people, procedures and independent protection.

DB HSE learning note: Independently prepared solely by Debjyoti Biswas for Level 6, Unit 3, Section 01. This explanation supports learning and is not an official OTHM interpretation.
Section 1.6.3

Common Reliability Calculations Used in Systems Reliability and Failure Tracing

Use operating, failure and repair data to create evidence. This section explains every symbol, formula, step, answer, limitation and workplace use for all nine calculations.

Independent DB HSE teaching content: prepared solely by Debjyoti Biswas for Level 6 · Unit 3 · Section 01.

1. CalculateUse valid data and correct units. 2. InterpretExplain what the answer means. 3. ApplyUse the evidence to improve control.

One scene · Nine techniques

The same master workplace scene will stay with you throughout Section 1.6.3. We will use the repairable pump, repair records, replaceable sensors, alarm chain and backup pumps to explain all nine techniques without changing the story.

Our master example · One workplace scene

The DB HSE Emergency-Protection System

Yes—we are going to use one connected workplace scene to explain all nine calculation techniques.

Imagine that you and I have entered this chemical warehouse. The protection system contains gas detection, a control panel, an alarm, an isolation valve, replaceable sensor cartridges and two fire-water pumps. We will not keep changing the story. We will simply ask nine different reliability questions about the relevant parts of this one system.

Master reliability example showing gas detection, alarm and control equipment, replaceable sensor cartridges, engineers and two parallel fire-water pumps in one industrial protection system
Our single master scene: one facility, one connected protection system and nine different reliability questions.
Part of our sceneInformation collectedWhy we need it
Repairable protection pump5,000 operating hours; 10 failures; 40 total repair hoursMTBF, MTTR, λ, availability, R(t) and F(t)
Required missionThe pump must operate for the next 100 hoursReliability and failure probability over time
Five replaceable sensor cartridgesLives: 1,800; 2,100; 1,900; 2,200; 2,000 hoursMTTF for non-repairable items
Series alarm chainDetector 98%; control panel 97%; alarm 99%Reliability when every required link must work
Two backup fire-water pumpsEach pump has 90% reliability for the stated demandParallel reliability when at least one independent pump must work

Watch how the same scene gives us nine different questions

1. MTBF

How long does the repairable pump operate, on average, between failures?

2. MTTR

How long does one pump repair take, on average?

3. λ

How frequently does the pump fail per operating hour?

4. Availability

What percentage of time is the pump ready for use?

5. R(t)

What is the chance that it completes 100 hours without failure?

6. F(t)

What is the chance that it fails during those 100 hours?

7. MTTF

What is the average life of a replaceable sensor cartridge?

8. Rs

Will the detector, control panel and alarm all work together?

9. Rp

Will at least one of the two independent fire pumps work?

Important: One scene does not mean one formula fits everything. We select the formula that matches the question, the equipment type and the data available.
Begin with the workplace problem

Why Do HSE Professionals Need These Calculations?

Statements such as “the machine fails frequently” or “the alarm is usually reliable” are opinions. Reliability calculations convert operating, failure and repair records into evidence that can be compared, investigated and improved.

Measure

Determine how often equipment fails, how long it runs and how long repairs take.

Trace

Use trends and component records to locate recurring failures and investigate their causes.

Improve

Verify whether maintenance, redesign, training, spare parts or redundancy improved performance.

Questions the calculations answer

  • How frequently does the equipment fail?
  • How long does it operate before failing?
  • How long is normally needed to repair it?
  • What proportion of time is it available?
  • What is the chance of completing a task without failure?
  • Which component repeats?
  • Did maintenance improve performance?

Safety-critical examples

Fire pumps, emergency alarms, pressure-relief systems, local exhaust ventilation, gas detectors, lifting equipment, emergency generators, rescue equipment and emergency shutdown systems.

The required reliability should reflect the consequence of failure and the availability of independent protection.

Interactive symbol guide

Understand Every Symbol Before Calculating

Select a symbol to see what it means and where it is used.

Total operating time — T

T represents the total time for which equipment actually operated. It must use the same time unit throughout the calculation, such as hours, days or cycles.

Live calculation laboratory

Explore Each Reliability Calculation

Choose a calculation, change the figures and examine how the result and interpretation respond.

Use the same easy method every time: ① Understand what is being measured → ② identify every symbol → ③ place the figures in the formula → ④ calculate → ⑤ explain what the answer means for the workplace.
1 of 9 1 of 9 calculations explored

Mean Time Between Failures — MTBF

Mean means average. MTBF is the average operating time between one failure and the next failure of repairable equipment.

MTBF = T ÷ N

T = total operating time
N = number of failures

Why it is relevant

  • Compares the reliability of machines or operating periods.
  • Supports preventive-maintenance planning.
  • Shows whether failures are becoming more frequent.
  • Provides evidence for investigating deterioration or repeated component failure.

Important interpretation

A higher MTBF is generally better. A falling MTBF—such as 1,000 → 600 → 300 hours—means failures are occurring more frequently and should be investigated.

Maintenance technicians checking vibration, bearing condition, lubrication and alignment on an industrial compressor
Example: A falling compressor MTBF is the clue. Technicians then inspect vibration, bearings, seals, lubrication, alignment, workload and maintenance history to find the cause.

Live MTBF calculator

See the complete worked example
  1. Record total operating time: T = 5,000 hours.
  2. Record failures: N = 10.
  3. Divide 5,000 by 10.
  4. MTBF = 500 hours.
  5. Interpretation: the machine operated for an average of 500 hours between failures.

MTBF trend inspector

What can make MTBF fall?

Equipment deterioration, ageing parts, ineffective maintenance, overloading, poor lubrication, incorrect operation, harsh conditions or repeated failure of the same component. Use maintenance records, failure reports and operating evidence to test these possibilities.

Read the complete picture

Combined Interpretation Dashboard

These figures describe different aspects of the same machine. One figure alone does not provide a complete reliability assessment.

MTBF500 hAverage operating time between failures
MTTR4 hAverage time needed to repair
Availability99.21%Time ready and operational
R(100)81.87%Chance of completing 100 hours without failure

Level 6 interpretation

The machine has high availability because failures can be repaired quickly, but its probability of completing 100 continuous hours without failure is only 81.87%. High availability must not be incorrectly presented as proof of high reliability or safety.

Interactive investigation sequence

Use Calculations for Failure Tracing

Step 1 of 6

Collect operating and failure data

Gather operating hours, downtime, repair duration, failure type, failed component, operating condition and maintenance history. Calculations are only as reliable as the data used.

Section 1.6.4

Using Reliability Calculations to Analyse System Performance

Calculation is only the beginning. Level 6 analysis explains what the figures show, why the change matters, what evidence is still needed and which management action is proportionate.

Independent DB HSE teaching content: prepared solely by Debjyoti Biswas; not produced by OTHM.

CalculateProduce valid indicators. InterpretCompare trends, targets and consequences. Decide & actPrioritise control and verify improvement.
The Level 6 questions

Move From a Number to a Defensible Judgement

Is it acceptable?

Compare the result with the performance target, manufacturer information, legal or organisational requirements and the safety function.

Is it changing?

Compare time periods. Decide whether MTBF is falling, failure rate or MTTR is rising, and whether availability or mission reliability is deteriorating.

What should happen?

Identify weak equipment, consider consequences, investigate causes, prioritise action and allocate competent people, time, spares and budget.

Questions an HSE professional should ask

  • Is the result acceptable for the required safety function?
  • Is performance improving, stable or deteriorating?
  • How does this period compare with earlier periods?
  • Which equipment or component is weakest?
  • What are the safety, health, environmental and business consequences?
  • Is further inspection or investigation required?
  • What management action and resources are justified?
  • Is the data accurate, complete and collected on a consistent basis?
Baseline, trend, comparison and target

Four Ways to Analyse Performance

1. Establish a baseline

A baseline is the starting performance against which future results are compared. Example: MTBF 1,000 h; MTTR 4 h; λ 0.001/h; availability 99.60%; R(100) 90.48%.

The same definition of failure, operating-time boundary, repair-time boundary and units must be used in every later comparison.

2. Analyse the trend

One result is a snapshot. A sequence reveals direction. Falling MTBF together with rising λ and MTTR is stronger evidence of deterioration than one isolated value.

MTBF: 1,000 → 800 → 600 → 400 h

3. Compare equipment

Compare like with like, then explain differences in age, condition, duty, workload, environment, design, operator competence, maintenance, materials and modifications.

4. Compare with a target

A fire pump may have a target availability of at least 99.5%, MTTR below 3 h and λ below 0.001/h. Actual values of 98.7%, 5 h and 0.002/h are adverse gaps requiring action.

IndicatorQ1Q2Q3Q4Interpretation
MTBF (h)1,000800600400Failures are becoming more frequent.
MTTR (h)4.04.55.07.0Recovery is becoming slower.
λ (failures/h)0.001000.001250.001670.00250The deterioration is consistent across indicators.

When should performance be treated as abnormal?

Do not judge a figure in isolation. Compare it with all relevant reference points:

  • The asset’s previous performance
  • Similar equipment performing the same duty
  • Manufacturer specifications and design limits
  • Internal standards and maintenance targets
  • Industry benchmarks and recognised good practice
  • The reliability required for the safety function

A statistically unusual result, a continuing adverse trend or any failure that threatens a critical protection function requires investigation—even when the percentage appears numerically high.

Live system-performance analyser

Compare Two Performance Periods

Change the data. The tool calculates MTBF, MTTR, λ, availability, R(t) and F(t), then explains the direction of change.

Period 1 — baseline

Period 2 — current

IndicatorPeriod 1Period 2Direction
Before and after reliability comparison A grouped bar chart comparing normalised reliability indicators.
Combine indicators and consequences

One Indicator Can Mislead

High availability can hide frequent failure

If MTBF = 500 h and MTTR = 2 h, availability is 99.60%. Yet, with λ = 0.002/h, R(100) is only 81.87%. Rapid repair creates high availability but does not make the system failure-free.

Availability ≠ reliability ≠ safety

Consequence changes the priority

Risk combines likelihood and consequence. The same failure probability may be tolerable for a non-critical printer but unacceptable for a gas detector, fire pump or emergency shutdown system.

IndicatorPump APump BMeaning
MTBF1,200 h450 hPump B fails more frequently.
MTTR3 h8 hPump B is slower to restore.
λ0.00083/h0.00222/hPump B has the higher failure rate.
Availability99.75%98.25%Pump B is less ready for use.

Prioritise for investigation when evidence shows:

  • High or rising failure rate
  • Falling MTBF
  • High or rising MTTR
  • Availability below target
  • Unacceptable F(t)
  • Repeated failure of one component
  • Safety-critical function
  • No effective backup
Section 1.6.5

Using Reliability Calculations in Failure Tracing

Calculations show that performance changed; they do not, by themselves, explain why. Failure tracing follows the evidence through the equipment, task, conditions, human factors and management system until controllable causes are identified.

Independent DB HSE teaching content: prepared solely by Debjyoti Biswas for Level 6 · Unit 3 · Section 01.

When to trace and what to collect

Start With a Clearly Defined Failure Event

Failure identification

States what failed, where and when: for example, “Fire Pump B failed to start during the weekly proof test at 09:20.”

Failure tracing

Examines how and why the failure developed, how it moved through the system and which immediate, underlying and root causes must be controlled.

Common triggers

  • Sudden or continuous reduction in MTBF
  • Rising λ or MTTR
  • Availability below target
  • Unacceptable failure probability
  • Repeated failure of one component
  • Large differences between similar assets
  • Failure soon after maintenance
  • Failure of a safety-critical item
  • Primary and backup equipment failing together
  • Deterioration in previously reliable equipment

Evidence to collect before concluding

  • Operating hours, cycles and duty
  • Failure date, time, type and alarms
  • Repair time, tests and replaced parts
  • Maintenance, inspection and proof-test history
  • Operators, contractors and shift conditions
  • Manufacturer information and design limits
  • Temperature, pressure, vibration and process data
  • Dust, moisture, corrosion and other environmental conditions
  • Photos and preserved damaged components
  • Permit-to-work and isolation records
  • Training, competence and handover records
  • Previous investigations and management-of-change records
Interactive reliability clues

Select the Result That Changed

Choose a reliability clueThe tool will suggest a focused investigation direction. It does not declare a root cause.
Interactive investigation route

Twelve Steps From Function to Verification

Select each step. A competent investigation may move back and forth as new evidence appears.

Cause structure

Do Not Stop at the First Technical Explanation

Immediate cause

The direct event closest to the failure: a bearing overheated and seized.

Underlying cause

The local condition that allowed it: insufficient lubrication and no condition warning.

Root / organisational cause

The management-system weakness: the maintenance system did not specify, schedule or verify lubrication and monitoring.

Five Whys — pump example

Problem: Pump stopped.
Why 1: Bearing seized.
Why 2: Bearing overheated.
Why 3: Lubrication was insufficient.
Why 4: The lubrication task was not completed.
Why 5: The maintenance system did not generate and verify the task.

Fault Tree Analysis — “fire pump fails to start”

Top event = Electrical OR Mechanical OR Control failure
  • Electrical: loss of supply, open protection, starter defect.
  • Mechanical: motor seizure, pump obstruction, coupling failure.
  • Control: sensor, logic, signal, set-point or interlock failure.

FTA helps test multiple credible paths instead of accepting the first visible defect.

Failure Modes and Effects Analysis — FMEA

FMEA is a forward-looking method. For each component or process step, identify the possible failure mode, its effect, its likely cause, existing controls and any further action. It helps prioritise what could fail before an incident occurs.

ItemFailure modePossible effectPossible causeExisting / further control
Pump sealLeakage or face damageLoss of containment and pump shutdownMisalignment, vibration or unsuitable materialAlignment verification, material review and vibration monitoring
Risk Priority Number: RPN = S × O × D

RPN is a prioritisation aid, not a universal risk-acceptance rule. Follow the organisation’s FMEA rating definitions; a catastrophic severity may require action even when the total RPN is not the highest.

Live Pareto failure-frequency tool

Find the Component Dominating the Failure Record

Example total = 12 failures. The tool sorts the components and calculates percentage and cumulative percentage.

ComponentFailuresShareCumulative
Broaden the investigation

Human, Organisational, Common-Cause and Hidden Failures

Human & organisational factors

Examine workload, fatigue, supervision, shift handover, competence, interface design, communication, production pressure, resources, contractor control, unclear roles and ignored warnings.

Common-cause failure

Parallel equipment may share electricity, control panel, fuel, cooling, room, procedure, maintenance error or defective component batch. Redundancy is valuable only when independence is credible.

Hidden failure

An alarm, emergency generator, relief valve, gas detector, interlock or shutdown may fail without being noticed until demanded. Inspection and proof testing must reveal dormant failure.

Full worked failure-tracing case

Recurring Pump-Seal Failure: Before and After

Initial evidence — 6,000 operating hours

12 failures, 60 repair hours; 8 of 12 failures involved the seal.

  • MTBF = 6,000 ÷ 12 = 500 h
  • MTTR = 60 ÷ 12 = 5 h
  • λ = 12 ÷ 6,000 = 0.002/h
  • A = 500 ÷ 505 × 100 = 99.01%
  • R(100) = e−0.2 = 81.87%

Tracing findings

Evidence showed excessive vibration and shaft misalignment. Seals had been replaced repeatedly without checking alignment; no vibration monitoring existed, and the maintenance procedure addressed replacement but not the failure cause.

Failure event → Seal damage → Vibration → Misalignment → Maintenance-system omission

Failure

Pump stopped because leakage exceeded the safe limit.

Mode & immediate cause

Seal faces damaged by excessive vibration.

Underlying & root cause

Misalignment plus a procedure that omitted alignment verification and vibration monitoring.

Corrective actions

Realign the pump, replace damaged parts, introduce vibration monitoring, revise the maintenance procedure, train maintainers, verify alignment after intervention and review similar pumps for the same weakness.

IndicatorBeforeAfter next 6,000 hChange
Failures12375% reduction
MTBF500 h2,000 hLonger between failures
MTTR5 h3 hFaster safe recovery
λ0.002/h0.0005/h75% reduction
Availability99.01%99.85%Improved readiness
R(100)81.87%95.12%Improved mission reliability

Verification: the recalculated results support the conclusion that the actions improved performance. Continue monitoring to confirm that the improvement is sustained and has not introduced another risk.

Evidence boundaries and records

What the Calculations Can and Cannot Prove

They can show

  • Frequency, average duration and direction of change
  • Comparison with another asset, period or target
  • Which component dominates the recorded failures
  • Whether performance improved after action

They cannot prove alone

  • The physical, human or organisational root cause
  • That the data definition and records are accurate
  • That a high availability figure means the system is safe
  • That a short-term improvement will continue

Failure-tracing record

Record the asset and required function; defined failure event; operating hours and duty; failed component and failure mode; immediate, underlying and root causes; evidence reviewed; risk and consequence; corrective actions, owner and due date; post-action calculations; proof of effectiveness; residual risk; and lessons shared with similar systems.

Practical workplace application

Reliability Case Studies

Compressor comparison

Compressor A operates for 12 months without failure. Compressor B requires repair every six weeks.

Annotated welding local exhaust ventilation case showing airflow and fume measurements before, during and after failure correction

LEV performance deterioration

Capture velocity falls from 0.52 m/s to 0.28 m/s while fume concentration rises from 1.2 mg/m³ to 4.8 mg/m³.

Knowledge check

Test Your Understanding

1. What does a falling MTBF normally indicate?
2. Which calculation measures average repair time?
3. Why can availability be high while reliability is lower?
4. What can defeat a parallel backup arrangement?
5. A pump operates 6,000 hours and fails 12 times. What is MTBF?
6. MTBF is stable but MTTR is increasing. Where should tracing focus first?
7. If R(100) = 81.87%, what is F(100)?
8. What is the defining feature of a series system?
9. How is Rp pronounced?
10. Can reliability calculations alone prove the root cause?
11. Seal 8, bearing 2, electrical 1, control 1: which component should receive focused tracing?
12. Why proof-test an emergency shutdown or alarm?
Answer all twelve questions, then check your score.
Unit 3 · Section 02

Understand the Strategies and Techniques of Risk Control

Welcome to the next part of our learning journey. Section 01 helped us identify, calculate and trace risk-related evidence. Section 02 asks the next management question: What should we do about the risk, and how can we justify that decision?

Learning outcomes: 2.1 Evaluate the use of common risk-management strategies. 2.2 Justify when to use risk avoidance, risk reduction, risk transfer, risk analysis, risk evaluation and risk review strategies. 2.3 Explain the development and characteristics of safe systems of work and safe operating procedures.

Let us begin with the simplest meaning

What Is Risk Control?

Imagine that we find an unguarded moving part on a machine.

The moving part is the hazard. The possibility and seriousness of someone being injured is the risk. The guard, isolation system or redesign used to prevent contact is the control.

Risk control means selecting, implementing and checking measures that eliminate the hazard or reduce the risk to an acceptable or required level.

Risk assessment is not the final product

A completed form does not protect anyone. Protection appears only when suitable controls are implemented, communicated, resourced, used and verified.

Easy memory: Find it → Understand it → Control it → Check it.

Hazard

Something with the potential to cause injury, ill health, damage or another unwanted outcome.

Risk

The combination of how likely harm is and how serious the consequence may be, considering exposure and existing controls.

Control

A measure that removes the hazard, prevents exposure, reduces likelihood, limits consequence or supports safe recovery.

Level 6 lens: Do not only name a strategy. Evaluate its strengths and limitations, then justify why it is suitable for the people, hazard, evidence, standards and workplace conditions involved.
Section 2.1

Evaluate the Use of Common Risk-Management Strategies

Evaluation means more than describing what a strategy is. Ask whether it is suitable, what evidence supports it, what benefit it offers, where it may fail and how its effectiveness will be checked.

1. SuitabilityDoes it fit the risk and context? 2. EvidenceWhat information supports the choice? 3. LimitationsWhat uncertainty or weakness remains?

The Risk-Assessment Process—Seven Clear Steps

Different organisations may group the stages differently. Here, we use the seven stages indicated for this learning outcome, followed by a continuous review loop. Select each step to hear the portal explain it.

Step 1 — Identify risks and hazardsExamine routine and non-routine work, equipment, substances, people, environment, foreseeable misuse, maintenance, change and emergencies. Use observation, worker consultation, incident data, instructions and technical information.
AssessBuild a suitable and sufficient picture. ActImplement the planned controls. ReviewCheck effectiveness and reassess change.

Control Standards, Action Plans and Priority

What is a risk-control standard?

It is the benchmark that tells us what level or type of control is required. Sources may include legislation, approved guidance, exposure limits, engineering codes, manufacturer instructions, industry good practice, internal rules and the hierarchy of controls.

Why it matters: A coloured risk score alone cannot decide whether a mandatory guard, ventilation system, permit or exposure limit is required.

What makes an action plan useful?

  • Specific control action—not “be careful”.
  • Named responsible owner and adequate resources.
  • Realistic completion date and interim protection.
  • Priority based on risk, standards, people and uncertainty.
  • Method to verify completion and effectiveness.

Priority is not “highest score only”

Give prompt attention to imminent danger, serious consequences, legal or control-standard gaps, many people exposed, vulnerable persons, ineffective controls and high uncertainty. A lower matrix score must not be used to postpone a non-negotiable requirement.

Generic, Specific and Dynamic Risk Assessments

These are not three levels of quality. They are three ways of matching the assessment to the work situation.

Assessment typeEasy meaningUse it when…Do not rely on it when…Main strength and limitation
GenericA baseline assessment for similar activities, hazards or locations.Work is routine, repeated and genuinely comparable; common controls can be standardised.People, equipment, substances, environment or task conditions differ materially.Strength: efficient and consistent. Limitation: may overlook local or individual differences.
SpecificAn assessment for one task, site, machine, substance, project or person.Work is unusual, complex, high risk, legally specific, non-routine or affected by vulnerability.A suitable generic assessment already covers truly identical low-risk work—although local checks are still needed.Strength: detailed and relevant. Limitation: requires time, information and competence.
DynamicA continuous, in-the-moment judgement as conditions change.Emergency response, rapidly changing work or an unexpected condition requires immediate reassessment.Work is planned and foreseeable. It must not replace a suitable formal assessment, method statement or permit.Strength: responds to reality. Limitation: time pressure and incomplete information may weaken judgement.

Interactive Assessment-Type Selector

Describe the situation. The tool will recommend a starting approach and explain why.

Now let us see the paperwork behind the words

What Do Generic, Specific and Dynamic Assessments Actually Look Like?

Think of these as three different lenses. A generic assessment establishes a reusable baseline. A specific assessment focuses that baseline on one real task, place, item, substance or person. A dynamic assessment keeps checking the live situation while work or an incident develops.

Safety adviser and warehouse supervisor reviewing a baseline assessment beside segregated loading-bay traffic
Generic — the repeatable baselineUseful when the activities, hazards and standard controls are genuinely similar. The local supervisor must still check that the real people, place, equipment and conditions match.
Competent team carrying out a pre-entry assessment at an industrial confined space with gas testing, ventilation and rescue arrangements
Specific — one real job in its real contextThis vessel entry needs named isolations, atmospheric tests, entrants, rescue arrangements, permit interfaces and a time-limited decision. A generic confined-space form is only background.
Industrial response team behind an exclusion barrier while a supervisor checks changing leak conditions with a detector and radio
Dynamic — the situation is changing nowThe leader continually observes, reassesses, controls, communicates and decides whether to proceed, withdraw or escalate. The later debrief must feed lessons back into formal assessments.
History note: These three formats were not invented together by one person. Generic and task-specific forms grew from systematic safety-management practice. Dynamic risk assessment was developed particularly through fire-service incident command for dangerous, unpredictable environments and was later applied more widely. It supplements planned assessment; it does not excuse foreseeable work from proper planning.

Interactive Assessment-Format Explorer

Select a type. The portal will show its purpose, its header, the columns a suitable form normally needs, a completed example and the final decision that must be recorded.

Core fields every formal assessment needs

  • Clear scope, boundaries, location, activity and version.
  • Assessor, competent contributors, workers consulted and approval.
  • Hazards and credible harm—not only a list of objects.
  • Who may be harmed, including contractors, visitors and vulnerable people.
  • Existing controls and evidence that they are really in place.
  • Risk judgement before further action, with the reasoning shown.
  • Further controls, owner, due date, interim protection and priority.
  • Residual risk, communication, verification and review triggers.

A blank box is not automatically a bad form

Good forms create space for sound thinking; they do not replace it. The assessor must walk the task, consult the people who understand the work, use relevant evidence and test assumptions.

Quality test: Could a competent supervisor read this record and understand what may happen, who is exposed, which controls must exist, what remains to be done and when work must stop?

Do Not Miss Long-Term Hazards to Health

An injury hazard may produce an immediate event. Many health hazards are quieter: exposure can accumulate and illness may appear months or years later. “Nothing happened today” is not evidence that the risk is controlled.

Noise—hearing damage can accumulateConsider sound level, exposure duration, work pattern, combined sources and susceptible workers. Use competent measurement where needed, compare with the applicable standard, control at source and review hearing-protection and health-surveillance evidence.

Exposure pattern

Frequency, duration, intensity, route, peaks, recovery time and combined exposure.

Evidence

Monitoring, sampling, health surveillance, absence records, worker reports and historical data.

Control

Prevent exposure at source. PPE and surveillance support control; they do not replace elimination or engineering where these are required.

Qualitative, Semi-Quantitative and Quantitative Assessment

The difference is how the risk is described and analysed—not how seriously the assessor takes it.

Words

Qualitative

Uses reasoned descriptions such as low, medium, high, unlikely or severe.

Useful for: straightforward work, screening, discussion and situations where numerical data would add little.

Limitation: categories may be subjective and different risks may appear equal.

Ordered scores

Semi-quantitative

Assigns ranked numbers to likelihood and severity, often calculating a matrix score such as L × S.

Useful for: consistent comparison and action planning across many hazards.

Limitation: numbers can create false precision; score boundaries and multiplication rules are organisational conventions.

Measured or modelled values

Quantitative

Uses numerical estimates of exposure, probability, frequency or consequence based on data and models.

Useful for: complex, high-consequence or technical decisions and comparison with numerical standards.

Limitation: needs valid data, competence and transparent assumptions; an exact-looking number may still be uncertain.

Important: Quantitative does not automatically mean “better”. Use the simplest method that is sufficiently reliable for the decision, consequence, uncertainty and applicable standard.
How did these methods develop?

From Professional Description to Ranked Scores and Probability Models

There is no honest single-name answer to “who invented risk assessment?” People have judged danger for as long as organised work has existed. Modern methods developed gradually as industries needed decisions that were more systematic, comparable and technically defensible.

Industrial safety team considering descriptive hazard evidence, a risk matrix and modelled technical evidence together
One decision may need three kinds of evidence

The methods are a progression of detail—not a competition

Qualitative thinking helps us name what may happen and judge its importance. Semi-quantitative scoring helps a group rank and compare many concerns using defined categories. Quantitative analysis estimates frequency, probability, exposure or consequence when the decision needs numerical evidence.

A major quantitative study still begins with qualitative questions: Which scenarios matter? What assumptions are credible? Who may be affected? A risk matrix still needs professional judgement. The methods often work together.

Qualitative assessment—description came firstThere is no single inventor or creation date. Early safety decisions were commonly deterministic and descriptive: experience, rules, testing and expert judgement were used to ask what could go wrong and what consequences could follow. Qualitative assessment remains useful for screening and straightforward work, especially when numerical data would not improve the decision. Its terms must be defined so “unlikely” or “major” means something consistent.

What the historical landmarks do—and do not—prove

System-safety programmes helped formalise ranked categories and risk matrices, but no single standard invented every semi-quantitative method. By 1971, NASA and aircraft manufacturers were using fault-tree tools. The 1975 US Reactor Safety Study, WASH-1400, directed by MIT professor Norman Rasmussen and AEC staff member Saul Levine, became a landmark probabilistic risk assessment. Later criticism of some numerical claims also taught an essential lesson: model structure, data gaps and uncertainty must be made visible.

Historical teaching sources: US Nuclear Regulatory Commission histories of the Reactor Safety Study and risk-informed regulation; NASA System Safety Handbook risk-matrix material; MIL-STD-882 system-safety practice.

Full formats + worked comparison + decision tools

The Qualitative, Semi-Quantitative and Quantitative Assessment Workshop

We will use one scene throughout: pedestrians and forklifts interact in a busy warehouse loading area. This makes the difference easy to see—the hazard does not change, but the depth and form of the analysis changes.

1. Qualitative formReasoned descriptors, evidence and a narrative decision.
Open qualitative form ↓
2. Semi-quantitative formDefined rankings combined in the interactive 5 × 5 matrix.
Open matrix ↓
3. Quantitative formModelled event frequency, uncertainty range and numerical criterion.
Open quantitative form ↓
Qualitative

Reasoned words

Judgement: “Collision risk is high because pedestrians frequently cross an active vehicle route and a collision could cause fatal injury.”

Decision value: Fast, understandable screening that identifies the urgent need for physical segregation.
Semi-quantitative

Defined ranks

Judgement: Likelihood 4 (likely) × severity 5 (catastrophic) = 20, “very high” under the example organisation’s approved matrix.

Decision value: Helps compare this issue with other recorded hazards and apply action rules consistently.
Quantitative

Measured or modelled values

Evidence: vehicle movements/hour, crossing frequency, near-miss rate, speed, exposure time, barrier reliability and predicted collision consequence.

Decision value: Tests route designs or investment options where reliable data and a suitable model exist.

Interactive Method-Format Explorer

Select a method to see the complete form structure. Notice that each higher-data method retains the basic hazard, people, controls, action and review fields.

Which Method Is Proportionate?

Answer five questions. The recommendation is a starting point for competent judgement—not an automatic approval.

Build a Defensible Qualitative Judgement

“High risk” alone is weak. Link the descriptor to exposure, consequence, controls, evidence and uncertainty.

Complete the evidence and let the portal build the reasoning chain.

Create One Complete Risk-Assessment Record

This tool mirrors the essential columns of a practical assessment register. It helps learners see that a risk rating sits inside a much larger management record.

Complete the owner and due date, then build the assessment row.
Complete interactive format 1 of 3

Interactive Qualitative Risk-Assessment Form

A qualitative assessment uses defined words and reasoned professional judgement. It does not multiply scores. The record must explain why the likelihood, consequence and overall priority descriptions fit the evidence.

A. Assessment identity and scope

B. Hazard, people and evidence

C. Initial qualitative judgement

D. Treatment, responsibility and residual judgement

Enter the assessment and action dates, then generate the complete qualitative record.

Example likelihood meanings

  • Rare: exceptional under the defined conditions.
  • Unlikely: foreseeable but not expected during normal activity.
  • Possible: could occur during the activity or assessment period.
  • Likely: expected to occur repeatedly unless controls improve.
  • Almost certain: occurs frequently or conditions make occurrence imminent.

Example consequence meanings

  • Minor: limited, short-term harm.
  • Moderate: treatment or restricted work may be required.
  • Serious: major injury or significant occupational ill health.
  • Major: life-changing harm or single fatality potential.
  • Catastrophic: multiple fatalities or major widespread impact.
Important: These descriptor definitions are teaching examples. An organisation must approve definitions appropriate to its activities and use them consistently. Legal requirements and recognised control standards override a convenient “low” description.

Interactive 5 × 5 Semi-Quantitative Risk Matrix

Compare the initial risk with the residual risk after proposed controls. The example bands are for learning only; an organisation must define and approve its own criteria.

Choose the ratings

Initial risk
20
Residual risk
10
1–4 Low5–9 Moderate10–16 High17–25 Very high

Solid outline: initial rating. Dashed white outline: residual rating.

Why might severity stay at 5?

A guard or interlock may make contact much less likely, but if contact still occurs the possible injury may remain catastrophic. Do not automatically reduce both numbers simply because controls were proposed. Rate the real effect of the selected controls and verify them.

A Simple Quantitative Comparison Tool

Measured value compared with an applicable limit

This demonstration calculates an exposure ratio. Both figures must use the same unit and come from a valid assessment.

Exposure ratio = Measured value ÷ Applicable limit

How to interpret carefully

A ratio of 0.70 means the measured value is 70% of the selected limit. It does not automatically mean there is no risk or that controls can be relaxed.

Check sampling quality, uncertainty, peak exposure, routes of exposure, combined substances, vulnerable people, the legal meaning of the limit and whether further reduction is required.

Quantitative Scenario-Frequency Calculator

This teaching model follows one event path. It estimates how often the defined harmful outcome may occur by combining an initiating-event frequency with conditional probabilities. It is useful for learning the logic of event trees; it is not a substitute for a validated QRA.

fharm = finitiator × P(exposure) × P(control failure) × P(harm | event)
Read the full formula aloud: “F sub harm equals F sub initiator, multiplied by the probability of exposure, multiplied by the probability of control failure, multiplied by the probability of harm given the event.”

How to Read Every Symbol—and Why It Is Used

fharmPronounced: “f sub harm.”

It means the estimated frequency of the defined harmful outcome, normally stated per year. It appears on the left because this is the answer the model is calculating.

finitiatorPronounced: “f sub initiator” or “initiating-event frequency.”

It means how many times the event that starts the harmful scenario occurs during a stated period, usually one year. The small word below f is a label—it is not multiplication.

PPronounced: “probability.”

P shows that the value inside the brackets is a chance from 0 to 1. For example, 0.25 means a 25% chance.

P(exposure)Pronounced: “probability of exposure.”

It means the chance that a person is present or exposed when the initiating event occurs. It is used because an initiating event cannot harm a person who is not in the exposure path.

P(control failure)Pronounced: “probability of control failure.”

It means the chance that the intended barrier, safeguard or protective control fails, is unavailable or does not stop the event path.

P(harm | event)Pronounced: “probability of harm given the event.”

The vertical bar | means “given that”. It asks: once the event has reached the exposed person, what is the chance of the defined harm?

×Pronounced: “multiplied by.”

Multiplication is used because the harmful outcome follows a sequence: the initiating event occurs, exposure exists, the control fails and harm follows. Each stage narrows the original frequency.

=Pronounced: “equals.”

It separates the answer on the left from the factors used to calculate it on the right. Both sides describe the same estimated harmful-outcome frequency.

per year · /yearPronounced: “per year.”

This is the unit, not a percentage. A result of 0.12 per year is a modelled event frequency; it does not mean a 12% annual risk unless a valid model specifically supports that interpretation.

Easy memory: Start events per year × chance of exposure × chance the control fails × chance harm follows = estimated harmful outcomes per year.

Four checks before trusting the output

  1. Scenario: Is the initiating event and harmful outcome defined without ambiguity?
  2. Dependence: Are the probabilities really independent, or can one common cause defeat several controls?
  3. Data: Are frequencies based on comparable equipment, tasks, people and operating conditions?
  4. Uncertainty: Would reasonable lower and upper assumptions materially change the decision?

Level 6 point: A calculated frequency is an estimate conditional on a model. Report units, source, time period, assumptions, uncertainty and sensitivity—not only the final number.

Complete interactive format 3 of 3

Interactive Quantitative Risk-Assessment Form

A quantitative assessment records more than a calculation. It defines the decision, scenario and model; gives every input a source and unit; shows uncertainty; compares the result with an approved criterion; and records treatment and verification.

A. Decision, scope and model boundary

B. Numerical model inputs

fharm = finitiator × P(person exposed) × P(protection fails) × P(harm | exposure)

The pronunciation and meaning of every symbol are explained in the symbol guide immediately above.

C. Assumptions, treatment and assurance

Enter the assessment date, review the numerical inputs and generate the complete quantitative record.

What the uncertainty factor means here

An uncertainty factor of 2 displays a teaching range from the central estimate ÷ 2 to the central estimate × 2. A real QRA may require probability distributions, confidence intervals, alternative models or structured expert judgement. The factor is a learning device—not a universal scientific rule.

What must be reviewed independently?

Scenario completeness, units, input provenance, relevance of data, dependencies and common causes, human-reliability assumptions, consequence model, uncertainty treatment, sensitivity, numerical criterion and whether the model is valid for the decision.

Decision rule: Do not compare a result with an invented criterion. The criterion must come from an applicable legal, regulatory, technical or approved organisational framework. Even a result below a criterion does not remove the duty to apply required good practice and reasonably practicable controls.

Build a Complete Risk-Control Action

An action without an owner, date, interim measure and verification method is only an intention. Complete the fields and create a practical action-plan entry.

Complete the fields, then create the action-plan entry.
Section 2.2

Justify When to Use Six Risk-Management Strategies

Avoidance, reduction and transfer are treatment responses. Analysis, evaluation and review are decision and assurance activities that help us choose, prioritise and verify treatment. In practice, a strong risk-management decision may combine several of them.

To justify means: state the chosen strategy, connect it to evidence and criteria, explain why it suits the context, recognise its limitations and state how it will be reviewed.

Meet the Six Strategies

Select a card. The portal will explain when to use it, when not to depend on it, and what a Level 6 justification should recognise.

Risk avoidance—remove the decision to be exposedUse when: the risk is intolerable, the activity is unnecessary, reliable control is not reasonably achievable, or a safer design or method can remove the hazard. Do not misuse it: cancelling one activity may shift risk elsewhere. Check the alternative for new hazards and operational consequences. Example: redesign a roof-level valve so routine operation can be completed at ground level.

When Each Strategy Is Most Relevant

StrategyUse when…Evidence neededKey limitation or warning
AvoidanceExposure can be removed by stopping, substituting, redesigning or choosing another objective.Severity, feasibility of alternatives, standards, lifecycle and risk-transfer effects.May create a different risk or sacrifice an essential activity; assess the replacement.
ReductionThe activity is necessary and the risk can be controlled using the hierarchy of controls.Control performance, human factors, maintenance, residual risk and verification.Administrative controls and PPE are vulnerable to failure; reduce at source where possible.
TransferSpecialist competence or financial sharing is appropriate—for example, a competent contractor or insurance.Competence, contract scope, interfaces, supervision, insurance and monitoring.Legal and ethical responsibility for protecting people is not simply transferred away.
AnalysisThe risk, causes, exposure, failure paths or options are uncertain or complex.Measurements, incidents, task information, models, assumptions and uncertainty.Do not delay obvious immediate controls while waiting for perfect data.
EvaluationAnalysed risk must be compared with legal, technical or organisational criteria to set priority and treatment.Approved criteria, control standards, consequence, affected groups and tolerance.A matrix colour cannot override a mandatory requirement or conceal uncertainty.
ReviewTime has passed or change, incident, failure, new evidence, worker concern or new requirements may affect validity.Inspection, monitoring, incidents, health data, assurance findings and change information.A scheduled annual review is insufficient when a trigger requires immediate reassessment.

Interactive Strategy Decision Lab

Choose a workplace situation and the strategy you think should lead the response. The tool will explain the strongest answer and supporting strategies.

Choose the scenario and your leading strategy, then check the decision.

Risk Reduction Must Follow the Hierarchy of Controls

When a risk cannot be avoided completely, begin with controls that act on the hazard and exposure pathway. Measures lower in the hierarchy usually depend more heavily on consistent human behaviour.

1 · Most effective

Eliminate

Remove the hazard from the work—for example, design out the need to enter a vessel.

2

Substitute

Replace it with a safer material, method, machine or energy source, then assess the substitute’s hazards.

3

Engineering

Isolate people through guarding, enclosure, segregation, automation, extraction or fail-safe design.

4

Administrative

Use planning, permits, procedures, competence, supervision, scheduling, signage and restricted access.

5 · Last line

PPE

Protect the individual when exposure remains. Select, fit, maintain and supervise its use; do not make it the automatic first answer.

Avoidance and elimination can overlap, but the emphasis differs: avoidance changes the decision or objective so the exposure is not undertaken; elimination removes the hazard or hazardous step from work that continues. In both cases, check that the alternative does not introduce a new serious risk.

Risk Transfer—The Point Learners Must Not Miss

What can be transferred or shared?

Some financial loss may be insured. Specialist work may be contracted to an organisation with suitable equipment and competence. Contract terms may allocate defined responsibilities.

What does not disappear?

The hazard remains until controlled. The client or employer must still select competent parties, provide information, coordinate interfaces, monitor work and meet applicable legal duties. A signature on a contract is not a physical control.

Easy example: Hiring a specialist crane contractor may transfer performance of the lift to specialist hands, but the site still has to coordinate exclusion zones, ground conditions, permits, communication and emergency arrangements.

Build a Level 6 Justification

Use the structure Decision → Because → Evidence → Limitation → Review. This produces a learning scaffold that you should explain in your own professional words.

Complete the evidence, limitation and review trigger, then build the explanation.

Analysis, Evaluation and Review—Do Not Mix Them Up

Risk analysisHow can harm occur? What are the causes, likelihood, consequences, controls and uncertainty? Risk evaluationHow does that analysed risk compare with criteria? Is more treatment required and how urgent is it? Risk reviewIs the assessment still valid and are the controls present, used and effective?

One simple example

Analysis: solvent-vapour measurements, duration and ventilation performance show the nature and level of exposure. Evaluation: the evidence is compared with applicable exposure criteria and good practice to decide whether treatment is adequate. Review: monitoring and reassessment confirm whether the new local exhaust ventilation continues to control exposure after process or maintenance changes.

Common Errors—and the Better Approach

Weak approachWhy it is weakBetter Level 6 approach
Copy a generic assessment without checking the site.Local hazards, people and conditions may be different.Use it as a baseline, then verify and adapt it before work.
Use a dynamic assessment for planned high-risk work.It avoids proper planning, consultation and control design.Complete a formal specific assessment, then use dynamic checks for real-time change.
Reduce both likelihood and severity scores automatically.The selected control may affect only one dimension.Explain how each control changes exposure, failure path or consequence.
Call a risk “low” because no one has yet been harmed.Absence of recorded harm is weak evidence, especially for latent health risks.Consider exposure data, potential severity, under-reporting and control reliability.
Transfer work and stop managing it.Contracting does not eliminate the hazard or all duties.Assess competence, coordinate, monitor and verify contractor controls.
Review only once a year.A change or failure can make the assessment invalid immediately.Use scheduled and event-triggered review.
Section 02 knowledge check

Can You Make and Justify the Decision?

1. When is a generic assessment most suitable?
2. What is the main limitation of a dynamic assessment?
3. A 5 × 5 likelihood–severity matrix is normally…
4. Why may long-term health hazards be underestimated?
5. Relocating a roof valve to ground level is mainly…
6. What does risk transfer NOT mean?
7. What is risk evaluation?
8. When should a risk assessment be reviewed?
9. Why might residual severity remain unchanged?
10. Which is the strongest Level 6 justification?
Answer all ten questions, then check your score.
Section 2.3

Explain the Development and Characteristics of Safe Systems of Work and Safe Operating Procedures

You have identified the hazard, assessed the risk and selected a strategy. The next question is practical: How will people complete the work safely, consistently and under control?

UnderstandDistinguish the documents and their purposes. DevelopTurn risk decisions into a workable system. AssureTrain, supervise, verify and review.
Your guided learning route

Learn One Control Layer at a Time

The portal starts with the wider safe system of work, develops it, then moves to the more focused safe operating procedure. Only after both are clear do we connect the document family and explore the complete permit-to-work lifecycle.

DB HSE International logo

DB HSE learning resource. Prepared solely by Debjyoti Biswas for teaching Unit 3, Section 2.3. This independent learning portal is not produced by OTHM and does not issue workplace authority.

The bridge from assessment to action

A Risk Assessment Decides What Must Be Controlled; the System of Work Decides How

Imagine a chemical-transfer pump that needs maintenance.

The assessment identifies hazardous chemical residue, stored pressure, electricity, moving parts, restricted access, contractors and possible conflict with nearby operations. That information is essential—but it does not yet tell the maintenance team exactly how the job will be prepared, authorised, completed, checked and handed back.

The safe system of work connects the people, equipment, controls, communication and sequence. The safe operating procedure gives the approved steps for a defined operation inside that wider system.

The paperwork test

A document does not make work safe merely because it has been signed. The controls must exist at the workplace, the people must understand them, and a responsible person must verify that they remain effective.

Easy memory: Assess → Design → Explain → Do → Check → Improve.

Level 6 lens: “Explain” requires a connected account of what an SSOW and SOP are, how they are developed, why each characteristic matters, when different strategies are used, and what may happen if the arrangements are weak.

One Scenario, Four Stages of Control

We will use the same chemical-transfer pump throughout Section 2.3. This allows you to see how one risk picture is converted into a complete working arrangement.

Technician beginning unsafe chemical-transfer pump maintenance without verified isolation, barriers, supervision or coordinated controls
1. The uncontrolled starting pointThe task has begun before the energy, chemical, access and coordination risks have been converted into verified controls.
Four industrial professionals consult beside a chemical-transfer pump, reviewing isolation points and a pre-job plan
2. Consultation and developmentThe operator, maintainer, supervisor and HSE professional combine task knowledge, hazard information and practical experience.
Technician and supervisor verify isolated, locked-out chemical-transfer pump maintenance within a controlled exclusion zone
3. Controlled executionIsolation, safe condition, barriers, tools, competence, PPE and supervision are present and verified before intrusive work begins.
Responsible people jointly verify a permit, isolation points and safe handover before chemical-transfer pump maintenance
4. Authorisation and handoverThe issuing and performing parties share the same understanding of the job, limits, precautions, status and return-to-service requirements.
Start here · the wider control system

What Is a Safe System of Work—SSOW?

Safe System of Work — SSOW

Pronounced: “S-S-O-W,” or simply “safe system of work.”

An SSOW is a deliberately organised method for completing work so that foreseeable hazards are controlled throughout preparation, execution, completion and foreseeable abnormal conditions.

It is a system because it joins people, plant, materials, environment, controls, responsibilities, communication, competence, supervision and review. It is not only a list of steps.

What an SSOW is not

It is not a risk-assessment form, a signature, a list of PPE, a copied method statement or an instruction to “be careful.” It must convert risk decisions into a realistic arrangement that people can understand and use.

Simple test: Could a competent team use this system to know who does what, in what order, with which controls, when to stop, and how the work returns safely to normal?

1 · Scope and boundaryTask, equipment, location, start and finish points, exclusions and conditions of use.
2 · People and authorityRoles, responsibilities, competence, supervision and stop-work authority.
3 · Hazards and evidenceRisk assessment, legal and technical standards, manuals, monitoring and incident learning.
4 · Controls and sequenceElimination or reduction measures, isolation, access, PPE, hold points and safe order.
5 · CommunicationBriefing, language and literacy needs, interfaces, handovers and changes in plant status.
6 · Abnormal conditionsStop rules, failed tests, loss of control, alarm, withdrawal, rescue and emergency escalation.
7 · Verification and handbackPhysical checks, completion, accounting for people/tools, reinstatement and acceptance.
8 · Assurance and reviewMonitoring, supervision, document control, feedback, investigation and review triggers.
SSOW format fieldWhat to recordWhy the field matters
Identity and controlTitle, number, owner, version, approval and review date.Prevents obsolete or unapproved instructions being used.
Purpose, scope and limitsActivity, plant, location, persons, conditions, interfaces and exclusions.Stops the system being applied outside the conditions it was designed for.
Roles and competenceWho plans, authorises, performs, supervises, verifies, hands back and reviews.Prevents gaps, duplication and unverified assumptions.
Hazards and control basisAssessment reference, standards, energy sources, substances, exposure and credible failures.Shows that the method is risk-based and technically supported.
Preparation and resourcesAccess, barriers, isolations, tools, staffing, communication, permits and PPE.Creates the conditions needed before work starts.
Safe sequence and hold pointsOrdered actions, responsible role, required result and checks before progression.Some controls only work when applied in the correct order.
Stop, abnormal and emergency rulesConditions that suspend work, safe state, escalation, rescue and recovery.Prevents unsafe improvisation when assumptions change.
Completion, handback and reviewInspection, reinstatement, records, acceptance, monitoring and review triggers.Controls the return to normal operation and captures learning.

Interactive Jargon Translator

Select any term. The portal will pronounce it, define it and explain why it matters.

Competent personPronounced: “kom-puh-tent person.” A person with suitable knowledge, training, skill and experience—and the ability to recognise their own limits—for the assigned work. A certificate alone does not prove competence for every situation.

Do Not Mix Up the Document Family

The documents are connected, but they do different jobs. The level of formality should be proportionate to the risk, complexity and need for coordination.

1 · AssessIdentify hazards, people, risk, standards and further controls.
2 · DesignConvert the assessment into an SSOW for the complete activity.
3 · InstructUse SOPs or method steps for defined operations inside the SSOW.
4 · AuthoriseUse PTW where specified high-risk work needs formal time-and-place control.
5 · PerformBrief, supervise, follow controls and monitor changing conditions.
6 · Hand backInspect, reinstate, communicate status and return plant to its owner.
7 · LearnKeep records, investigate deviations and review every affected layer.
Risk assessmentWhat could cause harm? SSOWHow will the complete activity be controlled? SOPWhat approved operating steps must be followed? Permit-to-workWho authorises this defined high-risk job, where and when?
Document or arrangementEasy purposeTypical useImportant limitation
Risk assessmentIdentifies hazards, people, existing controls, risk and further action.Before deciding the safe method and whenever relevant change occurs.A completed form does not implement the controls.
SSOWCoordinates the whole method, people, controls and interfaces.Where risks require a defined safe way of working, particularly complex or significant activities.It fails if impractical, unknown, unsupervised or not followed.
SOPStandardises the safe steps for a defined operation.Routine or repeated operation, inspection, start-up, shutdown, cleaning or maintenance.It cannot predict every abnormal condition; stop and escalation rules are needed.
Method statementDescribes how a particular job or project stage will be carried out.Construction, installation, maintenance and contractor work.A generic copied statement may not match the real site or sequence.
JSA/JHABreaks a job into steps, hazards and controls.Task planning and workforce discussion.Step-by-step analysis must still consider interactions and emergencies.
Permit-to-work — PTWFormally authorises specified work, location, time and precautions.Defined high-risk or tightly coordinated work under site rules.A permit is not a guarantee of safety and does not replace risk assessment.
ChecklistConfirms that required checks were completed.Pre-start, inspection, handover and verification.Ticking boxes without observation gives false assurance.
Emergency procedureExplains response when control is lost or conditions become unsafe.Credible abnormal and emergency situations.It must be resourced, communicated and tested—not only filed.

Tool: Which Arrangement Does This Job Need?

Development is a lifecycle, not a typing exercise

Fifteen Stages for Developing an Effective Safe System of Work

The stages are grouped into five phases. Select a phase to explore what must happen and why.

Phase 1 — Understand the real work1. Define the task, purpose, location, boundaries and normal/abnormal conditions. 2. Gather legislation, technical information, manufacturer instructions, incident history and existing procedures. 3. Observe the job and consult the people who perform, supervise and maintain it. The aim is to understand work as done—not only work as imagined.
Define the task and boundaries. State what is included, excluded, where it happens, what must be achieved and when the system applies.
Gather reliable information. Use applicable law, guidance, designs, manuals, safety data, previous assessments, monitoring and incident learning.
Observe and consult. Involve workers, supervisors, contractors, engineers and specialists who understand the real task and foreseeable shortcuts.
Identify hazards and credible failure paths. Include people, plant, substances, energy, environment, human factors, simultaneous operations and emergencies.
Analyse and evaluate risk. Understand causes, exposure, likelihood, consequence, uncertainty and applicable criteria.
Select the risk strategy. Avoid where possible; otherwise reduce using the hierarchy. Use specialist transfer, analysis, evaluation and review appropriately.
Design the control sequence. Put controls in the order they must exist, identify hold points and define stop-work conditions.
Assign roles and competence. State who prepares, authorises, performs, supervises, verifies, hands over and reviews.
Plan abnormal and emergency conditions. Explain safe shutdown, alarm, withdrawal, containment, rescue and escalation where relevant.
Write the SSOW and supporting SOPs. Use clear language, diagrams and workplace terminology. Remove ambiguity and unnecessary complexity.
Walk through and test. A competent team checks the sequence against the real workplace before full implementation.
Approve and control the document. Record owner, approver, version, date, review date and controlled availability.
Communicate, train and confirm understanding. Adapt for language, literacy, learning needs and role-specific competence.
Implement and supervise. Provide time, equipment, staffing and authority to stop when conditions differ.
Monitor, learn and improve. Observe work, inspect controls, investigate deviation, consult users and revise after relevant triggers.
Now focus on a defined operation

What Is a Safe Operating Procedure—SOP?

Safe Operating Procedure — SOP

Pronounced: “S-O-P,” or “safe operating procedure.”

An SOP is an approved, controlled and repeatable set of instructions for performing a defined operation safely and consistently. It tells the authorised user what conditions must exist, what to do in sequence, what result to confirm and when to stop.

Typical SOPs cover start-up, normal operation, sampling, cleaning, inspection, safe shutdown, isolation preparation, testing and return to service.

How it fits inside the SSOW

The SSOW coordinates the complete job—including teams, interfaces, permits, isolations, emergency arrangements and handback. The SOP standardises one defined operation within that system.

A pump-maintenance SSOW may refer to separate SOPs for shutdown, electrical isolation, line draining, gas testing and controlled recommissioning.

Important boundary: An SOP supports consistent work; it does not make an unsuitable task safe, replace risk assessment, authorise permit-controlled work or remove the need to stop when actual conditions differ.
SOP format fieldWhat it should containQuality question
Document identityTitle, equipment or process, number, version, owner, approver and review date.Can the user confirm this is the current approved procedure?
Purpose and scopeIntended result, authorised users, operating range, location and exclusions.Is it clear when the SOP applies—and when it does not?
Responsibilities and competenceOperator, supervisor, verifier, specialist and required authorisation.Does each person understand their role and limit of authority?
PrerequisitesPlant state, permits, isolations, tools, inspections, guards, ventilation and PPE.What must be true before Step 1?
Ordered stepsOne clear action per step, responsible role, location, setting, safe limit and expected result.Can the action and its successful outcome be observed?
Warnings and hold pointsCritical hazards, prohibited actions and mandatory verification before continuing.Are the most safety-critical instructions easy to find?
Operating limitsPressure, temperature, concentration, speed, time or other acceptance criteria.Does the user know when a result is outside the safe range?
Stop and escalationUnexpected state, failed check, alarm, leak, defect, safe shutdown and person to contact.Does the SOP prevent improvisation?
Completion and recordsFinal checks, housekeeping, status communication, log entries and retained evidence.Can another person confirm the operation ended safely?
Review and changeScheduled date plus triggers such as incident, modification, feedback or new evidence.Will the SOP remain aligned with the real process?
Develop the instruction from real work

How Is an SOP Developed?

Select each phase. The five phases contain ten connected stages: define → observe → assess → sequence → write → validate → approve → train → use → review.

Stages 1–2 — Define and observeDefine the operation, purpose, users, plant, boundaries and safe result. Then observe competent people performing the real task and ask where variation, delay, confusion, error or abnormal conditions occur.
1. Define. State the operation, intended result, equipment, users, range and exclusions.
2. Observe. Watch the job as actually performed and consult operators, maintainers and supervisors.
3. Assess. Link hazards, exposure, failure modes, limits and controls to the wider assessment and SSOW.
4. Sequence. Put preparation, operation, verification, shutdown and completion into a safe order.
5. Write. Use direct actions, familiar terms, diagrams where useful, measurable limits and clear stop rules.
6. Validate. Competent users walk through or trial the draft in controlled conditions and report ambiguity.
7. Approve. The authorised owner accepts the technical basis and controls the version.
8. Train. Explain the purpose, critical steps, limits, abnormal response and evidence of competence.
9. Use and supervise. Make the current SOP accessible, provide resources and verify real application.
10. Review. Learn from change, deviation, incident, user feedback, audit, monitoring and scheduled review.

Characteristics of a Strong SSOW and SOP

A good document must be technically correct and usable by the people who depend on it. These characteristics are evidence of quality—not decorative features.

CharacteristicEasy meaningWhy it mattersWarning sign
Risk-basedControls come from a suitable assessment and required standards.The system addresses credible harm rather than copying another job.The procedure mentions PPE but not the main energy or exposure source.
Task-specificIt matches the actual plant, people, place and conditions.Local differences can change the failure path.Wrong equipment number, location, substance or isolation point.
ProportionateDetail and formality reflect risk and complexity.Too little detail leaves gaps; excessive paperwork hides critical controls.A simple task has 50 pages, while a major intervention has one vague paragraph.
Clear and sequentialActions are unambiguous and in the correct order.Sequence can determine whether energy or exposure is controlled.“Make safe” without stating who, how or how safety is verified.
PracticalControls can be applied with available time, access, tools and resources.Impossible instructions encourage deviation and workarounds.The required test point cannot be reached safely.
ParticipativePeople who understand the work contribute to development and review.Worker knowledge reveals practical hazards and foreseeable shortcuts.Written remotely without observing or discussing the job.
Role-definedAuthority, responsibility and handover are explicit.Prevents gaps, overlaps and assumptions.Everyone believes someone else verified the isolation.
Competence-basedRequired knowledge, skill, experience and supervision are stated.The same instruction may not be safe for an inexperienced person.“Trained person” is stated but competence is never checked.
Human-centredIt considers workload, fatigue, usability, communication and predictable error.Controls must work in real human conditions.Critical information is buried, contradictory or unreadable.
Inclusive and accessibleUsers can find, read and understand it.Language, literacy, disability or unfamiliar terminology can affect safe use.Only one complex-language copy exists away from the workplace.
Abnormal-condition readyIt states when to stop, withdraw, isolate, escalate or use emergency arrangements.People must not improvise when normal conditions disappear.No instruction for a leak, failed test or unexpected pressure.
Controlled and currentOnly the approved version is available and changes are traceable.Obsolete instructions may conflict with modified plant or controls.Different versions are posted at the same workplace.
VerifiedCritical controls are checked before reliance.An assumed control may be absent, failed or incorrectly applied.The permit is signed without a field check.
Monitored and reviewedUse and effectiveness are checked over time and after triggers.Work, people, equipment and evidence change.Repeated deviations are normalised without investigation.

When a detailed written system is normally needed

  • Significant or high-risk work
  • Complex or non-routine tasks
  • Several people, teams or contractors
  • Critical sequence, isolation or verification
  • Permit-controlled work
  • Serious foreseeable abnormal conditions

When simplicity may be appropriate

Straightforward low-risk work may be controlled through concise instruction, training and normal supervision. Simplicity must come from low complexity—not from ignoring significant hazards. The arrangement still needs to be understood and effective.

How the Six Risk Strategies Shape the System of Work

These strategies are not six competing documents. They influence different decisions during development, authorisation and assurance.

StrategyWhen it is usedPump-maintenance applicationWhat the SSOW or SOP must show
AvoidanceThe exposure or activity can be removed or redesigned.Use remote condition monitoring to avoid unnecessary intrusive inspection.Why the hazardous step is no longer required and whether the alternative creates new risk.
ReductionNecessary work continues under stronger controls.Isolate, depressurise, drain, purge, verify, segregate and supervise.Control hierarchy, sequence, responsibilities, verification and residual risk.
TransferSpecialist competence, equipment or financial sharing is appropriate.Use a competent specialist for seal replacement or hazardous cleaning.Selection, information, coordination, interfaces, monitoring and retained duties.
AnalysisCauses, exposure, failure paths or uncertainty require deeper understanding.Analyse chemical residue, pressure, isolation effectiveness and previous failures.Evidence, assumptions and how findings affected the method.
EvaluationEvidence must be compared with criteria to decide adequacy and authorisation.Compare proposed precautions with legal, technical, manufacturer and site requirements.Acceptance criteria, decision authority and unresolved gaps.
ReviewTime, change, incident, feedback or failed control may affect validity.Revise after a leak, near miss, plant modification, contractor concern or recurring deviation.Review triggers, owner, evidence, revised version and communication.
Essential distinction: analysis, evaluation and review support the decision; they do not physically isolate energy or prevent exposure. Transfer may allocate some work or financial impact, but it does not make the hazard or all responsibilities disappear.

Tool: Build a Combined Risk-Strategy Route

Select one or more strategies and let the portal test whether the combination fits the scenario.
Interactive document workshop

Build an Educational Safe System of Work Draft

Adjust every field to your own workplace example. The output helps you understand the structure; it must be reviewed by competent people against the real workplace and applicable requirements before use.

Adjust the fields, then generate a structured learning draft.

Build an Educational Safe Operating Procedure Format

An SOP should tell the right person what to do, in what order, under which conditions, and when to stop. It should not ask the user to make undefined safety decisions during a critical step.

Complete the procedure fields, then generate the format.

Tool: Put the Pump-Maintenance Controls in a Defensible Order

Use the arrow buttons to move each step. The “correct” route is the approved teaching sequence for this example; a real installation may require a different, technically validated sequence.

    Move the steps, then check whether critical preparation, verification and recommissioning controls appear in the right order.
    Complete control-of-work concept

    Permit-to-Work—PTW: Meaning, Purpose, Issue, Use, Handback and Closure

    What is PTW?

    Pronounced: “P-T-W,” meaning permit-to-work.

    A PTW is a formal, time-limited system for authorising specified work at a defined place, on identified plant or equipment, by named or competent parties, subject to stated precautions and conditions.

    It communicates an agreement: this exact work may proceed, within this exact boundary, while these verified conditions remain true.

    What PTW does not mean

    A permit is not a risk assessment, an instruction to begin automatically, a substitute for isolation, a certificate that danger has disappeared, or a transfer of all responsibility to the worker.

    Issuing the paper alone does not make a job safe. The assessment, SSOW, competence, communication, physical controls, field checks, supervision and stop-work response must all function.

    Requested / draftThe job is being scoped and assessed. It is not permission to start.
    Issued / accepted / activeThe authorised issuer and performing party have verified, communicated and accepted the permit conditions.
    SuspendedWork has stopped and the permit is not active—for example after alarm, shift change, changed condition or control loss.
    Completed / handed back / closedThe work party declares completion; the area is checked, plant status is handed back, and the permit is formally cancelled or closed.

    Why is a permit-to-work used?

    A PTW creates disciplined communication where mistakes in plant identity, isolation, timing, coordination or handover could cause serious harm. It defines ownership, prevents incompatible simultaneous activities, records critical precautions, controls the period of work and manages the return to normal operation.

    Work categoryWhy formal control may be neededTypical linked controls or certificates
    Hot workFlame, arc, spark or heat may ignite flammable material or damage adjacent systems.Gas testing, area preparation, fire protection, fire watch and post-work monitoring.
    Confined-space or vessel entryAtmosphere, engulfment, restricted access, energy and rescue hazards can change rapidly.Isolation certificate, atmospheric test, ventilation, entry log, attendant and rescue plan.
    Line breaking / hazardous containmentOpening pipework or equipment can release pressure, temperature, toxic, corrosive or flammable material.Process isolation, drain/vent/purge, decontamination, test and line-break controls.
    Electrical or mechanical workContact, arc, unexpected start, gravity, pressure or stored energy may be fatal.Isolation plan, lockout/tagout, prove-dead or zero-energy verification and controlled reinstatement.
    Excavation / ground disturbanceUnderground services, collapse, water, contaminated ground and vehicle interaction may be present.Service drawings, detection and marking, trial holes, shoring, access and inspection.
    Other site-defined workWork at height, lifting, radiography, roof access, energised testing or unusual simultaneous work may need coordination.Site-specific permits, certificates, exclusion zones, specialist plans and interfaces.
    Do not assume permit categories are identical everywhere. The organisation’s control-of-work procedure and applicable legal/technical requirements determine which work needs a permit, who may issue or accept it, and which supporting certificates are required.

    Who is involved?

    Typical roleEasy meaningMain responsibility
    Area / operating authorityThe person controlling the plant or area.Confirms operating status, interfaces and whether the area can be released and later accepted back.
    Permit issuer / issuing authorityThe competent authorised person who issues the permit.Checks scope, assessment, precautions, isolations, conflicts, validity and field conditions before authorising.
    Performing authority / permit receiverThe person accepting the permit for the work party.Understands the permit, briefs the team, keeps within boundaries, monitors conditions and stops when conditions change.
    Isolating authorityThe person controlling required energy or process isolations.Applies, records, proves and later removes isolation under the approved process.
    Authorised gas testerA competent person approved to test atmosphere.Uses suitable equipment, records results and understands limits, frequency and conditions of testing.
    Work partyThe persons carrying out the authorised task.Attend the briefing, follow the SSOW/SOP and permit, protect controls and report change or uncertainty.
    Permit / SIMOPS coordinatorThe person who sees the whole work picture.Prevents conflicts between permits, operations, contractors, isolations and emergency arrangements.

    Terminology varies: a site may use different role names. The essential point is that authority, competence, accountability, communication and handover cannot be vague.

    The Full PTW Lifecycle—14 Phases

    Select a phase to see what must happen, why it matters and what evidence should exist.

    Phase 1 — Request and planDescribe the proposed job, reason, location, equipment, work order, timing, people and likely interactions. The request starts planning; it is not permission to begin.

    What should a complete permit form contain?

    Permit fieldInformation requiredControl purpose
    IdentificationPermit number/type, work order, exact plant/equipment tag, location and description.Prevents work on the wrong item or outside the authorised task.
    ValidityIssue date/time, start, expiry, shift and any rules for extension or revalidation.Stops an old permit being treated as continuing permission.
    Supporting documentsRisk assessment, SSOW, SOP/method, drawings, certificates and rescue/emergency plans.Connects authorisation to the technical control basis.
    Hazards and interfacesEnergy, substances, atmosphere, access, environment, nearby work and SIMOPS.Makes foreseeable interactions visible to both parties.
    Isolations and testsIsolation points/certificate, lock and tag references, drain/vent/purge, test type, result, time and tester.Provides traceable evidence of critical plant preparation.
    PrecautionsBarriers, ventilation, fire controls, access, tools, PPE, monitoring and prohibited actions.Defines conditions that must remain in place.
    Emergency and communicationAlarm, withdrawal, rescue, contact, stop-work rule, briefing and handover method.Supports response when normal assumptions fail.
    Authorisation and acceptanceIssuer and receiver names/signatures, date/time and declarations of understanding.Confirms that authority and shared understanding are explicit.
    Suspension / extension / handoverReason, safe state, new conditions, outgoing/incoming parties and revalidation.Prevents work continuing across a change without control.
    Completion and handbackWork complete/incomplete, people/tools cleared, guards restored, plant status, inspection and acceptance.Controls transfer back to operations and reinstatement.
    Cancellation and recordsClosure time, permit cancellation, linked documents closed, defects/actions and retained record.Ends authority clearly and preserves evidence for audit and learning.

    Interactive: Is the Permit Ready for Issue?

    Tick only items verified by evidence. A high total cannot compensate for a missing critical condition.

    Do not issue yet. Confirm each condition using the real worksite and approved control-of-work procedure.

    Interactive: Shift, Alarm and Change Decision

    Choose the event, then decide whether the permit remains controlled, is suspended, needs revalidation or can proceed to handback.

    Interactive: Build a Complete PTW Learning Form

    Complete every field. This produces a teaching draft—not a workplace permit or authorisation.

    Complete and check the form, then generate the structured learning draft.

    Interactive: Does the Work Need Formal Permit Control?

    A permit-to-work is a formal authorisation and communication system for defined work. Select the features that apply. Site procedures and applicable law make the final decision.

    Interactive: Stress-Test the Working Arrangement

    Rate each condition from 1 (weak) to 5 (strong). This is a learning diagnostic, not a risk-acceptance formula.

    4
    4
    3
    4
    5
    3

    PTW Knowledge Check

    1. What is the main purpose of PTW?
    2. What does issuing a permit automatically do?
    3. What must happen before issue?
    4. The job continues into a new shift. What is required?
    5. What if plant conditions or work scope change?
    6. What completes the PTW lifecycle?
    Answer all six questions, then check your understanding of the full permit lifecycle.

    Tools: SSOW Quality Diagnostic and Review Trigger

    Tool 1: Check the Quality of an Existing SSOW

    Select only what is genuinely present and effective in the system you are reviewing.

    Tool 2: What Should Trigger Review?

    Choose the trigger and current status, then decide whether work should stop, be restricted or continue under the approved system.

    Weak Wording Versus Defensible Wording

    Weak wordingWhy it failsStronger wording principle
    “Make the pump safe.”No person, isolation method, condition or verification is defined.Name the equipment, energy sources, authorised role, approved isolation method and verification requirement.
    “Wear proper PPE.”“Proper” is undefined and PPE may not control the main hazard.Select PPE from the assessment after applying higher-order controls; state type, limitation and checks.
    “Be careful when opening.”It transfers responsibility to behaviour without controlling stored pressure or residue.Prevent opening until depressurisation, drainage, safe condition and authorisation are verified.
    “Experienced workers only.”Experience is not defined or verified.State required authorisation, task knowledge, skill, experience and supervision.
    “In an emergency, act accordingly.”No stop, alarm, withdrawal or escalation route is given.Define credible abnormal conditions and the immediate response expected from each role.
    “Review annually.”Waits for a date even after a change, failure or near miss.Use both scheduled and event-triggered review.
    Assessment-ready learning support

    How to Explain 2.3 at Level 6

    A strong explanation normally contains

    1. Clear definitions of SSOW and SOP.
    2. The relationship with risk assessment and supporting documents.
    3. A connected development lifecycle.
    4. Characteristics explained with reasons—not only listed.
    5. Application of the six risk strategies.
    6. How risk assessment, SSOW, SOP, PTW, isolation and handback connect.
    7. The PTW lifecycle from request and field verification to suspension, handback and closure.
    8. A suitable workplace example.
    9. Limitations, implementation and review arrangements.

    Use this paragraph structure

    Point → Meaning → How developed/applied → Why it matters → Workplace example → Consequence if missing → Review.

    Write in your own professional words and relate the explanation to a genuine or realistic workplace. A list of headings alone does not fully satisfy “explain.”

    Practice question

    Explain how a safe system of work and supporting safe operating procedures should be developed for intrusive maintenance of a chemical-transfer pump. Your response should address consultation, risk strategies, control sequence, competence, human factors, permit interfaces, abnormal conditions, implementation and review.

    Section 2.3 knowledge check

    Can You Convert a Risk Decision Into Safe Work?

    1. What best describes an SSOW?
    2. What is the main purpose of an SOP?
    3. What should development begin with?
    4. Why is worker consultation important?
    5. What does a permit-to-work NOT do?
    6. What is a hold point?
    7. Intrusive pump maintenance must continue. What is normally the leading strategy?
    8. Which is a human-factor characteristic?
    9. Which is an event-triggered review?
    10. Which best demonstrates the command word “explain”?
    Answer all ten questions, then check your score and explanations.
    Unit 3 · Section 03

    Understand the Models of Loss Causation, Analysis of Loss Data and the Importance of Incident Investigation

    Section 03 follows one connected learning journey: understand why an event happened, test what the data reveals, assess who needs to receive the report, and explain how investigation converts evidence into prevention.

    Current learning route: 3.1 and 3.2 provide the causal and quantitative foundation. Sections 3.3 and 3.4 apply that foundation to reporting decisions, evidence-led investigation and verified risk reduction.

    Learning outcomes covered: 3.1 Outline loss-causation theories and techniques. 3.2 Justify quantitative methods in analysing loss data. 3.3 Assess the needs and impacts of reporting loss events. 3.4 Explain the importance and impact of incident investigations.

    How every part connects

    One Event · Four Connected Questions

    Do not study these outcomes as four separate chapters. Each part produces the evidence or decision needed by the next one.

    Event → immediate control → report and record → analyse the data → preserve evidence → investigate causes → improve controls → verify effectiveness → strengthen future learning
    Use the learning line: a report makes the event visible; investigation tests the evidence; cause-linked action reduces risk; verification shows whether the change works in real operations.
    Start with the language

    What Do “Loss” and “Causation” Mean?

    Loss means an unwanted outcome that removes or damages something of value. It may include injury, ill health, death, environmental harm, property damage, production interruption, legal exposure, financial cost, lost information or damaged trust.

    Causation means the way conditions, decisions, actions, failures and interactions combine to produce an event or outcome.

    An incident normally has more than one relevant cause. The final action or failed component may be easy to see, but a Level 6 analysis asks what shaped that action, why the control was absent or ineffective, and what management-system conditions allowed the vulnerability to remain.

    The central learning rule

    A model organises thinking; it does not manufacture evidence.

    Investigators must still preserve the scene, gather reliable evidence, test competing explanations, consult involved people, identify controls and verify corrective action.

    Immediate cause

    The unsafe act, condition, energy transfer or equipment state directly connected to the unwanted event.

    Ask: What directly triggered or enabled the contact?

    Underlying cause

    The job, workplace or organisational factor that allowed the immediate condition or action to arise.

    Ask: What influenced the work and weakened control?

    Root cause

    A deeper management-system or organisational failing whose correction can prevent a wider class of recurrence.

    Ask: Why did the system create, accept or fail to detect the vulnerability?

    Interactive Section 03 Jargon Translator

    Select a term to hear how it is pronounced and understand why it matters.

    LossPronounced: “loss.” The harmful or unwanted consequence of an event. Loss can affect people, environment, assets, operations, finance, compliance or reputation.
    One event · several ways to understand it

    Our Master Case: Forklift Contact With a Process Line

    A forklift enters a congested transfer-area route, passes a damaged low barrier and contacts a valve manifold. A small chemical release occurs. The alarm operates, workers withdraw and the response team isolates the area. No model should be used to blame the driver or to assume a cause before evidence is gathered.

    Industrial loading area before a forklift incident, showing interacting conditions around the route and process pipe
    1. Before the eventLook for conditions: layout, visibility, barrier condition, workload, traffic interaction, supervision and competing demands.
    Controlled aftermath of a forklift contact with a process line while workers withdraw and responders secure the area
    2. The loss eventSeparate the hazardous contact from its consequences. Notice which prevention, detection, mitigation and emergency barriers worked or failed.
    Multidisciplinary investigation team reviewing incident evidence, controls and loss data
    3. Investigation and learningCombine physical evidence, people’s accounts, documents, data and models. Test explanations before deciding actions.
    Evidence sourceInitial findingWhat must still be tested?
    CCTV and scene measurementsPallets narrowed the route; forklift contacted the low pipe barrier.Actual speed, visibility, pedestrian interaction and why storage entered the route.
    Inspection recordsBarrier damage had been recorded twice but not permanently repaired.Risk classification, escalation, ownership, resources and closure verification.
    Driver and worker interviewsPeak dispatch created queuing and radio interruptions.Work-as-done, production pressure, route rules, competence and normal adaptations.
    Alarm and response logDetection and emergency isolation limited the release.Alarm timing, response reliability, exposure and opportunities to strengthen recovery.
    Six-month loss dataVehicle–route near-miss reports increased, especially during peak dispatch.Reporting quality, exposure hours, location clustering and whether risk really increased.
    Learning outcome 3.1
    Central question: Why did the event happen?

    Outline a Range of Loss-Causation Theories and Techniques

    Outline means present the principal features and show the basic structure, purpose and application of each theory or technique. A Level 6 response should also distinguish models, use a relevant example and recognise important limitations.

    3.1
    Historical pattern studies

    Accident and Incident Ratio Studies

    Ratio studies arrange recorded events by consequence level. They helped organisations recognise that low-consequence events and near misses can reveal control weaknesses before a major loss occurs.

    Teaching illustration of H. W. Heinrich and Frank E. Bird Jr. beside historical safety records
    AI-generated teaching illustration—not an archival photograph. Heinrich is shown on the left; Bird on the right.

    Who, when, why—and what changed?

    1931 — HeinrichH. W. Heinrich published Industrial Accident Prevention: A Scientific Approach and presented the historical 1:29:300 relationship.
    1959–1965 — BirdFrank E. Bird Jr.’s Lukens Steel study examined about 90,000 incidents and reported a 1:100:500 pattern.
    1969 studyBird and colleagues analysed 1,753,498 accident reports from 297 companies and produced the widely taught 1:10:30:600 ratio.
    Modern useThe triangle is now best treated as a prompt for reporting and investigation—not a law, prediction or promise that minor-event reduction will automatically prevent catastrophe.

    Why it was created: to show that the serious injury at the top is only part of the recorded experience. Lower-consequence events may provide more frequent opportunities to discover exposure and weak controls.

    How thinking changed: Bird widened the categories to include property damage and near misses. Later research showed that the shape changes with definitions, industry, severity threshold and reporting practice, and that fatal and non-fatal events can follow different causal pathways.

    Historical context: DNV tribute to Frank Bird. Critical evidence: Salminen, Saari, Saarela and Räsänen (1992) and Marshall, Hirmas and Singer (2018).

    Heinrich’s historical triangle

    Often presented as 1 major injury : 29 minor injuries : 300 no-injury accidents. Read the colon “:” as “to”: one major injury to 29 minor injuries to 300 no-injury accidents. It was derived from historical insurance and accident records and promoted attention to the larger body of less-serious events.

    Bird’s historical ratio

    Often presented as 1 serious or major injury : 10 minor injuries : 30 property-damage events : 600 near-miss incidents. It broadened attention to damage and no-loss events. The figures describe Bird’s historical dataset; they do not calculate the probability of the next accident.

    Why use a ratio study? Use it to test whether people report near misses, identify recurring event types and decide where investigation may prevent loss. Do not multiply 600 near misses and claim that one major injury must follow. A ratio summarises a dataset; it is not a countdown.
    Essential caution: These historical ratios are not universal laws, probability predictions or targets. Event definitions, industry, hazard type, reporting culture and data quality change the observed pattern. Major-accident hazards may develop without a large visible base of minor personal injuries.
    Potential benefitWhy it helpsLimitation to explain
    Encourages near-miss reportingWeak signals can reveal exposure and failing controls before serious harm.More reports can mean better trust and reporting—not necessarily worsening safety.
    Supports preventionRecurring lower-level events can direct inspection and improvement.Preventing minor slips does not automatically control a low-frequency catastrophic process event.
    Communicates scale simplyThe triangle is memorable and helps introduce proactive learning.Simplicity may hide different causal pathways and consequence mechanisms.
    Provides trend categoriesOrganisations can compare reporting levels and event types over time.Changed definitions, workforce hours or reporting systems can create a false trend.
    Challenges injury-only thinkingDamage and near misses can contain valuable control information.A ratio does not replace investigation, risk assessment or barrier assurance.

    Interactive Tool: See How Reporting Culture Changes the Triangle

    The “true opportunities for learning” remain constant in this teaching example. Adjust the percentage that gets reported. Notice how the visible triangle changes even when the underlying events do not.

    25%
    70%
    90%
    Adjust a reporting slider to reveal the observed ratio.
    From a chain to a system

    Bird’s Loss-Causation Model and Multi-Causality

    Bird’s expanded domino approach connects management control with basic causes, immediate causes, the incident and the final loss. Removing or strengthening an earlier “domino” can interrupt the sequence.

    Where it came from

    Bird developed accident-prevention thinking beyond the earlier person-centred domino sequence. The model placed lack of management control at the beginning and widened “injury” into loss, including harm to people, property, environment and production. Bird’s later work with George Germain was published in Practical Loss Control Leadership in 1985.

    How it developed

    The sequence was adapted using energy-exchange concepts associated with William Haddon. Read it left to right to explain how loss developed; work from right to left during investigation to ask what control should have interrupted each step.

    Lack of controlWeak standards, responsibilities, planning, monitoring or correction within the management system.
    Basic causesJob factors and personal factors that shape exposure and performance.
    Immediate causesObservable substandard acts and conditions directly connected with the event.
    Incident or contactThe transfer of energy, substance or force—or loss of control—that creates the event.
    LossInjury, illness, environmental harm, damage, interruption or another unwanted consequence.
    Applied to our case: overdue corrective maintenance and weak closure assurance → damaged barrier and congested route → forklift enters a vulnerable path → vehicle contacts the manifold → small release and operational loss.
    Level 6 caution: A domino picture can appear too linear. Real events often involve feedback, several simultaneous pathways, successful controls and changing conditions. Use the five stages as organising headings, then use multi-causality to avoid forcing the evidence into one neat chain.

    Multi-causality: More Than One Path Can Matter

    Multi-causality rejects the idea that one unsafe act is normally a complete explanation. Several conditions may combine, interact or increase one another’s effect. Causes can exist at task, equipment, environmental, individual, supervisory and organisational levels.

    Immediate examples

    • Forklift contacts low barrier.
    • Route is narrowed by pallet storage.
    • Valve manifold remains exposed to vehicle energy.

    Underlying examples

    • Peak traffic and pedestrian interaction were not reassessed.
    • Temporary storage became normal.
    • Damaged barrier repair was delayed.

    Root examples

    • Defect priority criteria ignored major-consequence potential.
    • No effective owner verified corrective-action closure.
    • Layout-change governance excluded operations and workforce evidence.

    Interactive Tool: Classify the Cause—Then Look Deeper

    Choose the best classification, then read the explanation.
    Reason’s model of organisational accident causation

    The Swiss Cheese Model

    James Reason described safety as several layers of defence, barrier and safeguard. Each layer can contain weaknesses—shown as “holes.” An adverse trajectory can pass through when weaknesses in different layers align.

    James Reason: from human error to organisational defences

    Who and when: British psychologist James Reason developed the model across the 1990s, beginning with ideas in Human Error (1990) and refining the familiar defence-layer image through the decade. Safety practitioner John Wreathall also influenced its development.

    Why: Reason wanted to move investigation beyond “the operator made an error.” The model asks how front-line actions combine with weaknesses created earlier by design, staffing, maintenance, supervision and organisational decisions.

    How it changed: diagrams and terminology evolved between 1990 and 2000. This matters: Swiss Cheese is a family of evolving explanations, not one frozen diagram. Later safety thinking also stresses that defences interact dynamically and that a simple line through holes must not replace evidence.

    Read the system approach in Reason (2000), Human error: models and management, and the historical critique in Larouzee and Le Coze (2020).

    Teaching illustration of James Reason and H. A. Watson with defence layers and logic diagrams
    AI-generated teaching illustration—not an archival photograph. Reason is shown on the left; Watson on the right.

    Active failure

    An action or omission close in time and place to the event, such as an incorrect control input or missed check. It may trigger the event, but it is rarely the whole explanation.

    Latent condition

    A deeper weakness created by design, staffing, maintenance, procurement, priorities, procedures, supervision or management decisions. It may remain hidden until combined with local conditions.

    Defence layerA safeguard intended to prevent the event, detect it or reduce its consequence.
    HoleA relevant weakness in that layer. A hole can move or change as conditions, resources and work practices change.
    TrajectoryThe path by which a hazard passes through aligned weaknesses and reaches a harmful outcome.
    Person approach versus system approach: “The driver made an error” closes the question too early. A system approach asks what task, equipment, environment and organisational conditions shaped the action, and which defences should have prevented or contained its effect.

    Interactive Barrier Alignment Simulator

    Mark a layer as failed or ineffective. The event pathway opens only when every selected defence in this simplified example has a relevant weakness.

    Five defence layers availableA single weakness need not produce loss if another independent and effective layer interrupts the pathway.

    Limitation: The cheese image is a communication model, not a detailed dynamic simulation. It can oversimplify interactions unless each layer, threat, dependency, owner and performance requirement is defined with evidence.

    Structured logic techniques

    Fault Tree Analysis and Event Tree Analysis

    Teaching illustration of James Reason and Bell Laboratories engineer H. A. Watson with early fault-tree logic
    AI-generated teaching illustration—not an archival photograph. H. A. Watson is represented on the right.

    How structured logic entered safety analysis

    Fault Tree Analysis: commonly traced to H. A. Watson at Bell Laboratories in 1962 during the Minuteman missile programme. It was created because complex systems could fail through combinations that a simple checklist might miss.

    Event Tree Analysis: developed as a forward, consequence-oriented partner to fault-tree reasoning. Probabilistic risk work such as the US Reactor Safety Study, WASH-1400 (1975), helped establish combined fault-tree and event-tree methods for complex high-hazard systems.

    What changed: the techniques expanded from defence and nuclear applications into aviation, process safety and other industries. Software can now evaluate very large trees, but the logic, data, dependencies and uncertainty still need competent human review.

    Authoritative background: NRC Fault Tree Handbook and NRC history of WASH-1400.

    Read this before any equation

    FTA and ETA Symbol Decoder + Logic-Gate Starter

    Symbols are a form of shorthand. They make a large analysis easier to read, but only after every symbol has been defined. Start by reading the symbol aloud, identify what it represents, check its unit or range, and then ask why it is present in the equation.

    The important difference between capital N and small n: capital N normally means the total number of valid opportunities, demands, tests or items in the denominator. Lowercase n normally means the number actually counted. Therefore, in P(A) = nA ÷ N, nA is the number of times event A occurred and N is the total number of relevant opportunities. Symbols are local conventions, however: in a “k-out-of-n” gate, lowercase n means the total number of components in that gate. Always read the symbol key written for the equation being used.

    First: know what every common mark means

    A, B, CPronounced: “event A, event B, event C.”

    Letters name defined events. A might mean “detector fails”; B might mean “isolation valve fails.” The letter has no meaning until the analyst defines it.

    NPronounced: “capital N.”

    The total number of relevant demands, tests, opportunities or observations. It is commonly the denominator. Example: N = 100 valid detector tests.

    nPronounced: “lowercase n.”

    A count from the defined set. Example: n = 3 recorded failures. A subscript makes the count more specific, such as nF for number of failures.

    nA, nFPronounced: “n sub A” and “n sub F.”

    The subscript is a label, not multiplication. nA counts event A; nF counts failures. Thus nF ÷ N means failures divided by all valid opportunities.

    P(A)Pronounced: “probability of A.”

    The probability that defined event A occurs on the stated demand, mission or period. Probability has no unit and must lie from 0 to 1 inclusive.

    p, qPronounced: “p” and “q.”

    p commonly represents success probability and q the complementary failure probability. When success and failure cover all possibilities, q = 1 − p and p + q = 1.

    piPronounced: “p sub i.”

    i is an index meaning “the particular item or branch being considered.” p1, p2 and p3 can represent three different input probabilities.

    f, fIPronounced: “frequency” and “f sub initiator.”

    f is an event frequency with a unit such as per year. fI is the initiating-event frequency used at the start of an event tree.

    T, tPronounced: “capital T” and “lowercase t.”

    T often represents total observed exposure time; t often represents the mission time being evaluated. The analyst must state the unit—hours, days or years.

    λPronounced: “lambda.”

    A failure rate, normally stated per unit time. A value of λ = 0.002 per hour means the assumed rate basis is 0.002 failures per operating hour—not a 0.2% certainty for every hour.

    ePronounced: “Euler’s number” or simply “e.”

    A mathematical constant approximately equal to 2.71828. In e−λt, it supports the exponential reliability model; it does not mean an event count.

    S / FPronounced: “success” and “failure.”

    ETA branches are often labelled S and F. These labels say whether the defined barrier performs its required function at that branch point.

    ∩ / ∪Pronounced: “intersection” and “union.”

    ∩ means AND—events occur together. ∪ means OR—at least one of the defined events occurs, including the possibility that both occur.

    P(B | A)Pronounced: “probability of B given A.”

    The vertical bar means “given that.” It is used when B’s probability is evaluated under the condition that A has already occurred.

    ΣPronounced: “capital sigma” or “sum.”

    Add the listed values. In ETA, Σfbranch means add the frequencies of the mutually exclusive terminal branches.

    ∏Pronounced: “capital pi” or “product.”

    Multiply the listed values. ∏pi means p1 × p2 × p3 and so on; it is not the circle constant π.

    =, ×, ÷Pronounced: “equals,” “multiplied by,” and “divided by.”

    Equals states that both sides represent the same value. Multiplication combines required path factors. Division creates a proportion or rate from a count and denominator.

    1 − pPronounced: “one minus p.”

    The complement of p. If p is the probability of success, 1 − p is the probability of failure only when the two states are mutually exclusive and cover all defined possibilities.

    [ ] and ( )Pronounced: “brackets” and “parentheses.”

    Calculate the expression inside them first. They group terms and prevent the calculation from being performed in the wrong order.

    0 ≤ P ≤ 1Pronounced: “probability is greater than or equal to zero and less than or equal to one.”

    Zero means impossible within the defined model; one means certain within that model. Multiply a decimal by 100 to express it as a percentage.

    Second: what is a logic gate and why do we need it?

    A logic gate is a rule that explains how input events combine to create an output event. It is not a physical gate and it does not prove causation by itself. It converts a full English sentence into an exact visual instruction, helping an FTA remain consistent when many failure paths are connected.

    OR gate — “at least one is enough.”
    If detector failure by itself can cause the output, or valve failure by itself can cause it, connect them through OR. A normal inclusive OR also allows both events to occur.
    AND gate — “all stated inputs are required.”
    If the output occurs only when the detector fails and the automatic isolation also fails, connect them through AND. AND describes a required combination, not automatically a time sequence.
    k-out-of-n voting gate.
    The output occurs when at least k of n similar channels meet the stated condition. A 2-out-of-3 trip system needs any two of its three channels. Here n means total channels inside this gate.
    Advanced gates.
    NOT, exclusive-OR, inhibit and priority-AND can express special logic or sequence conditions. Use them only when their exact meaning is required and defined; most introductory FTAs begin with AND and OR.
    Event AEvent BA AND BA OR BPlain meaning
    0 — does not occur0 — does not occur00No input occurred.
    01 — occurs01Only B occurred; OR opens, AND does not.
    1 — occurs001Only A occurred; OR opens, AND does not.
    1111Both occurred; both rules are satisfied.
    FTA versus ETA: FTA mainly uses logic gates to reason backwards from a top event. ETA normally does not join causes through AND/OR gates; it moves forwards from one initiating event and divides into conditional success/failure branches. We multiply along one ETA path and add mutually exclusive end-state paths when they belong to the same outcome group.

    Interactive Logic-Rule Coach

    Choose the English statement you need to represent. The coach will identify the logic and explain how the symbols are used.

    Choose a statement to decode its logic.
    Five-pass formula-reading habit: (1) name the outcome on the left of “=”; (2) pronounce every symbol; (3) replace each symbol with its defined value and unit; (4) calculate brackets first, then multiplication/division, then addition/subtraction; and (5) explain what the answer means, its assumptions and what decision it supports.
    Begin before the tree

    Where Does the Probability Come From?

    FTA and ETA do not create probability simply because a box is drawn. Every input needs a defined event, population or equipment item, time or demand basis, data source and uncertainty. The first question is therefore not “Which formula shall I use?” It is “What exactly does this number describe?”

    P(A) = nA ÷ NPronounced: “probability of A equals number of A events divided by total relevant opportunities.”

    Use this empirical estimate when each opportunity is clearly defined. If a detector failed 3 of 100 valid proof-test demands, the observed failure-on-demand estimate is 3 ÷ 100 = 0.03, or 3%.

    q = 1 − pPronounced: “q equals one minus p.”

    If p is success probability, q is failure probability. A barrier that succeeds with probability 0.90 has a complementary failure probability of 1 − 0.90 = 0.10.

    P(B | A)Pronounced: “probability of B given A.”

    The vertical bar means given that. An ETA branch asks for the chance of the next success or failure after the initiating event and earlier branch conditions have occurred.

    F(t) = 1 − e−λtPronounced: “F of t equals one minus e to the power minus lambda t.”

    Under a justified constant-rate exponential model, this estimates the probability of at least one failure by mission time t. It is not suitable automatically for ageing, repair, dependence or changing conditions.

    f = n ÷ TPronounced: “frequency equals number of events divided by exposure time.”

    Frequency carries a unit such as events per year. Probability has no unit and stays between 0 and 1. A frequency of 0.5 per year must not automatically be called a 50% annual probability.

    P(A ∩ B)Pronounced: “probability of A intersection B.”

    This means A and B occur together. The general rule is P(A ∩ B) = P(A) × P(B | A). It becomes P(A) × P(B) only when independence is defensible.

    Possible sourceHow the value may be obtainedEssential quality question
    Operating or test dataDefined failures ÷ valid demands, or events ÷ exposure time.Are definitions, equipment, conditions and reporting sufficiently comparable?
    Reliability modelUse a justified distribution, failure rate and mission time—for example 1 − e−λt.Do the model assumptions fit ageing, repair, maintenance and operating conditions?
    Fault-tree calculationA barrier-failure probability used in an ETA may itself come from a detailed FTA.Were dependencies, shared utilities, human actions and common causes represented?
    Expert judgement or analogous dataElicit and document a defensible estimate or range when direct data are sparse.Are the experts, evidence, assumptions, bias controls and uncertainty traceable?

    Interactive Probability Source Calculator

    Choose one method. The portal will use only the fields required for that method and explain the units and assumptions.

    Choose a probability-source method.
    Level 6 interpretation: A calculated value is not automatically valid evidence. State the event definition, denominator, period, units, data provenance, assumptions, missing information and uncertainty. Use sensitivity analysis to see whether a reasonable change in an uncertain input changes the decision.

    Authoritative learning references: US NRC glossary definitions for fault trees and event trees, US NRC explanation of probabilistic risk assessment and NASA system-safety learning on logic models and probability.

    Fault Tree Analysis — FTA

    Pronounced: “F-T-A.” A Fault Tree Analysis is a structured, top-down method. Start with one precisely defined unwanted top event, then reason backwards to identify the equipment failures, human failures, external events and combinations capable of producing it.

    1. Define top event2. Set boundary3. Ask how4. Connect gates5. Check + quantify
    Think of FTA as asking: “The unwanted event is here. What must have happened beneath it?” A basic event is the lowest event taken forward for data or action. An intermediate event is created by lower events. A minimal cut set is a smallest combination of basic events sufficient to produce the top event.
    Top event: vehicle reaches hazardous manifold
    OR
    Failure path A: vehicle enters prohibited zone
    Failure path B: impact barrier fails when demanded
    Enter independent teaching probabilities.
    How the two gates differ: With an AND gate, every stated input is needed, so independent probabilities are multiplied. With an OR gate, any input can produce the output. Exact addition must remove the overlap where both occur; the complement method does this automatically. Choose the gate from the real causal logic—never choose it merely to obtain a preferred answer.
    Open the FTA equation and symbol guide
    P(A)Pronounced “P of A.” Probability that input event A occurs during the defined mission or period.
    P(B)Pronounced “P of B.” Probability that input event B occurs during the same basis.
    ×Pronounced “multiplied by.” For independent events, AND probability is P(A) × P(B).
    1 − P(A)Pronounced “one minus P of A.” The complement: probability that A does not occur.
    AND gateP(A ∩ B) = P(A) × P(B) when A and B are independent. The symbol ∩ is pronounced “intersection.”
    OR gateP(A ∪ B) = 1 − [1 − P(A)][1 − P(B)]. The symbol ∪ is pronounced “union.”
    General AND ruleP(A ∩ B) = P(A) × P(B | A). Use the conditional value when B’s chance changes because A occurred.
    General OR ruleP(A ∪ B) = P(A) + P(B) − P(A ∩ B). Subtract the intersection because simple addition counts “both” twice.
    n-input ANDP(top) = ∏pi for independent required inputs. ∏ is capital pi and means multiply the listed probabilities.
    n-input ORP(top) = 1 − ∏(1 − pi) for independent alternative inputs.
    Why calculate it? To identify combinations that contribute most to the top event, test design alternatives and prioritise reliability improvement. If inputs share a power supply, environment or maintenance error, independence is false and simple multiplication may understate risk.

    Real trees require validated logic, common-cause and dependency checks, suitable data, minimal cut sets and sensitivity analysis.

    Event Tree Analysis — ETA

    Pronounced: “E-T-A.” An Event Tree Analysis is a structured, forward or inductive method. Start with a defined initiating event, then move forwards through the conditional success or failure of safeguards, operator actions and recovery measures until every modelled path reaches an end state.

    1. Define initiator2. Order barriers3. Split branches4. Name end states5. Calculate + sum
    Think of ETA as asking: “The initiating event has occurred. What may happen next?” Every branch probability is conditional on arriving at that branch. The terminal paths should be mutually exclusive, and together they should cover the defined possibilities.
    Initiating event: contact causes leak
    Detection succeeds / fails
    Isolation succeeds / fails
    Enter an initiating frequency and conditional barrier probabilities.
    Two calculations occur: first multiply probabilities along a path to obtain the conditional branch probability. Then multiply that branch probability by the initiating-event frequency to obtain a branch frequency with units such as per year.
    Open the ETA equation and symbol guide
    fIPronounced “f sub initiator” or “initiating-event frequency.” Expected initiating events per stated period, normally one year.
    pDPronounced “p sub D.” Conditional probability that detection succeeds after the initiating event.
    qD = 1 − pDPronounced “q sub D equals one minus p sub D.” Probability that detection fails.
    pIPronounced “p sub I.” Conditional probability that isolation succeeds on the branch where it is demanded.
    qI = 1 − pIPronounced “q sub I equals one minus p sub I.” Conditional probability that isolation fails when demanded.
    PbranchPronounced “P sub branch.” Multiply the appropriate success or failure probability at every branch point on the selected path.
    fbranchPronounced “f sub branch.” fI multiplied by every conditional success or failure probability along that branch.
    ΣfbranchPronounced “sum of the branch frequencies.” Σ is capital sigma and means add all mutually exclusive terminal branches.
    Uncontrolled branch: funcontrolled = fI × (1 − pD) × (1 − pI). We calculate it to estimate how often a defined consequence pathway may occur and to see which barrier improvement changes the outcome most.

    ETA outputs are only as sound as the initiating frequency, branch definitions, conditional probabilities and dependency assumptions.

    FeatureFault treeEvent tree
    DirectionBackward from a defined unwanted top event.Forward from a defined initiating event.
    Main questionWhat combinations can cause this event?What outcomes can follow as barriers succeed or fail?
    LogicAND, OR and other gates combine causal events.Branches represent conditional success/failure pathways.
    Input basisBasic-event probabilities or frequencies on a consistent mission, demand or time basis.Initiating-event frequency plus conditional branch probabilities.
    Typical resultQualitative cut sets and, when quantified, top-event probability or frequency.Conditional path probabilities and frequencies for defined end states.
    Useful forComplex failure logic, critical combinations and design weaknesses.Escalation, mitigation, consequence pathways and outcome frequency.
    LimitationA poor top-event definition or missed dependency creates false confidence.Too many branches become difficult; dynamic interactions may be simplified.
    Connect causes, controls and consequences

    The Bowtie Model

    Bowtie combines fault-tree thinking on the left and event-tree thinking on the right. It places the top event—the moment control over the hazard is lost—in the centre.

    A collective industrial technique—not one person’s invention

    There is no single uncontested Bowtie inventor or creation date. The visual method evolved collectively from fault-tree and event-tree thinking and was progressively adopted in high-hazard industries. It became popular because a multidisciplinary team could see the complete threat–control–loss pathway on one page.

    Why it is used: to connect each threat to a preventive barrier, define the loss-of-control top event, connect consequences to mitigating barriers and make critical-control ownership visible.

    How it changed: modern practice adds escalation factors, escalation controls, barrier owners, performance standards and assurance evidence. A decorative “bowtie picture” without these elements is not enough for control management.

    See the UK Government Bowtie overview and the Office of Rail and Road’s health-risk application.

    Teaching illustration of a multidisciplinary industrial safety team collaboratively developing a Bowtie barrier map
    AI-generated teaching illustration—not an archival photograph. The team represents collective Bowtie development; no single inventor is implied.
    HazardThreats + preventionTop eventMitigation + recoveryConsequences
    ThreatCongested vehicle route
    Preventive barrierStorage exclusion and route inspection
    ThreatDamaged impact protection
    Preventive barrierEngineered barrier and verified repair
    TOP EVENT
    Vehicle contacts manifold and containment is lost
    Mitigating barrierLeak detection and alarm
    ConsequenceWorker chemical exposure
    Mitigating barrierEmergency isolation, exclusion and response
    ConsequenceEnvironmental release and interruption
    HazardA source with the potential to cause harm, such as hazardous chemical inventory or moving vehicles.
    ThreatA credible cause that could release control of the hazard. It belongs on the left of the top event.
    Top eventThe first moment control is lost—not the threat and not the final injury.
    Preventive barrierActs before the top event to stop a threat from causing loss of control.
    Mitigating barrierActs after the top event to reduce escalation or consequence.
    ConsequenceA credible harmful outcome to people, health, environment, assets or operations.

    Escalation factor

    A condition that can defeat or weaken a barrier—for example poor lighting reduces the reliability of a visual route check.

    Escalation-factor control

    A control that protects the main barrier—for example lighting inspection and emergency lighting support route visibility.

    Interactive Bowtie Draft Builder

    Complete the fields and build a barrier-focused teaching record.
    Understand behaviour in context

    Behavioural Root-Cause Analysis

    Behavioural RCA examines what a person did and the conditions that made the behaviour understandable or likely. It should not become a search for someone to blame.

    Teaching illustration of B. F. Skinner observing behaviour in a historical laboratory and a safety team analysing barriers
    AI-generated teaching illustration—not an archival photograph. B. F. Skinner is represented on the left; no single inventor of behavioural RCA is claimed.

    Where the ABC idea came from

    Who and when: there is no single creator of “behavioural root-cause analysis.” Its ABC structure grew from twentieth-century behavioural science. Psychologist B. F. Skinner’s work on operant conditioning helped explain how consequences can strengthen or weaken behaviour.

    Why safety practitioners use it: repeated behaviour often makes sense when the antecedents are clear and the immediate consequence is easier, faster or socially accepted. ABC analysis makes these influences visible.

    How it changed: modern human-factors practice rejects a behaviour-only explanation. ABC evidence should be combined with task design, competence, equipment, workload, supervision, leadership, culture and organisational controls. This prevents “the worker chose badly” from becoming the false root cause.

    Describe factsDefine behaviourFind antecedentsTest consequencesCorrect the system
    A — AntecedentWhat existed before the behaviour? Instructions, signals, layout, targets, training, workload, tools, norms and supervision.
    B — BehaviourWhat observable action occurred? Describe it neutrally and specifically—do not label a person “careless.”
    C — ConsequenceWhat followed the behaviour immediately or repeatedly? Time saved, praise, delay avoided, discomfort, correction—or no response at all.
    Easy rule: Behaviour is something observable that a person did or did not do. “Careless,” “complacent” and “poor attitude” are interpretations, not observable behaviours. Ask what evidence supports them—and what system conditions made the action likely.
    Example: “The driver used the narrowed route” is the behaviour. Antecedents may include pallets in the approved lane, dispatch pressure and an accepted local workaround. The immediate consequence may have been faster completion on many previous occasions. The corrective action must address the system that reinforced the behaviour—not only tell the driver to be careful.

    Interactive ABC and System-Factor Builder

    Describe behaviour neutrally, then connect it to antecedents, consequences and system conditions.

    Boundary: A fair system approach does not mean that every action is acceptable. Deliberate reckless conduct may require a just and proportionate response, but the investigation must still examine supervision, controls and organisational context.

    Tool: Which Causation Technique Should Lead?

    Complex investigations often combine techniques. Select the main question to see a defensible starting point.

    Choose the investigation question and system complexity.
    3.1 creates the causal questions.
    The next step is to test whether valid loss data reveals frequency, severity, distribution or change that supports further enquiry.
    Continue to 3.2
    Learning outcome 3.2
    Central question: What do the numbers reveal?

    Justify the Use of Quantitative Methods in Analysing Loss Data

    Quantitative analysis converts valid counts, exposure and consequences into comparable measures. The Level 6 requirement is not only to calculate; it is to justify why a numerical method is suitable, explain assumptions and limitations, interpret the result, and connect it to investigation and control.

    3.2
    Count correctly before calculating

    From Raw Loss Records to Decision-Useful Evidence

    CountHow many defined events occurred?
    RateHow many events occurred for a stated amount of exposure?
    SeverityHow much consequence, such as lost time, followed?
    DistributionWhere, when, to whom and through which mechanism did loss occur?

    Why a count can mislead

    Site A records 8 cases and Site B records 5. Site A may appear worse, but if it worked four times as many hours, its exposure-normalised rate may be lower.

    Why a rate can also mislead

    A single event can move a small workforce’s rate sharply. A low rate may also reflect under-reporting, changed classification, outsourced exposure or chance—not strong control.

    Symbols Used in the Calculations

    NPronounced: “capital N.” Number of cases matching the chosen definition.
    HPronounced: “capital H.” Total exposure hours actually worked for the population and period.
    WPronounced: “capital W.” Average workers or full-time-equivalent workers in the defined population.
    CPronounced: “capital C.” Existing cases: new and continuing cases present during the prevalence reference period.
    PPronounced: “capital P.” Population at risk of the defined ill-health condition.
    KPronounced: “capital K,” or “base multiplier.” The agreed standardising base, such as 1,000 workers, 100,000 people or 1,000,000 hours.
    FRPronounced: “F-R,” or “frequency rate.” Defined new events per stated number of exposure hours.
    IRPronounced: “I-R,” or “incidence rate.” Defined new cases per stated worker or population base.
    PrevPronounced: “prevalence.” Existing new and continuing ill-health cases per stated population base.
    DPronounced: “capital D.” Number of lost, restricted or otherwise defined consequence days.
    SRPronounced: “S-R,” or “severity rate.” SR = (D × K) ÷ H under the selected convention.
    D̄Pronounced: “D bar.” Average defined days per case: D̄ = D ÷ N.
    pPronounced: “lower-case p.” Probability of a defined event or branch success; its value lies from 0 to 1.
    q = 1 − pPronounced: “q equals one minus p.” Complementary probability that the defined event does not occur.
    ΣPronounced: “capital sigma,” or “sum.” Add all listed values that follow the symbol.
    x̄Pronounced: “x bar.” Arithmetic mean: add the observations and divide by their number.
    sPronounced: “lower-case s.” Sample standard deviation: typical spread of observations around the sample mean.
    SEPronounced: “S-E,” or “standard error.” Estimated sampling variability of a statistic; in the teaching tool, SE = s ÷ √n.
    CI₉₅Pronounced: “95 percent confidence interval.” An interval generated by a method expected to cover the population value in about 95% of repeated samples.
    Δ%Pronounced: “delta percent.” Percentage change between comparable periods.
    MA₃Pronounced: “M-A sub three.” Three-period moving average used to reduce short-term fluctuation.
    Definitions before arithmetic: Decide what counts as a case, which workers and hours are included, how contractors are treated, which period applies, and which base multiplier is required. Never compare rates that use different definitions without adjustment.
    Live calculation laboratory

    Calculate and Interpret Loss Rates

    Flexible Accident / Incident Frequency-Rate Calculator

    Why calculate it? A count alone ignores how long people were exposed. Frequency rate converts defined events into a common hours-worked base, supporting more defensible comparison.

    FR = (N × Kh) ÷ H
    The calculated rate will appear here.
    How to say and use every sign
    FR“F-R” or “frequency rate”: the result.
    =“equals”: the left side has the same calculated value as the right side.
    ( )“brackets”: complete the multiplication inside first.
    N“capital N”: defined new events in the period.
    דmultiplied by”: multiply N by the hours base Kh.
    ÷“divided by”: divide by exposure hours H.

    The 200,000-hour convention is used in US OSHA/BLS incidence rates. A one-million-hour base is common for some international frequency measures. Confirm the required definition before comparison.

    Severity and Average-Consequence Calculator

    Why calculate it? Frequency tells how often events occur; severity shows the amount of defined consequence relative to exposure. Average days per case describes the typical recorded consequence per relevant case.

    SR = (D × Kh) ÷ H   |   D̄ = D ÷ N
    Severity results will appear here.
    How to say and use every sign
    SR“S-R” or “severity rate”: defined lost days per chosen hours base.
    D̄“D bar”: average defined days per relevant case.
    |“and separately”: the vertical line separates two related calculations; it is not division.

    “Lost day” and severity conventions differ. Record calendar/workday rules, caps, fatalities and restricted work consistently.

    Accident Incidence-Rate Calculator

    Why calculate it? When reliable hours are unavailable or the required convention uses workers, incidence expresses new defined accidents or cases per standard number of workers.

    IR = (N × Kw) ÷ W
    The worker-based incidence rate will appear here.

    Do not confuse incidence with frequency: incidence uses workers or people; frequency uses hours worked. “New” means the case started within the reference period.

    Ill-Health Prevalence-Rate Calculator

    Why calculate it? Prevalence estimates the burden of a condition now—both new and continuing cases—so an organisation can plan health controls, surveillance, support and resources.

    Prev = (C × Kp) ÷ P
    The prevalence rate will appear here.

    Prevalence is not incidence: prevalence counts all qualifying existing cases; ill-health incidence counts only new cases. Long-latency occupational disease may reflect exposures from many years earlier.

    Interactive Tool: Which Rate Should I Use?

    Choose the decision question and available denominator.

    Formula conventions vary. Examples follow official explanations from the US Bureau of Labor Statistics, International Labour Organization, and UK HSE ill-health statistics guidance.

    Tool: Compare Two Sites Fairly

    Counts answer “how many?” Rates help answer “how many for the amount of exposure?” Use the same event definition, period and multiplier.

    Site A

    Site B

    The exposure-normalised comparison will appear here.
    Find direction and concentration

    Trend, Moving Average and Pareto Analysis

    Six-Period Trend Explorer

    Enter six comparable monthly rates separated by commas.

    Δ% = [(new − old) ÷ old] × 100   |   MA₃ = (xt + xt−1 + xt−2) ÷ 3
    Enter six rates and analyse the direction.

    Why calculate it? Δ (“delta”) shows relative endpoint change. MA₃ smooths short-term fluctuation by averaging the current and previous two periods. Neither proves the reason for change.

    Pareto Priority Explorer

    Pareto analysis orders categories from largest to smallest so the team can see where recorded loss is concentrated.

    Category % = (nc ÷ N) × 100   |   Cumulative % = Σ category %
    Adjust a count to update priority and cumulative contribution.
    What every symbol means
    nc“n sub c”: count in one category.
    N“capital N”: total count across all displayed categories.
    Σ“capital sigma”: add percentages as you move down the ranked categories.
    Interpretation rule: A trend or Pareto chart tells you where to look. It does not prove causation. Investigate exposure, reporting practice, severity potential, control performance and the narratives behind the categories.
    Interpret loss data visually

    Histogram, Pie Chart and Line Graph Laboratory

    Different charts answer different questions. A chart must match the data structure; attractive graphics do not repair weak definitions or incomplete records.

    Interactive Chart Explorer

    Choose a chart type and update the values.

    Choose by question—not appearance

    HistogramShows the distribution of numerical observations grouped into ordered, usually equal-width intervals called bins. Bars touch because the scale is continuous. Example: days lost per case.
    Pie chartShows how mutually exclusive categories make up one whole at one point or period. Use few categories and show the denominator. Example: recorded event mechanisms this year.
    Line graphShows an ordered series, normally time. It helps reveal direction, cycles and unusual change. Keep definitions, exposure and intervals comparable.
    Important distinction: a category chart with separate bars is a bar chart, not a histogram. A histogram needs a numerical scale split into intervals. The tool begins with lost-day intervals so the histogram use is valid.
    Interpretation sentence: “The chart shows ___, across ___, using ___ as the denominator. The main pattern is ___. However, ___ may affect validity; therefore we should investigate ___ before deciding ___.”
    Do not confuse precision with truth

    Statistical Variability, Distributions, Sampling and Data Validity

    Statistical variability means observations and sample results naturally differ. Validity asks whether the measure actually represents the decision question. A precise calculation from biased or incorrectly classified data is still misleading.

    Descriptive Statistics and Approximate CI Explorer

    Enter numerical observations such as lost days per case. This tool describes the sample; it does not certify a population model.

    Enter at least two observations.
    Open every equation and symbol
    x̄ = Σx ÷ n“x bar equals sum of x divided by n.” The arithmetic mean.
    MedianThe middle ordered value; for an even n, average the two middle values. It is less affected by extreme values than the mean.
    Range = max − minLargest observation minus smallest observation.
    s = √[Σ(x − x̄)² ÷ (n − 1)]Sample standard deviation. √ is “square root”; ² is “squared.”
    SE = s ÷ √nEstimated standard error of the sample mean.
    Approx. CI₉₅ = x̄ ± 1.96 × SE“x bar plus or minus 1.96 times standard error.” A teaching normal approximation, not suitable for every dataset.

    What a distribution can reveal

    SymmetricalValues spread similarly around the centre; mean and median may be close.
    Right-skewedMany small values with a few large losses. The mean may exceed the median—common with days lost or costs.
    Clusters or multiple peaksMay indicate different workgroups, tasks, mechanisms or reporting systems that should not be pooled without examination.

    Representative sample: a sample should reflect the population relevant to the question. Every important subgroup needs a fair chance of inclusion. A large convenience sample can still be biased.

    Sampling a population: define the target population, build a suitable sampling frame, select participants or records using a defensible method, record non-response and compare the achieved sample with the population.

    UK HSE explains sampling error and 95% confidence intervals in its Labour Force Survey guidance. See also the ONS guide to uncertainty.

    Interactive Sample and Validity Check

    The teaching validity check will appear here.

    This is a structured warning tool, not a formal sample-size calculation or statistical certification.

    Common data errors to test

    ErrorEasy meaningExample
    CoverageSome of the population cannot appear in the data.Night shift and contractors are missing.
    SelectionThe inclusion method favours certain people or records.Only volunteers answer a wellbeing survey.
    Non-responseSelected participants do not respond, and may differ from responders.Workers with symptoms do not trust confidentiality.
    MeasurementThe question, instrument or observer produces inaccurate values.Different clinics use different symptom questions.
    ClassificationThe same case is placed in different categories.Restricted work is recorded as first aid at one site.
    Duplicate / missingA case is counted twice or not counted.One injury exists in two systems without a unique ID.
    DenominatorThe exposure base excludes relevant work.Contractor cases included, contractor hours excluded.
    ProcessingEntry, coding, formula or transfer is wrong.Hours are entered as 40,000 instead of 400,000.
    Time lagThe measured harm appears long after exposure.Current respiratory disease reflects earlier dust exposure.
    Reporting cultureTrust and rules change what becomes visible.A reporting campaign increases near-miss counts.
    The command word is “justify”

    Why Use Quantitative Methods—and Why Not Use Them Alone?

    Reason for using numbersDecision valueCondition or limitation
    Normalise exposureRates allow more defensible comparison across differently sized populations or periods.Definitions, hours, workforce scope and base multiplier must match.
    Detect changeTime-series analysis can show sustained deterioration, improvement, seasonality or unusual variation.Short runs, rare events and changed reporting can produce unstable signals.
    Measure disease burdenIll-health prevalence estimates all qualifying existing cases and supports surveillance, control and resource planning.Long latency, diagnostic access, worker turnover and healthy-worker effects can disconnect current prevalence from current exposure.
    Prioritise investigationPareto and distribution analysis identify categories, locations or activities contributing most recorded loss.Frequency must be considered with credible severity and major-hazard potential.
    Express uncertaintySpread, standard error and confidence intervals show that a sample estimate is not an exact population truth.The calculation depends on a defensible sample, measurement quality and an appropriate statistical model.
    Evaluate interventionBefore/after measures can test whether performance changed following control.Control for exposure, operational change, reporting, regression to the mean and other influences.
    Communicate performanceDefined indicators support dashboards, accountability and resource decisions.Targets can encourage under-reporting or classification manipulation if poorly designed.
    Model event pathwaysFTA, ETA and quantitative risk methods estimate the contribution of failure combinations and barriers.Models contain assumptions, dependencies and uncertainty; precision is not certainty.

    Small numbers

    One event can double a rate in a small workforce. Use longer periods, confidence intervals or pooled evidence where appropriate.

    Under-reporting

    A “good” rate may reflect low trust or restricted definitions. Triangulate with audits, surveys, health data and workforce evidence.

    Lagging-only bias

    Injury data describes realised outcomes. Add leading evidence about exposure, critical controls, defects and corrective-action quality.

    Severity randomness

    Similar events can produce very different harm. Do not assume low historical injury means low potential consequence.

    Changing denominator

    Overtime, contractors, shutdowns and outsourcing change exposure. Record the population and hours consistently.

    Metric fixation

    Managing the number rather than the risk can distort behaviour. Indicators must serve learning and control—not replace them.

    Interactive Level 6 Justification Builder

    Build a connected justification: decision → method → evidence → value → limitation → complementary evidence → action.
    Learning outcome 3.3
    Central question: Who needs to know, when and why?

    Assess the Needs and Impacts of Reporting Loss Events

    Reporting is the controlled movement of information from the person who sees a loss event to the people who can protect, classify, investigate, notify, learn and act. To assess reporting at Level 6, weigh its need, urgency, benefits, adverse impacts and the consequences of silence, then reach a justified judgement.

    3.3
    Dr. Sarel Du Plessis in a black suit with folded arms
    Guest practitioner · short session overview

    Dr. Sarel Du Plessis

    A brief practitioner perspective before the learner journey

    DB HSE OTHM Level 6 in OSH Approved Tutor CEO & Founder · HSAFE Risk Solutions UAE · KSA · International Projects

    Guest contribution: Dr. Sarel will give a short practical overview of the Section 3.3 and 3.4 summary table and learning line—how a loss event moves from reporting into evidence-led investigation.

    The detailed explanations, examples and interactive tools below remain independent DB HSE learner resources for self-directed study and class discussion.

    3.3 · Make the event visibleWhat must be reported, to whom, how quickly, for which decision and with what impact?
    3.4 · Convert evidence into preventionHow does investigation test causes, improve controls and verify risk reduction?
    3.3 in one minute · See the process

    From a Workplace Signal to Safe Decisions

    What 3.3 is really asking: do not merely say that reporting is important. Judge what information is needed, who needs it, how urgently it must travel, what decision it enables, what burdens or risks it creates, and why the selected route is proportionate.
    Workers provide care, control a forklift incident area and make a factual digital loss-event report
    Practical reporting scene: a small contained outcome can still reveal high potential. Care, isolation, factual recording and competent escalation happen together—without declaring blame.
    1 · Care before paperworkFirst aid, withdrawal and containment take priority over form completion.
    2 · Record facts, not faultState what was observed, measured or reported; label estimates, hypotheses and unknowns.
    3 · Assess actual and credible potentialA minor injury or small release may expose failure of a critical control.
    4 · Choose the route and audienceInternal escalation, legal notification, insurer and client routes are separate decisions.
    5 · Close the loopAcknowledge the reporter, track action and explain what changed.
    Worker signal → emergency response → supervisor or control point → competent HSE triage → internal and applicable external routes → investigation → action → feedback to the reporter and workforce
    Start with the distinction

    What Is Loss-Event Reporting?

    Core definition: Reporting is the prompt and formal communication of an incident, accident, near miss, hazard or unsafe condition to the responsible person or authority. It makes the event officially visible so that immediate action, investigation, corrective action and prevention can begin.
    What the first report should contain

    Give enough reliable information for safe action

    • Date, time and exact location
    • People and equipment involved
    • Observable facts—without assumptions or blame
    • Actual injury, damage, loss or disruption
    • Potential consequences—what could credibly have happened
    • Immediate controls, such as stopping work, isolation, first aid or restricted access
    • Evidence requiring protection, including photographs, CCTV, documents, damaged equipment and witness details
    • The person reporting the event and the time it was reported
    Simple workplace example

    Oil spill beside Machine 3

    At 10:15 a.m., a worker slipped on an oil spill beside Machine 3 and sustained a minor ankle injury. Work was immediately stopped, first aid was provided and the area was isolated.

    The event could have caused a serious head injury or contact with moving machinery. Photographs, CCTV footage, the worker’s footwear, maintenance records and witness details were preserved for investigation.

    Important boundary: Reporting is not the full investigation. It is the first accurate notification that enables the organisation to control the situation, meet applicable legal requirements, investigate causes, take corrective action and prevent recurrence.

    Loss event

    A broad management term for an event that caused—or credibly could have caused—injury, ill health, environmental harm, asset damage, production interruption, financial loss or harm to trust and reputation.

    Report

    The first communication that makes the event visible. It should begin with observable facts, actual and potential consequence, immediate controls and the evidence that may need protection.

    Record

    The controlled, traceable entry kept for follow-up, analysis and required retention. One event may produce several witness or injury reports, but should have one master event identity with linked records.

    Notify

    A formal communication to a regulator, emergency service, insurer, client or other body when a current legal, permit or contractual rule requires it. Notification does not replace internal reporting or investigation.

    Essential rule: Internal reporting should be broader than statutory notification. “Not externally reportable” never means “do not record, assess or learn.” Near misses, control failures and unsafe conditions may provide warning before serious loss occurs.

    Six Communication Actions—Do Not Treat Them as One Form

    1. Emergency alertWarn people and activate rescue, isolation, spill, fire or medical arrangements. This is immediate operational communication.Typical timing: seconds to minutes
    2. Internal notificationTell the supervisor, control room, HSE lead or other defined responsible role that an event has occurred.Typical timing: without avoidable delay
    3. Factual recordCreate a traceable entry containing what is known, what remains unknown, the actual outcome, credible potential and immediate controls.Typical timing: as soon as the situation permits
    4. External notificationThe legally or contractually responsible person checks and completes any regulator, environmental, insurer, client or permit route.Typical timing: according to the current applicable rule
    5. Investigation updateCorrect or supplement the first record when verified evidence changes the sequence, classification or required action.Typical timing: controlled and version-traceable
    6. Learning feedbackTell the reporter and affected workforce what was decided, what changed and how effectiveness will be checked.Typical timing: throughout action and closure
    Practical 3.3 focus

    Reporting quality is measured by the safe decisions it enables

    Do not ask only whether a form was submitted. Ask whether reliable information reached the right people in time to protect, preserve, decide and learn—and whether the reporter received meaningful feedback.

    What a Level 6 assessment must weigh

    • Actual and credible potential loss
    • Stakeholder need and urgency
    • Internal and possible external routes
    • Benefits, burden and unintended effects
    • Privacy, fairness and reporting trust
    • Consequences of delay or silence

    Write the First Report Before You Know the Cause

    Information statusHow to write itMaster-case example
    Verified factState what a reliable source, measurement or record confirms.“The alarm log records activation at 07:42:18.”
    Direct observationName what the observer saw, heard or did without converting it into a cause.“The operator observed liquid inside the contained sump.”
    EstimateLabel the value and its basis; do not present it as a measurement.“The initial spill estimate is 12–18 litres based on bund level.”
    Witness recollectionAttribute the account and preserve its wording for later corroboration.“A witness recalled that pallets reduced route width before the contact.”
    UnknownState the gap and the evidence needed to resolve it.“The pressure immediately before contact is not yet confirmed; historian data is being preserved.”
    Hypothesis—not a report conclusionRecord it only as an explanation to test against supporting and conflicting evidence.“Barrier degradation may have influenced the outcome; inspection and defect records are required.”

    Interactive First-Report Language Coach

    Select a statement to see whether it is factual enough for an initial report or has moved prematurely into blame, minimisation or causal certainty.

    Choose a statement, then test whether it separates observation from investigation.

    What Should the System Be Able to Receive?

    PeopleFatality, injury, first aid, exposure, suspected or diagnosed work-related ill health.
    PotentialNear miss, dangerous occurrence, high-potential event or critical-control failure.
    Assets and operationsFire, damage, defective plant, material loss, downtime, quality or information loss.
    Wider effectsRelease, pollution, harm to the public, contractor event, supply disruption or loss of trust.
    Recognise the signal before choosing the route

    What Kind of Loss Event Is This?

    Classification gives a shared starting language. It does not decide the cause, prove legal notification or limit an event to one label. The same event may involve several loss types—for example, an injury, equipment damage, a process-safety failure and business interruption.

    A worker reports a near miss to a supervisor while the process area is cordoned and evidence is protected
    Practical example: the immediate outcome may be small, yet a damaged barrier and vulnerable process line show credible potential for greater loss. Reporting makes the signal visible while care, control and evidence protection continue.
    Actual outcomeWhat injury, damage, release, interruption or other loss actually occurred?
    Credible potentialWhat could reasonably have happened if one control or circumstance were different?
    Primary and linked typesChoose the main event type, then link every other affected loss category.
    Route comes laterInternal classification is broader than statutory notification and should never be used to avoid a competent legal screen.
    Loss-event typePlain-language meaningExample signalFirst reporting need
    Injury or fatalityPhysical harm ranging from first aid to death.Cut, fracture, burn or fatal contact.Care, scene control, factual record and competent notification screen.
    Occupational illness or diseaseAcute or gradual work-related harm to health.Dermatitis, hearing loss, vibration symptoms or occupational asthma.Support the person, protect health data and examine exposure history.
    Near missAn unplanned event with no actual loss but credible potential for harm.A suspended load falls into an empty exclusion zone.Capture the warning before evidence and memory disappear.
    Dangerous occurrenceA specified or serious event indicating major risk, whether or not someone was hurt.Collapse, plant failure or loss of containment.Escalate potential severity and screen the current legal definition.
    Property or equipment damageDamage to buildings, plant, vehicles, tools or materials.Forklift bends a process barrier.Make safe, preserve the damaged item and assess hidden risk.
    Fire or explosionUncontrolled combustion, ignition, deflagration or explosion.Small electrical cabinet fire.Emergency control, competent technical escalation and scene preservation.
    Environmental releaseActual or threatened harm to land, air, water or biodiversity.Chemical enters a surface-water drain.Contain, identify pathways and screen environmental notification duties.
    Process-safety eventLoss or degradation of containment or control in a hazardous process.Relief device lifts or isolation fails.Assess barrier performance and major-event potential, not injury count alone.
    Security eventThreat, intrusion, violence, theft, sabotage or information-security loss affecting safe operations.Unauthorised entry to a restricted process area.Protect people and evidence while using the defined security route.
    Business interruptionLoss of production, service, supply, access or continuity.A line stops for six hours after a collision.Coordinate safe recovery without allowing production pressure to weaken controls.
    Reputational or financial lossLoss of trust, claims, penalties, direct cost or wider commercial harm.Public concern after a visible release.Use verified facts, controlled communication and never conceal safety evidence.

    Interactive Loss-Event Classification Coach

    Choose the clearest first signal. The coach shows the linked categories and the first safe reporting priority.

    Choose a signal to distinguish its primary classification, linked losses, potential and first reporting need.
    Why, who, when and how

    The Needs a Reporting System Must Meet

    NeedWho needs the information?Decision or action enabledIf the event stays hidden
    Protect life and contain escalationWorkers, supervisor, emergency and occupational-health teamsFirst aid, evacuation, isolation, treatment, welfare and scene controlExposure can continue and evidence may disappear.
    Classify and escalateCompetent HSE, operations and senior leadersAssess actual and credible potential consequence; set investigation level and interim controlsA “minor outcome” can conceal a major control failure.
    Meet applicable dutiesNamed responsible person, regulator, insurer, client or permit authorityScreen current legal, contractual, environmental and sector-specific rulesDeadlines, evidence, compensation or formal duties may be missed.
    Support affected peopleInjured or exposed people, families, worker representatives and HRCare, communication, rehabilitation, fair treatment and lawful record handlingPeople may lose trust, support or access to a valid claim.
    Investigate and learnInvestigation team, technical specialists and workforcePreserve evidence, test causes, identify effective actions and share learningSimilar work can repeat the same conditions.
    Analyse performanceManagers, assurance functions and governance bodiesIdentify clusters, recurring control weakness, response delay and action qualityRates and trends become biased and falsely reassuring.
    1. Detect and care.
    Raise the alarm, obtain treatment and prevent escalation. Reporting never takes priority over immediate life safety.
    2. Make safe and alert.
    Use the emergency route and notify the responsible supervisor or control point without avoidable delay.
    3. Triage actual and potential loss.
    Ask what happened and what could credibly have happened if one condition were different.
    4. Create one master record.
    Give the event a unique identity; link people, witnesses, injuries and duplicate reports without double-counting.
    5. Preserve facts and evidence.
    Record time, place, work, people, equipment, substances, immediate controls, witnesses and evidence sources.
    6. Screen external routes.
    A competent person checks current legal, environmental, sector, insurer, client and permit requirements.
    7. Investigate proportionately.
    Use actual outcome, credible potential, recurrence likelihood, uncertainty and complexity to set the response.
    8. Assign and track action.
    Name owners, dates, interim safeguards and verification evidence.
    9. Feed back and support.
    Acknowledge the reporter, protect confidentiality, update affected people and share anonymised lessons.
    10. Analyse and review.
    Validate the dataset, assess repeated patterns and test whether action reduced risk.

    Current Great Britain illustration—not a global decision rule

    Under RIDDOR, the legally responsible person—not every witness—submits defined reports. Deaths, specified injuries, qualifying public injuries and dangerous occurrences generally require notification without delay and a report received within 10 days; qualifying over-seven-day incapacity is reported within 15 days; certain diagnosed occupational diseases are reported when the diagnosis is received.

    Always check the current rule, definitions and jurisdiction. Environmental, transport, major-hazard, insurer and client routes are separate.

    Design for work as actually done

    Provide simple mobile, telephone and paper routes; plain language; accessibility and multilingual support; contractor access; confidential routes; non-retaliation; prompt acknowledgement; and visible follow-up.

    A technically complete form that workers fear, cannot access or never hear back from is not an effective reporting system.

    One event · different legitimate information needs

    Who Needs Which Information—and Why?

    Good reporting does not send the entire file to everybody. It gives each authorised stakeholder the minimum reliable information needed for a defined decision, at the right time and through the right route. Urgent operational facts may travel quickly; identifiable health, witness or legal material may require much tighter access.

    StakeholderInformation needed and whyUrgency and routeConfidentiality boundaryDecision enabled
    Worker or affected personCare, exposure, immediate risk, support, next update and how their information will be used.Immediate face-to-face or emergency route; documented follow-up.Respect dignity; restrict health and identifying detail.Treatment, withdrawal, support and informed participation.
    Supervisor or line managerFacts, location, people at risk, controls taken and remaining danger.Without avoidable delay by alarm, radio, phone or direct report.Operational facts only unless more detail is necessary.Stop work, isolate, resource response and escalate.
    HSE functionActual and potential loss, evidence, repeated signals, exposure and applicable routes.Prompt internal system or emergency escalation.Role-based access; separate operational and clinical detail.Triage, investigation level, legal screen and learning.
    Senior management or boardMaterial risk, critical-control failure, trend, assurance gap and action ownership.Immediate for major risk; governed dashboard for trends.Use aggregated or anonymised data unless identity is essential.Resources, risk appetite response and governance challenge.
    Worker representatives or trade unionsRelevant facts, worker impact, controls, investigation participation and action.As soon as practicable through agreed consultation arrangements.Share lawfully; anonymise personal detail where appropriate.Represent workers, protect evidence and test practicality.
    Occupational-health teamExposure, symptoms, work history and relevant clinical information.Prompt confidential referral appropriate to health risk.Clinical detail stays within authorised health-data controls.Assessment, support, fitness advice and health surveillance.
    HR or legal functionNecessary employment, welfare, process, evidence and duty information.According to seriousness and formal-process need.Do not assume every investigation record is legally privileged.Fair process, support, claims management and legal advice.
    InsurerPolicy-defined notice, verified facts, loss estimate and preserved evidence.Within policy conditions through the specified claims route.Disclose only what is lawful, relevant and required.Coverage, claim handling, expertise and loss control.
    Client or principal contractorInterface risk, affected work, immediate controls and contract-defined notice.According to emergency and contractual arrangements.Contractual access does not justify unnecessary health detail.Coordinate site control, work interfaces and assurance.
    RegulatorInformation required by the current applicable legal framework.By the legally defined route and deadline.Use the official route; preserve the original record.Regulatory oversight, enquiry and enforcement where appropriate.
    Emergency servicesHazard, location, people, substances, access and live escalation risk.Immediately through the emergency channel.Life-safety need governs; avoid irrelevant personal detail.Rescue, medical, fire, evacuation and area control.
    Environmental authoritySubstance, amount, pathway, receptor, controls and monitoring evidence.According to the release and current environmental rule.Protect investigation integrity without delaying required notice.Containment, environmental protection and regulatory response.
    Family or public, where appropriateAccurate, humane and verified information about impact and public protection.Through an authorised, coordinated communication lead.Next-of-kin, privacy and investigation needs come before publicity.Support, public protection and trustworthy communication.

    Interactive Stakeholder Information Coach

    Choose a stakeholder to see the information, urgency, route, confidentiality boundary and decision enabled.

    Which Reporting Route Fits the Situation?

    RouteBest useBenefitLimitation or riskGood control
    Verbal reportImmediate local warning or simple access for workers.Fast and inclusive.Can be forgotten, changed or untraceable.Confirm important facts in the master record.
    Telephone or emergency notificationUrgent off-site or specialist response.Rapid two-way clarification.Wrong number, incomplete note or no audit trail.Use a defined call tree and log time, recipient and advice.
    Paper formLow-connectivity work or accessible local backup.Simple and portable.Delay, handwriting, loss and duplicate entry.Unique ID, secure transfer and timely digital reconciliation.
    Digital reporting systemStructured records, workflow, analysis and feedback.Traceability, required fields and trend data.Access, form fatigue, poor taxonomy or false precision.Short first report, role-based access and usability testing.
    Anonymous or confidential channelFear, sensitive conduct or serious trust barrier.Can reveal otherwise hidden risk.Harder clarification and possible misuse.Explain the difference between anonymous and confidential; protect non-retaliation.
    Toolbox talk or shift handoverImmediate shared operational awareness.Reaches the workgroup and connects context.Not a substitute for an event record or formal notification.Record the event once and document safety-critical handover.
    Statutory notificationA current legal reporting trigger.Meets the formal route and enables independent oversight.Narrow criteria can be mistaken for the whole internal system.Competent screen, current official form, deadline and retained record.

    Interactive Reporting-Route Comparator

    Choose a situation to see the primary route, supporting record, benefit and limitation.

    Leading and Lagging Information—Read Both

    Leading information

    Shows the conditions and activities that may influence future performance: hazard reports, critical-control failures, response time, reporting trust, investigation quality, overdue actions and verification results.

    Use: intervene before serious loss. Limit: activity counts do not prove the activity was effective.

    Lagging information

    Shows outcomes that have already occurred: injuries, diagnosed disease, releases, fires, damage, lost time, cost and business interruption.

    Use: understand loss experience and consequences. Limit: low counts can reflect luck, low exposure or under-reporting rather than strong control.

    Assessment means weighing both sides

    Positive, Adverse and Unintended Impacts

    Reporting featurePotential positive impactPossible adverse or unintended impactHow to manage the tension
    Immediate escalationFaster treatment, containment and evidence protectionOperational interruption and pressure on emergency resourcesUse predefined thresholds, trained roles and proportional escalation.
    Open near-miss reportingEarlier warning and worker voice before severe harmRecorded numbers may rise and be wrongly treated as poorer performanceInterpret volume with potential severity, trust, exposure, repeat rate and action closure.
    Detailed data collectionStronger investigation, patterns, claims and control decisionsAdministrative burden, duplication and sensitive-data riskCollect what is necessary, link duplicate records, validate fields and restrict access.
    External notificationCompliance, independent oversight and wider learningScrutiny, cost, enforcement, claim or reputation concernUse a competent legal screen; do not conceal an event to avoid consequences.
    Fair accountabilityTrust, learning and willingness to reportAn organisation may fear that “no blame” means no standardsSeparate honest error and reporting from deliberate or reckless conduct; decide fairly from evidence.
    Performance indicatorsTrend visibility, prioritisation and governance attentionTargets can drive under-reporting, reclassification or metric fixationBalance lagging outcomes with reporting trust, critical-control, response and action-quality indicators.

    Fear and blame

    Retaliation, ridicule or automatic discipline suppresses information. Reporting itself must never be treated as the offence.

    Production pressure

    If stopping work threatens targets or income, hazards and symptoms may remain invisible. Leaders must make safe reporting practicable.

    No feedback

    A silent system teaches workers that reporting changes nothing. Acknowledge, update, close and show what was improved.

    Complexity and exclusion

    Long forms, language barriers and contractor-only pathways distort the dataset. Offer accessible routes and one taxonomy.

    Latency and privacy

    Work-related disease may emerge slowly, while health data is sensitive. Enable delayed reports and use lawful, minimal, secure access.

    Counting confusion

    Report count, event count and affected-person count are different. Use a master event ID and state the unit being analysed.

    Real evidence: when an environmental loss was not reported

    In a published Environment Agency case, bleach reached a surface-water drain and killed more than 800 fish. The company did not report the incident because it wrongly assumed the drain led to the foul sewer. The case illustrates why drainage knowledge, rapid escalation and factual reporting matter; assumptions can enlarge both environmental and enforcement consequences.

    Read the official Environment Agency case

    A dataset reflects the system that produced it

    Can You Trust the Reporting Picture?

    Event data is not a transparent window into reality. What gets reported, how it is classified, who feels safe to speak and how duplicates are handled all shape the picture. Level 6 assessment therefore examines both the event information and the reporting system.

    Under-reporting

    Real events stay invisible because of fear, inconvenience, production pressure, uncertain definitions, inaccessible routes or lack of feedback. The dataset can look reassuring while exposure continues.

    Over-reporting

    Duplicate records, overly broad categories or repeated logging of the same condition inflate counts and workload. Do not discourage signals; link them to one master event and keep the original reports traceable.

    Selective reporting

    Some events, workers, contractors, shifts or sites appear while others disappear because incentives and access are unequal. Comparisons become systematically biased.

    Visibility change

    A new trusted route can increase reports even while risk control improves. Examine exposure, credible potential, repeat pathways, response quality and worker trust before declaring deterioration.

    Nine Common Reporting Biases

    Fear of discipline

    People hide honest errors or hazards when reporting itself is treated as misconduct.

    Blame culture

    Labels such as “careless” replace factual reporting and discourage useful detail.

    Production incentives

    Bonuses, league tables or zero-event targets can reward silence or reclassification.

    Contractor vulnerability

    People may fear removal, lost work or commercial penalty if they report.

    Normalisation of deviance

    Repeated unsafe conditions become accepted as normal and no longer feel reportable.

    Language and literacy

    Complex forms exclude people or strip meaning from their account.

    Severity bias

    Only injuries are noticed while high-potential near misses and control failures disappear.

    Duplicate reporting

    Several accounts are counted as several events instead of linked evidence.

    Delayed reporting

    Memory, scene condition and electronic evidence degrade before the event becomes visible.

    Eight Quality Tests for Reportable Data

    AccuracyFacts match reliable sources and errors are corrected traceably.
    CompletenessRequired facts, potential, controls and evidence gaps are not silently omitted.
    TimelinessInformation arrives soon enough for care, control, preservation and required action.
    RelevanceCollected information supports a legitimate safety, welfare, legal or learning decision.
    ConsistencyDefinitions, categories, units and periods are applied in the same way.
    TraceabilitySource, change, owner, event identity and decision history can be followed.
    ValidityThe field measures what it claims to measure and the event meets the stated definition.
    ConfidentialityPersonal and sensitive information is lawful, necessary, secure and role-restricted.

    Interactive Reporting-Distortion Diagnostic

    Select the conditions that genuinely exist. The tool identifies how the dataset may be distorted and which system response is needed.

    Select the real system conditions; the diagnosis will explain likely bias, risk and priority response.
    If reporting works

    Organisational value

    • Earlier control of exposure and escalation
    • More reliable trends and resource priorities
    • Better compliance, claims and insurance management
    • Stronger trust, participation and organisational learning
    • Reduced recurrence through visible, verified action
    If information is mishandled

    Adverse or unintended impact

    • Privacy breach, premature blame or reputational harm
    • Defensive reporting, fatigue and excessive bureaucracy
    • Distorted statistics and unnecessary circulation
    • Continued exposure, evidence loss and recurrence
    • Possible enforcement, claim, contractual and trust consequences
    Worked Level 6 “assess” example:

    The forklift–process-line event needs prompt reporting because a worker experienced symptoms, a chemical barrier was challenged and the credible potential was much greater than the contained outcome. The report enables care, scene control, evidence preservation, legal screening and investigation; it also supplies a signal about interface risk across similar routes. Those benefits must be weighed against downtime, administrative burden, staff anxiety and the risk of unnecessary disclosure of health information. The burdens are manageable through one master event record, factual language, role-based access, fair accountability and timely feedback. Non-reporting would leave the damaged barrier and route encroachment hidden, bias performance data and risk recurrence with more serious consequences. The justified judgement is therefore to report and escalate without avoidable delay, while limiting sensitive information and verifying that the resulting action reduces the risk.

    Assess the system, not only the event count

    How Healthy Is the Reporting System?

    A high number of reports is not automatically evidence of poor safety. It can indicate stronger visibility and trust. Read volume alongside exposure, potential severity, recurrence, response quality, action closure and workforce evidence.

    Reporting delayMedian time from discovery to the correct internal control point.
    AcknowledgementPercentage of reporters told promptly that the concern was received.
    High-potential screenPercentage of credible high-potential events triaged within target.
    Reporter feedbackPercentage receiving an update on decision, action and closure.
    Overdue workInvestigations and corrective actions beyond their agreed dates.
    Repeat eventsRecurrence of the same pathway, control failure or equivalent exposure.
    Data qualityMissing fields, duplicates, inconsistent classification and denominator errors.
    Worker trustWhether people expect fair treatment, confidentiality and useful follow-up.

    Practical case: when there is no “first 60 minutes”

    Three technicians separately report numbness and tingling after months of using vibrating tools. A useful system must accept a gradual health signal without forcing a false single-event time or an unproven diagnosis. It should support the workers, record symptoms and work history, link related reports without disclosing unnecessary health details, review exposure and tool-maintenance evidence, involve competent occupational-health support and screen current reporting duties.

    Assessment point: the need for visibility and prevention is strong, but identifiable health information requires restricted, lawful handling. Delay, fear or fragmented records can hide a group exposure; uncontrolled circulation can harm privacy and trust.

    Return to the master case

    The First 60 Minutes After the Forklift–Process-Line Event

    The release is small because the alarm and emergency isolation work. One contractor reports eye irritation; the forklift driver is shaken; liquid reaches a contained drainage sump; the route is closed for six hours. Previous near misses and barrier defects exist in separate informal messages.

    Safe learning boundary: This fictional scenario and its tools are educational—not an emergency, legal-reporting, disciplinary or restart system. In real work, follow emergency and organisational procedures, obtain competent jurisdiction-specific advice, and do not enter names, medical details, employee IDs, photographs or confidential evidence into this public portal.
    TimeDefensible actionWhy it matters
    0–5 minutesAlarm, withdrawal, emergency isolation, first aid and spill responseCare and containment come before form completion.
    5–15 minutesSupervisor alert, head count, exposure and environmental checks, scene boundaryActual harm and escalation routes may still be developing.
    15–30 minutesCreate master event ID; capture factual first reports; identify CCTV, alarm, plant and witness evidenceEarly information is perishable, but conclusions must not be invented.
    30–60 minutesCompetent triage of potential severity; statutory/contract/insurer screen; appoint investigation lead and interim controlsA limited actual outcome does not make the failed vehicle–chemical interface low risk.

    Interactive Reporting-Route and Impact Assessor

    This teaching tool recommends an internal response and legal-screen priority. It does not determine whether a statutory report is required.

    Choose the facts, then assess urgency, internal escalation and the need for a competent external-reporting screen.

    Interactive Report-Quality Diagnostic

    Select only what the proposed first report genuinely contains.

    0 / 8 quality elements selected. A report should be easy to make and strong enough to trigger safe action.

    Interactive Loss-Event Report Builder

    Complete the factual fields to create a learning report that can support triage and investigation.

    Interactive 3.3 “Assess” Response Planner

    Connect need, impact, mitigation, consequence and judgement—do not merely list advantages.

    Interactive Report-to-Investigation Handover Check

    A first report does not need to contain the final cause. It must give the investigation enough reliable starting information and protect what may disappear.

    0 / 8 handover elements selected. Check what the investigation team can safely rely on and what still needs urgent protection.
    3.3 makes the event visible and starts the response.
    A first report records what is known; 3.4 tests the evidence to explain how and why the event became possible.
    Continue to 3.4
    Learning outcome 3.4
    Central question: How does evidence become prevention?

    Explain the Importance and Impact of Incident Investigations

    An investigation is a structured, evidence-led process for understanding what happened, how and why it happened, which controls succeeded or failed, and what must change. To explain at Level 6, connect the process to its purpose, workplace application and consequences—not simply describe a form.

    3.4
    Earlier learning bridge · 3.4 → 4.1 → 4.2

    From Evidence to Action—and From Action to a Reliable System

    Today’s three outcomes form one continuous journey. We first investigate the event to understand what happened and why. We then manage the organisational response from immediate control to verified closure. Finally, we use policy, responsibilities and records to make that good practice consistent every time.

    “Do not treat investigation, incident management and policy as three separate subjects. Follow one event all the way through. Ask why it happened, decide what the organisation must do, and then build the system that makes the right response repeatable. Finding a cause is not the finish line—verified prevention is.” — Debjyoti Biswas, FIIRSM, CertIOSH
    3.4 · InvestigateTest the evidence, reconstruct the sequence and explain immediate, underlying and organisational causes.
    4.1 · ManageCoordinate life safety, control, escalation, reporting, investigation, recovery, action and verified closure.
    4.2 · GovernDefine the policies, roles, procedures and trustworthy records that make the response repeatable.
    One event → tested evidence → coordinated response → governed learning → safer work
    3.4 in one minute · Follow the evidence

    A Report Starts the Question—Investigation Tests the Answer

    What 3.4 is really asking: explain how and why a proportionate investigation preserves and tests evidence, reconstructs the sequence, examines control performance, identifies causes, selects stronger action and verifies that risk has genuinely reduced.
    A multidisciplinary team photographs, measures and reviews evidence at a controlled forklift incident scene
    Practical investigation scene: the team documents the barrier and route before avoidable change, compares physical and digital evidence, and listens to people without forcing a blame story.
    1 · Preserve what can disappearScene position, measurements, electronic logs, memories and temporary conditions are perishable.
    2 · CorroborateCompare people, place, plant, paper and digital evidence rather than relying on one source.
    3 · Test alternativesAsk what evidence would exist if a different explanation were true.
    4 · Examine barriers and systemsLook beyond the immediate act to design, workload, change, maintenance and ownership.
    5 · Verify the changeAn action is not closed until field evidence shows that it controls the demonstrated risk.
    Observe → preserve → log → corroborate → reconstruct → analyse causes and barriers → control → implement → verify → share learning
    Learning, not a search for someone to blame

    What Is an Incident Investigation?

    It is

    A planned examination of reliable physical, documentary, digital, organisational and human evidence. It reconstructs the event, tests alternative explanations, identifies immediate, underlying and root causes, and produces proportionate action.

    It is not

    A blame interview, an assumption that the last person caused everything, a report written to defend a predetermined conclusion, or a list of retraining actions unsupported by evidence.

    Purpose test: A strong investigation increases the organisation’s ability to prevent recurrence and reduce wider risk. A beautifully formatted report that leaves failed controls unchanged has not achieved that purpose.
    Practical 3.4 focus

    Begin with evidence, not a predetermined explanation

    The first story may be plausible, but a professional investigation tests it against physical facts, records, independent accounts and credible alternatives. The work is not complete when the report is signed; it must show that the risk has been reduced.

    Use this discipline

    Observe → preserve → corroborate → analyse → control → verify. Record uncertainty honestly. A finding can be useful without pretending that every detail is known.

    Why Investigation Is Important

    Prevention

    Finds control weaknesses before the same pathway—or a related pathway—produces greater harm.

    Understanding

    Separates evidence from assumption and connects immediate events to job, organisational and system conditions.

    Control assurance

    Shows which barriers worked, failed, were missing or were never verified.

    People and trust

    Demonstrates concern, involves workers fairly and supports those affected when communication is honest and respectful.

    Governance

    Provides evidence for legal, insurance, leadership, worker-representation and resource decisions.

    Organisational learning

    Transfers lessons to similar assets, tasks, contractors, sites and design standards—not just the event location.

    Set the Terms of Reference Before the Team Starts

    Terms of reference means the written mandate that tells the investigation team what it is authorised and expected to do. It protects focus and fairness without fixing the answer in advance.

    ElementQuestion it must answerWhy it matters
    Purpose and scopeWhat event, period, activities, interfaces and credible consequences are included?Prevents an investigation that is either superficial or unmanageably broad.
    Authority and independenceWho commissions the work, who may secure evidence, and is the lead sufficiently free from the decisions being examined?Allows access and challenge while managing conflicts of interest.
    Team and competenceWhich operational, technical, HSE, human-factors and worker perspectives are required?No single specialist sees the whole socio-technical system.
    Evidence interfacesHow will evidence be protected, logged and coordinated with regulators, police, insurers or equipment specialists where relevant?Avoids loss, contamination, unlawful access or conflict with another authority.
    People, privacy and communicationHow will affected people be supported, consulted and updated, and who may access identifiable health or witness information?Protects dignity, trust, lawful handling and evidence quality.
    Timetable and outputsWhat interim controls, updates, report, actions, owners and effectiveness review are required?Links investigation to timely risk reduction rather than report production alone.

    Who Does What? Investigation Governance and Roles

    One person may hold more than one role in a small review, but authority, competence, independence and conflicts of interest must still be explicit.

    SponsorSets the mandate, resources and authority; does not predetermine the finding.
    Investigation leadPlans the work, protects fairness, tests evidence and integrates the team’s reasoning.
    Evidence custodianLogs identity, condition, access, storage and transfer so integrity is traceable.
    Operations or technical specialistExplains plant, process and work context without marking their own assumptions as fact.
    Human-factors adviserExamines task demands, design, workload, communication and organisational influences.
    Worker representativeProvides workforce insight, supports consultation and challenges impractical conclusions.
    Action ownerAccepts a cause-linked action, resources it and supplies implementation evidence.
    Independent reviewerTests whether evidence supports the findings and whether verification is credible.
    Conflict-of-interest safeguard: operational knowledge is valuable, but a person should not have unchecked authority to investigate and approve conclusions about decisions for which they were directly responsible. Use transparent declarations, multidisciplinary challenge and independent review proportionate to the event.
    Investigate the potential, not only the outcome

    Choose a Proportionate Investigation

    LevelTypical triggerPeople and depthOutput
    Local learning reviewLow actual and credible potential, understood local issue, no sign of wider control failureCompetent supervisor with worker involvement; factual sequence and local control checkRecorded learning, proportionate action and closure evidence
    Formal investigationMedical treatment, material loss, repeat event, uncertain causes, important barrier failure or significant potentialIndependent-enough multidisciplinary team with HSE and technical competenceEvidence pack, timeline, causal analysis, prioritised actions and management review
    Major / specialist investigationFatality, life-changing harm, major release, multiple people, catastrophic potential, complex system or public/regulatory significanceSenior mandate, specialist and worker-representative input, controlled interfaces with authoritiesFormal terms of reference, rigorous evidence governance, systemic learning and executive assurance
    Proportionality rule: Base depth on the worst credible consequence and recurrence potential as well as actual harm. A near miss caused by failure of a critical barrier may deserve more investigation than a minor injury with a simple, well-understood cause.

    Interactive Investigation-Level Selector

    Assess actual loss, credible potential, recurrence and complexity.
    Proportionate does not mean superficial

    How Deep Should the Investigation Go?

    The investigation level is a governance choice about authority, independence, competence and depth. Organisations use different names, but the six levels below help learners see the available range. The correct choice depends on the event and jurisdiction—not the label alone.

    Possible levelTypical purposeTeam and independenceTypical output
    Supervisor-level investigationPrompt learning from a low-potential, well-understood local event.Competent supervisor with the people doing the work.Facts, local causes, action, owner and closure evidence.
    Preliminary investigationEstablish enough evidence to classify risk, preserve material and decide whether a fuller investigation is needed.Supervisor or HSE lead with suitable technical support.Initial sequence, potential, evidence map, interim controls and escalation recommendation.
    Full internal investigationExamine a significant, repeated, uncertain or multi-factor organisational event.Multidisciplinary team sufficiently independent from the decisions examined.Terms of reference, evidence pack, causal and barrier analysis, approved actions and verification plan.
    Independent investigationProvide additional credibility or specialist challenge after severe, sensitive or governance-significant loss.External or organisationally separate lead with no material conflict.Independent findings, systemic recommendations and assurance to senior governance.
    Joint employer–worker investigationCombine duty-holder knowledge with workforce experience and representation.Employer and worker representatives with agreed access, competence and evidence safeguards.Shared fact-finding, practical recommendations and stronger workforce learning.
    Regulatory or criminal investigationEstablish facts under statutory authority and determine regulatory or criminal matters.Competent authority using its legal powers; organisational teams must not obstruct it.Official evidence, decisions and possible enforcement or prosecution process.
    Parallel-process rule: an internal safety investigation may continue only with appropriate coordination. Preserve evidence, obtain competent advice, respect authority directions and avoid contaminating witnesses or representing an internal conclusion as a legal verdict.

    Eight Selection Questions

    Actual severityWhat harm or loss occurred?
    Credible potentialWhat was the worst realistic outcome?
    RecurrenceHas the pathway or warning appeared before?
    Regulatory significanceMay a formal authority route or evidence interface apply?
    Public or environmental impactCould people off site, ecosystems or public confidence be affected?
    ComplexityDo technical, human, contractor or organisational factors interact?
    Control failureDid a critical prevention or mitigation barrier fail or go unverified?
    Learning valueCould findings improve equivalent work, sites or standards?

    What Competence Does the Team Need?

    Technical knowledge

    Understands the plant, substances, task, design limits and relevant control standards.

    Investigation method

    Can scope, plan, reconstruct, analyse and communicate without forcing a preferred answer.

    Interviewing skill

    Uses open, non-leading questions and supports affected people fairly.

    Evidence management

    Protects identity, integrity, source, access, transfer and uncertainty.

    Legal awareness

    Recognises when specialist advice or coordination with an authority is needed.

    Independence

    Declares conflicts and can challenge decisions without improper pressure.

    Worker participation

    Brings work-as-done knowledge and tests whether recommendations are usable.

    Team integration

    Combines disciplines, resolves evidence conflict and communicates limitations.

    A complete evidence-to-improvement route

    The Incident-Investigation Lifecycle

    1. Care, control and preserve.
    Protect people and environment, prevent escalation and disturb the scene only for safety, rescue or essential control.
    2. Appoint and scope.
    Set a competent, sufficiently independent team; define authority, boundaries, interfaces, timescale and communication.
    3. Plan the evidence.
    Identify what is perishable, who must be consulted, what expertise is required and how evidence integrity will be recorded.
    4. Gather information.
    Secure photographs, measurements, parts, records, electronic data, conditions and fair accounts from involved people.
    5. Build the sequence.
    Place verified events on a timeline; label fact, inference, conflict, uncertainty and missing evidence.
    6. Analyse causes and controls.
    Test immediate, underlying and root causes; successful, failed and absent barriers; human and organisational influences.
    7. Test alternatives.
    Ask what evidence would be expected if another explanation were true. Do not stop at the first plausible story.
    8. Select risk controls.
    Use the hierarchy, address system causes, define interim safeguards and avoid over-reliance on warnings or retraining.
    9. Report and communicate.
    Explain evidence, reasoning, uncertainty, lessons, actions, owners and dates; protect personal information.
    10. Implement and verify.
    Track completion, test whether the action works in practice and check equivalent risks across the organisation.
    HSE learning structure: HSG245 groups the investigation into four connected stages—gather information, analyse it, identify suitable risk-control measures, and plan/implement action. The ten steps above make those four stages operational for learners. HSG245 remains a valuable investigation framework, but any historical references it contains to statutory reporting must not replace the current HSE RIDDOR pages and the organisation’s competent legal check.
    Four HSG245 stages · fifteen practical steps

    How Does the Full Investigation Process Connect?

    HSE’s HSG245 gives a clear four-stage learning structure. The more detailed lifecycle below does not replace it; it shows where the practical governance, evidence, communication and verification activities sit.

    HSG245 stagePlain-language questionPractical activities in this portalOutput
    1. Gather informationWhat happened, in what conditions, and what evidence exists?Immediate response, scene control, level selection, team appointment, terms of reference, evidence plan and collection.Protected evidence, factual sequence, gaps and uncertainties.
    2. Analyse the informationHow and why did the event become possible?Timeline reconstruction; causal, barrier, human-factors and organisational analysis; alternative explanations.Supported immediate, underlying and root/system findings.
    3. Identify suitable risk-control measuresWhat change will address the demonstrated causes?Generate options, apply the hierarchy, consider similar exposure and define success measures.Prioritised interim and permanent controls linked to findings.
    4. Develop and implement an action planWho will do what, by when, and how will we know it worked?Approval, communication, ownership, resources, implementation, effectiveness verification and closure.Completed and verified risk reduction with transferred learning.

    The Fifteen-Stage Route—from Notification to Closure

    1. Immediate response.
    Care for people, raise alarms, contain escalation and preserve life before evidence.
    2. Scene control.
    Set a safe boundary, record necessary changes and prevent avoidable loss or contamination.
    3. Select the investigation level.
    Use actual harm, credible potential, recurrence, significance, complexity and learning value.
    4. Appoint the team.
    Match technical, method, people, legal and evidence competence; declare conflicts.
    5. Agree terms of reference.
    Set purpose, authority, scope, interfaces, outputs, communication and timetable without fixing the answer.
    6. Create the evidence plan.
    Prioritise perishable sources, authorisations, specialist tests, access and storage.
    7. Collect and log evidence.
    Secure people, place, plant, records and process evidence with traceable identity and condition.
    8. Reconstruct the timeline.
    Arrange verified events from the last normal point; show conflicts, gaps and uncertainty.
    9. Analyse causes and barriers.
    Select suitable methods; examine immediate, underlying, systemic and recovery factors.
    10. Develop recommendations.
    Link each action to an established cause, use the hierarchy and define effectiveness evidence.
    11. Review and approve the report.
    Test evidence, reasoning, fairness, limitations, actions and governance acceptance.
    12. Communicate findings.
    Support affected people, brief authorised stakeholders and share usable, privacy-protected learning.
    13. Implement action.
    Resource named owners, manage dependencies and retain interim controls until permanent change is safe.
    14. Verify effectiveness.
    Observe real work, measure control performance, consult users and check for unintended risk.
    15. Formally close and transfer learning.
    Independent-enough evidence confirms the action works and equivalent operations have been considered.
    Currency boundary: HSG245 is official HSE guidance on investigation method, not legislation. It was published in 2004, so its historical statutory-reporting references must not be used instead of current RIDDOR guidance.
    Evidence before explanation

    Gather, Protect and Test the Evidence

    Evidence familyExamples in the master caseQuality question
    PeopleDriver, contractor, operators, supervisor, maintenance planner, dispatcher and worker representativeWas the account gathered promptly, respectfully and without leading or blame?
    Position and placeRoute width, pallet positions, lighting, sight lines, floor marks, spill boundary and weather/shift conditionsWas the scene photographed and measured before avoidable change?
    Plant, parts and substancesForklift, low barrier, valve manifold, damaged fixings, chemical and alarm/isolation equipmentIs identity, condition and chain of custody traceable?
    Paper and electronic recordsRisk assessment, layout, training, inspections, defect tickets, change records, CCTV, alarms, radio and maintenance historyIs the record authentic, complete and understood in context?
    Performance and organisational dataNear-miss trends, dispatch volume, overtime, backlog, action closure, contractor interfaces and previous learningAre definitions, periods, denominators and possible under-reporting understood?

    The Minimum Evidence Log

    Evidence ID → description → original source or location → collector → date and time → condition → storage or transfer → relevance → access restriction

    Record any unavoidable scene change, copy, conversion, repair, test or transfer. “Chain of custody” means the documented history of who controlled an item and what happened to it; the level of formality should match the event and any regulator, police, legal or insurance interface.

    Formal-process boundary: a safety investigation must not interfere with a regulator, police or other authorised process. Agree evidence access and preservation arrangements with the relevant authority, protect personal information and avoid statements that pre-judge legal liability.

    Four Tests for a Defensible Finding

    1. Relevance and causal link

    Does the evidence help explain the sequence, control performance or conditions that made the event possible—or is it merely interesting background?

    2. Corroboration

    Is the conclusion supported by more than one reliable source, such as measurement plus records or an account plus digital evidence?

    3. Alternatives and conflict

    What other explanation could fit, and what supporting or conflicting evidence would be expected if it were true?

    4. Confidence and limitations

    What is known, inferred, disputed or missing, and how strongly can the team state the conclusion?

    Interactive Finding-Quality Coach

    Choose a statement and test its causal support, alternatives and limitations.

    Interview for learning

    Explain purpose and support; choose a suitable place; ask open questions; invite the person’s own sequence; distinguish observation from inference; explore what made sense at the time; confirm understanding; and allow correction.

    Avoid contamination

    Do not ask “Why were you careless?” or circulate a shared story before accounts are captured. A better prompt is: “Please take me through what you saw, heard and did, from the last normal point.”

    Interactive Evidence-Preservation Diagnostic

    0 / 8 evidence safeguards selected. Prioritise perishable evidence without delaying care or emergency control.
    A memory aid—not a substitute for an evidence plan

    What Evidence Should the Team Seek?

    A multidisciplinary investigation team reviews a plan, a tagged component, digital timeline and protected photographs
    Evidence-led teamwork: different disciplines compare the scene, plant, records and people’s accounts. No single photograph, document or recollection is treated as the whole story.
    Secure before interpretingPreserve identity, position, condition, time and access before testing a theory.
    Use several familiesPhysical, digital, documentary, measurement and human evidence can support—or challenge—each other.
    Reliability is contextualA reliable source may still answer only part of the question; authenticity does not automatically prove causation.
    State uncertainty“Not yet known” is more professional than invented certainty.

    The “5 Ps” Evidence Check

    The five Ps are a DB HSE learning aid for remembering evidence families. HSG245 discusses physical, verbal and written evidence but does not formally prescribe this mnemonic.

    PeopleInvolved workers, witnesses, supervisors, specialists, affected people and representatives.
    Position / placeLayout, distances, lines of sight, marks, conditions, location and changes to the scene.
    Parts / plantEquipment, components, substances, settings, damage, protection and test results.
    PaperworkRisk assessments, procedures, permits, training, inspection, maintenance, design and change records.
    Processes / proceduresHow work was planned, resourced, supervised and actually performed; organisational and contractor interfaces.

    Evidence Reliability Ladder—Use with Caution

    Evidence typeTypical strengthQuestion before relying on it
    Contemporaneous physical evidenceDirect condition, position, damage or residue close to event time.Was the scene changed, item identified and condition protected?
    Digital or system recordTime-stamped CCTV, alarm, historian, access or communication data.Are clocks synchronised, data complete and export authentic?
    Controlled documentApproved design, procedure, inspection or maintenance record.Was it current, available and actually used in the work?
    Measurement or testQuantified distance, force, exposure, condition or performance.Was the method suitable, calibrated, repeatable and representative?
    Witness recollectionExplains perception, sequence, context and why actions made sense.How soon was it gathered, was the question neutral, and what corroborates it?
    Opinion or assumptionCan generate a hypothesis or identify specialist questions.What evidence would support or disprove it? Never relabel it as fact.
    Important: This is not a rigid legal ranking. A precise digital timestamp may be misleading if clocks differ; a witness may be the only source of task context; physical damage may show what happened but not why. Strong findings explain source, relevance, corroboration, conflict and limitation.

    Interactive Evidence-Source Coach

    Choose an evidence source to see what it can show, what it cannot prove and how to corroborate it.
    Apply 3.1, 3.2 and 3.3

    From Evidence to Causes, Controls and Action

    LevelEvidence-based findingControl implication
    ImmediateThe forklift entered the narrowed route and contacted a low barrier that did not prevent contact with the manifold.Restore control of the route and equipment integrity before restart.
    Underlying job factorsPallet storage reduced clearance; peak dispatch, visibility and radio interruption shaped the task; temporary adaptations had become normal.Redesign storage and traffic flow; review capacity, communication and work scheduling.
    Underlying organisational factorsRepeated barrier defects were recorded without effective risk escalation, ownership or closure verification.Repair the defect-management, prioritisation and assurance process across comparable barriers.
    Root / systemicWarehouse, production and maintenance changes were managed separately, so the combined vehicle–chemical interface was not owned as one critical risk.Create accountable interface ownership, change control and critical-control performance standards.
    Successful recoveryDetection, alarm, withdrawal and emergency isolation limited the actual release.Preserve and verify these controls; do not let success hide the preventive-control failure.

    Barrier Performance Review—Link Back to 3.1

    Use Bowtie and Swiss Cheese thinking without simply redrawing the models. For each barrier, ask what it was meant to do, whether a demand occurred, what evidence shows performance and why any weakness developed or persisted.

    BarrierIntended functionEvidence and statusCause-linked actionVerification
    Traffic-route clearanceKeep vehicles separated from process equipmentPallets reduced usable width; barrier failed before contactRedesign storage and physically control the clear zonePeak-shift measurement, observation and encroachment trend
    Vehicle restraint barrierPrevent vehicle reach to the manifoldLow, damaged barrier did not resist the demandEngineer a rated segregation system and protect foundationsDesign acceptance, inspection and controlled impact criteria
    Defect escalationPrioritise and close safety-critical damageRepeat defects existed without verified risk ownershipSet criticality, owner, deadline, interim control and escalationBacklog audit and sample field verification of closed defects
    Alarm and isolationLimit consequence after loss of containmentWorked and restricted the actual releasePreserve, maintain and test the successful recovery controlsFunction test, response-time evidence and drill observation

    Interactive Cause-to-Control Builder

    Connect evidence to causes, then match the action to the system weakness and define verification.
    Choose the method to fit the question

    Which Investigation Method Fits the Evidence?

    A method is a disciplined way to organise questions; it is not a machine that automatically discovers “the root cause.” Use the simplest method that can represent the event fairly, and combine techniques when different questions require different views.

    MethodBest question or useStrengthLimitation
    Five WhysWhy did a relatively simple, evidence-rich problem persist?Fast and accessible; pushes beyond the immediate event.One questioning path can oversimplify, reflect facilitator bias or stop too early.
    Fishbone / IshikawaWhich categories of contributing factors should the team explore?Encourages broad multidisciplinary brainstorming.A list of possibilities does not establish sequence, evidence or causation.
    Timeline analysisWhat happened before, during and after the event?Makes gaps, conflicts and parallel actions visible.Chronology alone does not explain why conditions existed.
    Change analysisWhat differed from an earlier safe state, design or expectation?Useful after modification, drift, staffing or process change.Can miss long-standing weaknesses that did not recently change.
    Barrier analysisWhich prevention, control and recovery barriers worked, failed or were missing?Connects evidence directly to risk-control performance.Weak barrier definitions or performance standards produce vague conclusions.
    Fault Tree Analysis (FTA)Which combinations of failures could produce a defined top event?Shows AND/OR logic and can support probability analysis.Depends on scope, independence assumptions and defensible input data.
    Event Tree Analysis (ETA)What outcomes can follow an initiating event as barriers succeed or fail?Shows consequence pathways and recovery importance.Branch dependence and omitted pathways can mislead.
    BowtieHow do threats, a top event, consequences and preventive/recovery barriers connect?Clear control-focused communication across disciplines.Can become a static picture and hide dynamic interactions or barrier degradation.
    Tripod BetaWhich failed barriers, preconditions and organisational influences shaped the event?Structured connection from active failures to latent conditions.Needs trained use and can be disproportionate for a simple local event.
    AcciMapHow did decisions and interactions across regulators, organisation, management and work contribute?Useful for complex multi-level systems and public interfaces.Resource intensive; broad maps still need evidence-based prioritisation.
    STAMPHow did inadequate constraints and feedback across a complex sociotechnical control system allow loss?Represents software, organisations, adaptation and non-linear interaction.Requires specialist competence and careful boundary definition; often excessive for simple events.

    Interactive Investigation-Method Selector

    Choose the question; the coach will suggest a primary method, a useful companion and a limitation to manage.

    Human Factors—Why Did the Action Make Sense at the Time?

    Human-factors analysis examines the interaction between people, task, equipment, environment and organisation. It does not excuse unsafe conduct; it avoids the weak assumption that naming an error explains the system that produced it.

    Fatigue and workload

    Hours, rest, pace, task switching, physical demand and cumulative workload.

    Competence and supervision

    Knowledge, practice, authorisation, coaching, availability and span of control.

    Communication

    Handover, language, radio, alarm meaning, coordination and shared situational awareness.

    Usability and design

    Controls, displays, access, visibility, procedure usability and error-tolerant design.

    Time and production pressure

    Targets, staffing, incentive, scheduling, backlog and conflicting priorities.

    Culture and adaptation

    What leaders reward, what is tolerated, how deviations become normal and whether concerns are acted on.

    Work as imagined versus work as done: “Work as imagined” is how a designer, manager or procedure expects the task to occur. “Work as done” is how people actually achieve the task under real constraints and variation. The gap is evidence to understand—not automatic proof of worker misconduct.

    Just Culture—Fair Accountability Without Automatic Blame

    Human error

    An unintended slip, lapse or mistake. Console and support the person, then improve design, conditions, barriers and learning.

    At-risk behaviour

    A choice where risk is underestimated, normalised or perceived as justified. Understand incentives and context; coach and redesign the system.

    Reckless behaviour

    A conscious, unjustifiable disregard of a substantial and obvious risk. A fair accountability process may be appropriate after evidence, due process and system context are considered.

    Fair accountability

    Apply consistent criteria, separate investigation from predetermined discipline, consider capacity and intent, protect reporting and avoid outcome bias.

    Boundary: This learning model is not a legal finding or automatic disciplinary formula. Do not classify conduct from outcome severity alone; use reliable evidence, employment procedures, representation rights and competent jurisdiction-specific advice.

    Interactive Just-Culture Reflection Coach

    Choose an evidence pattern to see a fair response and the system questions that must still be asked.

    Legal and Ethical Safeguards

    ConfidentialityRestrict identifiable health, witness and personal information to authorised need.
    Data protectionUse a lawful purpose, minimum necessary data, secure access and controlled retention.
    Legal privilegeObtain legal advice where needed; do not assume labelling an internal document makes it privileged.
    Worker representationInvolve representatives proportionately and early while preserving evidence and privacy.
    Evidence disclosureKeep material authentic, traceable and available for lawful authority or process requirements.
    Parallel investigationCoordinate with police, regulators or other authorities; never obstruct or contaminate their work.
    Non-retaliationProtect honest reporting and participation; concerns about conduct require a separate fair process.
    Respect and supportCommunicate with affected people humanely and avoid avoidable re-traumatisation or public exposure.
    The impact depends on investigation quality

    What Changes When Investigation Is Done Well—or Badly?

    AreaStrong investigation impactWeak or blame-led impact
    RiskSystem weaknesses and equivalent exposures are controlled.Superficial causes remain and recurrence becomes likely.
    WorkersPeople are supported, heard and more willing to share evidence.Fear, distress, silence and adversarial accounts increase.
    DecisionsResources target higher-order controls with owners and verification.Action lists favour reminders, retraining and paperwork without risk reduction.
    Evidence and complianceReasoning is traceable and reporting, claims and governance use reliable facts.Delay, contamination, inconsistency or concealment undermines confidence.
    Learning cultureSuccessful and failed barriers are shared across similar work.The report is filed, lessons stay local and organisational memory decays.
    BusinessRecurrence, downtime, loss cost and uncertainty can reduce over time.Investigation consumes time but returns little value; repeat loss enlarges cost and reputation damage.
    Balanced impact: A rigorous investigation uses time, specialist skill and operational attention; it may uncover uncomfortable failures or require expensive change. Those are real impacts. They should be managed through proportional scope, support, confidentiality and prioritisation—not by weakening the search for preventable causes.
    Closure has two gates: first verify implementation—the agreed action exists, is owned and has been introduced. Then verify effectiveness—the control works during real operating conditions, addresses the demonstrated cause, transfers where needed and has not created a new risk. Passing the first gate alone is administrative completion, not prevention.

    A Defensible Investigation Report Structure

    1 · Mandate and scopeEvent, purpose, authority, team, conflicts, boundaries and limitations.
    2 · Verified sequenceFacts, sources, scene changes, uncertainties and disputed points.
    3 · AnalysisImmediate, underlying and systemic causes; successful, failed and absent barriers.
    4 · ConclusionsFindings supported by evidence, alternatives considered and confidence stated.
    5 · Risk controlsCause-linked interim and permanent action using the hierarchy of control.
    6 · DeliveryNamed owners, resources, dates, dependencies and escalation of delay.
    7 · VerificationImplementation evidence, effectiveness measure, reviewer and review date.
    8 · Learning and recordsCommunication, privacy, equivalent-risk review and controlled retention.

    Example 30–60–90 Day Verification Plan

    0–30 days · Stabilise

    Confirm interim controls, protect people, assign owners, preserve evidence and stop uncontrolled recurrence while permanent action is designed.

    31–60 days · Implement

    Introduce engineering and system changes, consult users, update linked documents and train only where competence genuinely forms part of the cause and control.

    61–90 days · Verify

    Observe real work, inspect barrier performance, test worker understanding, review repeat signals and confirm equivalent sites or tasks received the learning.

    The dates are a teaching example, not a universal deadline. Urgent controls must not wait, and complex engineering work may require a longer governed plan. Verification criteria should be decided when the action is agreed—not invented at closure.

    Interactive Corrective-Action Quality Coach

    Evaluate control strength, ownership, interim safety and verification.

    Interactive Investigation-Quality Diagnostic

    0 / 10 investigation-quality elements selected.

    Interactive 3.4 “Explain” Response Builder

    Explain what, how, why, application, impact and verification in one connected argument.
    A finding is not yet prevention.
    Before 3.4 is complete, test whether the recommended action was implemented, works in real conditions and transfers to equivalent risks.
    Verify action before closure
    Completion is not the same as effectiveness

    When Is Corrective Action Truly Closed?

    An engineer, operator and worker representative measure a new engineered segregation barrier and review verification evidence
    Verified closure: the engineered barrier exists, but the team also measures the installation, checks the real interface and records evidence. Closure requires proof of risk reduction—not a tick beside “completed.”
    Implementation evidenceDesign approval, installation record, inspection and operating-document update show the change exists.
    Effectiveness evidenceObservation, measurements, barrier tests, repeat-event data and worker feedback show it controls the cause in real work.
    Transfer evidenceEquivalent routes, assets, contractors and sites were screened and treated where needed.
    Unintended effectsThe change did not create restricted access, new trapping points, unsafe workarounds or emergency-response delay.

    Eight Quality Tests for Every Recommendation

    Cause-linkedAddresses a finding supported by evidence.
    Specific and measurableDefines the change and the observable success condition.
    Hierarchy alignedPrefers elimination, engineering or robust system control where practicable.
    OwnedNames an accountable person with authority to deliver.
    Time-boundSets a justified date and escalation for delay.
    ResourcedProvides budget, people, technical input and controlled dependencies.
    TransferableReviews similar operations and shared causes.
    VerifiableDefines who will test effectiveness, how and when.

    Interactive Recommendation-Quality Gate

    Select only the qualities that the proposed action genuinely demonstrates.

    0 / 8 recommendation qualities selected. Completion and effectiveness are separate gates.

    Completed Example—from First Notification to Verified Closure

    StageWhat the team didEvidence or decision produced
    1. Initial notificationThe operator raised the alarm after forklift contact beside manifold V-12; first aid, withdrawal and isolation followed.Time, place, people, actual symptoms, credible potential and immediate controls entered under one event ID.
    2. Scene and evidenceThe area was cordoned; barrier position, pallet layout, marks and route width were photographed and measured; CCTV and alarm data were secured.Traceable physical, digital, documentary and human evidence with scene changes recorded.
    3. Level and teamHigh credible chemical-release potential, repeated barrier defects and cross-department interfaces justified a full internal investigation.Multidisciplinary team with worker representation, technical skill and independent review.
    4. ReconstructionThe team synchronised CCTV, alarm and dispatch records with separate accounts from the driver, contractor and operators.A tested timeline from the last normal point, including gaps and uncertainty.
    5. AnalysisTimeline and barrier analysis showed route encroachment, degraded physical protection, weak defect escalation and divided interface ownership.Immediate, underlying and systemic findings; alarm and isolation recorded as successful recovery controls.
    6. RecommendationsThe team selected rated segregation, controlled storage, critical-defect escalation and accountable interface change control.Cause-linked actions using stronger controls rather than reminders alone.
    7. Approval and communicationEvidence, limits, owners, dates and learning were reviewed; affected people and the workforce received privacy-protected updates.Approved report, action plan, interim controls and feedback record.
    8. ImplementationEngineered segregation and a clear pedestrian route were installed; storage limits and defect workflow were changed.Design acceptance, installation evidence, updated controls and owner sign-off.
    9. Effectiveness verificationAn independent-enough reviewer measured clearances, observed peak work, sampled defect closure and consulted users over three months.Barrier performance, no repeated encroachment, timely critical-defect escalation and no harmful workaround.
    10. Closure and transferComparable process routes and contractor interfaces were screened; lessons and standards were transferred.Formal closure based on verified risk reduction—not report signature alone.
    Worked Level 6 “explain” example:

    Incident investigation is important because it converts a reported event into tested knowledge and preventive action. In the forklift case, early scene control and evidence preservation protected measurements, CCTV, alarm data and people’s accounts before they changed. A multidisciplinary team then reconstructed the sequence and used barrier analysis to connect the contact not only to route encroachment, but also to damaged protection, weak defect escalation and divided ownership of the vehicle–process interface. This matters because action based only on the driver’s final movement would leave the organisational conditions intact. Engineering segregation, controlled storage and accountable critical-defect management were therefore linked to established causes. Their impact became credible only when implementation and field effectiveness were verified across real peak-shift work and similar routes. A strong investigation can reduce recurrence, improve trust and direct resources to system controls; a weak or blame-led investigation can suppress future reports, preserve failed barriers and allow a more serious event to recur.

    3.4 now closes the prevention loop.
    Evidence was tested, causes were connected to action and effectiveness was verified. Bring the complete reasoning into one Level 6 response.
    Practise all four outcomes
    Assessment-ready learning support

    How to Answer 3.1–3.4 at Level 6

    3.1 — Outline

    For each theory or technique: state its name and origin/context, describe its principal structure, explain its direction or logic, apply it briefly to an incident, and state one useful feature and limitation.

    Suggested structure: Name → main idea → components → application → usefulness → limitation.

    3.2 — Justify

    Identify the decision, explain why the selected quantitative method fits, show the data and formula, interpret the result, discuss reliability and limitations, combine it with qualitative evidence, and state the resulting action and review.

    Suggested structure: Decision → method → evidence → calculation → interpretation → limitation → complementary evidence → action.

    3.3 — Assess

    Weigh the reporting need, affected stakeholders, urgency, route and benefit against burden, privacy, trust and other adverse effects. Examine what non-reporting would cause, show how tensions can be controlled, and reach an evidence-based judgement.

    Suggested structure: Trigger → need → route → benefit → adverse impact → mitigation → consequence of silence → judgement.

    3.4 — Explain

    State what investigation is, show how evidence becomes causal understanding and control, explain why each stage matters, apply it to a realistic event, and connect strong or weak investigation quality to people, risk and organisational outcomes.

    Suggested structure: Meaning → process → reason → workplace application → positive impact → consequence if weak → verification.

    Integrated practice task

    Using the forklift–process-line case: outline the causation theories and techniques; justify suitable quantitative methods; assess the need and impact of internal and possible external reporting; and explain how a proportionate, fair investigation would convert evidence into verified risk reduction.

    Section 03 knowledge check

    Can You Connect Models, Data and Investigation?

    1. What is the best use of an incident triangle?
    2. What is the first domino in Bird’s expanded loss-causation sequence?
    3. What does multi-causality emphasise?
    4. What do the holes in Swiss Cheese represent?
    5. Which technique starts with a top event and reasons backwards?
    6. What does ETA mainly explore?
    7. What sits at the centre of a Bowtie?
    8. Which is a sound behavioural RCA approach?
    9. Why divide cases by hours worked?
    10. What is the strongest Level 6 conclusion?
    11. An event is not externally notifiable. What follows?
    12. Why assess credible potential as well as actual harm?
    13. Near-miss reports rise after a trusted mobile route launches. What is the best interpretation?
    14. Which event may justify a deep investigation?
    15. Which is the strongest first interview prompt?
    16. When is an investigation action truly closed?
    Answer all sixteen questions, then check your score and explanations.
    Unit 3 · Section 04

    Understand Processes and Strategies to Manage Health and Safety Incidents in an Organisation

    Section 04 moves from understanding individual loss events to managing and assuring the complete organisational response. Learners follow an incident from readiness and first response through evidence, investigation, corrective action, lawful record maintenance, safe recovery, evaluation and verified organisational learning.

    Learning outcomes covered: 4.1 Outline the critical stages for managing incidents in the organisation. 4.2 Outline organisational policies to identify, investigate, report and record incidents. 4.3 Explain how to maintain records of incidents to meet regulatory and statutory requirements. 4.4 Evaluate an organisational process for managing health and safety incidents.

    Independent DB HSE learning resource: prepared solely by Debjyoti Biswas to support OTHM Level 6 Unit 3 teaching. It is not produced or endorsed by OTHM or Ofqual.
    Do not study the criteria as isolated boxes

    How 3.1–3.4 and 4.1–4.4 Work Together

    A mature incident system needs every part. Causation models organise possible explanations; loss data tests patterns; reporting makes events visible; investigation turns evidence into findings; incident management coordinates the response; policy makes good practice repeatable; records preserve defensible proof; and evaluation determines whether the system really works.

    3.1 · Explain whyUse causation theories and structured logic to examine how controls failed.
    3.2 · Test the patternUse valid quantitative information to understand frequency, severity and trends.
    3.3 · Make it visibleReport reliable facts to the people who must act.
    3.4 · Learn from evidenceInvestigate causes, controls and organisational conditions.
    4.1 · Coordinate actionManage the complete sequence from alarm to verified closure.
    4.2 · Make it repeatableUse policy, responsibility, procedure and records to govern every incident.
    4.3 · Preserve proofMaintain lawful, reliable, secure and retrievable incident records.
    4.4 · Test the systemEvaluate design, delivery and results; then prioritise improvement.
    The simple relationship: 3.3 and 3.4 explain two essential activities—reporting and investigating. Section 4 places those activities inside a wider organisational system that also protects life, controls escalation, assigns authority, supports recovery, verifies actions and retains trustworthy records.
    OUTLINE
    EXPLAIN
    EVALUATE
    How the command words change the answer

    4.1–4.2 outline: identify the principal stages or policy features and show what each involves. 4.3 explain: connect each recordkeeping method to the legal or regulatory requirement it satisfies and the consequence of failure. 4.4 evaluate: compare evidence against a benchmark, make a balanced judgement and propose prioritised, verifiable improvement.

    Plain language before technical detail

    What Is Incident Management?

    Incident management is the coordinated organisational process used to prepare for, respond to, control, report, investigate, recover from and learn from an event that caused—or could have caused—injury, ill health, damage, environmental harm or operational loss.
    Choose a termEvery abbreviation and technical term is explained before learners use it.
    One realistic case · Three connected moments

    Master Case: Forklift, Racking and Chemical Release

    A reversing forklift strikes warehouse racking and damages a cleaning-chemical container. One employee experiences eye irritation, a contractor narrowly avoids falling material, liquid moves towards a drain, CCTV is available, and management wants the area reopened quickly. The employee later reports skin symptoms.

    Why this case is useful: it involves workers, a contractor, occupational health, operations, maintenance, environment, management and possible external authorities. The same event now runs through all four criteria: manage it in 4.1, govern it in 4.2, maintain its records in 4.3 and evaluate the process in 4.4.
    Learning outcome 4.1
    Central question: What must the organisation do—from the first warning until evidence shows the incident is properly closed?

    Outline the Critical Stages for Managing Incidents in the Organisation

    The stages form a controlled sequence, but some activities happen in parallel. Emergency response, internal escalation, evidence protection and any legally required notification must be coordinated rather than forced into a rigid queue.

    4.1
    Before Stage 1 · Readiness is the foundation

    An Organisation Cannot Invent Its Incident System During the Emergency

    Foresee credible events

    Use risk assessments, emergency planning, loss history and specialist advice to consider fires, releases, vehicle events, violence, structural failure, occupational-health events and other credible scenarios.

    Prepare competent roles

    Define incident control, first aid, evacuation, technical isolation, occupational health, communications and regulatory-notification authority.

    Provide resources

    Maintain alarms, communications, emergency equipment, spill control, access information, contact lists, evidence kits and alternative work arrangements.

    Train, exercise and review

    Practise arrangements, record learning, correct gaps and update plans after change. A paper plan that nobody can use is not readiness.

    HSE principle: emergency procedures should define foreseeable actions, competent people and roles; they should be recorded, rehearsed and reviewed. Work must not resume while serious danger remains.
    The full incident-management cycle

    Four Phases · Ten Critical Stages

    A

    Respond and stabilise

    Recognise, raise the alarm, protect life, establish control and prevent escalation.

    B

    Preserve and report

    Protect the scene, capture initial facts, classify the event and activate required routes.

    C

    Investigate and improve

    Gather and analyse evidence, select controls and assign governed corrective actions.

    D

    Recover and learn

    Authorise safe restart, support people, verify effectiveness, share learning and close.

    Recognise → protect life → command and contain → preserve and report → classify and escalate → investigate → control and act → recover → verify → learn and close
    Interactive stage explorer

    Select a Stage to See Its Purpose, Owner and Output

    Stage 1 · Recognise and activate

    Purpose: make the event visible quickly enough for a proportionate response.

    • Main actions: stop unsafe work, raise the alarm, identify location and immediate danger.
    • Typical owner: first person aware, supervisor or control point.
    • Output: verified alert and activated response route.
    Apply the first stages

    First 15 Minutes: Choose Immediate Actions and Investigation Depth

    First 15 Minutes Simulator

    Choose the safest coordinated actions, then test the response.

    Actual vs Credible-Potential Triage

    This learning tool selects organisational investigation depth. It does not decide legal reportability.

    The current example requires a proportionate team investigation because credible potential is greater than actual harm.
    Safety first · Evidence second · Verification before restart

    Can We Restart Safely? Preserve Evidence and Test Readiness

    Scene-Preservation Decision Coach

    Select a condition to decide how safety and evidence should be balanced.

    Safe-Restart Gate

    Select only what has been demonstrated with evidence.

    0 / 10 restart conditions confirmed. Do not restart.
    Important distinction: emergency stabilisation allows investigation to begin safely. It does not automatically authorise normal operation. Restart is a separate risk decision requiring defined evidence and accountable approval.
    4.1 gives us the sequence.
    4.2 now defines the policies, responsibilities, mandatory rules and controlled records that make the sequence reliable across shifts, departments, sites and contractors.
    Continue to 4.2
    Learning outcome 4.2
    Central question: What must the organisation formally require so that incidents are recognised, investigated, reported and recorded consistently?

    Outline Organisational Policies to Identify, Investigate, Report and Record Incidents

    Policy converts good intentions into approved expectations. It defines scope, responsibilities, authority and mandatory rules; linked procedures, forms and records then make those rules operational and demonstrable.

    4.2
    A frequent assessment and workplace confusion

    Policy, Procedure, Plan, Form and Record Are Not the Same

    PolicyApproved intent, principles, scope, responsibilities and mandatory organisational rules.
    ProcedureThe operational sequence: who does what, when, how and where the matter escalates.
    PlanPrepared arrangements for a foreseeable event, site, project or emergency.
    Form or templateStructured prompts used to gather consistent information.
    RecordCompleted evidence showing what happened, what was decided and what was done.

    Document Classifier

    A blank incident form is a tool; the completed form becomes a record. Neither is the policy itself.
    One umbrella policy · Several linked arrangements

    Incident Reporting, Investigation and Organisational Learning Policy

    An organisation may use one umbrella policy rather than four artificial policies. The essential test is whether the controlled system clearly covers identification, investigation, reporting and recording.

    1 · Purpose and commitmentPrevention, welfare, objective learning, compliance and continual improvement.
    2 · ScopeEmployees, agency workers, contractors, visitors, public, remote work and shared workplaces.
    3 · DefinitionsConsistent taxonomy for accidents, ill health, near misses, dangerous occurrences and high-potential events.
    4 · ResponsibilitiesNamed accountability from the worker and supervisor through competent specialists and senior leadership.
    5 · Reporting routesAccessible channels, required information, internal times, escalation and an alternative route.
    6 · Investigation rulesProportionate level, competence, objectivity, worker involvement, scope and time expectations.
    7 · Evidence controlScene, physical, digital, documentary, witness and occupational-health evidence.
    8 · External interfacesEmergency services, regulators, insurers, clients and environmental authorities where applicable.
    9 · Records and privacyMinimum fields, case ID, access control, medical confidentiality, retention and audit trail.
    10 · Corrective actionCause-linked actions, owners, resources, deadlines, interim controls and effectiveness checks.
    11 · Recovery and learningRestart authority, feedback, wider transfer and support for affected people.
    12 · GovernanceApproval, version control, competence, consultation, audit, review dates and change triggers.
    4.2 in organisational practice · DB HSE teaching architecture

    Organisational Incident Policy Suite

    OTHM asks learners to outline organisational policies used to identify, investigate, report and record incidents. A defensible organisation can meet those functions through one controlled umbrella policy supported by specialist policies and procedures. The eight-policy suite below is a practical learning model; OTHM does not prescribe these exact document titles.

    Why use a suite? One document states commitment and authority, while linked documents give enough detail for classification, reporting, investigation, evidence, notification, records, corrective action and shared-workplace coordination. A smaller organisation may combine documents, but it must not lose any of those control functions.
    Policy sets the ruleProcedure gives the stepsForm prompts informationRecord proves what happenedAssurance tests performanceReview improves the system
    Policy 1 of 8

    Incident Identification and Classification Policy

    Defines which events enter the organisational incident system and how their actual and credible potential consequences are classified.

    Mandatory rulesInclude injury, occupational ill health, near miss, dangerous occurrence, high-potential event and relevant property, environmental or operational loss.
    Accountable roleSystem owner with operational managers and competent health-and-safety support.
    Supporting procedureEvent taxonomy, classification matrix and escalation method.
    Required recordInitial event record plus the classification decision and its evidence.
    Monitoring and reviewSample classifications, examine reclassification and compare recurring event categories.
    Master-case applicationThe eye irritation, contractor near miss, chemical release and credible falling-stock potential enter one connected case.
    Interactive policy owner coach

    Which Policy Owns the Problem?

    Select a realistic failure. The coach identifies the primary policy and the supporting links needed to prevent a gap between documents.

    A problem may have one primary policy owner and several supporting policies. Select a situation to see the complete route.
    Boundary for the next criteria: 4.2 should establish the policy rules, responsibilities, procedures and records. Detailed statutory record-maintenance requirements belong mainly to 4.3, while evaluation of whether the process works belongs mainly to 4.4. The suite introduces those links without replacing the later analysis.
    The exact words in criterion 4.2

    Four Policy Pillars

    01Identify

    The policy explains what must enter the incident system.

    • Injury and occupational ill health
    • Near miss and dangerous occurrence
    • Property, environmental and operational loss
    • Delayed symptoms and diagnoses
    • High-potential and repeated control failures
    • Signals from alarms, inspections, monitoring and complaints

    02Investigate

    The policy defines how learning will be obtained fairly and proportionately.

    • Investigation levels and escalation criteria
    • Competent, authorised and sufficiently objective team
    • Worker and representative involvement
    • Terms of reference, evidence and time expectations
    • Immediate, underlying and root/systemic causes
    • Cause-linked action and effectiveness verification

    03Report

    The policy defines information routes and time expectations.

    • Who reports, what, to whom and how
    • Immediate emergency escalation
    • Alternative route if normal management is involved
    • Good-faith reporting without retaliation
    • Authorised external notification
    • Feedback to reporters and affected workers

    04Record

    The policy defines trustworthy, retrievable and protected evidence.

    • Unique event ID and contemporaneous facts
    • Actual and potential consequences
    • Notifications, evidence and investigation links
    • Decisions, actions, owners and deadlines
    • Effectiveness and closure evidence
    • Access, retention, version control and secure disposal
    ILO distinction: reporting normally means a worker or other person informs the organisation; recording means the employer creates and retains the organisational record; notification means an authorised duty holder sends required information to an external authority.
    Policy must allocate authority—not merely mention departments

    Who Is Responsible for What?

    RolePrincipal responsibilityEvidence or decision
    Worker or witnessProtect immediate safety and report promptly through an accessible route.Initial factual signal.
    SupervisorActivate response, make the area safe, receive the report and escalate.Initial event record and controls.
    Incident controllerCoordinate priorities, resources, communications and external emergency interface.Incident log and controlled response.
    Competent H&S personTriage actual and potential severity, advise on investigation level and screen legal duties.Classification and escalation decision.
    Responsible statutory reporterSubmit any required external report on behalf of the duty holder.Notification reference and retained copy.
    Investigation lead/teamPreserve and test evidence, analyse causes and propose controls.Findings and recommendations.
    Worker representativeContribute workforce knowledge and support fair learning.Consultation and challenge.
    Occupational health / HRSupport people and control sensitive health and welfare information.Restricted health/welfare record.
    Action ownerImplement assigned action and provide completion evidence.Implementation record.
    Authorised senior managerResource serious cases, decide restart/closure and review organisational implications.Risk acceptance, restart and closure approval.
    4.1 sequence + 4.2 accountability

    Who Owns the Next Decision? Responsibility and Handover Board

    Incident management can fail between stages even when each person performs one task well. Select both the accountable role and the evidence or decision that must be handed forward.

    Critical handoverAccountable decision ownerRequired output or evidenceStatus
    Alarm and initial escalationNot checked
    Command and stabilisationNot checked
    Investigation mobilisationNot checked
    Required external notificationNot checked
    Safe restart authorisationNot checked
    Effectiveness verification and closureNot checked
    1. Alarm
    2. Control
    3. Investigation
    4. Notification
    5. Restart
    6. Closure
    Choose an owner and a traceable output for every handover. The tool will identify exactly where organisational control breaks.
    Interactive policy assurance

    Policy X-Ray, Reporting Route and Record-Quality Tools

    Test whether the policy is complete, usable, accountable and capable of producing trustworthy evidence.

    Policy X-Ray

    Select only the provisions actually present in the organisation's policy.

    0 / 14 policy provisions confirmed.

    Internal Report vs External Notification

    Every event enters the internal system; only defined events enter a particular statutory route.
    Great Britain example: current RIDDOR duties apply to defined events and are submitted by the responsible person. They do not replace internal reporting, emergency calls, insurer notification, environmental reporting or local requirements. International learners must check their jurisdiction.

    Record-Quality Coach

    Good records separate facts, potential, analysis and sensitive information.

    Umbrella Policy Builder

    Complete the six policy elements to create an educational policy summary.
    Operational chain within the complete 3.4 → 4.4 route

    From Finding to Prevention: Learning-Loop Mapper

    A 3.4 finding becomes useful only when 4.1 converts it into the right organisational action and 4.2 governs that action through policy, responsibility and a required record. Section 4.3 then maintains that record as defensible evidence, while 4.4 evaluates whether the complete chain delivered effective prevention.

    3.4 · FindingChoose an investigation finding
    4.1 · ActionChoose the organisational stage
    4.2 · GovernanceChoose the policy and record
    Build the complete chain. A correct action today is not yet a reliable organisational system.
    Trust, privacy and governance

    What Makes an Incident Policy Work in Practice?

    Fair reporting culture

    Protect good-faith reporting, avoid premature blame, explain how information will be used and provide feedback. Accountability decisions may still occur through a separate fair process.

    Worker participation

    Workers and representatives provide work-as-done knowledge, identify practical barriers and help test whether corrective actions are usable.

    Privacy by design

    Collect only necessary personal information, separate detailed health records, apply role-based access, use anonymised learning where identity is unnecessary and dispose securely under the retention schedule.

    Management assurance

    Monitor delayed actions, repeat events, reporting routes, investigation quality, effectiveness checks, competence and policy-review triggers—not only injury totals.

    Policy review triggers: a serious or repeated incident; evidence that reporting is being suppressed; legal or organisational change; new technology or work pattern; audit findings; contractor-interface failure; emergency exercise learning; or the scheduled review date.
    The 4.2 → 4.3 handover

    Policy Sets the Rule—Maintained Records Prove What Happened

    4.2 · GovernThe policy defines what must be identified, investigated, reported and recorded; who is accountable; and which controls apply.
    4.3 · Maintain proofThe organisation creates, verifies, protects, updates, retrieves, retains and lawfully disposes of the resulting records.
    4.4 · EvaluateThose records become evidence for judging whether the process is compliant, consistently followed and effective.
    The relationship in one sentence: 4.2 tells people what the system requires; 4.3 preserves trustworthy evidence of decisions and actions; 4.4 tests that evidence to decide whether the system actually works.
    Today’s masterclass · 4.3 → 4.4
    FIIRSM · Fellow Member of IIRSM CertIOSH

    From Incident Records to Organisational Assurance

    Debjyoti Biswas, FIIRSM, CertIOSH

    Fellow Member of IIRSM · CEO & Founder, DB HSE INTERNATIONAL

    “An incident record is not simply paperwork. It is legal evidence, organisational memory and the foundation for proving whether lessons have genuinely prevented recurrence. Today, we move from maintaining defensible records to evaluating whether the complete incident-management process actually works.”
    Learning outcome 4.3
    Central question: How does an organisation keep incident records lawful, trustworthy, secure and retrievable throughout their lifecycle?

    Explain How to Maintain Records of Incidents to Meet Regulatory and Statutory Requirements

    At Level 6, do more than name a form or retention period. Explain the applicable requirement, how the organisation maintains the record, why that method satisfies the duty, and what legal, operational or learning consequence follows if the record cannot be trusted.

    4.3
    Begin with the meaning of “maintain”

    Incident Records Are Evidence, Memory and Accountability

    Maintaining incident records means creating accurate records promptly; linking all related evidence and decisions; protecting integrity, confidentiality and availability; updating them through a controlled audit trail; retrieving them when lawfully needed; retaining them for the correct period; applying any legal hold; and disposing of them securely when authorised.

    A record may support immediate control, statutory reporting, worker welfare, investigation, enforcement, claims, trend analysis, corrective-action tracking and organisational learning. One purpose must not destroy another—for example, a learning summary can be anonymised while the controlled legal record remains complete.

    Important distinction: statutory duties arise from legislation. Regulatory requirements may include enforceable rules, licence or sector conditions and regulator expectations. Guidance helps organisations comply but is not automatically legislation. Internal policy may lawfully require more than the statutory minimum.
    Safety professional linking an incident form, digital evidence, sealed evidence item and secure records storage in a warehouse office
    Visualise the complete record. The form, original evidence, access controls, correction history, decisions and closure proof remain connected through one traceable incident identity.
    One event · a controlled family of records

    What Must Stay Connected in the Incident File?

    Initial event recordTime, place, people, equipment, observable facts, actual and potential consequences, immediate controls and reporter details.
    Incident register or accident bookControlled index entry showing the event identity, classification and route into the organisation’s system.
    Notification evidenceDecision, authorised submitter, regulator or external recipient, date, method, acknowledgement and any reference number.
    Evidence registerPhotographs, CCTV, equipment, samples, documents, witness accounts, source, collector, custody and access history.
    Investigation fileScope, competence, chronology, analysis, findings, uncertainties, approvals and links from evidence to conclusions.
    Corrective-action recordCause-linked actions, interim controls, owner, resources, priority, due date, completion evidence and escalation.
    Recovery and restart recordRisk reassessment, inspection, worker briefing, conditions, authorisation, restrictions and monitoring plan.
    Health and welfare recordNecessary operational facts linked to separately protected occupational-health or HR information where detailed health data are required.
    Effectiveness and closure recordFollow-up measures, observations, workforce feedback, recurrence review, wider learning, residual risk and accountable closure decision.
    Control principle: use one master incident reference and a clear index. Do not copy sensitive data into every document. Link controlled records so authorised users can reconstruct who knew what, what they decided, why they decided it and whether the action worked.
    Record maintenance is a lifecycle—not filing at the end

    Create → Verify → Protect → Use → Retain → Dispose

    Create promptlyCapture contemporaneous facts and source details before memory, system data or evidence changes.
    Classify and indexAssign a unique ID, event type, actual/potential severity, jurisdiction and responsible owner.
    Verify qualityCheck completeness, factual language, dates, identity, source, duplication, uncertainty and required statutory fields.
    Protect integrityPreserve originals, use controlled copies, permissions, backups and evidence custody appropriate to the risk.
    Update transparentlyNever silently overwrite. Record the original, correction, reason, author, approval and date/time.
    Use and disclose lawfullyShare the minimum necessary information with authorised internal and external recipients.
    Retrieve reliablyKeep the file indexed, readable and available for investigation, regulator requests, audit, claims and learning.
    Retain or holdApply the longest applicable legal, sector, insurance, contractual or litigation-hold requirement.
    Dispose securelyAuthorise, document and securely destroy records only when every retention purpose and hold has ended.
    Translate the law into a controlled record rule

    Build a Legal Recordkeeping Profile for Every Jurisdiction

    A global policy should not hard-code one country’s rules for every site. Each organisation needs a current legal register or jurisdiction profile that converts applicable requirements into an operational record rule.

    Profile fieldQuestion the organisation must answerControl evidence
    AuthorityWhich legislation, regulation, licence, regulator, court or sector rule applies?Current legal register, version/date and competent review.
    Duty holderWho must create, notify, sign, keep or produce the record?Named accountable role and authorised deputy.
    Trigger and timeWhich event, diagnosis, incapacity or consequence activates the duty—and by when?Classification decision, deadline control and escalation log.
    Required contentWhich facts, identities, injury/diagnosis details, circumstances or submission references are mandatory?Controlled form fields and completeness check.
    Place, format and accessWhere must the record be kept, in what readable format, and who may inspect it?Repository, role permissions, retrieval test and disclosure log.
    Retention and disposalWhat is the minimum period, when does it start, what extends it, and how is disposal authorised?Retention schedule, legal-hold process and destruction certificate.

    Great Britain worked example—use current rules, not memory

    Reportable under RIDDOR

    The responsible person screens the event against current RIDDOR categories and timing. The external report and acknowledgement are linked to the incident file; completing the investigation is not a reason to miss the notification deadline.

    Record required under RIDDOR 2013 Regulation 12

    Records cover reportable deaths, injuries, dangerous occurrences and occupational diseases, plus over-three-consecutive-day worker incapacity even though that latter category alone is not externally reportable. Required particulars must be kept for at least three years from the date the record was made.

    Recordable is wider than reportable

    The accident book, internal policy, insurer, client, environmental or sector rules may require records when RIDDOR notification does not. “Not RIDDOR-reportable” never means “delete the event.”

    Current-law control

    HSG245 remains useful investigation guidance, but it was published in 2004 and contains historic RIDDOR references. Use current HSE and legislation pages for legal decisions, and use the correct separate regime for Northern Ireland.

    Jurisdiction caution: the examples above explain Great Britain only. International learners must identify current national, sectoral, environmental, data-protection, social-security, insurance, client and contractual requirements for their organisation. A consultation or proposal does not change the law unless formal provisions come into force.
    Trustworthy, not merely complete

    Integrity, Privacy, Retention and Legal Holds

    Eight tests of a defensible record

    • Accurate: facts reflect the best available evidence.
    • Complete: mandatory fields, attachments and decisions are present.
    • Timely: created and updated without avoidable delay.
    • Objective: observation is separated from inference and blame.
    • Traceable: source, author, date/time and event ID are clear.
    • Controlled: access, versions, copies and custody are governed.
    • Retrievable: readable records can be found throughout retention.
    • Auditable: changes, disclosures, approvals and disposal leave evidence.

    Correction without destroying history

    Preserve the original entry; add the corrected information; state why it changed; identify the person making and approving the change; time-stamp it; and notify affected decision makers when the correction changes classification, notification, welfare or action.

    Never: erase the original, backdate a record, fabricate certainty, copy an altered file over the only original, or change evidence after a regulator or legal hold has attached.

    Privacy and worker health data

    In UK data-protection terms, injury and health information is special-category data. Identify a lawful basis and special-category condition; collect only what is necessary; separate detailed medical information; restrict access; encrypt or lock it; and use anonymised learning where identity is unnecessary.

    Retention is a reasoned decision

    Do not invent one universal period such as “keep everything for six years.” Apply the specific legal minimum, then test longer sector, exposure, insurance, contractual, claim-limitation, safeguarding and litigation-hold needs. Retain no longer than justified once every duty and hold has ended.

    Interactive 4.3 record laboratory

    Make the Recordkeeping Decisions

    1Record, Report or Notify?

    Separate the internal record, internal report and any authorised external notification.

    2Legal Recordkeeping Profile Builder

    Complete the eight fields to translate a legal duty into a usable record rule.

    3Incident Record Integrity Test

    The record has a clear start. Test whether it remains trustworthy throughout its lifecycle.

    4Retention and Access Decision

    Apply purpose, minimum period, legal hold, necessity and authorised access together.

    5Audit-Trail Correction Challenge

    A defensible record can change when evidence changes—but its history must remain visible.
    The 4.3 → 4.4 handover

    Records Are the Evidence—Evaluation Is the Judgement

    Without reliable 4.3 records, 4.4 cannot distinguish a genuinely effective process from a well-written policy or a reassuring story. Evaluation now compares expected practice with what the records, people, observations and outcomes actually show.
    Learning outcome 4.4
    Central question: Is the incident-management process legally compliant, consistently delivered and effective—and what must improve first?

    Evaluate an Organisational Process for Managing Health and Safety Incidents

    Evaluation is not a description of the procedure and not an opinion about whether it “looks good.” It requires a benchmark, relevant evidence, balanced strengths and weaknesses, an explicit judgement, consequences and prioritised recommendations whose effectiveness can later be verified.

    4.4
    Move from activity to assurance

    Does the Process Work in Design, Delivery and Results?

    Evaluation is the systematic comparison of the organisation’s incident-management process and evidence against defined legal, policy, guidance and risk-based criteria in order to judge adequacy, implementation and effectiveness, then determine proportionate improvement.
    MonitoringTracks routine indicators and exceptions continuously—for example reporting delay or overdue actions.
    AuditChecks conformance against defined criteria through planned sampling, interviews, records and observation.
    EvaluationCombines compliance, implementation, quality and outcome evidence to reach a balanced judgement about value and risk.
    BenchmarkEvidenceFindingJudgementImprovement + verification
    Multidisciplinary workplace team evaluating incident evidence, process stages and corrective-action performance in an industrial meeting room
    Evaluation is multidisciplinary. Records show formal evidence; workers and contractors explain work as done; observation tests control performance; and trend data show whether improvement lasts.
    Use HSG245 as an assurance benchmark—not as current law

    Evaluate Every Link from Scene Preservation to Shared Learning

    HSG245-linked stageEvaluation questionExamples of evidencePossible process failure
    Preserve the sceneDid urgent safety action occur without avoidable destruction or loss of evidence?Scene log, photographs, isolation record, CCTV copy, disturbance log.Evidence overwritten, moved without record or inaccessible.
    Note people and equipmentCan the organisation identify exposure, witnesses, assets and relevant operating conditions?People/equipment register, shift list, contractor record, asset history.Contractors omitted; equipment state or health exposure disconnected.
    Report the eventWas the event visible quickly to everyone who had to protect, decide or notify?Time-stamped report, acknowledgement, escalation and notification logs.Delay, suppression, wrong recipient or missed external duty.
    Decide investigation needDid actual and credible potential risk determine proportionate depth?Triage score, classification evidence, scope and team appointment.Minor actual injury hides high-potential systemic failure.
    Gather informationWas evidence sufficient, reliable, diverse and traceable?5 Ps evidence map, interviews, documents, digital data, health evidence.Single narrative, missing source, bias or no work-as-done evidence.
    Analyse informationDid analysis test barriers and immediate, underlying and root/systemic causes?Timeline, barrier analysis, change analysis, cause logic and uncertainties.“Human error” becomes the stopping point; unsupported root cause.
    Identify controlsDo proposed controls address supported causes at an appropriate hierarchy level?Option appraisal, risk assessment, hierarchy and human-factors review.Training and reminders substitute for engineering or system change.
    Implement action planAre actions resourced, owned, prioritised, tracked and verified?Owners, due dates, interim controls, completion and effectiveness evidence.Action closed on invoice, attendance sheet or assertion alone.
    Share learningDid relevant sites, assets, contractors and workers receive usable learning and change?Targeted communication, transfer review, updated systems and feedback.A bulletin is sent but equivalent risks remain unchanged elsewhere.
    Current-law warning: HSG245 provides a strong prevention methodology, but it was published in 2004. Evaluate legal compliance against current applicable legislation and regulator guidance; do not treat historic reporting references as current statutory rules.
    Five tests drawn from the indicative-content purpose

    What Value Should Investigation and Incident Management Create?

    Legal complianceCorrect duty holder, route, timing, record and evidence can be demonstrated.
    Reliable dataInformation is complete enough for learning, trend analysis and defensible decisions.
    Causal depthAnalysis goes beyond the immediate act to barriers, conditions and organisational factors.
    PreventionControls address causes, transfer across similar risks and show sustained effectiveness.
    Morale and trustPeople report, participate and receive feedback without premature blame or unfair treatment.
    Evaluation principle: a process can meet a deadline yet fail to learn; produce a long report yet lack causal depth; close every action yet fail to prevent recurrence; or show low incident numbers because people no longer trust the reporting system.
    Triangulate before judging

    Three Evidence Layers and Four Defensible Ratings

    Design · work as imaginedPolicy, procedure, legal register, forms, roles, competence standards, escalation matrices and planned controls.
    Delivery · work as doneIncident files, interviews, observations, contractor experience, system logs, samples and actual handovers.
    Results · difference achievedTimeliness, record quality, recurrence, control performance, overdue actions, worker confidence and wider learning.
    EffectiveCriteria are met consistently and outcome evidence shows risk reduction with no material critical gap.
    Partially effectiveImportant strengths exist, but inconsistent delivery or incomplete evidence weakens assurance.
    IneffectiveMaterial failures create legal, safety, learning or trust risk despite some activity.
    Insufficient evidenceA conclusion would be speculation; further records, interviews, observation or follow-up data are required.
    Critical-control rule: do not hide a missed statutory notification, unsafe restart, destroyed evidence or repeated high-potential event inside an average score. Critical failures require explicit judgement and urgent action.
    Mixed evidence · balanced evaluation

    Evaluate the Forklift and Chemical Incident Process

    The organisation’s procedure looks complete. The incident file, interviews and follow-up data reveal a more complicated picture.

    Strength · immediate controlThe alarm was raised, first aid was provided, the area was isolated and the drain was protected.
    Strength · initial documentationAn incident ID, initial report and updated risk assessment exist, with a restart decision recorded.
    Gap · visibility and triageThe report was entered eight hours later and classified mainly by minor actual harm rather than serious credible potential.
    Gap · evidence integrityRelevant CCTV was overwritten and the contractor was not included in the first evidence plan.
    Gap · independence and causal depthThe operational manager under production pressure led the review; “driver inattention” became the main conclusion.
    Gap · control qualityTraining and signs were completed, but vehicle–pedestrian separation, racking protection and earlier warnings were not fully addressed.
    Gap · recovery and health linkageWork restarted before varied-condition verification; delayed skin symptoms were stored separately and did not trigger reclassification.
    Gap · resultsA similar route conflict occurred two months later, and contractor survey feedback shows low confidence in reporting follow-up.
    Balanced judgement: the process is partially effective but materially weak. Immediate response and basic documentation are credible strengths, but evidence loss, weak potential-risk triage, limited causal depth, lower-order controls, premature restart and recurrence prevent a conclusion of effectiveness. Priorities are time-sensitive evidence preservation, risk-based triage, independent-enough investigation, cause-linked higher-order controls, integrated health updates and defined effectiveness verification.
    Measure the health of the system—not only injury totals

    A Balanced Incident-Management Dashboard

    Reporting latencyEvent occurrence or discovery → useful internal visibility.
    High-potential escalationCorrectly recognised and escalated high-potential events.
    Evidence preservationTime-sensitive sources protected before loss or overwrite.
    Investigation qualityEvidence-to-finding logic, causal depth and worker involvement.
    Action traceabilityFindings linked to actions, owners, deadlines and controls.
    Overdue critical actionsAge, risk exposure and escalation—not just total overdue count.
    Effectiveness verifiedActions tested in real work after implementation.
    Recurrence / precursor trendSimilar event, near-miss and control-failure patterns.
    Statutory accuracyCorrect decisions, fields, timing, record and audit trail.
    Worker confidenceCan people report, participate and expect feedback without unfair blame?
    Learning reachEquivalent sites, tasks, assets and contractors assessed.
    Record retrievabilityComplete file available, readable and access-controlled when tested.
    Metric caution: zero recorded incidents may mean excellent control—or hidden underreporting. Faster closure may mean better governance—or superficial investigation. Interpret every metric with exposure, severity, potential, quality sampling and workforce evidence.
    Interactive 4.4 assurance laboratory

    Test the Evidence, Diagnose the Metric and Prioritise Improvement

    1HSG245 Assurance Review

    Select the stage, evidence strength and consequence; then make an explicit rating.

    2Evidence Triangulation Lab

    Claim: “Corrective actions are effective because every action is marked complete.” Which independent evidence would you test?

    Strong evaluation triangulates records, observation, people and outcomes.

    3Metric Detective

    Ask what the measure includes, what changed, what is missing and what other evidence would confirm the meaning.

    4Improvement Prioritiser

    Priority should reflect legal exposure, credible harm, recurrence, evidence loss, trust and the strength of interim control.
    4.4 evaluation tool · implemented is not the same as effective

    Completed or Effective? Closure-Evidence Dashboard

    Use comparable performance evidence, real-work observation, workforce input and independent-enough verification to decide whether an action is merely installed or genuinely ready for closure.

    Evidence picture

    Baseline
    6.0 / 1,000
    Follow-up
    1.0 / 1,000
    Assess whether the apparent improvement is valid, sufficiently tested and free from harmful side effects.
    Learning boundary: six observed conditions is used only as an educational completeness prompt in this dashboard. It is not a universal legal or statistical threshold. Real verification criteria must be risk-based, defined in advance and appropriate to the organisation.
    Assessment-ready learning support

    How to Answer 4.1–4.4 at Level 6

    4.1 · Outline the stages

    Present the incident lifecycle in a logical sequence. For each major stage, state its purpose, principal action, responsibility and output or handover.

    Structure: Stage → what happens → why it matters → owner → output → brief case application.

    4.2 · Outline the policies

    Describe the umbrella policy and its four pillars. Include mandatory rules, responsible roles, supporting procedures/records and how the policy is assured.

    Structure: Policy requirement → main rules → accountability → controlled document/record → brief application.

    4.3 · Explain record maintenance

    Identify the statutory or regulatory requirement, then explain how the record is created, verified, protected, corrected, retrieved, retained and disposed of so the duty is met.

    Structure: Requirement → maintenance method → why it satisfies the duty → incident-file application → consequence of failure.

    4.4 · Evaluate the process

    Use a clear benchmark and triangulated evidence to judge design, delivery and results. Balance strengths and weaknesses, state the consequence and prioritise verifiable improvement.

    Structure: Benchmark → evidence → finding → judgement → consequence → recommendation → verification.

    Interactive Level 6 Command-Word Coach

    Complete the framework; the coach will adapt the paragraph to Outline, Explain or Evaluate.

    Common Weak Answers

    Only a list“Report, investigate, close” does not outline what the stages involve.
    Reporting equals RIDDORInternal reporting is broader than one statutory notification regime.
    Evidence before lifeRescue, first aid and immediate hazard control take priority.
    Actual harm onlyA fortunate outcome can conceal a high-potential control failure.
    Policy equals formA template cannot replace approved rules, authority and responsibility.
    Close on completionAction completion is different from verified effectiveness.
    One retention period for everythingApply the specific legal minimum, other justified purposes and any legal hold.
    Description called evaluation4.4 needs a benchmark, evidence, explicit judgement, consequence and prioritised improvement.
    Section 04 knowledge assessment

    Check the Complete Incident-Management System

    1. What has priority when evidence and rescue conflict?
    2. What should determine investigation level?
    3. Must external notification wait for the full investigation?
    4. Which document defines mandatory organisational rules?
    5. A serious near miss is not RIDDOR-reportable. What should happen?
    6. Who should submit a required RIDDOR report?
    7. How should sensitive health information be handled?
    8. When may normal work restart?
    9. What does a completed incident form become?
    10. Why involve workers in investigation and learning?
    11. When is corrective action truly closed?
    12. Which is the strongest use of “outline”?
    13. What does “maintain incident records” mean?
    14. What is the minimum RIDDOR Regulation 12 record period in Great Britain?
    15. How should a material record error be corrected?
    16. What is the strongest approach to detailed worker-health data?
    17. Which response genuinely evaluates a process?
    18. What does “zero incidents” prove by itself?
    19. When is an incident-management process best rated effective?
    20. How should a missed statutory notification be treated in evaluation?
    Answer all twenty questions, then check your score and explanations.
    Authority and jurisdiction boundary

    What This Section Is Built On

    Official OTHM specificationSets criteria 4.1–4.4 and identifies reporting/recording practice, the importance of investigation and the HSG245 sequence in its indicative content.
    HSE HSG245Supports gathering, analysing, identifying controls and implementing an action plan. Its investigation method remains useful, but its historic RIDDOR references are superseded.
    HSE emergency proceduresSupports planned roles, alarms, first aid, evacuation, isolation, training and the rule not to resume while serious danger remains.
    HSE policy structureUses statement of intent, responsibilities and arrangements as a practical policy architecture.
    RIDDOR 2013 Regulation 12Primary Great Britain legislation for records of defined reportable events, diagnoses and over-three-day incapacity, including the minimum three-year record period.
    HSE RIDDOR record guidanceExplains the information that must be retained, where records are kept and how recordkeeping differs from external reporting.
    ICO worker injury recordsSupports necessary, proportionate and access-controlled handling of worker health and injury information. The ICO page currently states that this guidance is under review following the Data (Use and Access) Act, so organisations must check the current law and guidance.
    ISO 45001 official FAQIllustrates response, correction, cause determination, corrective action and decisions about safe return to use. ISO 45001 is voluntary and not legislation.
    Legal caution: RIDDOR is a Great Britain example. International learners must identify applicable national, sectoral, environmental, client and contractual duties. Internal reporting should remain broader than the minimum statutory-notification threshold.
    Final Unit 3 stage · Assignment preparation masterclass

    Turn Unit 3 Learning into 17 Clear Pieces of Assessment Evidence

    This is the bridge between knowing Risk and Incident Management and demonstrating it at Level 6. Follow the exact task, cover every assessment criterion, apply it to one clearly identified organisation and support each judgement with reliable evidence.

    4 tasksTwo essays · Two reports
    17 ACsEvery criterion must be met
    1,000 eachEach task has its own 900–1,100 word range
    Pass / FailNo compensation for a missed AC
    Y/617/7540Official Unit 3 code
    Level 6Risk and Incident Management
    8 credits80 TQT · 30 GLH
    Word document.doc or .docx · Harvard references
    The complete route

    Do the Exact Work the Brief Requires

    The assignment is not one general discussion about safety. It is a controlled evidence journey. Each task has its own genre, word limit and assessment criteria.

    1 · DecodeRead the whole task, then let each AC command verb set the depth.
    2 · ContextualiseName one organisation, site, industry and legal jurisdiction.
    3 · MapCreate one visible heading or clearly traceable section for every AC.
    4 · ResearchUse internal evidence plus academic, regulatory and professional sources.
    5 · DraftWrite in essay or report form exactly as directed—within the task limit.
    6 · AnalyseConnect theory, organisational practice, evidence, limitations and judgement.
    7 · VerifyCheck all 17 ACs, all six named strategies and all four Task 4 criteria.
    8 · AuthenticateCite accurately, confirm your own work and sign the authenticity statement.
    Level 6 standard: Build clear links between theory and practice, use independent learning and a range of contextually relevant sources, engage critically and analytically, reach evidence-based conclusions and present a well-structured argument.
    Word-count control
    Each task states 1,000 words and allows 10% either side: 900–1,100 words per task. Do not rely on a combined total if one task is outside its own range.
    Mandatory safeguard — do not miss AC 4.4. The Task 4 heading maps 4.1–4.4, but the printed bullet list contains only 4.1–4.3. A separate evaluation of the organisation’s incident-management process is still compulsory.
    Exact task-by-task navigator

    Plan Every Assessment Criterion Before Writing

    Select a task. The word figures below are a recommended planning budget—not an OTHM-mandated distribution. Adjust them while keeping the task between 900 and 1,100 words.

    Task 1 of 4

    Risk Identification, Assessment and Reliability

    Essay1,000 wordsAC 1.1–1.6

    Task instruction: Detail the processes and strategies used to identify and analyse risk and how this influences risk management within the organisation.

    ACWhat the learner must demonstrateUseful organisational evidenceSuggested words
    1.1 OutlineEssential internal and external information sources used to identify hazards and assess risks, plus their organisational relevance.Incident/near-miss data, sickness and damage trends, maintenance records, inspections, HSE/OSHA, ILO, WHO, professional and trade guidance.
    1.2 ExplainHow hazard-identification techniques work, how they are used and why they suit the organisation.Observation, consultation, task analysis, inspections, audits, HAZOP, failure tracing, FTA or other proportionate techniques.
    1.3 ExplainThe end-to-end risk-assessment system and how risk is evaluated—not only a five-step list.Scope, competence, people at risk, method, likelihood/severity, control standards, prioritised action, approval, review and limitations.
    1.4 ExplainHow the organisation measures, monitors and reports hazards. Address all three actions.Exposure measurements, inspections, leading/reactive indicators, reporting channels, escalation, frequency, responsibility and dashboards.
    1.5 ExplainHow risk-assessment records are controlled to meet applicable regulatory and statutory requirements.Jurisdiction, required content, owner, approval, revision, accessibility, retention, privacy, review triggers and audit trail.
    1.6 ExplainActual calculations used to analyse and improve system reliability and trace failure, followed by interpretation and action.Relevant worked examples such as MTBF, MTTR, failure rate, availability or R(t); data source, result, meaning, limitation, improvement and verification.
    StructureIntroduction, synthesis showing influence on risk-management decisions, and a concise conclusion.Organisation/site scope and a clear line of argument across 1.1–1.6.
    Planned Task 1 total1,000 words · within range
    Command-word discipline

    The Verb Controls the Depth of the Answer

    Read the full question, then use the assessment criterion’s verb. “Detail,” “carry out” and other wording in the task introduction do not reduce the demand of the AC.

    OutlineGive a concise but complete summary of the essential features. Not a bare list.
    ExplainShow how and why, make connections and apply knowledge coherently to the organisation.
    EvaluateUse criteria and evidence, analyse strengths and limitations, then reach a justified verdict.
    JustifyState the choice, support it with evidence and show why it is appropriate in context.
    AssessWeigh evidence, options and consequences, then conclude on effectiveness or validity.
    A strong Level 6 paragraph: Point → organisational context → reliable evidence/source → analysis of meaning → limitation or alternative → reasoned judgement/action. Not every paragraph needs the same formula, but unsupported description is rarely enough.
    Recommended structure—not an OTHM-mandated template

    Essay and Report Are Not the Same

    Tasks 1 and 3 · Essay

    1. Specific title
    2. Introduction: organisation, scope and argument
    3. Connected thematic sections mapped to the ACs
    4. Critical discussion linking theory, evidence and practice
    5. Conclusion answering the task—normally no new evidence
    6. Harvard reference list

    Tasks 2 and 4 · Report

    1. Title and clear organisational scope
    2. Contents page if the centre template expects one
    3. Introduction / purpose / method
    4. Numbered findings mapped to the ACs
    5. Evidence-based conclusion
    6. Prioritised recommendations where justified
    7. Harvard reference list
    Clarify locally: The uploaded brief does not state whether references, tables, appendices or contents count toward the word limit, or whether the four tasks are submitted separately or in one Word file. Follow the centre’s issued submission template and instructions.
    Interactive class planning tool

    Define the Organisation Before Researching

    A learner may use their own workplace or another organisation they can evidence. Use one consistent context and distinguish real evidence from reasonable recommendations.

    Your assignment context map

    Complete the fields to create a planning map. The tool will not write assignment prose.

    Evidence and academic integrity

    Research Widely—Write Independently

    Build a balanced evidence set

    • Current applicable legislation and regulator guidance
    • Academic books and peer-reviewed research
    • Recognised standards and professional-body guidance
    • Controlled organisational policies, records and data
    • Credible workplace observations or interviews, where permitted

    Protect authenticity

    • Write the work yourself and keep a traceable research trail
    • Cite every borrowed idea, model, table, figure or quotation
    • Paraphrase genuinely—not by changing a few words
    • Never invent organisational data, legislation or references
    • Follow the centre’s current rules on permitted AI assistance and disclosure
    • Sign the required statement of authenticity
    This masterclass is a preparation guide, not a model answer. Learners must select their own evidence, make their own judgements and produce their own work. Copying portal wording as an assignment response may create authenticity and plagiarism concerns.
    Referral-prevention clinic

    Nine Common Ways a Learner Can Miss the Standard

    Missing an ACGood work elsewhere does not compensate for an unaddressed criterion.
    Describing instead of evaluating2.1 and 4.4 need evidence, strengths, limitations and a verdict.
    No “when” in 2.2All six strategies need actual use and circumstances that justify use.
    No calculation in 1.6Narrative alone may not demonstrate calculations, reliability and failure tracing.
    Task 3 narrowed to investigationLoss causation, quantitative data and reporting remain separately assessed.
    AC 4.4 omittedThe brief’s missing bullet does not remove the criterion.
    Generic organisationName the site/context and apply claims to credible policies, evidence and constraints.
    Wrong jurisdictionDo not present Great Britain law as universally applicable to another country.
    Weak sourcing or authenticityUnverifiable references, copied wording or invented data undermine the evidence.
    Final 17-criterion self-audit

    Nothing Is Complete Until Every Criterion Is Visible

    Tick only after the draft contains traceable evidence at the required command-word depth. The checklist measures coverage, not assessor approval.

    0 of 17 criteria evidenced. Begin with the task plan—do not tick from memory.

    Submission-control checklist

    0 of 12 submission controls complete.

    Authority and version control

    Use the Brief and the Current Specification Together

    The uploaded assignment brief is dated January 2021 and sets the four task genres, assessment-criterion mapping and 1,000-word limit for each task. The December 2023 OTHM qualification specification confirms the same 17 Unit 3 assessment criteria, but it is not the source of the task genres or word limits shown here. The centre must confirm that the issued assignment brief is the authorised version for the learner’s cohort.

    Independent resource status: This assignment-preparation masterclass was prepared by Debjyoti Biswas, FIIRSM, CertIOSH, for DB HSE International, OTHM Approved Training Centre DC2201634. It supports teaching and does not replace the centre-issued assignment brief, assessor judgement or OTHM requirements.

    Use Authoritative Information

    This independent DB HSE explanation supports learning. Workplace decisions must use current applicable legislation, exposure limits, approved organisational criteria and competent specialist advice where required.

    Annotated Source Guide for Sections 3.3–4.4

    OTHM Level 6 official specificationSets assessment criteria 3.3–4.4. It is the qualification source; the expanded topics, professional-practice framework and tools here are independent DB HSE teaching interpretation.
    HSE HSG245Official Great Britain guidance—not legislation—supporting investigation depth, evidence, cause analysis, risk controls and action planning. Published in 2004; use current RIDDOR pages for statutory reporting.
    Current HSE RIDDOR guidanceOfficial Great Britain legal guidance for defined statutory notifications. It is narrower than internal reporting and does not automatically notify insurers, clients or environmental authorities.
    HSE 2026 RIDDOR consultationOfficial proposal record, closed 7 July 2026. Proposals are not current law unless and until formal changes come into force.
    ILO recording and notification codeInternational, non-binding 1996 code of practice supporting prevention-focused recording and notification. National law and local arrangements remain controlling.
    ISO/TC 283 ISO 45001 FAQOfficial voluntary-standard guidance supporting response, correction, cause determination, corrective action and effectiveness review. ISO 45001 is not legislation; the portal paraphrases rather than reproduces copyrighted standard text.
    HSE worker-representative guidanceGreat Britain guidance explaining prompt, proportionate involvement in investigations and the need to protect evidence and personal information.
    ICO worker injury-record guidanceOfficial UK data-protection guidance supporting necessary, proportionate and access-controlled handling of identifiable health information. The page currently notes that it is under review following the Data (Use and Access) Act; it is guidance, not the legislation itself.
    Independent resource status: Sections 1.6.3–1.6.5, 2.1–2.3, 3.1–3.4 and 4.1–4.4 are prepared solely by Debjyoti Biswas for DB HSE International to support OTHM Level 6, Unit 3 teaching. They are not produced by OTHM or Ofqual and do not replace the official specification, assessment guidance or applicable workplace requirements.