Welcome to Unit 3 · Sections 01–04 + Assignment Brief
Welcome to Unit 3 of the OTHM Level 6 International Diploma in Occupational Health and Safety. This unit develops complete risk-and-incident management judgement—from understanding system reliability and selecting effective controls to learning from loss events, maintaining defensible records and evaluating whether the organisation truly improves. The final assignment stage then shows learners how to convert that knowledge into all 17 required pieces of assessment evidence.
See Risk Clearly · Control It Systematically · Learn from Every Event · Prove Improvement
The four sections form one connected professional journey. Learners move from understanding how systems fail, through choosing and designing controls, to investigating what happened and assuring that the complete incident-management process is lawful, reliable and effective.
Why it matters: Unit 3 connects prevention before an event with disciplined learning after it. A Level 6 practitioner must not only identify what went wrong, but decide what the organisation should change, demonstrate that the change was implemented and verify that it reduced risk.
Systems Reliability and Failure Tracing
- 1.6.3Common reliability calculations
- 1.6.4Analyse system performance
- 1.6.5Use evidence to trace failures
- ToolsSystem explainer, nine calculators, analysis and tracing activities
Strategies and Techniques of Risk Control
- 2.1Evaluate common risk-management strategies
- 2.2Justify when each strategy should be used
- 2.3Develop SSOWs and SOPs
- ToolsAssessment formats, risk matrix, strategy planner, SSOW and SOP builders
Loss Causation, Reporting and Investigation
- 3.1Completed foundation: loss-causation theories and techniques
- 3.2Completed foundation: quantitative loss-data analysis
- 3.3Who needs to know, when and why? Assess loss-event reporting
- 3.4How does evidence become prevention? Explain incident investigation
- ModelsBird, multi-causality, Swiss Cheese, FTA, ETA, Bowtie and behavioural RCA
- ToolsCause, data, reporting, evidence, investigation and action-quality laboratories
Managing Incidents Through Organisational Systems
- 4.1What must happen from the first alarm to verified closure?
- 4.2Which policies make identification, investigation, reporting and recording consistent?
- 4.3How are incident records kept lawful, trustworthy, secure and retrievable?
- 4.4How do we evaluate whether the complete incident-management process really works?
- ToolsResponse, policy, record-integrity, legal-profile, assurance, metrics and evaluation laboratories
Unit 3 Assignment Preparation Masterclass
- Task 1Essay · AC 1.1–1.6 · Risk identification, assessment and reliability
- Task 2Report · AC 2.1–2.3 · Risk control, SSOW and SOP
- Task 3Essay · AC 3.1–3.4 · Loss causation, data, reporting and investigation
- Task 4Report · AC 4.1–4.4 · Incident systems, records and evaluation
- ToolsExact AC map, live word budgets, organisation planner and command-verb clinic
- Assurance17-criterion coverage audit, academic-integrity rules and submission controls
Systems Reliability and Failure Tracing
Welcome to Section 01. Begin with system reliability, complete all nine calculations in 1.6.3, analyse what the results mean in 1.6.4, and use the evidence to trace failures in 1.6.5.
FIIRSM · CertIOSH · CEO & Founder, DB HSE International
Meet Your Tutor
Debjyoti Biswas, FIIRSM, CertIOSH is an Approved Tutor of the OTHM Level 6 International Diploma in Occupational Health and Safety Qualification, a Fellow Member of IIRSM and a CertIOSH Member. His teaching approach connects Level 6 knowledge with practical workplace application, evidence-based judgement, leadership and professional development.
What Is Reliability?
Let us begin with one easy question:
If you ask a machine, alarm, pump or complete safety system to do its job, how confident are you that it will perform correctly without failing?
That confidence—expressed as a probability—is reliability.
The exact job the item must perform. A fire pump must deliver water; a gas detector must detect gas.
The real environment and duty: temperature, pressure, dust, workload, operating mode and maintenance condition.
The period for which successful operation is required—for example, a 100-hour mission or an emergency demand.
The item completes the required function without stopping, leaking, giving a false signal or becoming unavailable.
The answer is normally written between 0 and 1, or as a percentage from 0% to 100%.
Now apply that idea to a complete system.
A system is a group of connected parts—equipment, power, controls, information, people and procedures—working together. System reliability asks whether the complete required function succeeds.
Explore One Safety Function—One Step at a Time
Imagine a gas-release protection system. Select each part to understand its job.
Component reliability
Asks whether one item—such as a sensor, bearing, seal or pump—will perform successfully.
System reliability
Asks whether the complete required function succeeds. Reliable individual parts do not automatically guarantee a reliable system because connections, power, controls, procedures and common causes also matter.
Not only quality
A well-made item may still be unreliable if it is used outside its design conditions or maintained poorly.
Not the same as availability
An item can fail frequently yet remain highly available when every repair is very fast.
Not proof of safety
Reliability supports risk decisions, but safety also depends on consequence, design, inspection, people, procedures and independent protection.
The DB HSE Emergency-Protection System
Yes—we are going to use one connected workplace scene to explain all nine calculation techniques.
Imagine that you and I have entered this chemical warehouse. The protection system contains gas detection, a control panel, an alarm, an isolation valve, replaceable sensor cartridges and two fire-water pumps. We will not keep changing the story. We will simply ask nine different reliability questions about the relevant parts of this one system.
| Part of our scene | Information collected | Why we need it |
|---|---|---|
| Repairable protection pump | 5,000 operating hours; 10 failures; 40 total repair hours | MTBF, MTTR, λ, availability, R(t) and F(t) |
| Required mission | The pump must operate for the next 100 hours | Reliability and failure probability over time |
| Five replaceable sensor cartridges | Lives: 1,800; 2,100; 1,900; 2,200; 2,000 hours | MTTF for non-repairable items |
| Series alarm chain | Detector 98%; control panel 97%; alarm 99% | Reliability when every required link must work |
| Two backup fire-water pumps | Each pump has 90% reliability for the stated demand | Parallel reliability when at least one independent pump must work |
Watch how the same scene gives us nine different questions
How long does the repairable pump operate, on average, between failures?
How long does one pump repair take, on average?
How frequently does the pump fail per operating hour?
What percentage of time is the pump ready for use?
What is the chance that it completes 100 hours without failure?
What is the chance that it fails during those 100 hours?
What is the average life of a replaceable sensor cartridge?
Will the detector, control panel and alarm all work together?
Will at least one of the two independent fire pumps work?
Why Do HSE Professionals Need These Calculations?
Statements such as “the machine fails frequently” or “the alarm is usually reliable” are opinions. Reliability calculations convert operating, failure and repair records into evidence that can be compared, investigated and improved.
Measure
Determine how often equipment fails, how long it runs and how long repairs take.
Trace
Use trends and component records to locate recurring failures and investigate their causes.
Improve
Verify whether maintenance, redesign, training, spare parts or redundancy improved performance.
Questions the calculations answer
- How frequently does the equipment fail?
- How long does it operate before failing?
- How long is normally needed to repair it?
- What proportion of time is it available?
- What is the chance of completing a task without failure?
- Which component repeats?
- Did maintenance improve performance?
Safety-critical examples
Fire pumps, emergency alarms, pressure-relief systems, local exhaust ventilation, gas detectors, lifting equipment, emergency generators, rescue equipment and emergency shutdown systems.
The required reliability should reflect the consequence of failure and the availability of independent protection.
Understand Every Symbol Before Calculating
Select a symbol to see what it means and where it is used.
Total operating time — T
T represents the total time for which equipment actually operated. It must use the same time unit throughout the calculation, such as hours, days or cycles.
Explore Each Reliability Calculation
Choose a calculation, change the figures and examine how the result and interpretation respond.
Mean Time Between Failures — MTBF
Mean means average. MTBF is the average operating time between one failure and the next failure of repairable equipment.
T = total operating time
N = number of failures
Why it is relevant
- Compares the reliability of machines or operating periods.
- Supports preventive-maintenance planning.
- Shows whether failures are becoming more frequent.
- Provides evidence for investigating deterioration or repeated component failure.
Important interpretation
A higher MTBF is generally better. A falling MTBF—such as 1,000 → 600 → 300 hours—means failures are occurring more frequently and should be investigated.
Live MTBF calculator
See the complete worked example
- Record total operating time: T = 5,000 hours.
- Record failures: N = 10.
- Divide 5,000 by 10.
- MTBF = 500 hours.
- Interpretation: the machine operated for an average of 500 hours between failures.
MTBF trend inspector
What can make MTBF fall?
Equipment deterioration, ageing parts, ineffective maintenance, overloading, poor lubrication, incorrect operation, harsh conditions or repeated failure of the same component. Use maintenance records, failure reports and operating evidence to test these possibilities.
Mean Time to Repair — MTTR
MTTR is the average time required to detect, diagnose, repair, test and safely return failed equipment to service.
Why it is relevant
- Evaluates maintenance response and recovery performance.
- Supports staffing, competence and spare-parts planning.
- Identifies access, diagnostic or permit-to-work delays.
- Helps reduce downtime without compromising safe isolation and testing.
Failure-tracing clue
If MTBF remains stable but MTTR increases, the machine is not necessarily failing more often; the organisation is taking longer to restore it. Investigate the repair process.
Live MTTR calculator
What can increase MTTR?
Slow diagnosis, unavailable spare parts, poor equipment access, insufficient competence, delayed authorisation, inadequate tools, weak procedures or complex testing and restart requirements.
Failure Rate — Lambda (λ)
Failure rate measures how frequently a system or component fails during its operating time. The Greek letter λ is pronounced “lambda”.
λ = failure rate
N = number of failures
T = total operating time
Relationship with MTBF
When a constant failure rate is assumed, higher MTBF normally means a lower failure rate.
Important limitation
The simple relationship λ = 1 ÷ MTBF assumes a reasonably constant failure rate. It may not represent early-life defects, rapid deterioration or age-related wear.
Live failure-rate calculator
Why express it per 1,000 hours?
A rate of 0.002 failures per hour may feel abstract. Multiplying it by 1,000 gives two failures per 1,000 operating hours, which is easier to communicate to managers and workers.
Where is failure rate used?
To compare equipment, monitor deterioration, establish inspection intervals, evaluate maintenance, estimate future failure probability and prioritise improvement. If a seal records 8 of 12 pump failures, trace seal selection, alignment, pressure, temperature, installation and maintenance.
Availability — A
Availability is the proportion of time that equipment is operational and ready to perform its required function.
A = availability
U = uptime
D = downtime
Why it is relevant
Availability is critical for fire pumps, emergency generators, gas detectors, alarms, ventilation and emergency communication systems that must be ready when required.
Availability is not reliability
A machine may fail regularly yet still show high availability if it is repaired quickly. Reliability asks whether it will operate without failure; availability asks whether it is ready for use.
Live availability calculator
How can availability be improved?
Increase MTBF; safely reduce MTTR; provide competent maintainers, fault detection and essential spares; improve preventive maintenance; remove repeated causes; and provide reliable, independently tested backup where justified.
Reliability Probability — R(t)
R(t) is the probability that a system performs its required function without failure for a specified time, t.
R(t) = reliability during time t
e = mathematical constant, approximately 2.718
λ = failure rate
t = required operating time
Why it is relevant
It estimates whether equipment can complete a continuous production period, lifting operation, emergency duty, shutdown task or inspection interval without failure.
Assumption
This exponential formula assumes a constant failure rate. It is an estimate based on recorded data, not a guarantee of future performance.
Live reliability calculator
Probability of Failure — F(t)
F(t), also called unreliability, is the probability that equipment will fail during a specified operating period.
F(t) = probability of failure
1 = total probability, equal to 100%
R(t) = probability of successful operation
Why it is relevant
Management can compare the probability of failure with risk tolerance and decide whether maintenance, inspection, standby equipment, replacement or additional controls are required.
Always check the total
Reliability and failure probability must add to 100%: R(t) + F(t) = 1.
Live failure-probability calculator
Connect this with the previous example
If R(100) = 81.87%, then F(100) = 100% − 81.87% = 18.13%. The machine therefore has an estimated 18.13% probability of failure during the 100-hour period.
What decisions can F(t) support?
Whether risk is acceptable; whether maintenance or inspection is needed before operation; whether independent standby equipment is required; whether replacement is justified; and whether emergency arrangements must be strengthened.
Mean Time to Failure — MTTF
MTTF is the average operating life of a non-repairable item—an item that is normally replaced after failure.
Why it is relevant
MTTF supports component selection, life-cycle planning, stock control, purchasing and the preventive replacement of fuses, disposable sensors, certain batteries and other non-repairable parts.
MTBF versus MTTF
MTBF is mainly for repairable equipment and measures time between failures. MTTF is mainly for non-repairable items and measures time until failure.
Live MTTF calculator
Examples of non-repairable items
Disposable sensors, fuses, certain electronic parts, light bulbs, single-use protective devices and some batteries. They are normally replaced, so MTTF—not MTBF—is the appropriate average life measure.
Reliability of a Series System — Rs
In a series system, every component must work for the complete system to succeed. If one component fails, the complete function fails.
Rs = complete series-system reliability
R1, R2, R3 = reliability of individual components
Workplace example
A safety alarm may require a sensor, control unit and audible alarm. All three must operate successfully. The system’s reliability will therefore be lower than the reliability of its best component.
Live series-system builder
Why does the overall value fall?
Each additional required component creates another point at which the system can fail. Improving the weakest component can have an important effect on the complete system.
Reliability of a Parallel System — Rp
Rp is pronounced “R sub p”. The subscript p means parallel. A parallel system contains backup or redundant components and continues to operate if at least one independent component remains functional.
Rp = parallel-system reliability
1 − R = probability that a component fails
Why it is relevant
Redundancy may improve the reliability of fire pumps, emergency generators, ventilation, alarms, communication systems and emergency shutdown arrangements.
Common-cause failure
Two backups are not fully independent if they share the same electrical supply, fuel source, control panel, cooling system, room or maintenance error.
Live parallel-system builder
Combined Interpretation Dashboard
These figures describe different aspects of the same machine. One figure alone does not provide a complete reliability assessment.
Level 6 interpretation
The machine has high availability because failures can be repaired quickly, but its probability of completing 100 continuous hours without failure is only 81.87%. High availability must not be incorrectly presented as proof of high reliability or safety.
Use Calculations for Failure Tracing
Collect operating and failure data
Gather operating hours, downtime, repair duration, failure type, failed component, operating condition and maintenance history. Calculations are only as reliable as the data used.
Move From a Number to a Defensible Judgement
Is it acceptable?
Compare the result with the performance target, manufacturer information, legal or organisational requirements and the safety function.
Is it changing?
Compare time periods. Decide whether MTBF is falling, failure rate or MTTR is rising, and whether availability or mission reliability is deteriorating.
What should happen?
Identify weak equipment, consider consequences, investigate causes, prioritise action and allocate competent people, time, spares and budget.
Questions an HSE professional should ask
- Is the result acceptable for the required safety function?
- Is performance improving, stable or deteriorating?
- How does this period compare with earlier periods?
- Which equipment or component is weakest?
- What are the safety, health, environmental and business consequences?
- Is further inspection or investigation required?
- What management action and resources are justified?
- Is the data accurate, complete and collected on a consistent basis?
Four Ways to Analyse Performance
1. Establish a baseline
A baseline is the starting performance against which future results are compared. Example: MTBF 1,000 h; MTTR 4 h; λ 0.001/h; availability 99.60%; R(100) 90.48%.
The same definition of failure, operating-time boundary, repair-time boundary and units must be used in every later comparison.
2. Analyse the trend
One result is a snapshot. A sequence reveals direction. Falling MTBF together with rising λ and MTTR is stronger evidence of deterioration than one isolated value.
3. Compare equipment
Compare like with like, then explain differences in age, condition, duty, workload, environment, design, operator competence, maintenance, materials and modifications.
4. Compare with a target
A fire pump may have a target availability of at least 99.5%, MTTR below 3 h and λ below 0.001/h. Actual values of 98.7%, 5 h and 0.002/h are adverse gaps requiring action.
| Indicator | Q1 | Q2 | Q3 | Q4 | Interpretation |
|---|---|---|---|---|---|
| MTBF (h) | 1,000 | 800 | 600 | 400 | Failures are becoming more frequent. |
| MTTR (h) | 4.0 | 4.5 | 5.0 | 7.0 | Recovery is becoming slower. |
| λ (failures/h) | 0.00100 | 0.00125 | 0.00167 | 0.00250 | The deterioration is consistent across indicators. |
When should performance be treated as abnormal?
Do not judge a figure in isolation. Compare it with all relevant reference points:
- The asset’s previous performance
- Similar equipment performing the same duty
- Manufacturer specifications and design limits
- Internal standards and maintenance targets
- Industry benchmarks and recognised good practice
- The reliability required for the safety function
A statistically unusual result, a continuing adverse trend or any failure that threatens a critical protection function requires investigation—even when the percentage appears numerically high.
Compare Two Performance Periods
Change the data. The tool calculates MTBF, MTTR, λ, availability, R(t) and F(t), then explains the direction of change.
Period 1 — baseline
Period 2 — current
| Indicator | Period 1 | Period 2 | Direction |
|---|
One Indicator Can Mislead
High availability can hide frequent failure
If MTBF = 500 h and MTTR = 2 h, availability is 99.60%. Yet, with λ = 0.002/h, R(100) is only 81.87%. Rapid repair creates high availability but does not make the system failure-free.
Consequence changes the priority
Risk combines likelihood and consequence. The same failure probability may be tolerable for a non-critical printer but unacceptable for a gas detector, fire pump or emergency shutdown system.
| Indicator | Pump A | Pump B | Meaning |
|---|---|---|---|
| MTBF | 1,200 h | 450 h | Pump B fails more frequently. |
| MTTR | 3 h | 8 h | Pump B is slower to restore. |
| λ | 0.00083/h | 0.00222/h | Pump B has the higher failure rate. |
| Availability | 99.75% | 98.25% | Pump B is less ready for use. |
Prioritise for investigation when evidence shows:
- High or rising failure rate
- Falling MTBF
- High or rising MTTR
- Availability below target
- Unacceptable F(t)
- Repeated failure of one component
- Safety-critical function
- No effective backup
Start With a Clearly Defined Failure Event
Failure identification
States what failed, where and when: for example, “Fire Pump B failed to start during the weekly proof test at 09:20.”
Failure tracing
Examines how and why the failure developed, how it moved through the system and which immediate, underlying and root causes must be controlled.
Common triggers
- Sudden or continuous reduction in MTBF
- Rising λ or MTTR
- Availability below target
- Unacceptable failure probability
- Repeated failure of one component
- Large differences between similar assets
- Failure soon after maintenance
- Failure of a safety-critical item
- Primary and backup equipment failing together
- Deterioration in previously reliable equipment
Evidence to collect before concluding
- Operating hours, cycles and duty
- Failure date, time, type and alarms
- Repair time, tests and replaced parts
- Maintenance, inspection and proof-test history
- Operators, contractors and shift conditions
- Manufacturer information and design limits
- Temperature, pressure, vibration and process data
- Dust, moisture, corrosion and other environmental conditions
- Photos and preserved damaged components
- Permit-to-work and isolation records
- Training, competence and handover records
- Previous investigations and management-of-change records
Select the Result That Changed
Twelve Steps From Function to Verification
Select each step. A competent investigation may move back and forth as new evidence appears.
Do Not Stop at the First Technical Explanation
Immediate cause
The direct event closest to the failure: a bearing overheated and seized.
Underlying cause
The local condition that allowed it: insufficient lubrication and no condition warning.
Root / organisational cause
The management-system weakness: the maintenance system did not specify, schedule or verify lubrication and monitoring.
Five Whys — pump example
Fault Tree Analysis — “fire pump fails to start”
- Electrical: loss of supply, open protection, starter defect.
- Mechanical: motor seizure, pump obstruction, coupling failure.
- Control: sensor, logic, signal, set-point or interlock failure.
FTA helps test multiple credible paths instead of accepting the first visible defect.
Failure Modes and Effects Analysis — FMEA
FMEA is a forward-looking method. For each component or process step, identify the possible failure mode, its effect, its likely cause, existing controls and any further action. It helps prioritise what could fail before an incident occurs.
| Item | Failure mode | Possible effect | Possible cause | Existing / further control |
|---|---|---|---|---|
| Pump seal | Leakage or face damage | Loss of containment and pump shutdown | Misalignment, vibration or unsuitable material | Alignment verification, material review and vibration monitoring |
RPN is a prioritisation aid, not a universal risk-acceptance rule. Follow the organisation’s FMEA rating definitions; a catastrophic severity may require action even when the total RPN is not the highest.
Find the Component Dominating the Failure Record
Example total = 12 failures. The tool sorts the components and calculates percentage and cumulative percentage.
| Component | Failures | Share | Cumulative |
|---|
Human, Organisational, Common-Cause and Hidden Failures
Human & organisational factors
Examine workload, fatigue, supervision, shift handover, competence, interface design, communication, production pressure, resources, contractor control, unclear roles and ignored warnings.
Common-cause failure
Parallel equipment may share electricity, control panel, fuel, cooling, room, procedure, maintenance error or defective component batch. Redundancy is valuable only when independence is credible.
Hidden failure
An alarm, emergency generator, relief valve, gas detector, interlock or shutdown may fail without being noticed until demanded. Inspection and proof testing must reveal dormant failure.
Recurring Pump-Seal Failure: Before and After
Initial evidence — 6,000 operating hours
12 failures, 60 repair hours; 8 of 12 failures involved the seal.
- MTBF = 6,000 ÷ 12 = 500 h
- MTTR = 60 ÷ 12 = 5 h
- λ = 12 ÷ 6,000 = 0.002/h
- A = 500 ÷ 505 × 100 = 99.01%
- R(100) = e−0.2 = 81.87%
Tracing findings
Evidence showed excessive vibration and shaft misalignment. Seals had been replaced repeatedly without checking alignment; no vibration monitoring existed, and the maintenance procedure addressed replacement but not the failure cause.
Failure
Pump stopped because leakage exceeded the safe limit.
Mode & immediate cause
Seal faces damaged by excessive vibration.
Underlying & root cause
Misalignment plus a procedure that omitted alignment verification and vibration monitoring.
Corrective actions
Realign the pump, replace damaged parts, introduce vibration monitoring, revise the maintenance procedure, train maintainers, verify alignment after intervention and review similar pumps for the same weakness.
| Indicator | Before | After next 6,000 h | Change |
|---|---|---|---|
| Failures | 12 | 3 | 75% reduction |
| MTBF | 500 h | 2,000 h | Longer between failures |
| MTTR | 5 h | 3 h | Faster safe recovery |
| λ | 0.002/h | 0.0005/h | 75% reduction |
| Availability | 99.01% | 99.85% | Improved readiness |
| R(100) | 81.87% | 95.12% | Improved mission reliability |
Verification: the recalculated results support the conclusion that the actions improved performance. Continue monitoring to confirm that the improvement is sustained and has not introduced another risk.
What the Calculations Can and Cannot Prove
They can show
- Frequency, average duration and direction of change
- Comparison with another asset, period or target
- Which component dominates the recorded failures
- Whether performance improved after action
They cannot prove alone
- The physical, human or organisational root cause
- That the data definition and records are accurate
- That a high availability figure means the system is safe
- That a short-term improvement will continue
Failure-tracing record
Record the asset and required function; defined failure event; operating hours and duty; failed component and failure mode; immediate, underlying and root causes; evidence reviewed; risk and consequence; corrective actions, owner and due date; post-action calculations; proof of effectiveness; residual risk; and lessons shared with similar systems.
Reliability Case Studies
Compressor comparison
Compressor A operates for 12 months without failure. Compressor B requires repair every six weeks.
LEV performance deterioration
Capture velocity falls from 0.52 m/s to 0.28 m/s while fume concentration rises from 1.2 mg/m³ to 4.8 mg/m³.
Test Your Understanding
Understand the Strategies and Techniques of Risk Control
Welcome to the next part of our learning journey. Section 01 helped us identify, calculate and trace risk-related evidence. Section 02 asks the next management question: What should we do about the risk, and how can we justify that decision?
Learning outcomes: 2.1 Evaluate the use of common risk-management strategies. 2.2 Justify when to use risk avoidance, risk reduction, risk transfer, risk analysis, risk evaluation and risk review strategies. 2.3 Explain the development and characteristics of safe systems of work and safe operating procedures.
What Is Risk Control?
Imagine that we find an unguarded moving part on a machine.
The moving part is the hazard. The possibility and seriousness of someone being injured is the risk. The guard, isolation system or redesign used to prevent contact is the control.
Risk control means selecting, implementing and checking measures that eliminate the hazard or reduce the risk to an acceptable or required level.
Risk assessment is not the final product
A completed form does not protect anyone. Protection appears only when suitable controls are implemented, communicated, resourced, used and verified.
Easy memory: Find it → Understand it → Control it → Check it.
Hazard
Something with the potential to cause injury, ill health, damage or another unwanted outcome.
Risk
The combination of how likely harm is and how serious the consequence may be, considering exposure and existing controls.
Control
A measure that removes the hazard, prevents exposure, reduces likelihood, limits consequence or supports safe recovery.
The Risk-Assessment Process—Seven Clear Steps
Different organisations may group the stages differently. Here, we use the seven stages indicated for this learning outcome, followed by a continuous review loop. Select each step to hear the portal explain it.
Control Standards, Action Plans and Priority
What is a risk-control standard?
It is the benchmark that tells us what level or type of control is required. Sources may include legislation, approved guidance, exposure limits, engineering codes, manufacturer instructions, industry good practice, internal rules and the hierarchy of controls.
Why it matters: A coloured risk score alone cannot decide whether a mandatory guard, ventilation system, permit or exposure limit is required.
What makes an action plan useful?
- Specific control action—not “be careful”.
- Named responsible owner and adequate resources.
- Realistic completion date and interim protection.
- Priority based on risk, standards, people and uncertainty.
- Method to verify completion and effectiveness.
Priority is not “highest score only”
Give prompt attention to imminent danger, serious consequences, legal or control-standard gaps, many people exposed, vulnerable persons, ineffective controls and high uncertainty. A lower matrix score must not be used to postpone a non-negotiable requirement.
Generic, Specific and Dynamic Risk Assessments
These are not three levels of quality. They are three ways of matching the assessment to the work situation.
| Assessment type | Easy meaning | Use it when… | Do not rely on it when… | Main strength and limitation |
|---|---|---|---|---|
| Generic | A baseline assessment for similar activities, hazards or locations. | Work is routine, repeated and genuinely comparable; common controls can be standardised. | People, equipment, substances, environment or task conditions differ materially. | Strength: efficient and consistent. Limitation: may overlook local or individual differences. |
| Specific | An assessment for one task, site, machine, substance, project or person. | Work is unusual, complex, high risk, legally specific, non-routine or affected by vulnerability. | A suitable generic assessment already covers truly identical low-risk work—although local checks are still needed. | Strength: detailed and relevant. Limitation: requires time, information and competence. |
| Dynamic | A continuous, in-the-moment judgement as conditions change. | Emergency response, rapidly changing work or an unexpected condition requires immediate reassessment. | Work is planned and foreseeable. It must not replace a suitable formal assessment, method statement or permit. | Strength: responds to reality. Limitation: time pressure and incomplete information may weaken judgement. |
Interactive Assessment-Type Selector
Describe the situation. The tool will recommend a starting approach and explain why.
What Do Generic, Specific and Dynamic Assessments Actually Look Like?
Think of these as three different lenses. A generic assessment establishes a reusable baseline. A specific assessment focuses that baseline on one real task, place, item, substance or person. A dynamic assessment keeps checking the live situation while work or an incident develops.
Interactive Assessment-Format Explorer
Select a type. The portal will show its purpose, its header, the columns a suitable form normally needs, a completed example and the final decision that must be recorded.
Core fields every formal assessment needs
- Clear scope, boundaries, location, activity and version.
- Assessor, competent contributors, workers consulted and approval.
- Hazards and credible harm—not only a list of objects.
- Who may be harmed, including contractors, visitors and vulnerable people.
- Existing controls and evidence that they are really in place.
- Risk judgement before further action, with the reasoning shown.
- Further controls, owner, due date, interim protection and priority.
- Residual risk, communication, verification and review triggers.
A blank box is not automatically a bad form
Good forms create space for sound thinking; they do not replace it. The assessor must walk the task, consult the people who understand the work, use relevant evidence and test assumptions.
Quality test: Could a competent supervisor read this record and understand what may happen, who is exposed, which controls must exist, what remains to be done and when work must stop?
Do Not Miss Long-Term Hazards to Health
An injury hazard may produce an immediate event. Many health hazards are quieter: exposure can accumulate and illness may appear months or years later. “Nothing happened today” is not evidence that the risk is controlled.
Exposure pattern
Frequency, duration, intensity, route, peaks, recovery time and combined exposure.
Evidence
Monitoring, sampling, health surveillance, absence records, worker reports and historical data.
Control
Prevent exposure at source. PPE and surveillance support control; they do not replace elimination or engineering where these are required.
Qualitative, Semi-Quantitative and Quantitative Assessment
The difference is how the risk is described and analysed—not how seriously the assessor takes it.
Qualitative
Uses reasoned descriptions such as low, medium, high, unlikely or severe.
Useful for: straightforward work, screening, discussion and situations where numerical data would add little.
Limitation: categories may be subjective and different risks may appear equal.
Semi-quantitative
Assigns ranked numbers to likelihood and severity, often calculating a matrix score such as L × S.
Useful for: consistent comparison and action planning across many hazards.
Limitation: numbers can create false precision; score boundaries and multiplication rules are organisational conventions.
Quantitative
Uses numerical estimates of exposure, probability, frequency or consequence based on data and models.
Useful for: complex, high-consequence or technical decisions and comparison with numerical standards.
Limitation: needs valid data, competence and transparent assumptions; an exact-looking number may still be uncertain.
From Professional Description to Ranked Scores and Probability Models
There is no honest single-name answer to “who invented risk assessment?” People have judged danger for as long as organised work has existed. Modern methods developed gradually as industries needed decisions that were more systematic, comparable and technically defensible.
The methods are a progression of detail—not a competition
Qualitative thinking helps us name what may happen and judge its importance. Semi-quantitative scoring helps a group rank and compare many concerns using defined categories. Quantitative analysis estimates frequency, probability, exposure or consequence when the decision needs numerical evidence.
A major quantitative study still begins with qualitative questions: Which scenarios matter? What assumptions are credible? Who may be affected? A risk matrix still needs professional judgement. The methods often work together.
What the historical landmarks do—and do not—prove
System-safety programmes helped formalise ranked categories and risk matrices, but no single standard invented every semi-quantitative method. By 1971, NASA and aircraft manufacturers were using fault-tree tools. The 1975 US Reactor Safety Study, WASH-1400, directed by MIT professor Norman Rasmussen and AEC staff member Saul Levine, became a landmark probabilistic risk assessment. Later criticism of some numerical claims also taught an essential lesson: model structure, data gaps and uncertainty must be made visible.
Historical teaching sources: US Nuclear Regulatory Commission histories of the Reactor Safety Study and risk-informed regulation; NASA System Safety Handbook risk-matrix material; MIL-STD-882 system-safety practice.
The Qualitative, Semi-Quantitative and Quantitative Assessment Workshop
We will use one scene throughout: pedestrians and forklifts interact in a busy warehouse loading area. This makes the difference easy to see—the hazard does not change, but the depth and form of the analysis changes.
Open qualitative form ↓ 2. Semi-quantitative formDefined rankings combined in the interactive 5 × 5 matrix.
Open matrix ↓ 3. Quantitative formModelled event frequency, uncertainty range and numerical criterion.
Open quantitative form ↓
Reasoned words
Judgement: “Collision risk is high because pedestrians frequently cross an active vehicle route and a collision could cause fatal injury.”
Defined ranks
Judgement: Likelihood 4 (likely) × severity 5 (catastrophic) = 20, “very high” under the example organisation’s approved matrix.
Measured or modelled values
Evidence: vehicle movements/hour, crossing frequency, near-miss rate, speed, exposure time, barrier reliability and predicted collision consequence.
Interactive Method-Format Explorer
Select a method to see the complete form structure. Notice that each higher-data method retains the basic hazard, people, controls, action and review fields.
Which Method Is Proportionate?
Answer five questions. The recommendation is a starting point for competent judgement—not an automatic approval.
Build a Defensible Qualitative Judgement
“High risk” alone is weak. Link the descriptor to exposure, consequence, controls, evidence and uncertainty.
Create One Complete Risk-Assessment Record
This tool mirrors the essential columns of a practical assessment register. It helps learners see that a risk rating sits inside a much larger management record.
Interactive Qualitative Risk-Assessment Form
A qualitative assessment uses defined words and reasoned professional judgement. It does not multiply scores. The record must explain why the likelihood, consequence and overall priority descriptions fit the evidence.
A. Assessment identity and scope
B. Hazard, people and evidence
C. Initial qualitative judgement
D. Treatment, responsibility and residual judgement
Example likelihood meanings
- Rare: exceptional under the defined conditions.
- Unlikely: foreseeable but not expected during normal activity.
- Possible: could occur during the activity or assessment period.
- Likely: expected to occur repeatedly unless controls improve.
- Almost certain: occurs frequently or conditions make occurrence imminent.
Example consequence meanings
- Minor: limited, short-term harm.
- Moderate: treatment or restricted work may be required.
- Serious: major injury or significant occupational ill health.
- Major: life-changing harm or single fatality potential.
- Catastrophic: multiple fatalities or major widespread impact.
Interactive 5 × 5 Semi-Quantitative Risk Matrix
Compare the initial risk with the residual risk after proposed controls. The example bands are for learning only; an organisation must define and approve its own criteria.
Choose the ratings
Solid outline: initial rating. Dashed white outline: residual rating.
Why might severity stay at 5?
A guard or interlock may make contact much less likely, but if contact still occurs the possible injury may remain catastrophic. Do not automatically reduce both numbers simply because controls were proposed. Rate the real effect of the selected controls and verify them.
A Simple Quantitative Comparison Tool
Measured value compared with an applicable limit
This demonstration calculates an exposure ratio. Both figures must use the same unit and come from a valid assessment.
How to interpret carefully
A ratio of 0.70 means the measured value is 70% of the selected limit. It does not automatically mean there is no risk or that controls can be relaxed.
Check sampling quality, uncertainty, peak exposure, routes of exposure, combined substances, vulnerable people, the legal meaning of the limit and whether further reduction is required.
Quantitative Scenario-Frequency Calculator
This teaching model follows one event path. It estimates how often the defined harmful outcome may occur by combining an initiating-event frequency with conditional probabilities. It is useful for learning the logic of event trees; it is not a substitute for a validated QRA.
How to Read Every Symbol—and Why It Is Used
It means the estimated frequency of the defined harmful outcome, normally stated per year. It appears on the left because this is the answer the model is calculating.
It means how many times the event that starts the harmful scenario occurs during a stated period, usually one year. The small word below f is a label—it is not multiplication.
P shows that the value inside the brackets is a chance from 0 to 1. For example, 0.25 means a 25% chance.
It means the chance that a person is present or exposed when the initiating event occurs. It is used because an initiating event cannot harm a person who is not in the exposure path.
It means the chance that the intended barrier, safeguard or protective control fails, is unavailable or does not stop the event path.
The vertical bar | means “given that”. It asks: once the event has reached the exposed person, what is the chance of the defined harm?
Multiplication is used because the harmful outcome follows a sequence: the initiating event occurs, exposure exists, the control fails and harm follows. Each stage narrows the original frequency.
It separates the answer on the left from the factors used to calculate it on the right. Both sides describe the same estimated harmful-outcome frequency.
This is the unit, not a percentage. A result of 0.12 per year is a modelled event frequency; it does not mean a 12% annual risk unless a valid model specifically supports that interpretation.
Four checks before trusting the output
- Scenario: Is the initiating event and harmful outcome defined without ambiguity?
- Dependence: Are the probabilities really independent, or can one common cause defeat several controls?
- Data: Are frequencies based on comparable equipment, tasks, people and operating conditions?
- Uncertainty: Would reasonable lower and upper assumptions materially change the decision?
Level 6 point: A calculated frequency is an estimate conditional on a model. Report units, source, time period, assumptions, uncertainty and sensitivity—not only the final number.
Interactive Quantitative Risk-Assessment Form
A quantitative assessment records more than a calculation. It defines the decision, scenario and model; gives every input a source and unit; shows uncertainty; compares the result with an approved criterion; and records treatment and verification.
A. Decision, scope and model boundary
B. Numerical model inputs
The pronunciation and meaning of every symbol are explained in the symbol guide immediately above.
C. Assumptions, treatment and assurance
What the uncertainty factor means here
An uncertainty factor of 2 displays a teaching range from the central estimate ÷ 2 to the central estimate × 2. A real QRA may require probability distributions, confidence intervals, alternative models or structured expert judgement. The factor is a learning device—not a universal scientific rule.
What must be reviewed independently?
Scenario completeness, units, input provenance, relevance of data, dependencies and common causes, human-reliability assumptions, consequence model, uncertainty treatment, sensitivity, numerical criterion and whether the model is valid for the decision.
Build a Complete Risk-Control Action
An action without an owner, date, interim measure and verification method is only an intention. Complete the fields and create a practical action-plan entry.
Meet the Six Strategies
Select a card. The portal will explain when to use it, when not to depend on it, and what a Level 6 justification should recognise.
When Each Strategy Is Most Relevant
| Strategy | Use when… | Evidence needed | Key limitation or warning |
|---|---|---|---|
| Avoidance | Exposure can be removed by stopping, substituting, redesigning or choosing another objective. | Severity, feasibility of alternatives, standards, lifecycle and risk-transfer effects. | May create a different risk or sacrifice an essential activity; assess the replacement. |
| Reduction | The activity is necessary and the risk can be controlled using the hierarchy of controls. | Control performance, human factors, maintenance, residual risk and verification. | Administrative controls and PPE are vulnerable to failure; reduce at source where possible. |
| Transfer | Specialist competence or financial sharing is appropriate—for example, a competent contractor or insurance. | Competence, contract scope, interfaces, supervision, insurance and monitoring. | Legal and ethical responsibility for protecting people is not simply transferred away. |
| Analysis | The risk, causes, exposure, failure paths or options are uncertain or complex. | Measurements, incidents, task information, models, assumptions and uncertainty. | Do not delay obvious immediate controls while waiting for perfect data. |
| Evaluation | Analysed risk must be compared with legal, technical or organisational criteria to set priority and treatment. | Approved criteria, control standards, consequence, affected groups and tolerance. | A matrix colour cannot override a mandatory requirement or conceal uncertainty. |
| Review | Time has passed or change, incident, failure, new evidence, worker concern or new requirements may affect validity. | Inspection, monitoring, incidents, health data, assurance findings and change information. | A scheduled annual review is insufficient when a trigger requires immediate reassessment. |
Interactive Strategy Decision Lab
Choose a workplace situation and the strategy you think should lead the response. The tool will explain the strongest answer and supporting strategies.
Risk Reduction Must Follow the Hierarchy of Controls
When a risk cannot be avoided completely, begin with controls that act on the hazard and exposure pathway. Measures lower in the hierarchy usually depend more heavily on consistent human behaviour.
Eliminate
Remove the hazard from the work—for example, design out the need to enter a vessel.
Substitute
Replace it with a safer material, method, machine or energy source, then assess the substitute’s hazards.
Engineering
Isolate people through guarding, enclosure, segregation, automation, extraction or fail-safe design.
Administrative
Use planning, permits, procedures, competence, supervision, scheduling, signage and restricted access.
PPE
Protect the individual when exposure remains. Select, fit, maintain and supervise its use; do not make it the automatic first answer.
Risk Transfer—The Point Learners Must Not Miss
What can be transferred or shared?
Some financial loss may be insured. Specialist work may be contracted to an organisation with suitable equipment and competence. Contract terms may allocate defined responsibilities.
What does not disappear?
The hazard remains until controlled. The client or employer must still select competent parties, provide information, coordinate interfaces, monitor work and meet applicable legal duties. A signature on a contract is not a physical control.
Build a Level 6 Justification
Use the structure Decision → Because → Evidence → Limitation → Review. This produces a learning scaffold that you should explain in your own professional words.
Analysis, Evaluation and Review—Do Not Mix Them Up
One simple example
Analysis: solvent-vapour measurements, duration and ventilation performance show the nature and level of exposure. Evaluation: the evidence is compared with applicable exposure criteria and good practice to decide whether treatment is adequate. Review: monitoring and reassessment confirm whether the new local exhaust ventilation continues to control exposure after process or maintenance changes.
Common Errors—and the Better Approach
| Weak approach | Why it is weak | Better Level 6 approach |
|---|---|---|
| Copy a generic assessment without checking the site. | Local hazards, people and conditions may be different. | Use it as a baseline, then verify and adapt it before work. |
| Use a dynamic assessment for planned high-risk work. | It avoids proper planning, consultation and control design. | Complete a formal specific assessment, then use dynamic checks for real-time change. |
| Reduce both likelihood and severity scores automatically. | The selected control may affect only one dimension. | Explain how each control changes exposure, failure path or consequence. |
| Call a risk “low” because no one has yet been harmed. | Absence of recorded harm is weak evidence, especially for latent health risks. | Consider exposure data, potential severity, under-reporting and control reliability. |
| Transfer work and stop managing it. | Contracting does not eliminate the hazard or all duties. | Assess competence, coordinate, monitor and verify contractor controls. |
| Review only once a year. | A change or failure can make the assessment invalid immediately. | Use scheduled and event-triggered review. |
Can You Make and Justify the Decision?
Learn One Control Layer at a Time
The portal starts with the wider safe system of work, develops it, then moves to the more focused safe operating procedure. Only after both are clear do we connect the document family and explore the complete permit-to-work lifecycle.
DB HSE learning resource. Prepared solely by Debjyoti Biswas for teaching Unit 3, Section 2.3. This independent learning portal is not produced by OTHM and does not issue workplace authority.
A Risk Assessment Decides What Must Be Controlled; the System of Work Decides How
Imagine a chemical-transfer pump that needs maintenance.
The assessment identifies hazardous chemical residue, stored pressure, electricity, moving parts, restricted access, contractors and possible conflict with nearby operations. That information is essential—but it does not yet tell the maintenance team exactly how the job will be prepared, authorised, completed, checked and handed back.
The safe system of work connects the people, equipment, controls, communication and sequence. The safe operating procedure gives the approved steps for a defined operation inside that wider system.
The paperwork test
A document does not make work safe merely because it has been signed. The controls must exist at the workplace, the people must understand them, and a responsible person must verify that they remain effective.
Easy memory: Assess → Design → Explain → Do → Check → Improve.
One Scenario, Four Stages of Control
We will use the same chemical-transfer pump throughout Section 2.3. This allows you to see how one risk picture is converted into a complete working arrangement.
What Is a Safe System of Work—SSOW?
Safe System of Work — SSOW
Pronounced: “S-S-O-W,” or simply “safe system of work.”
An SSOW is a deliberately organised method for completing work so that foreseeable hazards are controlled throughout preparation, execution, completion and foreseeable abnormal conditions.
It is a system because it joins people, plant, materials, environment, controls, responsibilities, communication, competence, supervision and review. It is not only a list of steps.
What an SSOW is not
It is not a risk-assessment form, a signature, a list of PPE, a copied method statement or an instruction to “be careful.” It must convert risk decisions into a realistic arrangement that people can understand and use.
Simple test: Could a competent team use this system to know who does what, in what order, with which controls, when to stop, and how the work returns safely to normal?
| SSOW format field | What to record | Why the field matters |
|---|---|---|
| Identity and control | Title, number, owner, version, approval and review date. | Prevents obsolete or unapproved instructions being used. |
| Purpose, scope and limits | Activity, plant, location, persons, conditions, interfaces and exclusions. | Stops the system being applied outside the conditions it was designed for. |
| Roles and competence | Who plans, authorises, performs, supervises, verifies, hands back and reviews. | Prevents gaps, duplication and unverified assumptions. |
| Hazards and control basis | Assessment reference, standards, energy sources, substances, exposure and credible failures. | Shows that the method is risk-based and technically supported. |
| Preparation and resources | Access, barriers, isolations, tools, staffing, communication, permits and PPE. | Creates the conditions needed before work starts. |
| Safe sequence and hold points | Ordered actions, responsible role, required result and checks before progression. | Some controls only work when applied in the correct order. |
| Stop, abnormal and emergency rules | Conditions that suspend work, safe state, escalation, rescue and recovery. | Prevents unsafe improvisation when assumptions change. |
| Completion, handback and review | Inspection, reinstatement, records, acceptance, monitoring and review triggers. | Controls the return to normal operation and captures learning. |
Interactive Jargon Translator
Select any term. The portal will pronounce it, define it and explain why it matters.
Do Not Mix Up the Document Family
The documents are connected, but they do different jobs. The level of formality should be proportionate to the risk, complexity and need for coordination.
| Document or arrangement | Easy purpose | Typical use | Important limitation |
|---|---|---|---|
| Risk assessment | Identifies hazards, people, existing controls, risk and further action. | Before deciding the safe method and whenever relevant change occurs. | A completed form does not implement the controls. |
| SSOW | Coordinates the whole method, people, controls and interfaces. | Where risks require a defined safe way of working, particularly complex or significant activities. | It fails if impractical, unknown, unsupervised or not followed. |
| SOP | Standardises the safe steps for a defined operation. | Routine or repeated operation, inspection, start-up, shutdown, cleaning or maintenance. | It cannot predict every abnormal condition; stop and escalation rules are needed. |
| Method statement | Describes how a particular job or project stage will be carried out. | Construction, installation, maintenance and contractor work. | A generic copied statement may not match the real site or sequence. |
| JSA/JHA | Breaks a job into steps, hazards and controls. | Task planning and workforce discussion. | Step-by-step analysis must still consider interactions and emergencies. |
| Permit-to-work — PTW | Formally authorises specified work, location, time and precautions. | Defined high-risk or tightly coordinated work under site rules. | A permit is not a guarantee of safety and does not replace risk assessment. |
| Checklist | Confirms that required checks were completed. | Pre-start, inspection, handover and verification. | Ticking boxes without observation gives false assurance. |
| Emergency procedure | Explains response when control is lost or conditions become unsafe. | Credible abnormal and emergency situations. | It must be resourced, communicated and tested—not only filed. |
Tool: Which Arrangement Does This Job Need?
Fifteen Stages for Developing an Effective Safe System of Work
The stages are grouped into five phases. Select a phase to explore what must happen and why.
What Is a Safe Operating Procedure—SOP?
Safe Operating Procedure — SOP
Pronounced: “S-O-P,” or “safe operating procedure.”
An SOP is an approved, controlled and repeatable set of instructions for performing a defined operation safely and consistently. It tells the authorised user what conditions must exist, what to do in sequence, what result to confirm and when to stop.
Typical SOPs cover start-up, normal operation, sampling, cleaning, inspection, safe shutdown, isolation preparation, testing and return to service.
How it fits inside the SSOW
The SSOW coordinates the complete job—including teams, interfaces, permits, isolations, emergency arrangements and handback. The SOP standardises one defined operation within that system.
A pump-maintenance SSOW may refer to separate SOPs for shutdown, electrical isolation, line draining, gas testing and controlled recommissioning.
| SOP format field | What it should contain | Quality question |
|---|---|---|
| Document identity | Title, equipment or process, number, version, owner, approver and review date. | Can the user confirm this is the current approved procedure? |
| Purpose and scope | Intended result, authorised users, operating range, location and exclusions. | Is it clear when the SOP applies—and when it does not? |
| Responsibilities and competence | Operator, supervisor, verifier, specialist and required authorisation. | Does each person understand their role and limit of authority? |
| Prerequisites | Plant state, permits, isolations, tools, inspections, guards, ventilation and PPE. | What must be true before Step 1? |
| Ordered steps | One clear action per step, responsible role, location, setting, safe limit and expected result. | Can the action and its successful outcome be observed? |
| Warnings and hold points | Critical hazards, prohibited actions and mandatory verification before continuing. | Are the most safety-critical instructions easy to find? |
| Operating limits | Pressure, temperature, concentration, speed, time or other acceptance criteria. | Does the user know when a result is outside the safe range? |
| Stop and escalation | Unexpected state, failed check, alarm, leak, defect, safe shutdown and person to contact. | Does the SOP prevent improvisation? |
| Completion and records | Final checks, housekeeping, status communication, log entries and retained evidence. | Can another person confirm the operation ended safely? |
| Review and change | Scheduled date plus triggers such as incident, modification, feedback or new evidence. | Will the SOP remain aligned with the real process? |
How Is an SOP Developed?
Select each phase. The five phases contain ten connected stages: define → observe → assess → sequence → write → validate → approve → train → use → review.
Characteristics of a Strong SSOW and SOP
A good document must be technically correct and usable by the people who depend on it. These characteristics are evidence of quality—not decorative features.
| Characteristic | Easy meaning | Why it matters | Warning sign |
|---|---|---|---|
| Risk-based | Controls come from a suitable assessment and required standards. | The system addresses credible harm rather than copying another job. | The procedure mentions PPE but not the main energy or exposure source. |
| Task-specific | It matches the actual plant, people, place and conditions. | Local differences can change the failure path. | Wrong equipment number, location, substance or isolation point. |
| Proportionate | Detail and formality reflect risk and complexity. | Too little detail leaves gaps; excessive paperwork hides critical controls. | A simple task has 50 pages, while a major intervention has one vague paragraph. |
| Clear and sequential | Actions are unambiguous and in the correct order. | Sequence can determine whether energy or exposure is controlled. | “Make safe” without stating who, how or how safety is verified. |
| Practical | Controls can be applied with available time, access, tools and resources. | Impossible instructions encourage deviation and workarounds. | The required test point cannot be reached safely. |
| Participative | People who understand the work contribute to development and review. | Worker knowledge reveals practical hazards and foreseeable shortcuts. | Written remotely without observing or discussing the job. |
| Role-defined | Authority, responsibility and handover are explicit. | Prevents gaps, overlaps and assumptions. | Everyone believes someone else verified the isolation. |
| Competence-based | Required knowledge, skill, experience and supervision are stated. | The same instruction may not be safe for an inexperienced person. | “Trained person” is stated but competence is never checked. |
| Human-centred | It considers workload, fatigue, usability, communication and predictable error. | Controls must work in real human conditions. | Critical information is buried, contradictory or unreadable. |
| Inclusive and accessible | Users can find, read and understand it. | Language, literacy, disability or unfamiliar terminology can affect safe use. | Only one complex-language copy exists away from the workplace. |
| Abnormal-condition ready | It states when to stop, withdraw, isolate, escalate or use emergency arrangements. | People must not improvise when normal conditions disappear. | No instruction for a leak, failed test or unexpected pressure. |
| Controlled and current | Only the approved version is available and changes are traceable. | Obsolete instructions may conflict with modified plant or controls. | Different versions are posted at the same workplace. |
| Verified | Critical controls are checked before reliance. | An assumed control may be absent, failed or incorrectly applied. | The permit is signed without a field check. |
| Monitored and reviewed | Use and effectiveness are checked over time and after triggers. | Work, people, equipment and evidence change. | Repeated deviations are normalised without investigation. |
When a detailed written system is normally needed
- Significant or high-risk work
- Complex or non-routine tasks
- Several people, teams or contractors
- Critical sequence, isolation or verification
- Permit-controlled work
- Serious foreseeable abnormal conditions
When simplicity may be appropriate
Straightforward low-risk work may be controlled through concise instruction, training and normal supervision. Simplicity must come from low complexity—not from ignoring significant hazards. The arrangement still needs to be understood and effective.
How the Six Risk Strategies Shape the System of Work
These strategies are not six competing documents. They influence different decisions during development, authorisation and assurance.
| Strategy | When it is used | Pump-maintenance application | What the SSOW or SOP must show |
|---|---|---|---|
| Avoidance | The exposure or activity can be removed or redesigned. | Use remote condition monitoring to avoid unnecessary intrusive inspection. | Why the hazardous step is no longer required and whether the alternative creates new risk. |
| Reduction | Necessary work continues under stronger controls. | Isolate, depressurise, drain, purge, verify, segregate and supervise. | Control hierarchy, sequence, responsibilities, verification and residual risk. |
| Transfer | Specialist competence, equipment or financial sharing is appropriate. | Use a competent specialist for seal replacement or hazardous cleaning. | Selection, information, coordination, interfaces, monitoring and retained duties. |
| Analysis | Causes, exposure, failure paths or uncertainty require deeper understanding. | Analyse chemical residue, pressure, isolation effectiveness and previous failures. | Evidence, assumptions and how findings affected the method. |
| Evaluation | Evidence must be compared with criteria to decide adequacy and authorisation. | Compare proposed precautions with legal, technical, manufacturer and site requirements. | Acceptance criteria, decision authority and unresolved gaps. |
| Review | Time, change, incident, feedback or failed control may affect validity. | Revise after a leak, near miss, plant modification, contractor concern or recurring deviation. | Review triggers, owner, evidence, revised version and communication. |
Tool: Build a Combined Risk-Strategy Route
Build an Educational Safe System of Work Draft
Adjust every field to your own workplace example. The output helps you understand the structure; it must be reviewed by competent people against the real workplace and applicable requirements before use.
Build an Educational Safe Operating Procedure Format
An SOP should tell the right person what to do, in what order, under which conditions, and when to stop. It should not ask the user to make undefined safety decisions during a critical step.
Tool: Put the Pump-Maintenance Controls in a Defensible Order
Use the arrow buttons to move each step. The “correct” route is the approved teaching sequence for this example; a real installation may require a different, technically validated sequence.
Permit-to-Work—PTW: Meaning, Purpose, Issue, Use, Handback and Closure
What is PTW?
Pronounced: “P-T-W,” meaning permit-to-work.
A PTW is a formal, time-limited system for authorising specified work at a defined place, on identified plant or equipment, by named or competent parties, subject to stated precautions and conditions.
It communicates an agreement: this exact work may proceed, within this exact boundary, while these verified conditions remain true.
What PTW does not mean
A permit is not a risk assessment, an instruction to begin automatically, a substitute for isolation, a certificate that danger has disappeared, or a transfer of all responsibility to the worker.
Issuing the paper alone does not make a job safe. The assessment, SSOW, competence, communication, physical controls, field checks, supervision and stop-work response must all function.
Why is a permit-to-work used?
A PTW creates disciplined communication where mistakes in plant identity, isolation, timing, coordination or handover could cause serious harm. It defines ownership, prevents incompatible simultaneous activities, records critical precautions, controls the period of work and manages the return to normal operation.
| Work category | Why formal control may be needed | Typical linked controls or certificates |
|---|---|---|
| Hot work | Flame, arc, spark or heat may ignite flammable material or damage adjacent systems. | Gas testing, area preparation, fire protection, fire watch and post-work monitoring. |
| Confined-space or vessel entry | Atmosphere, engulfment, restricted access, energy and rescue hazards can change rapidly. | Isolation certificate, atmospheric test, ventilation, entry log, attendant and rescue plan. |
| Line breaking / hazardous containment | Opening pipework or equipment can release pressure, temperature, toxic, corrosive or flammable material. | Process isolation, drain/vent/purge, decontamination, test and line-break controls. |
| Electrical or mechanical work | Contact, arc, unexpected start, gravity, pressure or stored energy may be fatal. | Isolation plan, lockout/tagout, prove-dead or zero-energy verification and controlled reinstatement. |
| Excavation / ground disturbance | Underground services, collapse, water, contaminated ground and vehicle interaction may be present. | Service drawings, detection and marking, trial holes, shoring, access and inspection. |
| Other site-defined work | Work at height, lifting, radiography, roof access, energised testing or unusual simultaneous work may need coordination. | Site-specific permits, certificates, exclusion zones, specialist plans and interfaces. |
Who is involved?
| Typical role | Easy meaning | Main responsibility |
|---|---|---|
| Area / operating authority | The person controlling the plant or area. | Confirms operating status, interfaces and whether the area can be released and later accepted back. |
| Permit issuer / issuing authority | The competent authorised person who issues the permit. | Checks scope, assessment, precautions, isolations, conflicts, validity and field conditions before authorising. |
| Performing authority / permit receiver | The person accepting the permit for the work party. | Understands the permit, briefs the team, keeps within boundaries, monitors conditions and stops when conditions change. |
| Isolating authority | The person controlling required energy or process isolations. | Applies, records, proves and later removes isolation under the approved process. |
| Authorised gas tester | A competent person approved to test atmosphere. | Uses suitable equipment, records results and understands limits, frequency and conditions of testing. |
| Work party | The persons carrying out the authorised task. | Attend the briefing, follow the SSOW/SOP and permit, protect controls and report change or uncertainty. |
| Permit / SIMOPS coordinator | The person who sees the whole work picture. | Prevents conflicts between permits, operations, contractors, isolations and emergency arrangements. |
Terminology varies: a site may use different role names. The essential point is that authority, competence, accountability, communication and handover cannot be vague.
The Full PTW Lifecycle—14 Phases
Select a phase to see what must happen, why it matters and what evidence should exist.
What should a complete permit form contain?
| Permit field | Information required | Control purpose |
|---|---|---|
| Identification | Permit number/type, work order, exact plant/equipment tag, location and description. | Prevents work on the wrong item or outside the authorised task. |
| Validity | Issue date/time, start, expiry, shift and any rules for extension or revalidation. | Stops an old permit being treated as continuing permission. |
| Supporting documents | Risk assessment, SSOW, SOP/method, drawings, certificates and rescue/emergency plans. | Connects authorisation to the technical control basis. |
| Hazards and interfaces | Energy, substances, atmosphere, access, environment, nearby work and SIMOPS. | Makes foreseeable interactions visible to both parties. |
| Isolations and tests | Isolation points/certificate, lock and tag references, drain/vent/purge, test type, result, time and tester. | Provides traceable evidence of critical plant preparation. |
| Precautions | Barriers, ventilation, fire controls, access, tools, PPE, monitoring and prohibited actions. | Defines conditions that must remain in place. |
| Emergency and communication | Alarm, withdrawal, rescue, contact, stop-work rule, briefing and handover method. | Supports response when normal assumptions fail. |
| Authorisation and acceptance | Issuer and receiver names/signatures, date/time and declarations of understanding. | Confirms that authority and shared understanding are explicit. |
| Suspension / extension / handover | Reason, safe state, new conditions, outgoing/incoming parties and revalidation. | Prevents work continuing across a change without control. |
| Completion and handback | Work complete/incomplete, people/tools cleared, guards restored, plant status, inspection and acceptance. | Controls transfer back to operations and reinstatement. |
| Cancellation and records | Closure time, permit cancellation, linked documents closed, defects/actions and retained record. | Ends authority clearly and preserves evidence for audit and learning. |
Interactive: Is the Permit Ready for Issue?
Tick only items verified by evidence. A high total cannot compensate for a missing critical condition.
Interactive: Shift, Alarm and Change Decision
Interactive: Build a Complete PTW Learning Form
Complete every field. This produces a teaching draft—not a workplace permit or authorisation.
Interactive: Does the Work Need Formal Permit Control?
A permit-to-work is a formal authorisation and communication system for defined work. Select the features that apply. Site procedures and applicable law make the final decision.
Interactive: Stress-Test the Working Arrangement
Rate each condition from 1 (weak) to 5 (strong). This is a learning diagnostic, not a risk-acceptance formula.
PTW Knowledge Check
Tools: SSOW Quality Diagnostic and Review Trigger
Tool 1: Check the Quality of an Existing SSOW
Tool 2: What Should Trigger Review?
Weak Wording Versus Defensible Wording
| Weak wording | Why it fails | Stronger wording principle |
|---|---|---|
| “Make the pump safe.” | No person, isolation method, condition or verification is defined. | Name the equipment, energy sources, authorised role, approved isolation method and verification requirement. |
| “Wear proper PPE.” | “Proper” is undefined and PPE may not control the main hazard. | Select PPE from the assessment after applying higher-order controls; state type, limitation and checks. |
| “Be careful when opening.” | It transfers responsibility to behaviour without controlling stored pressure or residue. | Prevent opening until depressurisation, drainage, safe condition and authorisation are verified. |
| “Experienced workers only.” | Experience is not defined or verified. | State required authorisation, task knowledge, skill, experience and supervision. |
| “In an emergency, act accordingly.” | No stop, alarm, withdrawal or escalation route is given. | Define credible abnormal conditions and the immediate response expected from each role. |
| “Review annually.” | Waits for a date even after a change, failure or near miss. | Use both scheduled and event-triggered review. |
How to Explain 2.3 at Level 6
A strong explanation normally contains
- Clear definitions of SSOW and SOP.
- The relationship with risk assessment and supporting documents.
- A connected development lifecycle.
- Characteristics explained with reasons—not only listed.
- Application of the six risk strategies.
- How risk assessment, SSOW, SOP, PTW, isolation and handback connect.
- The PTW lifecycle from request and field verification to suspension, handback and closure.
- A suitable workplace example.
- Limitations, implementation and review arrangements.
Use this paragraph structure
Point → Meaning → How developed/applied → Why it matters → Workplace example → Consequence if missing → Review.
Write in your own professional words and relate the explanation to a genuine or realistic workplace. A list of headings alone does not fully satisfy “explain.”
Explain how a safe system of work and supporting safe operating procedures should be developed for intrusive maintenance of a chemical-transfer pump. Your response should address consultation, risk strategies, control sequence, competence, human factors, permit interfaces, abnormal conditions, implementation and review.
Can You Convert a Risk Decision Into Safe Work?
Understand the Models of Loss Causation, Analysis of Loss Data and the Importance of Incident Investigation
Section 03 follows one connected learning journey: understand why an event happened, test what the data reveals, assess who needs to receive the report, and explain how investigation converts evidence into prevention.
Current learning route: 3.1 and 3.2 provide the causal and quantitative foundation. Sections 3.3 and 3.4 apply that foundation to reporting decisions, evidence-led investigation and verified risk reduction.
Learning outcomes covered: 3.1 Outline loss-causation theories and techniques. 3.2 Justify quantitative methods in analysing loss data. 3.3 Assess the needs and impacts of reporting loss events. 3.4 Explain the importance and impact of incident investigations.
One Event · Four Connected Questions
Do not study these outcomes as four separate chapters. Each part produces the evidence or decision needed by the next one.
Why did it happen?
Use theories and techniques to organise immediate, underlying, systemic and barrier-related causes.
Output → causal questions and hypotheses 3.2What do the numbers reveal?
Use valid rates, severity, graphs and statistics to identify patterns and justify attention.
Output → quantitative signals and limitations 3.3Who needs to know, when and why?
Assess reporting needs, internal and external routes, stakeholders, impacts, privacy and culture.
Output → timely visibility and escalation 3.4How does evidence become prevention?
Explain proportionate investigation, evidence testing, causal findings, stronger action and verification.
Output → learning, controls and assuranceWhat Do “Loss” and “Causation” Mean?
Loss means an unwanted outcome that removes or damages something of value. It may include injury, ill health, death, environmental harm, property damage, production interruption, legal exposure, financial cost, lost information or damaged trust.
Causation means the way conditions, decisions, actions, failures and interactions combine to produce an event or outcome.
An incident normally has more than one relevant cause. The final action or failed component may be easy to see, but a Level 6 analysis asks what shaped that action, why the control was absent or ineffective, and what management-system conditions allowed the vulnerability to remain.
The central learning rule
A model organises thinking; it does not manufacture evidence.
Investigators must still preserve the scene, gather reliable evidence, test competing explanations, consult involved people, identify controls and verify corrective action.
Immediate cause
The unsafe act, condition, energy transfer or equipment state directly connected to the unwanted event.
Ask: What directly triggered or enabled the contact?
Underlying cause
The job, workplace or organisational factor that allowed the immediate condition or action to arise.
Ask: What influenced the work and weakened control?
Root cause
A deeper management-system or organisational failing whose correction can prevent a wider class of recurrence.
Ask: Why did the system create, accept or fail to detect the vulnerability?
Interactive Section 03 Jargon Translator
Select a term to hear how it is pronounced and understand why it matters.
Our Master Case: Forklift Contact With a Process Line
A forklift enters a congested transfer-area route, passes a damaged low barrier and contacts a valve manifold. A small chemical release occurs. The alarm operates, workers withdraw and the response team isolates the area. No model should be used to blame the driver or to assume a cause before evidence is gathered.
| Evidence source | Initial finding | What must still be tested? |
|---|---|---|
| CCTV and scene measurements | Pallets narrowed the route; forklift contacted the low pipe barrier. | Actual speed, visibility, pedestrian interaction and why storage entered the route. |
| Inspection records | Barrier damage had been recorded twice but not permanently repaired. | Risk classification, escalation, ownership, resources and closure verification. |
| Driver and worker interviews | Peak dispatch created queuing and radio interruptions. | Work-as-done, production pressure, route rules, competence and normal adaptations. |
| Alarm and response log | Detection and emergency isolation limited the release. | Alarm timing, response reliability, exposure and opportunities to strengthen recovery. |
| Six-month loss data | Vehicle–route near-miss reports increased, especially during peak dispatch. | Reporting quality, exposure hours, location clustering and whether risk really increased. |
Accident and Incident Ratio Studies
Ratio studies arrange recorded events by consequence level. They helped organisations recognise that low-consequence events and near misses can reveal control weaknesses before a major loss occurs.
Who, when, why—and what changed?
Why it was created: to show that the serious injury at the top is only part of the recorded experience. Lower-consequence events may provide more frequent opportunities to discover exposure and weak controls.
How thinking changed: Bird widened the categories to include property damage and near misses. Later research showed that the shape changes with definitions, industry, severity threshold and reporting practice, and that fatal and non-fatal events can follow different causal pathways.
Historical context: DNV tribute to Frank Bird. Critical evidence: Salminen, Saari, Saarela and Räsänen (1992) and Marshall, Hirmas and Singer (2018).
Heinrich’s historical triangle
Often presented as 1 major injury : 29 minor injuries : 300 no-injury accidents. Read the colon “:” as “to”: one major injury to 29 minor injuries to 300 no-injury accidents. It was derived from historical insurance and accident records and promoted attention to the larger body of less-serious events.
Bird’s historical ratio
Often presented as 1 serious or major injury : 10 minor injuries : 30 property-damage events : 600 near-miss incidents. It broadened attention to damage and no-loss events. The figures describe Bird’s historical dataset; they do not calculate the probability of the next accident.
| Potential benefit | Why it helps | Limitation to explain |
|---|---|---|
| Encourages near-miss reporting | Weak signals can reveal exposure and failing controls before serious harm. | More reports can mean better trust and reporting—not necessarily worsening safety. |
| Supports prevention | Recurring lower-level events can direct inspection and improvement. | Preventing minor slips does not automatically control a low-frequency catastrophic process event. |
| Communicates scale simply | The triangle is memorable and helps introduce proactive learning. | Simplicity may hide different causal pathways and consequence mechanisms. |
| Provides trend categories | Organisations can compare reporting levels and event types over time. | Changed definitions, workforce hours or reporting systems can create a false trend. |
| Challenges injury-only thinking | Damage and near misses can contain valuable control information. | A ratio does not replace investigation, risk assessment or barrier assurance. |
Interactive Tool: See How Reporting Culture Changes the Triangle
The “true opportunities for learning” remain constant in this teaching example. Adjust the percentage that gets reported. Notice how the visible triangle changes even when the underlying events do not.
Bird’s Loss-Causation Model and Multi-Causality
Bird’s expanded domino approach connects management control with basic causes, immediate causes, the incident and the final loss. Removing or strengthening an earlier “domino” can interrupt the sequence.
Where it came from
Bird developed accident-prevention thinking beyond the earlier person-centred domino sequence. The model placed lack of management control at the beginning and widened “injury” into loss, including harm to people, property, environment and production. Bird’s later work with George Germain was published in Practical Loss Control Leadership in 1985.
How it developed
The sequence was adapted using energy-exchange concepts associated with William Haddon. Read it left to right to explain how loss developed; work from right to left during investigation to ask what control should have interrupted each step.
Multi-causality: More Than One Path Can Matter
Multi-causality rejects the idea that one unsafe act is normally a complete explanation. Several conditions may combine, interact or increase one another’s effect. Causes can exist at task, equipment, environmental, individual, supervisory and organisational levels.
Immediate examples
- Forklift contacts low barrier.
- Route is narrowed by pallet storage.
- Valve manifold remains exposed to vehicle energy.
Underlying examples
- Peak traffic and pedestrian interaction were not reassessed.
- Temporary storage became normal.
- Damaged barrier repair was delayed.
Root examples
- Defect priority criteria ignored major-consequence potential.
- No effective owner verified corrective-action closure.
- Layout-change governance excluded operations and workforce evidence.
Interactive Tool: Classify the Cause—Then Look Deeper
The Swiss Cheese Model
James Reason described safety as several layers of defence, barrier and safeguard. Each layer can contain weaknesses—shown as “holes.” An adverse trajectory can pass through when weaknesses in different layers align.
James Reason: from human error to organisational defences
Who and when: British psychologist James Reason developed the model across the 1990s, beginning with ideas in Human Error (1990) and refining the familiar defence-layer image through the decade. Safety practitioner John Wreathall also influenced its development.
Why: Reason wanted to move investigation beyond “the operator made an error.” The model asks how front-line actions combine with weaknesses created earlier by design, staffing, maintenance, supervision and organisational decisions.
How it changed: diagrams and terminology evolved between 1990 and 2000. This matters: Swiss Cheese is a family of evolving explanations, not one frozen diagram. Later safety thinking also stresses that defences interact dynamically and that a simple line through holes must not replace evidence.
Read the system approach in Reason (2000), Human error: models and management, and the historical critique in Larouzee and Le Coze (2020).
Active failure
An action or omission close in time and place to the event, such as an incorrect control input or missed check. It may trigger the event, but it is rarely the whole explanation.
Latent condition
A deeper weakness created by design, staffing, maintenance, procurement, priorities, procedures, supervision or management decisions. It may remain hidden until combined with local conditions.
Interactive Barrier Alignment Simulator
Mark a layer as failed or ineffective. The event pathway opens only when every selected defence in this simplified example has a relevant weakness.
Limitation: The cheese image is a communication model, not a detailed dynamic simulation. It can oversimplify interactions unless each layer, threat, dependency, owner and performance requirement is defined with evidence.
Fault Tree Analysis and Event Tree Analysis
How structured logic entered safety analysis
Fault Tree Analysis: commonly traced to H. A. Watson at Bell Laboratories in 1962 during the Minuteman missile programme. It was created because complex systems could fail through combinations that a simple checklist might miss.
Event Tree Analysis: developed as a forward, consequence-oriented partner to fault-tree reasoning. Probabilistic risk work such as the US Reactor Safety Study, WASH-1400 (1975), helped establish combined fault-tree and event-tree methods for complex high-hazard systems.
What changed: the techniques expanded from defence and nuclear applications into aviation, process safety and other industries. Software can now evaluate very large trees, but the logic, data, dependencies and uncertainty still need competent human review.
Authoritative background: NRC Fault Tree Handbook and NRC history of WASH-1400.
FTA and ETA Symbol Decoder + Logic-Gate Starter
Symbols are a form of shorthand. They make a large analysis easier to read, but only after every symbol has been defined. Start by reading the symbol aloud, identify what it represents, check its unit or range, and then ask why it is present in the equation.
First: know what every common mark means
Letters name defined events. A might mean “detector fails”; B might mean “isolation valve fails.” The letter has no meaning until the analyst defines it.
The total number of relevant demands, tests, opportunities or observations. It is commonly the denominator. Example: N = 100 valid detector tests.
A count from the defined set. Example: n = 3 recorded failures. A subscript makes the count more specific, such as nF for number of failures.
The subscript is a label, not multiplication. nA counts event A; nF counts failures. Thus nF ÷ N means failures divided by all valid opportunities.
The probability that defined event A occurs on the stated demand, mission or period. Probability has no unit and must lie from 0 to 1 inclusive.
p commonly represents success probability and q the complementary failure probability. When success and failure cover all possibilities, q = 1 − p and p + q = 1.
i is an index meaning “the particular item or branch being considered.” p1, p2 and p3 can represent three different input probabilities.
f is an event frequency with a unit such as per year. fI is the initiating-event frequency used at the start of an event tree.
T often represents total observed exposure time; t often represents the mission time being evaluated. The analyst must state the unit—hours, days or years.
A failure rate, normally stated per unit time. A value of λ = 0.002 per hour means the assumed rate basis is 0.002 failures per operating hour—not a 0.2% certainty for every hour.
A mathematical constant approximately equal to 2.71828. In e−λt, it supports the exponential reliability model; it does not mean an event count.
ETA branches are often labelled S and F. These labels say whether the defined barrier performs its required function at that branch point.
∩ means AND—events occur together. ∪ means OR—at least one of the defined events occurs, including the possibility that both occur.
The vertical bar means “given that.” It is used when B’s probability is evaluated under the condition that A has already occurred.
Add the listed values. In ETA, Σfbranch means add the frequencies of the mutually exclusive terminal branches.
Multiply the listed values. ∏pi means p1 × p2 × p3 and so on; it is not the circle constant π.
Equals states that both sides represent the same value. Multiplication combines required path factors. Division creates a proportion or rate from a count and denominator.
The complement of p. If p is the probability of success, 1 − p is the probability of failure only when the two states are mutually exclusive and cover all defined possibilities.
Calculate the expression inside them first. They group terms and prevent the calculation from being performed in the wrong order.
Zero means impossible within the defined model; one means certain within that model. Multiply a decimal by 100 to express it as a percentage.
Second: what is a logic gate and why do we need it?
A logic gate is a rule that explains how input events combine to create an output event. It is not a physical gate and it does not prove causation by itself. It converts a full English sentence into an exact visual instruction, helping an FTA remain consistent when many failure paths are connected.
If detector failure by itself can cause the output, or valve failure by itself can cause it, connect them through OR. A normal inclusive OR also allows both events to occur.
If the output occurs only when the detector fails and the automatic isolation also fails, connect them through AND. AND describes a required combination, not automatically a time sequence.
The output occurs when at least k of n similar channels meet the stated condition. A 2-out-of-3 trip system needs any two of its three channels. Here n means total channels inside this gate.
NOT, exclusive-OR, inhibit and priority-AND can express special logic or sequence conditions. Use them only when their exact meaning is required and defined; most introductory FTAs begin with AND and OR.
| Event A | Event B | A AND B | A OR B | Plain meaning |
|---|---|---|---|---|
| 0 — does not occur | 0 — does not occur | 0 | 0 | No input occurred. |
| 0 | 1 — occurs | 0 | 1 | Only B occurred; OR opens, AND does not. |
| 1 — occurs | 0 | 0 | 1 | Only A occurred; OR opens, AND does not. |
| 1 | 1 | 1 | 1 | Both occurred; both rules are satisfied. |
Interactive Logic-Rule Coach
Choose the English statement you need to represent. The coach will identify the logic and explain how the symbols are used.
Where Does the Probability Come From?
FTA and ETA do not create probability simply because a box is drawn. Every input needs a defined event, population or equipment item, time or demand basis, data source and uncertainty. The first question is therefore not “Which formula shall I use?” It is “What exactly does this number describe?”
Use this empirical estimate when each opportunity is clearly defined. If a detector failed 3 of 100 valid proof-test demands, the observed failure-on-demand estimate is 3 ÷ 100 = 0.03, or 3%.
If p is success probability, q is failure probability. A barrier that succeeds with probability 0.90 has a complementary failure probability of 1 − 0.90 = 0.10.
The vertical bar means given that. An ETA branch asks for the chance of the next success or failure after the initiating event and earlier branch conditions have occurred.
Under a justified constant-rate exponential model, this estimates the probability of at least one failure by mission time t. It is not suitable automatically for ageing, repair, dependence or changing conditions.
Frequency carries a unit such as events per year. Probability has no unit and stays between 0 and 1. A frequency of 0.5 per year must not automatically be called a 50% annual probability.
This means A and B occur together. The general rule is P(A ∩ B) = P(A) × P(B | A). It becomes P(A) × P(B) only when independence is defensible.
| Possible source | How the value may be obtained | Essential quality question |
|---|---|---|
| Operating or test data | Defined failures ÷ valid demands, or events ÷ exposure time. | Are definitions, equipment, conditions and reporting sufficiently comparable? |
| Reliability model | Use a justified distribution, failure rate and mission time—for example 1 − e−λt. | Do the model assumptions fit ageing, repair, maintenance and operating conditions? |
| Fault-tree calculation | A barrier-failure probability used in an ETA may itself come from a detailed FTA. | Were dependencies, shared utilities, human actions and common causes represented? |
| Expert judgement or analogous data | Elicit and document a defensible estimate or range when direct data are sparse. | Are the experts, evidence, assumptions, bias controls and uncertainty traceable? |
Interactive Probability Source Calculator
Choose one method. The portal will use only the fields required for that method and explain the units and assumptions.
Authoritative learning references: US NRC glossary definitions for fault trees and event trees, US NRC explanation of probabilistic risk assessment and NASA system-safety learning on logic models and probability.
Fault Tree Analysis — FTA
Pronounced: “F-T-A.” A Fault Tree Analysis is a structured, top-down method. Start with one precisely defined unwanted top event, then reason backwards to identify the equipment failures, human failures, external events and combinations capable of producing it.
Open the FTA equation and symbol guide
Real trees require validated logic, common-cause and dependency checks, suitable data, minimal cut sets and sensitivity analysis.
Event Tree Analysis — ETA
Pronounced: “E-T-A.” An Event Tree Analysis is a structured, forward or inductive method. Start with a defined initiating event, then move forwards through the conditional success or failure of safeguards, operator actions and recovery measures until every modelled path reaches an end state.
Open the ETA equation and symbol guide
ETA outputs are only as sound as the initiating frequency, branch definitions, conditional probabilities and dependency assumptions.
| Feature | Fault tree | Event tree |
|---|---|---|
| Direction | Backward from a defined unwanted top event. | Forward from a defined initiating event. |
| Main question | What combinations can cause this event? | What outcomes can follow as barriers succeed or fail? |
| Logic | AND, OR and other gates combine causal events. | Branches represent conditional success/failure pathways. |
| Input basis | Basic-event probabilities or frequencies on a consistent mission, demand or time basis. | Initiating-event frequency plus conditional branch probabilities. |
| Typical result | Qualitative cut sets and, when quantified, top-event probability or frequency. | Conditional path probabilities and frequencies for defined end states. |
| Useful for | Complex failure logic, critical combinations and design weaknesses. | Escalation, mitigation, consequence pathways and outcome frequency. |
| Limitation | A poor top-event definition or missed dependency creates false confidence. | Too many branches become difficult; dynamic interactions may be simplified. |
The Bowtie Model
Bowtie combines fault-tree thinking on the left and event-tree thinking on the right. It places the top event—the moment control over the hazard is lost—in the centre.
A collective industrial technique—not one person’s invention
There is no single uncontested Bowtie inventor or creation date. The visual method evolved collectively from fault-tree and event-tree thinking and was progressively adopted in high-hazard industries. It became popular because a multidisciplinary team could see the complete threat–control–loss pathway on one page.
Why it is used: to connect each threat to a preventive barrier, define the loss-of-control top event, connect consequences to mitigating barriers and make critical-control ownership visible.
How it changed: modern practice adds escalation factors, escalation controls, barrier owners, performance standards and assurance evidence. A decorative “bowtie picture” without these elements is not enough for control management.
See the UK Government Bowtie overview and the Office of Rail and Road’s health-risk application.
Vehicle contacts manifold and containment is lost
Escalation factor
A condition that can defeat or weaken a barrier—for example poor lighting reduces the reliability of a visual route check.
Escalation-factor control
A control that protects the main barrier—for example lighting inspection and emergency lighting support route visibility.
Interactive Bowtie Draft Builder
Behavioural Root-Cause Analysis
Behavioural RCA examines what a person did and the conditions that made the behaviour understandable or likely. It should not become a search for someone to blame.
Where the ABC idea came from
Who and when: there is no single creator of “behavioural root-cause analysis.” Its ABC structure grew from twentieth-century behavioural science. Psychologist B. F. Skinner’s work on operant conditioning helped explain how consequences can strengthen or weaken behaviour.
Why safety practitioners use it: repeated behaviour often makes sense when the antecedents are clear and the immediate consequence is easier, faster or socially accepted. ABC analysis makes these influences visible.
How it changed: modern human-factors practice rejects a behaviour-only explanation. ABC evidence should be combined with task design, competence, equipment, workload, supervision, leadership, culture and organisational controls. This prevents “the worker chose badly” from becoming the false root cause.
Interactive ABC and System-Factor Builder
Boundary: A fair system approach does not mean that every action is acceptable. Deliberate reckless conduct may require a just and proportionate response, but the investigation must still examine supervision, controls and organisational context.
Tool: Which Causation Technique Should Lead?
Complex investigations often combine techniques. Select the main question to see a defensible starting point.
The next step is to test whether valid loss data reveals frequency, severity, distribution or change that supports further enquiry.
From Raw Loss Records to Decision-Useful Evidence
Why a count can mislead
Site A records 8 cases and Site B records 5. Site A may appear worse, but if it worked four times as many hours, its exposure-normalised rate may be lower.
Why a rate can also mislead
A single event can move a small workforce’s rate sharply. A low rate may also reflect under-reporting, changed classification, outsourced exposure or chance—not strong control.
Symbols Used in the Calculations
Calculate and Interpret Loss Rates
Flexible Accident / Incident Frequency-Rate Calculator
Why calculate it? A count alone ignores how long people were exposed. Frequency rate converts defined events into a common hours-worked base, supporting more defensible comparison.
How to say and use every sign
The 200,000-hour convention is used in US OSHA/BLS incidence rates. A one-million-hour base is common for some international frequency measures. Confirm the required definition before comparison.
Severity and Average-Consequence Calculator
Why calculate it? Frequency tells how often events occur; severity shows the amount of defined consequence relative to exposure. Average days per case describes the typical recorded consequence per relevant case.
How to say and use every sign
“Lost day” and severity conventions differ. Record calendar/workday rules, caps, fatalities and restricted work consistently.
Accident Incidence-Rate Calculator
Why calculate it? When reliable hours are unavailable or the required convention uses workers, incidence expresses new defined accidents or cases per standard number of workers.
Do not confuse incidence with frequency: incidence uses workers or people; frequency uses hours worked. “New” means the case started within the reference period.
Ill-Health Prevalence-Rate Calculator
Why calculate it? Prevalence estimates the burden of a condition now—both new and continuing cases—so an organisation can plan health controls, surveillance, support and resources.
Prevalence is not incidence: prevalence counts all qualifying existing cases; ill-health incidence counts only new cases. Long-latency occupational disease may reflect exposures from many years earlier.
Interactive Tool: Which Rate Should I Use?
Formula conventions vary. Examples follow official explanations from the US Bureau of Labor Statistics, International Labour Organization, and UK HSE ill-health statistics guidance.
Tool: Compare Two Sites Fairly
Counts answer “how many?” Rates help answer “how many for the amount of exposure?” Use the same event definition, period and multiplier.
Site A
Site B
Trend, Moving Average and Pareto Analysis
Six-Period Trend Explorer
Enter six comparable monthly rates separated by commas.
Why calculate it? Δ (“delta”) shows relative endpoint change. MA₃ smooths short-term fluctuation by averaging the current and previous two periods. Neither proves the reason for change.
Pareto Priority Explorer
Pareto analysis orders categories from largest to smallest so the team can see where recorded loss is concentrated.
What every symbol means
Histogram, Pie Chart and Line Graph Laboratory
Different charts answer different questions. A chart must match the data structure; attractive graphics do not repair weak definitions or incomplete records.
Interactive Chart Explorer
Choose by question—not appearance
Statistical Variability, Distributions, Sampling and Data Validity
Statistical variability means observations and sample results naturally differ. Validity asks whether the measure actually represents the decision question. A precise calculation from biased or incorrectly classified data is still misleading.
Descriptive Statistics and Approximate CI Explorer
Enter numerical observations such as lost days per case. This tool describes the sample; it does not certify a population model.
Open every equation and symbol
What a distribution can reveal
Representative sample: a sample should reflect the population relevant to the question. Every important subgroup needs a fair chance of inclusion. A large convenience sample can still be biased.
Sampling a population: define the target population, build a suitable sampling frame, select participants or records using a defensible method, record non-response and compare the achieved sample with the population.
UK HSE explains sampling error and 95% confidence intervals in its Labour Force Survey guidance. See also the ONS guide to uncertainty.
Interactive Sample and Validity Check
This is a structured warning tool, not a formal sample-size calculation or statistical certification.
Common data errors to test
| Error | Easy meaning | Example |
|---|---|---|
| Coverage | Some of the population cannot appear in the data. | Night shift and contractors are missing. |
| Selection | The inclusion method favours certain people or records. | Only volunteers answer a wellbeing survey. |
| Non-response | Selected participants do not respond, and may differ from responders. | Workers with symptoms do not trust confidentiality. |
| Measurement | The question, instrument or observer produces inaccurate values. | Different clinics use different symptom questions. |
| Classification | The same case is placed in different categories. | Restricted work is recorded as first aid at one site. |
| Duplicate / missing | A case is counted twice or not counted. | One injury exists in two systems without a unique ID. |
| Denominator | The exposure base excludes relevant work. | Contractor cases included, contractor hours excluded. |
| Processing | Entry, coding, formula or transfer is wrong. | Hours are entered as 40,000 instead of 400,000. |
| Time lag | The measured harm appears long after exposure. | Current respiratory disease reflects earlier dust exposure. |
| Reporting culture | Trust and rules change what becomes visible. | A reporting campaign increases near-miss counts. |
Why Use Quantitative Methods—and Why Not Use Them Alone?
| Reason for using numbers | Decision value | Condition or limitation |
|---|---|---|
| Normalise exposure | Rates allow more defensible comparison across differently sized populations or periods. | Definitions, hours, workforce scope and base multiplier must match. |
| Detect change | Time-series analysis can show sustained deterioration, improvement, seasonality or unusual variation. | Short runs, rare events and changed reporting can produce unstable signals. |
| Measure disease burden | Ill-health prevalence estimates all qualifying existing cases and supports surveillance, control and resource planning. | Long latency, diagnostic access, worker turnover and healthy-worker effects can disconnect current prevalence from current exposure. |
| Prioritise investigation | Pareto and distribution analysis identify categories, locations or activities contributing most recorded loss. | Frequency must be considered with credible severity and major-hazard potential. |
| Express uncertainty | Spread, standard error and confidence intervals show that a sample estimate is not an exact population truth. | The calculation depends on a defensible sample, measurement quality and an appropriate statistical model. |
| Evaluate intervention | Before/after measures can test whether performance changed following control. | Control for exposure, operational change, reporting, regression to the mean and other influences. |
| Communicate performance | Defined indicators support dashboards, accountability and resource decisions. | Targets can encourage under-reporting or classification manipulation if poorly designed. |
| Model event pathways | FTA, ETA and quantitative risk methods estimate the contribution of failure combinations and barriers. | Models contain assumptions, dependencies and uncertainty; precision is not certainty. |
Small numbers
One event can double a rate in a small workforce. Use longer periods, confidence intervals or pooled evidence where appropriate.
Under-reporting
A “good” rate may reflect low trust or restricted definitions. Triangulate with audits, surveys, health data and workforce evidence.
Lagging-only bias
Injury data describes realised outcomes. Add leading evidence about exposure, critical controls, defects and corrective-action quality.
Severity randomness
Similar events can produce very different harm. Do not assume low historical injury means low potential consequence.
Changing denominator
Overtime, contractors, shutdowns and outsourcing change exposure. Record the population and hours consistently.
Metric fixation
Managing the number rather than the risk can distort behaviour. Indicators must serve learning and control—not replace them.
Interactive Level 6 Justification Builder
Why Incident Investigation Remains Essential
Protect people, prevent escalation and avoid disturbing evidence unnecessarily.
Define scope and competence; obtain physical evidence, photographs, measurements, documents, records, data and fair accounts from involved people.
Distinguish fact, interpretation and uncertainty. Check inconsistencies and alternative explanations.
Use suitable models to identify immediate, underlying and root causes, successful controls and failed or missing defences.
Prioritise higher-order and system-level action; avoid relying only on reminders or retraining.
Name owners and dates, manage interim risk, confirm effectiveness and share relevant learning.
Check similar equipment, tasks, sites and procedures so the organisation learns beyond the single event.
Now assess how events enter that dataset, who needs the information, and how reporting quality or culture changes the picture.
From a Workplace Signal to Safe Decisions
What Is Loss-Event Reporting?
Give enough reliable information for safe action
- Date, time and exact location
- People and equipment involved
- Observable facts—without assumptions or blame
- Actual injury, damage, loss or disruption
- Potential consequences—what could credibly have happened
- Immediate controls, such as stopping work, isolation, first aid or restricted access
- Evidence requiring protection, including photographs, CCTV, documents, damaged equipment and witness details
- The person reporting the event and the time it was reported
Oil spill beside Machine 3
At 10:15 a.m., a worker slipped on an oil spill beside Machine 3 and sustained a minor ankle injury. Work was immediately stopped, first aid was provided and the area was isolated.
The event could have caused a serious head injury or contact with moving machinery. Photographs, CCTV footage, the worker’s footwear, maintenance records and witness details were preserved for investigation.
Loss event
A broad management term for an event that caused—or credibly could have caused—injury, ill health, environmental harm, asset damage, production interruption, financial loss or harm to trust and reputation.
Report
The first communication that makes the event visible. It should begin with observable facts, actual and potential consequence, immediate controls and the evidence that may need protection.
Record
The controlled, traceable entry kept for follow-up, analysis and required retention. One event may produce several witness or injury reports, but should have one master event identity with linked records.
Notify
A formal communication to a regulator, emergency service, insurer, client or other body when a current legal, permit or contractual rule requires it. Notification does not replace internal reporting or investigation.
Six Communication Actions—Do Not Treat Them as One Form
Reporting quality is measured by the safe decisions it enables
Do not ask only whether a form was submitted. Ask whether reliable information reached the right people in time to protect, preserve, decide and learn—and whether the reporter received meaningful feedback.
What a Level 6 assessment must weigh
- Actual and credible potential loss
- Stakeholder need and urgency
- Internal and possible external routes
- Benefits, burden and unintended effects
- Privacy, fairness and reporting trust
- Consequences of delay or silence
Write the First Report Before You Know the Cause
| Information status | How to write it | Master-case example |
|---|---|---|
| Verified fact | State what a reliable source, measurement or record confirms. | “The alarm log records activation at 07:42:18.” |
| Direct observation | Name what the observer saw, heard or did without converting it into a cause. | “The operator observed liquid inside the contained sump.” |
| Estimate | Label the value and its basis; do not present it as a measurement. | “The initial spill estimate is 12–18 litres based on bund level.” |
| Witness recollection | Attribute the account and preserve its wording for later corroboration. | “A witness recalled that pallets reduced route width before the contact.” |
| Unknown | State the gap and the evidence needed to resolve it. | “The pressure immediately before contact is not yet confirmed; historian data is being preserved.” |
| Hypothesis—not a report conclusion | Record it only as an explanation to test against supporting and conflicting evidence. | “Barrier degradation may have influenced the outcome; inspection and defect records are required.” |
Interactive First-Report Language Coach
Select a statement to see whether it is factual enough for an initial report or has moved prematurely into blame, minimisation or causal certainty.
What Should the System Be Able to Receive?
What Kind of Loss Event Is This?
Classification gives a shared starting language. It does not decide the cause, prove legal notification or limit an event to one label. The same event may involve several loss types—for example, an injury, equipment damage, a process-safety failure and business interruption.
| Loss-event type | Plain-language meaning | Example signal | First reporting need |
|---|---|---|---|
| Injury or fatality | Physical harm ranging from first aid to death. | Cut, fracture, burn or fatal contact. | Care, scene control, factual record and competent notification screen. |
| Occupational illness or disease | Acute or gradual work-related harm to health. | Dermatitis, hearing loss, vibration symptoms or occupational asthma. | Support the person, protect health data and examine exposure history. |
| Near miss | An unplanned event with no actual loss but credible potential for harm. | A suspended load falls into an empty exclusion zone. | Capture the warning before evidence and memory disappear. |
| Dangerous occurrence | A specified or serious event indicating major risk, whether or not someone was hurt. | Collapse, plant failure or loss of containment. | Escalate potential severity and screen the current legal definition. |
| Property or equipment damage | Damage to buildings, plant, vehicles, tools or materials. | Forklift bends a process barrier. | Make safe, preserve the damaged item and assess hidden risk. |
| Fire or explosion | Uncontrolled combustion, ignition, deflagration or explosion. | Small electrical cabinet fire. | Emergency control, competent technical escalation and scene preservation. |
| Environmental release | Actual or threatened harm to land, air, water or biodiversity. | Chemical enters a surface-water drain. | Contain, identify pathways and screen environmental notification duties. |
| Process-safety event | Loss or degradation of containment or control in a hazardous process. | Relief device lifts or isolation fails. | Assess barrier performance and major-event potential, not injury count alone. |
| Security event | Threat, intrusion, violence, theft, sabotage or information-security loss affecting safe operations. | Unauthorised entry to a restricted process area. | Protect people and evidence while using the defined security route. |
| Business interruption | Loss of production, service, supply, access or continuity. | A line stops for six hours after a collision. | Coordinate safe recovery without allowing production pressure to weaken controls. |
| Reputational or financial loss | Loss of trust, claims, penalties, direct cost or wider commercial harm. | Public concern after a visible release. | Use verified facts, controlled communication and never conceal safety evidence. |
Interactive Loss-Event Classification Coach
Choose the clearest first signal. The coach shows the linked categories and the first safe reporting priority.
The Needs a Reporting System Must Meet
| Need | Who needs the information? | Decision or action enabled | If the event stays hidden |
|---|---|---|---|
| Protect life and contain escalation | Workers, supervisor, emergency and occupational-health teams | First aid, evacuation, isolation, treatment, welfare and scene control | Exposure can continue and evidence may disappear. |
| Classify and escalate | Competent HSE, operations and senior leaders | Assess actual and credible potential consequence; set investigation level and interim controls | A “minor outcome” can conceal a major control failure. |
| Meet applicable duties | Named responsible person, regulator, insurer, client or permit authority | Screen current legal, contractual, environmental and sector-specific rules | Deadlines, evidence, compensation or formal duties may be missed. |
| Support affected people | Injured or exposed people, families, worker representatives and HR | Care, communication, rehabilitation, fair treatment and lawful record handling | People may lose trust, support or access to a valid claim. |
| Investigate and learn | Investigation team, technical specialists and workforce | Preserve evidence, test causes, identify effective actions and share learning | Similar work can repeat the same conditions. |
| Analyse performance | Managers, assurance functions and governance bodies | Identify clusters, recurring control weakness, response delay and action quality | Rates and trends become biased and falsely reassuring. |
Raise the alarm, obtain treatment and prevent escalation. Reporting never takes priority over immediate life safety.
Use the emergency route and notify the responsible supervisor or control point without avoidable delay.
Ask what happened and what could credibly have happened if one condition were different.
Give the event a unique identity; link people, witnesses, injuries and duplicate reports without double-counting.
Record time, place, work, people, equipment, substances, immediate controls, witnesses and evidence sources.
A competent person checks current legal, environmental, sector, insurer, client and permit requirements.
Use actual outcome, credible potential, recurrence likelihood, uncertainty and complexity to set the response.
Name owners, dates, interim safeguards and verification evidence.
Acknowledge the reporter, protect confidentiality, update affected people and share anonymised lessons.
Validate the dataset, assess repeated patterns and test whether action reduced risk.
Current Great Britain illustration—not a global decision rule
Under RIDDOR, the legally responsible person—not every witness—submits defined reports. Deaths, specified injuries, qualifying public injuries and dangerous occurrences generally require notification without delay and a report received within 10 days; qualifying over-seven-day incapacity is reported within 15 days; certain diagnosed occupational diseases are reported when the diagnosis is received.
Always check the current rule, definitions and jurisdiction. Environmental, transport, major-hazard, insurer and client routes are separate.
Design for work as actually done
Provide simple mobile, telephone and paper routes; plain language; accessibility and multilingual support; contractor access; confidential routes; non-retaliation; prompt acknowledgement; and visible follow-up.
A technically complete form that workers fear, cannot access or never hear back from is not an effective reporting system.
RIDDOR 2013 remains the current Great Britain reporting framework taught here. HSE’s consultation, which closed on 7 July 2026, proposed possible changes to definitions, dangerous occurrences, occupational diseases and the online form; a consultation proposal is not itself a change in law. Before a real decision, the responsible person must check the current HSE category, route and deadline.
Also keep the information routes separate: a RIDDOR submission does not automatically notify an insurer, client or environmental authority. Identifiable injury and health information requires restricted, lawful handling; detailed health records should not be circulated simply because operational teams need the event facts.
HSE 2026 RIDDOR consultation record—proposals, not current law · ICO guidance on sickness and injury records
Who Needs Which Information—and Why?
Good reporting does not send the entire file to everybody. It gives each authorised stakeholder the minimum reliable information needed for a defined decision, at the right time and through the right route. Urgent operational facts may travel quickly; identifiable health, witness or legal material may require much tighter access.
| Stakeholder | Information needed and why | Urgency and route | Confidentiality boundary | Decision enabled |
|---|---|---|---|---|
| Worker or affected person | Care, exposure, immediate risk, support, next update and how their information will be used. | Immediate face-to-face or emergency route; documented follow-up. | Respect dignity; restrict health and identifying detail. | Treatment, withdrawal, support and informed participation. |
| Supervisor or line manager | Facts, location, people at risk, controls taken and remaining danger. | Without avoidable delay by alarm, radio, phone or direct report. | Operational facts only unless more detail is necessary. | Stop work, isolate, resource response and escalate. |
| HSE function | Actual and potential loss, evidence, repeated signals, exposure and applicable routes. | Prompt internal system or emergency escalation. | Role-based access; separate operational and clinical detail. | Triage, investigation level, legal screen and learning. |
| Senior management or board | Material risk, critical-control failure, trend, assurance gap and action ownership. | Immediate for major risk; governed dashboard for trends. | Use aggregated or anonymised data unless identity is essential. | Resources, risk appetite response and governance challenge. |
| Worker representatives or trade unions | Relevant facts, worker impact, controls, investigation participation and action. | As soon as practicable through agreed consultation arrangements. | Share lawfully; anonymise personal detail where appropriate. | Represent workers, protect evidence and test practicality. |
| Occupational-health team | Exposure, symptoms, work history and relevant clinical information. | Prompt confidential referral appropriate to health risk. | Clinical detail stays within authorised health-data controls. | Assessment, support, fitness advice and health surveillance. |
| HR or legal function | Necessary employment, welfare, process, evidence and duty information. | According to seriousness and formal-process need. | Do not assume every investigation record is legally privileged. | Fair process, support, claims management and legal advice. |
| Insurer | Policy-defined notice, verified facts, loss estimate and preserved evidence. | Within policy conditions through the specified claims route. | Disclose only what is lawful, relevant and required. | Coverage, claim handling, expertise and loss control. |
| Client or principal contractor | Interface risk, affected work, immediate controls and contract-defined notice. | According to emergency and contractual arrangements. | Contractual access does not justify unnecessary health detail. | Coordinate site control, work interfaces and assurance. |
| Regulator | Information required by the current applicable legal framework. | By the legally defined route and deadline. | Use the official route; preserve the original record. | Regulatory oversight, enquiry and enforcement where appropriate. |
| Emergency services | Hazard, location, people, substances, access and live escalation risk. | Immediately through the emergency channel. | Life-safety need governs; avoid irrelevant personal detail. | Rescue, medical, fire, evacuation and area control. |
| Environmental authority | Substance, amount, pathway, receptor, controls and monitoring evidence. | According to the release and current environmental rule. | Protect investigation integrity without delaying required notice. | Containment, environmental protection and regulatory response. |
| Family or public, where appropriate | Accurate, humane and verified information about impact and public protection. | Through an authorised, coordinated communication lead. | Next-of-kin, privacy and investigation needs come before publicity. | Support, public protection and trustworthy communication. |
Interactive Stakeholder Information Coach
Which Reporting Route Fits the Situation?
| Route | Best use | Benefit | Limitation or risk | Good control |
|---|---|---|---|---|
| Verbal report | Immediate local warning or simple access for workers. | Fast and inclusive. | Can be forgotten, changed or untraceable. | Confirm important facts in the master record. |
| Telephone or emergency notification | Urgent off-site or specialist response. | Rapid two-way clarification. | Wrong number, incomplete note or no audit trail. | Use a defined call tree and log time, recipient and advice. |
| Paper form | Low-connectivity work or accessible local backup. | Simple and portable. | Delay, handwriting, loss and duplicate entry. | Unique ID, secure transfer and timely digital reconciliation. |
| Digital reporting system | Structured records, workflow, analysis and feedback. | Traceability, required fields and trend data. | Access, form fatigue, poor taxonomy or false precision. | Short first report, role-based access and usability testing. |
| Anonymous or confidential channel | Fear, sensitive conduct or serious trust barrier. | Can reveal otherwise hidden risk. | Harder clarification and possible misuse. | Explain the difference between anonymous and confidential; protect non-retaliation. |
| Toolbox talk or shift handover | Immediate shared operational awareness. | Reaches the workgroup and connects context. | Not a substitute for an event record or formal notification. | Record the event once and document safety-critical handover. |
| Statutory notification | A current legal reporting trigger. | Meets the formal route and enables independent oversight. | Narrow criteria can be mistaken for the whole internal system. | Competent screen, current official form, deadline and retained record. |
Interactive Reporting-Route Comparator
Leading and Lagging Information—Read Both
Leading information
Shows the conditions and activities that may influence future performance: hazard reports, critical-control failures, response time, reporting trust, investigation quality, overdue actions and verification results.
Use: intervene before serious loss. Limit: activity counts do not prove the activity was effective.
Lagging information
Shows outcomes that have already occurred: injuries, diagnosed disease, releases, fires, damage, lost time, cost and business interruption.
Use: understand loss experience and consequences. Limit: low counts can reflect luck, low exposure or under-reporting rather than strong control.
Positive, Adverse and Unintended Impacts
| Reporting feature | Potential positive impact | Possible adverse or unintended impact | How to manage the tension |
|---|---|---|---|
| Immediate escalation | Faster treatment, containment and evidence protection | Operational interruption and pressure on emergency resources | Use predefined thresholds, trained roles and proportional escalation. |
| Open near-miss reporting | Earlier warning and worker voice before severe harm | Recorded numbers may rise and be wrongly treated as poorer performance | Interpret volume with potential severity, trust, exposure, repeat rate and action closure. |
| Detailed data collection | Stronger investigation, patterns, claims and control decisions | Administrative burden, duplication and sensitive-data risk | Collect what is necessary, link duplicate records, validate fields and restrict access. |
| External notification | Compliance, independent oversight and wider learning | Scrutiny, cost, enforcement, claim or reputation concern | Use a competent legal screen; do not conceal an event to avoid consequences. |
| Fair accountability | Trust, learning and willingness to report | An organisation may fear that “no blame” means no standards | Separate honest error and reporting from deliberate or reckless conduct; decide fairly from evidence. |
| Performance indicators | Trend visibility, prioritisation and governance attention | Targets can drive under-reporting, reclassification or metric fixation | Balance lagging outcomes with reporting trust, critical-control, response and action-quality indicators. |
Fear and blame
Retaliation, ridicule or automatic discipline suppresses information. Reporting itself must never be treated as the offence.
Production pressure
If stopping work threatens targets or income, hazards and symptoms may remain invisible. Leaders must make safe reporting practicable.
No feedback
A silent system teaches workers that reporting changes nothing. Acknowledge, update, close and show what was improved.
Complexity and exclusion
Long forms, language barriers and contractor-only pathways distort the dataset. Offer accessible routes and one taxonomy.
Latency and privacy
Work-related disease may emerge slowly, while health data is sensitive. Enable delayed reports and use lawful, minimal, secure access.
Counting confusion
Report count, event count and affected-person count are different. Use a master event ID and state the unit being analysed.
Real evidence: when an environmental loss was not reported
In a published Environment Agency case, bleach reached a surface-water drain and killed more than 800 fish. The company did not report the incident because it wrongly assumed the drain led to the foul sewer. The case illustrates why drainage knowledge, rapid escalation and factual reporting matter; assumptions can enlarge both environmental and enforcement consequences.
Can You Trust the Reporting Picture?
Event data is not a transparent window into reality. What gets reported, how it is classified, who feels safe to speak and how duplicates are handled all shape the picture. Level 6 assessment therefore examines both the event information and the reporting system.
Under-reporting
Real events stay invisible because of fear, inconvenience, production pressure, uncertain definitions, inaccessible routes or lack of feedback. The dataset can look reassuring while exposure continues.
Over-reporting
Duplicate records, overly broad categories or repeated logging of the same condition inflate counts and workload. Do not discourage signals; link them to one master event and keep the original reports traceable.
Selective reporting
Some events, workers, contractors, shifts or sites appear while others disappear because incentives and access are unequal. Comparisons become systematically biased.
Visibility change
A new trusted route can increase reports even while risk control improves. Examine exposure, credible potential, repeat pathways, response quality and worker trust before declaring deterioration.
Nine Common Reporting Biases
Fear of discipline
People hide honest errors or hazards when reporting itself is treated as misconduct.
Blame culture
Labels such as “careless” replace factual reporting and discourage useful detail.
Production incentives
Bonuses, league tables or zero-event targets can reward silence or reclassification.
Contractor vulnerability
People may fear removal, lost work or commercial penalty if they report.
Normalisation of deviance
Repeated unsafe conditions become accepted as normal and no longer feel reportable.
Language and literacy
Complex forms exclude people or strip meaning from their account.
Severity bias
Only injuries are noticed while high-potential near misses and control failures disappear.
Duplicate reporting
Several accounts are counted as several events instead of linked evidence.
Delayed reporting
Memory, scene condition and electronic evidence degrade before the event becomes visible.
Eight Quality Tests for Reportable Data
Interactive Reporting-Distortion Diagnostic
Select the conditions that genuinely exist. The tool identifies how the dataset may be distorted and which system response is needed.
Organisational value
- Earlier control of exposure and escalation
- More reliable trends and resource priorities
- Better compliance, claims and insurance management
- Stronger trust, participation and organisational learning
- Reduced recurrence through visible, verified action
Adverse or unintended impact
- Privacy breach, premature blame or reputational harm
- Defensive reporting, fatigue and excessive bureaucracy
- Distorted statistics and unnecessary circulation
- Continued exposure, evidence loss and recurrence
- Possible enforcement, claim, contractual and trust consequences
The forklift–process-line event needs prompt reporting because a worker experienced symptoms, a chemical barrier was challenged and the credible potential was much greater than the contained outcome. The report enables care, scene control, evidence preservation, legal screening and investigation; it also supplies a signal about interface risk across similar routes. Those benefits must be weighed against downtime, administrative burden, staff anxiety and the risk of unnecessary disclosure of health information. The burdens are manageable through one master event record, factual language, role-based access, fair accountability and timely feedback. Non-reporting would leave the damaged barrier and route encroachment hidden, bias performance data and risk recurrence with more serious consequences. The justified judgement is therefore to report and escalate without avoidable delay, while limiting sensitive information and verifying that the resulting action reduces the risk.
How Healthy Is the Reporting System?
A high number of reports is not automatically evidence of poor safety. It can indicate stronger visibility and trust. Read volume alongside exposure, potential severity, recurrence, response quality, action closure and workforce evidence.
Practical case: when there is no “first 60 minutes”
Three technicians separately report numbness and tingling after months of using vibrating tools. A useful system must accept a gradual health signal without forcing a false single-event time or an unproven diagnosis. It should support the workers, record symptoms and work history, link related reports without disclosing unnecessary health details, review exposure and tool-maintenance evidence, involve competent occupational-health support and screen current reporting duties.
Assessment point: the need for visibility and prevention is strong, but identifiable health information requires restricted, lawful handling. Delay, fear or fragmented records can hide a group exposure; uncontrolled circulation can harm privacy and trust.
The First 60 Minutes After the Forklift–Process-Line Event
The release is small because the alarm and emergency isolation work. One contractor reports eye irritation; the forklift driver is shaken; liquid reaches a contained drainage sump; the route is closed for six hours. Previous near misses and barrier defects exist in separate informal messages.
| Time | Defensible action | Why it matters |
|---|---|---|
| 0–5 minutes | Alarm, withdrawal, emergency isolation, first aid and spill response | Care and containment come before form completion. |
| 5–15 minutes | Supervisor alert, head count, exposure and environmental checks, scene boundary | Actual harm and escalation routes may still be developing. |
| 15–30 minutes | Create master event ID; capture factual first reports; identify CCTV, alarm, plant and witness evidence | Early information is perishable, but conclusions must not be invented. |
| 30–60 minutes | Competent triage of potential severity; statutory/contract/insurer screen; appoint investigation lead and interim controls | A limited actual outcome does not make the failed vehicle–chemical interface low risk. |
Interactive Reporting-Route and Impact Assessor
This teaching tool recommends an internal response and legal-screen priority. It does not determine whether a statutory report is required.
Interactive Report-Quality Diagnostic
Select only what the proposed first report genuinely contains.
Interactive Loss-Event Report Builder
Interactive 3.3 “Assess” Response Planner
Interactive Report-to-Investigation Handover Check
A first report does not need to contain the final cause. It must give the investigation enough reliable starting information and protect what may disappear.
A first report records what is known; 3.4 tests the evidence to explain how and why the event became possible.
From Evidence to Action—and From Action to a Reliable System
Today’s three outcomes form one continuous journey. We first investigate the event to understand what happened and why. We then manage the organisational response from immediate control to verified closure. Finally, we use policy, responsibilities and records to make that good practice consistent every time.
“Do not treat investigation, incident management and policy as three separate subjects. Follow one event all the way through. Ask why it happened, decide what the organisation must do, and then build the system that makes the right response repeatable. Finding a cause is not the finish line—verified prevention is.” — Debjyoti Biswas, FIIRSM, CertIOSH
A Report Starts the Question—Investigation Tests the Answer
What Is an Incident Investigation?
It is
A planned examination of reliable physical, documentary, digital, organisational and human evidence. It reconstructs the event, tests alternative explanations, identifies immediate, underlying and root causes, and produces proportionate action.
It is not
A blame interview, an assumption that the last person caused everything, a report written to defend a predetermined conclusion, or a list of retraining actions unsupported by evidence.
Begin with evidence, not a predetermined explanation
The first story may be plausible, but a professional investigation tests it against physical facts, records, independent accounts and credible alternatives. The work is not complete when the report is signed; it must show that the risk has been reduced.
Use this discipline
Observe → preserve → corroborate → analyse → control → verify. Record uncertainty honestly. A finding can be useful without pretending that every detail is known.
Why Investigation Is Important
Prevention
Finds control weaknesses before the same pathway—or a related pathway—produces greater harm.
Understanding
Separates evidence from assumption and connects immediate events to job, organisational and system conditions.
Control assurance
Shows which barriers worked, failed, were missing or were never verified.
People and trust
Demonstrates concern, involves workers fairly and supports those affected when communication is honest and respectful.
Governance
Provides evidence for legal, insurance, leadership, worker-representation and resource decisions.
Organisational learning
Transfers lessons to similar assets, tasks, contractors, sites and design standards—not just the event location.
Set the Terms of Reference Before the Team Starts
Terms of reference means the written mandate that tells the investigation team what it is authorised and expected to do. It protects focus and fairness without fixing the answer in advance.
| Element | Question it must answer | Why it matters |
|---|---|---|
| Purpose and scope | What event, period, activities, interfaces and credible consequences are included? | Prevents an investigation that is either superficial or unmanageably broad. |
| Authority and independence | Who commissions the work, who may secure evidence, and is the lead sufficiently free from the decisions being examined? | Allows access and challenge while managing conflicts of interest. |
| Team and competence | Which operational, technical, HSE, human-factors and worker perspectives are required? | No single specialist sees the whole socio-technical system. |
| Evidence interfaces | How will evidence be protected, logged and coordinated with regulators, police, insurers or equipment specialists where relevant? | Avoids loss, contamination, unlawful access or conflict with another authority. |
| People, privacy and communication | How will affected people be supported, consulted and updated, and who may access identifiable health or witness information? | Protects dignity, trust, lawful handling and evidence quality. |
| Timetable and outputs | What interim controls, updates, report, actions, owners and effectiveness review are required? | Links investigation to timely risk reduction rather than report production alone. |
Who Does What? Investigation Governance and Roles
One person may hold more than one role in a small review, but authority, competence, independence and conflicts of interest must still be explicit.
Choose a Proportionate Investigation
| Level | Typical trigger | People and depth | Output |
|---|---|---|---|
| Local learning review | Low actual and credible potential, understood local issue, no sign of wider control failure | Competent supervisor with worker involvement; factual sequence and local control check | Recorded learning, proportionate action and closure evidence |
| Formal investigation | Medical treatment, material loss, repeat event, uncertain causes, important barrier failure or significant potential | Independent-enough multidisciplinary team with HSE and technical competence | Evidence pack, timeline, causal analysis, prioritised actions and management review |
| Major / specialist investigation | Fatality, life-changing harm, major release, multiple people, catastrophic potential, complex system or public/regulatory significance | Senior mandate, specialist and worker-representative input, controlled interfaces with authorities | Formal terms of reference, rigorous evidence governance, systemic learning and executive assurance |
Interactive Investigation-Level Selector
How Deep Should the Investigation Go?
The investigation level is a governance choice about authority, independence, competence and depth. Organisations use different names, but the six levels below help learners see the available range. The correct choice depends on the event and jurisdiction—not the label alone.
| Possible level | Typical purpose | Team and independence | Typical output |
|---|---|---|---|
| Supervisor-level investigation | Prompt learning from a low-potential, well-understood local event. | Competent supervisor with the people doing the work. | Facts, local causes, action, owner and closure evidence. |
| Preliminary investigation | Establish enough evidence to classify risk, preserve material and decide whether a fuller investigation is needed. | Supervisor or HSE lead with suitable technical support. | Initial sequence, potential, evidence map, interim controls and escalation recommendation. |
| Full internal investigation | Examine a significant, repeated, uncertain or multi-factor organisational event. | Multidisciplinary team sufficiently independent from the decisions examined. | Terms of reference, evidence pack, causal and barrier analysis, approved actions and verification plan. |
| Independent investigation | Provide additional credibility or specialist challenge after severe, sensitive or governance-significant loss. | External or organisationally separate lead with no material conflict. | Independent findings, systemic recommendations and assurance to senior governance. |
| Joint employer–worker investigation | Combine duty-holder knowledge with workforce experience and representation. | Employer and worker representatives with agreed access, competence and evidence safeguards. | Shared fact-finding, practical recommendations and stronger workforce learning. |
| Regulatory or criminal investigation | Establish facts under statutory authority and determine regulatory or criminal matters. | Competent authority using its legal powers; organisational teams must not obstruct it. | Official evidence, decisions and possible enforcement or prosecution process. |
Eight Selection Questions
What Competence Does the Team Need?
Technical knowledge
Understands the plant, substances, task, design limits and relevant control standards.
Investigation method
Can scope, plan, reconstruct, analyse and communicate without forcing a preferred answer.
Interviewing skill
Uses open, non-leading questions and supports affected people fairly.
Evidence management
Protects identity, integrity, source, access, transfer and uncertainty.
Legal awareness
Recognises when specialist advice or coordination with an authority is needed.
Independence
Declares conflicts and can challenge decisions without improper pressure.
Worker participation
Brings work-as-done knowledge and tests whether recommendations are usable.
Team integration
Combines disciplines, resolves evidence conflict and communicates limitations.
The Incident-Investigation Lifecycle
Protect people and environment, prevent escalation and disturb the scene only for safety, rescue or essential control.
Set a competent, sufficiently independent team; define authority, boundaries, interfaces, timescale and communication.
Identify what is perishable, who must be consulted, what expertise is required and how evidence integrity will be recorded.
Secure photographs, measurements, parts, records, electronic data, conditions and fair accounts from involved people.
Place verified events on a timeline; label fact, inference, conflict, uncertainty and missing evidence.
Test immediate, underlying and root causes; successful, failed and absent barriers; human and organisational influences.
Ask what evidence would be expected if another explanation were true. Do not stop at the first plausible story.
Use the hierarchy, address system causes, define interim safeguards and avoid over-reliance on warnings or retraining.
Explain evidence, reasoning, uncertainty, lessons, actions, owners and dates; protect personal information.
Track completion, test whether the action works in practice and check equivalent risks across the organisation.
How Does the Full Investigation Process Connect?
HSE’s HSG245 gives a clear four-stage learning structure. The more detailed lifecycle below does not replace it; it shows where the practical governance, evidence, communication and verification activities sit.
| HSG245 stage | Plain-language question | Practical activities in this portal | Output |
|---|---|---|---|
| 1. Gather information | What happened, in what conditions, and what evidence exists? | Immediate response, scene control, level selection, team appointment, terms of reference, evidence plan and collection. | Protected evidence, factual sequence, gaps and uncertainties. |
| 2. Analyse the information | How and why did the event become possible? | Timeline reconstruction; causal, barrier, human-factors and organisational analysis; alternative explanations. | Supported immediate, underlying and root/system findings. |
| 3. Identify suitable risk-control measures | What change will address the demonstrated causes? | Generate options, apply the hierarchy, consider similar exposure and define success measures. | Prioritised interim and permanent controls linked to findings. |
| 4. Develop and implement an action plan | Who will do what, by when, and how will we know it worked? | Approval, communication, ownership, resources, implementation, effectiveness verification and closure. | Completed and verified risk reduction with transferred learning. |
The Fifteen-Stage Route—from Notification to Closure
Care for people, raise alarms, contain escalation and preserve life before evidence.
Set a safe boundary, record necessary changes and prevent avoidable loss or contamination.
Use actual harm, credible potential, recurrence, significance, complexity and learning value.
Match technical, method, people, legal and evidence competence; declare conflicts.
Set purpose, authority, scope, interfaces, outputs, communication and timetable without fixing the answer.
Prioritise perishable sources, authorisations, specialist tests, access and storage.
Secure people, place, plant, records and process evidence with traceable identity and condition.
Arrange verified events from the last normal point; show conflicts, gaps and uncertainty.
Select suitable methods; examine immediate, underlying, systemic and recovery factors.
Link each action to an established cause, use the hierarchy and define effectiveness evidence.
Test evidence, reasoning, fairness, limitations, actions and governance acceptance.
Support affected people, brief authorised stakeholders and share usable, privacy-protected learning.
Resource named owners, manage dependencies and retain interim controls until permanent change is safe.
Observe real work, measure control performance, consult users and check for unintended risk.
Independent-enough evidence confirms the action works and equivalent operations have been considered.
Gather, Protect and Test the Evidence
| Evidence family | Examples in the master case | Quality question |
|---|---|---|
| People | Driver, contractor, operators, supervisor, maintenance planner, dispatcher and worker representative | Was the account gathered promptly, respectfully and without leading or blame? |
| Position and place | Route width, pallet positions, lighting, sight lines, floor marks, spill boundary and weather/shift conditions | Was the scene photographed and measured before avoidable change? |
| Plant, parts and substances | Forklift, low barrier, valve manifold, damaged fixings, chemical and alarm/isolation equipment | Is identity, condition and chain of custody traceable? |
| Paper and electronic records | Risk assessment, layout, training, inspections, defect tickets, change records, CCTV, alarms, radio and maintenance history | Is the record authentic, complete and understood in context? |
| Performance and organisational data | Near-miss trends, dispatch volume, overtime, backlog, action closure, contractor interfaces and previous learning | Are definitions, periods, denominators and possible under-reporting understood? |
The Minimum Evidence Log
Record any unavoidable scene change, copy, conversion, repair, test or transfer. “Chain of custody” means the documented history of who controlled an item and what happened to it; the level of formality should match the event and any regulator, police, legal or insurance interface.
Four Tests for a Defensible Finding
Does the evidence help explain the sequence, control performance or conditions that made the event possible—or is it merely interesting background?
Is the conclusion supported by more than one reliable source, such as measurement plus records or an account plus digital evidence?
What other explanation could fit, and what supporting or conflicting evidence would be expected if it were true?
What is known, inferred, disputed or missing, and how strongly can the team state the conclusion?
Interactive Finding-Quality Coach
Interview for learning
Explain purpose and support; choose a suitable place; ask open questions; invite the person’s own sequence; distinguish observation from inference; explore what made sense at the time; confirm understanding; and allow correction.
Avoid contamination
Do not ask “Why were you careless?” or circulate a shared story before accounts are captured. A better prompt is: “Please take me through what you saw, heard and did, from the last normal point.”
Interactive Evidence-Preservation Diagnostic
What Evidence Should the Team Seek?
The “5 Ps” Evidence Check
The five Ps are a DB HSE learning aid for remembering evidence families. HSG245 discusses physical, verbal and written evidence but does not formally prescribe this mnemonic.
Evidence Reliability Ladder—Use with Caution
| Evidence type | Typical strength | Question before relying on it |
|---|---|---|
| Contemporaneous physical evidence | Direct condition, position, damage or residue close to event time. | Was the scene changed, item identified and condition protected? |
| Digital or system record | Time-stamped CCTV, alarm, historian, access or communication data. | Are clocks synchronised, data complete and export authentic? |
| Controlled document | Approved design, procedure, inspection or maintenance record. | Was it current, available and actually used in the work? |
| Measurement or test | Quantified distance, force, exposure, condition or performance. | Was the method suitable, calibrated, repeatable and representative? |
| Witness recollection | Explains perception, sequence, context and why actions made sense. | How soon was it gathered, was the question neutral, and what corroborates it? |
| Opinion or assumption | Can generate a hypothesis or identify specialist questions. | What evidence would support or disprove it? Never relabel it as fact. |
Interactive Evidence-Source Coach
From Evidence to Causes, Controls and Action
| Level | Evidence-based finding | Control implication |
|---|---|---|
| Immediate | The forklift entered the narrowed route and contacted a low barrier that did not prevent contact with the manifold. | Restore control of the route and equipment integrity before restart. |
| Underlying job factors | Pallet storage reduced clearance; peak dispatch, visibility and radio interruption shaped the task; temporary adaptations had become normal. | Redesign storage and traffic flow; review capacity, communication and work scheduling. |
| Underlying organisational factors | Repeated barrier defects were recorded without effective risk escalation, ownership or closure verification. | Repair the defect-management, prioritisation and assurance process across comparable barriers. |
| Root / systemic | Warehouse, production and maintenance changes were managed separately, so the combined vehicle–chemical interface was not owned as one critical risk. | Create accountable interface ownership, change control and critical-control performance standards. |
| Successful recovery | Detection, alarm, withdrawal and emergency isolation limited the actual release. | Preserve and verify these controls; do not let success hide the preventive-control failure. |
Barrier Performance Review—Link Back to 3.1
Use Bowtie and Swiss Cheese thinking without simply redrawing the models. For each barrier, ask what it was meant to do, whether a demand occurred, what evidence shows performance and why any weakness developed or persisted.
| Barrier | Intended function | Evidence and status | Cause-linked action | Verification |
|---|---|---|---|---|
| Traffic-route clearance | Keep vehicles separated from process equipment | Pallets reduced usable width; barrier failed before contact | Redesign storage and physically control the clear zone | Peak-shift measurement, observation and encroachment trend |
| Vehicle restraint barrier | Prevent vehicle reach to the manifold | Low, damaged barrier did not resist the demand | Engineer a rated segregation system and protect foundations | Design acceptance, inspection and controlled impact criteria |
| Defect escalation | Prioritise and close safety-critical damage | Repeat defects existed without verified risk ownership | Set criticality, owner, deadline, interim control and escalation | Backlog audit and sample field verification of closed defects |
| Alarm and isolation | Limit consequence after loss of containment | Worked and restricted the actual release | Preserve, maintain and test the successful recovery controls | Function test, response-time evidence and drill observation |
Interactive Cause-to-Control Builder
Which Investigation Method Fits the Evidence?
A method is a disciplined way to organise questions; it is not a machine that automatically discovers “the root cause.” Use the simplest method that can represent the event fairly, and combine techniques when different questions require different views.
| Method | Best question or use | Strength | Limitation |
|---|---|---|---|
| Five Whys | Why did a relatively simple, evidence-rich problem persist? | Fast and accessible; pushes beyond the immediate event. | One questioning path can oversimplify, reflect facilitator bias or stop too early. |
| Fishbone / Ishikawa | Which categories of contributing factors should the team explore? | Encourages broad multidisciplinary brainstorming. | A list of possibilities does not establish sequence, evidence or causation. |
| Timeline analysis | What happened before, during and after the event? | Makes gaps, conflicts and parallel actions visible. | Chronology alone does not explain why conditions existed. |
| Change analysis | What differed from an earlier safe state, design or expectation? | Useful after modification, drift, staffing or process change. | Can miss long-standing weaknesses that did not recently change. |
| Barrier analysis | Which prevention, control and recovery barriers worked, failed or were missing? | Connects evidence directly to risk-control performance. | Weak barrier definitions or performance standards produce vague conclusions. |
| Fault Tree Analysis (FTA) | Which combinations of failures could produce a defined top event? | Shows AND/OR logic and can support probability analysis. | Depends on scope, independence assumptions and defensible input data. |
| Event Tree Analysis (ETA) | What outcomes can follow an initiating event as barriers succeed or fail? | Shows consequence pathways and recovery importance. | Branch dependence and omitted pathways can mislead. |
| Bowtie | How do threats, a top event, consequences and preventive/recovery barriers connect? | Clear control-focused communication across disciplines. | Can become a static picture and hide dynamic interactions or barrier degradation. |
| Tripod Beta | Which failed barriers, preconditions and organisational influences shaped the event? | Structured connection from active failures to latent conditions. | Needs trained use and can be disproportionate for a simple local event. |
| AcciMap | How did decisions and interactions across regulators, organisation, management and work contribute? | Useful for complex multi-level systems and public interfaces. | Resource intensive; broad maps still need evidence-based prioritisation. |
| STAMP | How did inadequate constraints and feedback across a complex sociotechnical control system allow loss? | Represents software, organisations, adaptation and non-linear interaction. | Requires specialist competence and careful boundary definition; often excessive for simple events. |
Interactive Investigation-Method Selector
Human Factors—Why Did the Action Make Sense at the Time?
Human-factors analysis examines the interaction between people, task, equipment, environment and organisation. It does not excuse unsafe conduct; it avoids the weak assumption that naming an error explains the system that produced it.
Fatigue and workload
Hours, rest, pace, task switching, physical demand and cumulative workload.
Competence and supervision
Knowledge, practice, authorisation, coaching, availability and span of control.
Communication
Handover, language, radio, alarm meaning, coordination and shared situational awareness.
Usability and design
Controls, displays, access, visibility, procedure usability and error-tolerant design.
Time and production pressure
Targets, staffing, incentive, scheduling, backlog and conflicting priorities.
Culture and adaptation
What leaders reward, what is tolerated, how deviations become normal and whether concerns are acted on.
Just Culture—Fair Accountability Without Automatic Blame
Human error
An unintended slip, lapse or mistake. Console and support the person, then improve design, conditions, barriers and learning.
At-risk behaviour
A choice where risk is underestimated, normalised or perceived as justified. Understand incentives and context; coach and redesign the system.
Reckless behaviour
A conscious, unjustifiable disregard of a substantial and obvious risk. A fair accountability process may be appropriate after evidence, due process and system context are considered.
Fair accountability
Apply consistent criteria, separate investigation from predetermined discipline, consider capacity and intent, protect reporting and avoid outcome bias.
Interactive Just-Culture Reflection Coach
Legal and Ethical Safeguards
What Changes When Investigation Is Done Well—or Badly?
| Area | Strong investigation impact | Weak or blame-led impact |
|---|---|---|
| Risk | System weaknesses and equivalent exposures are controlled. | Superficial causes remain and recurrence becomes likely. |
| Workers | People are supported, heard and more willing to share evidence. | Fear, distress, silence and adversarial accounts increase. |
| Decisions | Resources target higher-order controls with owners and verification. | Action lists favour reminders, retraining and paperwork without risk reduction. |
| Evidence and compliance | Reasoning is traceable and reporting, claims and governance use reliable facts. | Delay, contamination, inconsistency or concealment undermines confidence. |
| Learning culture | Successful and failed barriers are shared across similar work. | The report is filed, lessons stay local and organisational memory decays. |
| Business | Recurrence, downtime, loss cost and uncertainty can reduce over time. | Investigation consumes time but returns little value; repeat loss enlarges cost and reputation damage. |
A Defensible Investigation Report Structure
Example 30–60–90 Day Verification Plan
0–30 days · Stabilise
Confirm interim controls, protect people, assign owners, preserve evidence and stop uncontrolled recurrence while permanent action is designed.
31–60 days · Implement
Introduce engineering and system changes, consult users, update linked documents and train only where competence genuinely forms part of the cause and control.
61–90 days · Verify
Observe real work, inspect barrier performance, test worker understanding, review repeat signals and confirm equivalent sites or tasks received the learning.
The dates are a teaching example, not a universal deadline. Urgent controls must not wait, and complex engineering work may require a longer governed plan. Verification criteria should be decided when the action is agreed—not invented at closure.
Interactive Corrective-Action Quality Coach
Interactive Investigation-Quality Diagnostic
Interactive 3.4 “Explain” Response Builder
Before 3.4 is complete, test whether the recommended action was implemented, works in real conditions and transfers to equivalent risks.
When Is Corrective Action Truly Closed?
Eight Quality Tests for Every Recommendation
Interactive Recommendation-Quality Gate
Select only the qualities that the proposed action genuinely demonstrates.
Completed Example—from First Notification to Verified Closure
| Stage | What the team did | Evidence or decision produced |
|---|---|---|
| 1. Initial notification | The operator raised the alarm after forklift contact beside manifold V-12; first aid, withdrawal and isolation followed. | Time, place, people, actual symptoms, credible potential and immediate controls entered under one event ID. |
| 2. Scene and evidence | The area was cordoned; barrier position, pallet layout, marks and route width were photographed and measured; CCTV and alarm data were secured. | Traceable physical, digital, documentary and human evidence with scene changes recorded. |
| 3. Level and team | High credible chemical-release potential, repeated barrier defects and cross-department interfaces justified a full internal investigation. | Multidisciplinary team with worker representation, technical skill and independent review. |
| 4. Reconstruction | The team synchronised CCTV, alarm and dispatch records with separate accounts from the driver, contractor and operators. | A tested timeline from the last normal point, including gaps and uncertainty. |
| 5. Analysis | Timeline and barrier analysis showed route encroachment, degraded physical protection, weak defect escalation and divided interface ownership. | Immediate, underlying and systemic findings; alarm and isolation recorded as successful recovery controls. |
| 6. Recommendations | The team selected rated segregation, controlled storage, critical-defect escalation and accountable interface change control. | Cause-linked actions using stronger controls rather than reminders alone. |
| 7. Approval and communication | Evidence, limits, owners, dates and learning were reviewed; affected people and the workforce received privacy-protected updates. | Approved report, action plan, interim controls and feedback record. |
| 8. Implementation | Engineered segregation and a clear pedestrian route were installed; storage limits and defect workflow were changed. | Design acceptance, installation evidence, updated controls and owner sign-off. |
| 9. Effectiveness verification | An independent-enough reviewer measured clearances, observed peak work, sampled defect closure and consulted users over three months. | Barrier performance, no repeated encroachment, timely critical-defect escalation and no harmful workaround. |
| 10. Closure and transfer | Comparable process routes and contractor interfaces were screened; lessons and standards were transferred. | Formal closure based on verified risk reduction—not report signature alone. |
Incident investigation is important because it converts a reported event into tested knowledge and preventive action. In the forklift case, early scene control and evidence preservation protected measurements, CCTV, alarm data and people’s accounts before they changed. A multidisciplinary team then reconstructed the sequence and used barrier analysis to connect the contact not only to route encroachment, but also to damaged protection, weak defect escalation and divided ownership of the vehicle–process interface. This matters because action based only on the driver’s final movement would leave the organisational conditions intact. Engineering segregation, controlled storage and accountable critical-defect management were therefore linked to established causes. Their impact became credible only when implementation and field effectiveness were verified across real peak-shift work and similar routes. A strong investigation can reduce recurrence, improve trust and direct resources to system controls; a weak or blame-led investigation can suppress future reports, preserve failed barriers and allow a more serious event to recur.
Evidence was tested, causes were connected to action and effectiveness was verified. Bring the complete reasoning into one Level 6 response.
How to Answer 3.1–3.4 at Level 6
3.1 — Outline
For each theory or technique: state its name and origin/context, describe its principal structure, explain its direction or logic, apply it briefly to an incident, and state one useful feature and limitation.
Suggested structure: Name → main idea → components → application → usefulness → limitation.
3.2 — Justify
Identify the decision, explain why the selected quantitative method fits, show the data and formula, interpret the result, discuss reliability and limitations, combine it with qualitative evidence, and state the resulting action and review.
Suggested structure: Decision → method → evidence → calculation → interpretation → limitation → complementary evidence → action.
3.3 — Assess
Weigh the reporting need, affected stakeholders, urgency, route and benefit against burden, privacy, trust and other adverse effects. Examine what non-reporting would cause, show how tensions can be controlled, and reach an evidence-based judgement.
Suggested structure: Trigger → need → route → benefit → adverse impact → mitigation → consequence of silence → judgement.
3.4 — Explain
State what investigation is, show how evidence becomes causal understanding and control, explain why each stage matters, apply it to a realistic event, and connect strong or weak investigation quality to people, risk and organisational outcomes.
Suggested structure: Meaning → process → reason → workplace application → positive impact → consequence if weak → verification.
Using the forklift–process-line case: outline the causation theories and techniques; justify suitable quantitative methods; assess the need and impact of internal and possible external reporting; and explain how a proportionate, fair investigation would convert evidence into verified risk reduction.
Can You Connect Models, Data and Investigation?
Understand Processes and Strategies to Manage Health and Safety Incidents in an Organisation
Section 04 moves from understanding individual loss events to managing and assuring the complete organisational response. Learners follow an incident from readiness and first response through evidence, investigation, corrective action, lawful record maintenance, safe recovery, evaluation and verified organisational learning.
Learning outcomes covered: 4.1 Outline the critical stages for managing incidents in the organisation. 4.2 Outline organisational policies to identify, investigate, report and record incidents. 4.3 Explain how to maintain records of incidents to meet regulatory and statutory requirements. 4.4 Evaluate an organisational process for managing health and safety incidents.
How 3.1–3.4 and 4.1–4.4 Work Together
A mature incident system needs every part. Causation models organise possible explanations; loss data tests patterns; reporting makes events visible; investigation turns evidence into findings; incident management coordinates the response; policy makes good practice repeatable; records preserve defensible proof; and evaluation determines whether the system really works.
EXPLAIN
EVALUATE
4.1–4.2 outline: identify the principal stages or policy features and show what each involves. 4.3 explain: connect each recordkeeping method to the legal or regulatory requirement it satisfies and the consequence of failure. 4.4 evaluate: compare evidence against a benchmark, make a balanced judgement and propose prioritised, verifiable improvement.
What Is Incident Management?
Master Case: Forklift, Racking and Chemical Release
A reversing forklift strikes warehouse racking and damages a cleaning-chemical container. One employee experiences eye irritation, a contractor narrowly avoids falling material, liquid moves towards a drain, CCTV is available, and management wants the area reopened quickly. The employee later reports skin symptoms.
An Organisation Cannot Invent Its Incident System During the Emergency
Foresee credible events
Use risk assessments, emergency planning, loss history and specialist advice to consider fires, releases, vehicle events, violence, structural failure, occupational-health events and other credible scenarios.
Prepare competent roles
Define incident control, first aid, evacuation, technical isolation, occupational health, communications and regulatory-notification authority.
Provide resources
Maintain alarms, communications, emergency equipment, spill control, access information, contact lists, evidence kits and alternative work arrangements.
Train, exercise and review
Practise arrangements, record learning, correct gaps and update plans after change. A paper plan that nobody can use is not readiness.
Four Phases · Ten Critical Stages
Respond and stabilise
Recognise, raise the alarm, protect life, establish control and prevent escalation.
Preserve and report
Protect the scene, capture initial facts, classify the event and activate required routes.
Investigate and improve
Gather and analyse evidence, select controls and assign governed corrective actions.
Recover and learn
Authorise safe restart, support people, verify effectiveness, share learning and close.
Select a Stage to See Its Purpose, Owner and Output
Stage 1 · Recognise and activate
Purpose: make the event visible quickly enough for a proportionate response.
- Main actions: stop unsafe work, raise the alarm, identify location and immediate danger.
- Typical owner: first person aware, supervisor or control point.
- Output: verified alert and activated response route.
First 15 Minutes: Choose Immediate Actions and Investigation Depth
First 15 Minutes Simulator
Actual vs Credible-Potential Triage
This learning tool selects organisational investigation depth. It does not decide legal reportability.
Can We Restart Safely? Preserve Evidence and Test Readiness
Scene-Preservation Decision Coach
Safe-Restart Gate
Select only what has been demonstrated with evidence.
4.2 now defines the policies, responsibilities, mandatory rules and controlled records that make the sequence reliable across shifts, departments, sites and contractors.
Policy, Procedure, Plan, Form and Record Are Not the Same
Document Classifier
Incident Reporting, Investigation and Organisational Learning Policy
An organisation may use one umbrella policy rather than four artificial policies. The essential test is whether the controlled system clearly covers identification, investigation, reporting and recording.
Organisational Incident Policy Suite
OTHM asks learners to outline organisational policies used to identify, investigate, report and record incidents. A defensible organisation can meet those functions through one controlled umbrella policy supported by specialist policies and procedures. The eight-policy suite below is a practical learning model; OTHM does not prescribe these exact document titles.
Incident Identification and Classification Policy
Defines which events enter the organisational incident system and how their actual and credible potential consequences are classified.
Which Policy Owns the Problem?
Select a realistic failure. The coach identifies the primary policy and the supporting links needed to prevent a gap between documents.
Four Policy Pillars
01Identify
The policy explains what must enter the incident system.
- Injury and occupational ill health
- Near miss and dangerous occurrence
- Property, environmental and operational loss
- Delayed symptoms and diagnoses
- High-potential and repeated control failures
- Signals from alarms, inspections, monitoring and complaints
02Investigate
The policy defines how learning will be obtained fairly and proportionately.
- Investigation levels and escalation criteria
- Competent, authorised and sufficiently objective team
- Worker and representative involvement
- Terms of reference, evidence and time expectations
- Immediate, underlying and root/systemic causes
- Cause-linked action and effectiveness verification
03Report
The policy defines information routes and time expectations.
- Who reports, what, to whom and how
- Immediate emergency escalation
- Alternative route if normal management is involved
- Good-faith reporting without retaliation
- Authorised external notification
- Feedback to reporters and affected workers
04Record
The policy defines trustworthy, retrievable and protected evidence.
- Unique event ID and contemporaneous facts
- Actual and potential consequences
- Notifications, evidence and investigation links
- Decisions, actions, owners and deadlines
- Effectiveness and closure evidence
- Access, retention, version control and secure disposal
Who Is Responsible for What?
| Role | Principal responsibility | Evidence or decision |
|---|---|---|
| Worker or witness | Protect immediate safety and report promptly through an accessible route. | Initial factual signal. |
| Supervisor | Activate response, make the area safe, receive the report and escalate. | Initial event record and controls. |
| Incident controller | Coordinate priorities, resources, communications and external emergency interface. | Incident log and controlled response. |
| Competent H&S person | Triage actual and potential severity, advise on investigation level and screen legal duties. | Classification and escalation decision. |
| Responsible statutory reporter | Submit any required external report on behalf of the duty holder. | Notification reference and retained copy. |
| Investigation lead/team | Preserve and test evidence, analyse causes and propose controls. | Findings and recommendations. |
| Worker representative | Contribute workforce knowledge and support fair learning. | Consultation and challenge. |
| Occupational health / HR | Support people and control sensitive health and welfare information. | Restricted health/welfare record. |
| Action owner | Implement assigned action and provide completion evidence. | Implementation record. |
| Authorised senior manager | Resource serious cases, decide restart/closure and review organisational implications. | Risk acceptance, restart and closure approval. |
Who Owns the Next Decision? Responsibility and Handover Board
Incident management can fail between stages even when each person performs one task well. Select both the accountable role and the evidence or decision that must be handed forward.
| Critical handover | Accountable decision owner | Required output or evidence | Status |
|---|---|---|---|
| Alarm and initial escalation | Not checked | ||
| Command and stabilisation | Not checked | ||
| Investigation mobilisation | Not checked | ||
| Required external notification | Not checked | ||
| Safe restart authorisation | Not checked | ||
| Effectiveness verification and closure | Not checked |
- Alarm
- Control
- Investigation
- Notification
- Restart
- Closure
Policy X-Ray, Reporting Route and Record-Quality Tools
Test whether the policy is complete, usable, accountable and capable of producing trustworthy evidence.
Policy X-Ray
Select only the provisions actually present in the organisation's policy.
Internal Report vs External Notification
Record-Quality Coach
Umbrella Policy Builder
From Finding to Prevention: Learning-Loop Mapper
A 3.4 finding becomes useful only when 4.1 converts it into the right organisational action and 4.2 governs that action through policy, responsibility and a required record. Section 4.3 then maintains that record as defensible evidence, while 4.4 evaluates whether the complete chain delivered effective prevention.
What Makes an Incident Policy Work in Practice?
Fair reporting culture
Protect good-faith reporting, avoid premature blame, explain how information will be used and provide feedback. Accountability decisions may still occur through a separate fair process.
Worker participation
Workers and representatives provide work-as-done knowledge, identify practical barriers and help test whether corrective actions are usable.
Privacy by design
Collect only necessary personal information, separate detailed health records, apply role-based access, use anonymised learning where identity is unnecessary and dispose securely under the retention schedule.
Management assurance
Monitor delayed actions, repeat events, reporting routes, investigation quality, effectiveness checks, competence and policy-review triggers—not only injury totals.
Policy Sets the Rule—Maintained Records Prove What Happened
From Incident Records to Organisational Assurance
Debjyoti Biswas, FIIRSM, CertIOSHFellow Member of IIRSM · CEO & Founder, DB HSE INTERNATIONAL
“An incident record is not simply paperwork. It is legal evidence, organisational memory and the foundation for proving whether lessons have genuinely prevented recurrence. Today, we move from maintaining defensible records to evaluating whether the complete incident-management process actually works.”
Incident Records Are Evidence, Memory and Accountability
A record may support immediate control, statutory reporting, worker welfare, investigation, enforcement, claims, trend analysis, corrective-action tracking and organisational learning. One purpose must not destroy another—for example, a learning summary can be anonymised while the controlled legal record remains complete.
What Must Stay Connected in the Incident File?
Create → Verify → Protect → Use → Retain → Dispose
Build a Legal Recordkeeping Profile for Every Jurisdiction
A global policy should not hard-code one country’s rules for every site. Each organisation needs a current legal register or jurisdiction profile that converts applicable requirements into an operational record rule.
| Profile field | Question the organisation must answer | Control evidence |
|---|---|---|
| Authority | Which legislation, regulation, licence, regulator, court or sector rule applies? | Current legal register, version/date and competent review. |
| Duty holder | Who must create, notify, sign, keep or produce the record? | Named accountable role and authorised deputy. |
| Trigger and time | Which event, diagnosis, incapacity or consequence activates the duty—and by when? | Classification decision, deadline control and escalation log. |
| Required content | Which facts, identities, injury/diagnosis details, circumstances or submission references are mandatory? | Controlled form fields and completeness check. |
| Place, format and access | Where must the record be kept, in what readable format, and who may inspect it? | Repository, role permissions, retrieval test and disclosure log. |
| Retention and disposal | What is the minimum period, when does it start, what extends it, and how is disposal authorised? | Retention schedule, legal-hold process and destruction certificate. |
Great Britain worked example—use current rules, not memory
Reportable under RIDDOR
The responsible person screens the event against current RIDDOR categories and timing. The external report and acknowledgement are linked to the incident file; completing the investigation is not a reason to miss the notification deadline.
Record required under RIDDOR 2013 Regulation 12
Records cover reportable deaths, injuries, dangerous occurrences and occupational diseases, plus over-three-consecutive-day worker incapacity even though that latter category alone is not externally reportable. Required particulars must be kept for at least three years from the date the record was made.
Recordable is wider than reportable
The accident book, internal policy, insurer, client, environmental or sector rules may require records when RIDDOR notification does not. “Not RIDDOR-reportable” never means “delete the event.”
Current-law control
HSG245 remains useful investigation guidance, but it was published in 2004 and contains historic RIDDOR references. Use current HSE and legislation pages for legal decisions, and use the correct separate regime for Northern Ireland.
Integrity, Privacy, Retention and Legal Holds
Eight tests of a defensible record
- Accurate: facts reflect the best available evidence.
- Complete: mandatory fields, attachments and decisions are present.
- Timely: created and updated without avoidable delay.
- Objective: observation is separated from inference and blame.
- Traceable: source, author, date/time and event ID are clear.
- Controlled: access, versions, copies and custody are governed.
- Retrievable: readable records can be found throughout retention.
- Auditable: changes, disclosures, approvals and disposal leave evidence.
Correction without destroying history
Preserve the original entry; add the corrected information; state why it changed; identify the person making and approving the change; time-stamp it; and notify affected decision makers when the correction changes classification, notification, welfare or action.
Privacy and worker health data
In UK data-protection terms, injury and health information is special-category data. Identify a lawful basis and special-category condition; collect only what is necessary; separate detailed medical information; restrict access; encrypt or lock it; and use anonymised learning where identity is unnecessary.
Retention is a reasoned decision
Do not invent one universal period such as “keep everything for six years.” Apply the specific legal minimum, then test longer sector, exposure, insurance, contractual, claim-limitation, safeguarding and litigation-hold needs. Retain no longer than justified once every duty and hold has ended.
Make the Recordkeeping Decisions
1Record, Report or Notify?
2Legal Recordkeeping Profile Builder
3Incident Record Integrity Test
4Retention and Access Decision
5Audit-Trail Correction Challenge
Records Are the Evidence—Evaluation Is the Judgement
Does the Process Work in Design, Delivery and Results?
Evaluate Every Link from Scene Preservation to Shared Learning
| HSG245-linked stage | Evaluation question | Examples of evidence | Possible process failure |
|---|---|---|---|
| Preserve the scene | Did urgent safety action occur without avoidable destruction or loss of evidence? | Scene log, photographs, isolation record, CCTV copy, disturbance log. | Evidence overwritten, moved without record or inaccessible. |
| Note people and equipment | Can the organisation identify exposure, witnesses, assets and relevant operating conditions? | People/equipment register, shift list, contractor record, asset history. | Contractors omitted; equipment state or health exposure disconnected. |
| Report the event | Was the event visible quickly to everyone who had to protect, decide or notify? | Time-stamped report, acknowledgement, escalation and notification logs. | Delay, suppression, wrong recipient or missed external duty. |
| Decide investigation need | Did actual and credible potential risk determine proportionate depth? | Triage score, classification evidence, scope and team appointment. | Minor actual injury hides high-potential systemic failure. |
| Gather information | Was evidence sufficient, reliable, diverse and traceable? | 5 Ps evidence map, interviews, documents, digital data, health evidence. | Single narrative, missing source, bias or no work-as-done evidence. |
| Analyse information | Did analysis test barriers and immediate, underlying and root/systemic causes? | Timeline, barrier analysis, change analysis, cause logic and uncertainties. | “Human error” becomes the stopping point; unsupported root cause. |
| Identify controls | Do proposed controls address supported causes at an appropriate hierarchy level? | Option appraisal, risk assessment, hierarchy and human-factors review. | Training and reminders substitute for engineering or system change. |
| Implement action plan | Are actions resourced, owned, prioritised, tracked and verified? | Owners, due dates, interim controls, completion and effectiveness evidence. | Action closed on invoice, attendance sheet or assertion alone. |
| Share learning | Did relevant sites, assets, contractors and workers receive usable learning and change? | Targeted communication, transfer review, updated systems and feedback. | A bulletin is sent but equivalent risks remain unchanged elsewhere. |
What Value Should Investigation and Incident Management Create?
Three Evidence Layers and Four Defensible Ratings
Evaluate the Forklift and Chemical Incident Process
The organisation’s procedure looks complete. The incident file, interviews and follow-up data reveal a more complicated picture.
A Balanced Incident-Management Dashboard
Test the Evidence, Diagnose the Metric and Prioritise Improvement
1HSG245 Assurance Review
2Evidence Triangulation Lab
Claim: “Corrective actions are effective because every action is marked complete.” Which independent evidence would you test?
3Metric Detective
4Improvement Prioritiser
Completed or Effective? Closure-Evidence Dashboard
Use comparable performance evidence, real-work observation, workforce input and independent-enough verification to decide whether an action is merely installed or genuinely ready for closure.
Evidence picture
How to Answer 4.1–4.4 at Level 6
4.1 · Outline the stages
Present the incident lifecycle in a logical sequence. For each major stage, state its purpose, principal action, responsibility and output or handover.
Structure: Stage → what happens → why it matters → owner → output → brief case application.
4.2 · Outline the policies
Describe the umbrella policy and its four pillars. Include mandatory rules, responsible roles, supporting procedures/records and how the policy is assured.
Structure: Policy requirement → main rules → accountability → controlled document/record → brief application.
4.3 · Explain record maintenance
Identify the statutory or regulatory requirement, then explain how the record is created, verified, protected, corrected, retrieved, retained and disposed of so the duty is met.
Structure: Requirement → maintenance method → why it satisfies the duty → incident-file application → consequence of failure.
4.4 · Evaluate the process
Use a clear benchmark and triangulated evidence to judge design, delivery and results. Balance strengths and weaknesses, state the consequence and prioritise verifiable improvement.
Structure: Benchmark → evidence → finding → judgement → consequence → recommendation → verification.
Interactive Level 6 Command-Word Coach
Common Weak Answers
Check the Complete Incident-Management System
What This Section Is Built On
Turn Unit 3 Learning into 17 Clear Pieces of Assessment Evidence
This is the bridge between knowing Risk and Incident Management and demonstrating it at Level 6. Follow the exact task, cover every assessment criterion, apply it to one clearly identified organisation and support each judgement with reliable evidence.
Do the Exact Work the Brief Requires
The assignment is not one general discussion about safety. It is a controlled evidence journey. Each task has its own genre, word limit and assessment criteria.
Each task states 1,000 words and allows 10% either side: 900–1,100 words per task. Do not rely on a combined total if one task is outside its own range.
Plan Every Assessment Criterion Before Writing
Select a task. The word figures below are a recommended planning budget—not an OTHM-mandated distribution. Adjust them while keeping the task between 900 and 1,100 words.
Risk Identification, Assessment and Reliability
Task instruction: Detail the processes and strategies used to identify and analyse risk and how this influences risk management within the organisation.
| AC | What the learner must demonstrate | Useful organisational evidence | Suggested words |
|---|---|---|---|
| 1.1 Outline | Essential internal and external information sources used to identify hazards and assess risks, plus their organisational relevance. | Incident/near-miss data, sickness and damage trends, maintenance records, inspections, HSE/OSHA, ILO, WHO, professional and trade guidance. | |
| 1.2 Explain | How hazard-identification techniques work, how they are used and why they suit the organisation. | Observation, consultation, task analysis, inspections, audits, HAZOP, failure tracing, FTA or other proportionate techniques. | |
| 1.3 Explain | The end-to-end risk-assessment system and how risk is evaluated—not only a five-step list. | Scope, competence, people at risk, method, likelihood/severity, control standards, prioritised action, approval, review and limitations. | |
| 1.4 Explain | How the organisation measures, monitors and reports hazards. Address all three actions. | Exposure measurements, inspections, leading/reactive indicators, reporting channels, escalation, frequency, responsibility and dashboards. | |
| 1.5 Explain | How risk-assessment records are controlled to meet applicable regulatory and statutory requirements. | Jurisdiction, required content, owner, approval, revision, accessibility, retention, privacy, review triggers and audit trail. | |
| 1.6 Explain | Actual calculations used to analyse and improve system reliability and trace failure, followed by interpretation and action. | Relevant worked examples such as MTBF, MTTR, failure rate, availability or R(t); data source, result, meaning, limitation, improvement and verification. | |
| Structure | Introduction, synthesis showing influence on risk-management decisions, and a concise conclusion. | Organisation/site scope and a clear line of argument across 1.1–1.6. |
Risk-Control Strategies, SSOWs and SOPs
Task instruction: Carry out an evaluation of the organisation’s strategies and techniques for controlling risk.
| AC | What the learner must demonstrate | Useful organisational evidence | Suggested words |
|---|---|---|---|
| 2.1 Evaluate | Judge the effectiveness of common strategies actually used: criteria, evidence, strengths, limitations, constraints and verdict. | Risk register/actions, control performance, residual risk, audits, incidents, worker feedback, legal/industry benchmarks and gaps. | |
| 2.2 Justify | Address all six separately and explain both actual use and when each is appropriate: avoidance, reduction, transfer, analysis, evaluation and review. | One or more credible decisions showing why the selected timing and strategy fit risk, feasibility, duty and organisational context. | |
| 2.3 Explain | Explain both the development process and characteristics of safe systems of work and safe operating procedures. | Risk information, consultation, task sequence, competence, roles, permits, controls, emergency arrangements, authorisation, communication, monitoring and review. | |
| Report frame | Introduction/scope, findings under numbered headings, conclusion and focused recommendations. | Keep recommendations connected to evaluated gaps; do not replace evaluation with a list. |
Loss Causation, Data, Reporting and Investigation
Task instruction: Assess the organisation’s approach to incident investigation.
| AC | What the learner must demonstrate | Useful organisational evidence | Suggested words |
|---|---|---|---|
| 3.1 Outline | A range of loss-causation theories and techniques and how they apply to organisational events. | Bird/domino, multi-causality, immediate/underlying/root causes, Swiss Cheese, FTA, ETA, Bowtie and behavioural RCA. | |
| 3.2 Justify | Why quantitative loss-data methods are useful and appropriate, with limitations and data-quality caution. | Frequency, incidence, severity or ill-health prevalence rates; trend charts; exposure denominator, sampling, variability, errors and interpretation. | |
| 3.3 Assess | Weigh the needs and impacts of reporting loss events, then reach an organisational conclusion. | Legal/internal learning need, early warning, morale/trust, resource and reputation effects, under-reporting, data quality, feedback and action. | |
| 3.4 Explain | How and why investigation matters and what impact it has on the organisation. | Compliance, evidence, causal learning, recurrence prevention, control quality, worker morale, cost, reputation and organisational performance. | |
| Structure | Introduction connecting the four ACs, a sustained evidence-led argument, and a concise conclusion. | Do not write only about investigation: 3.1–3.3 are separately assessed. |
Incident Management, Records and Process Evaluation
Task instruction: Outline the organisation’s processes and strategies to manage health and safety incidents.
| AC | What the learner must demonstrate | Useful organisational evidence | Suggested words |
|---|---|---|---|
| 4.1 Outline | The critical end-to-end stages for managing incidents. | Preserve scene, people/equipment, report, triage investigation depth, gather/analyse, identify controls, implement actions, share learning and verify closure. | |
| 4.2 Outline | Organisational policies for all four functions: identify, investigate, report and record. | Policy purpose, responsibilities, reporting routes, investigation authority, escalation, record ownership, action closure and communication. | |
| 4.3 Explain | How incident records are maintained to meet applicable regulatory and statutory requirements. | Jurisdiction, duty holder, required information, classification, notification links, integrity, privacy, correction, retention, legal hold, access and disposal. | |
| 4.4 Evaluate | Judge whether an actual organisational incident-management process works in design, delivery and results. | Defined benchmark, representative evidence, strengths, weaknesses, consequences, overall verdict, prioritised improvements and verification method. | |
| Report frame | Introduction/scope, findings, overall conclusion and prioritised recommendations. | Keep 4.4 visible as a separate evaluative heading even though the brief’s bullet list omits it. |
The Verb Controls the Depth of the Answer
Read the full question, then use the assessment criterion’s verb. “Detail,” “carry out” and other wording in the task introduction do not reduce the demand of the AC.
Essay and Report Are Not the Same
Tasks 1 and 3 · Essay
- Specific title
- Introduction: organisation, scope and argument
- Connected thematic sections mapped to the ACs
- Critical discussion linking theory, evidence and practice
- Conclusion answering the task—normally no new evidence
- Harvard reference list
Tasks 2 and 4 · Report
- Title and clear organisational scope
- Contents page if the centre template expects one
- Introduction / purpose / method
- Numbered findings mapped to the ACs
- Evidence-based conclusion
- Prioritised recommendations where justified
- Harvard reference list
Define the Organisation Before Researching
A learner may use their own workplace or another organisation they can evidence. Use one consistent context and distinguish real evidence from reasonable recommendations.
Your assignment context map
Complete the fields to create a planning map. The tool will not write assignment prose.
Research Widely—Write Independently
Build a balanced evidence set
- Current applicable legislation and regulator guidance
- Academic books and peer-reviewed research
- Recognised standards and professional-body guidance
- Controlled organisational policies, records and data
- Credible workplace observations or interviews, where permitted
Protect authenticity
- Write the work yourself and keep a traceable research trail
- Cite every borrowed idea, model, table, figure or quotation
- Paraphrase genuinely—not by changing a few words
- Never invent organisational data, legislation or references
- Follow the centre’s current rules on permitted AI assistance and disclosure
- Sign the required statement of authenticity
Nine Common Ways a Learner Can Miss the Standard
Nothing Is Complete Until Every Criterion Is Visible
Tick only after the draft contains traceable evidence at the required command-word depth. The checklist measures coverage, not assessor approval.
0 of 17 criteria evidenced. Begin with the task plan—do not tick from memory.
Submission-control checklist
0 of 12 submission controls complete.
Use the Brief and the Current Specification Together
The uploaded assignment brief is dated January 2021 and sets the four task genres, assessment-criterion mapping and 1,000-word limit for each task. The December 2023 OTHM qualification specification confirms the same 17 Unit 3 assessment criteria, but it is not the source of the task genres or word limits shown here. The centre must confirm that the issued assignment brief is the authorised version for the learner’s cohort.
Independent resource status: This assignment-preparation masterclass was prepared by Debjyoti Biswas, FIIRSM, CertIOSH, for DB HSE International, OTHM Approved Training Centre DC2201634. It supports teaching and does not replace the centre-issued assignment brief, assessor judgement or OTHM requirements.
Use Authoritative Information
This independent DB HSE explanation supports learning. Workplace decisions must use current applicable legislation, exposure limits, approved organisational criteria and competent specialist advice where required.