A work safety dashboard in good health is a reassuring sight. Lost-time injuries are trending down, the latest audit scored well, and every work health and safety performance indicator on the board is green. The natural response is relief, and perhaps a quiet sense that the organisation has its risks under control.
The watermelon effect is a warning about that moment when a watermelon is green on the outside and red on the inside. Applied to work safety, it describes a set of indicators that look healthy on the surface while the controls that prevent serious harm are quietly weakening underneath.
Key points
- Work health and safety performance indicators are the measures a business uses to judge whether its safety is under control. Chosen poorly, they create the watermelon effect: green on the surface, red underneath.
- A low injury rate is a lagging indicator of mostly minor harm. It does not predict the low-frequency, high-consequence events that kill people.
- Better indicators measure the health of the controls that prevent serious harm, pairing a leading and a lagging measure for each critical control.
- Three questions strip the problem to its core: do we understand what can go wrong, do we know the controls are in place, and do we have assurance they are working.
- A small set of verifiable leading indicators tells a board more than a large dashboard of convenient ones.
The term began in information technology, where service dashboards glowed green while customers sat through outages. The consultancy ABB later borrowed it for process safety. The pattern it names belongs to any organisation that judges its own safety by indicators, which is to say almost all of them.
When the board is green, is the organisation measuring what keeps people safe, or measuring what is easy to keep green?
What the watermelon effect describes
The watermelon effect is rarely deliberate. It is a failure of measurement and assurance, and very few organisations set out to hide anything. Most simply measure the wrong things, or measure the right things in a way that flatters.
The green skin is the reported performance: low injury frequency, high audit scores, a tidy set of indicators. The red flesh is the real condition of the controls that stand between routine operations and a fatality. The two can drift a long way apart before anyone notices, because the surface numbers keep saying that all is well.
That gap matters most where the consequences are largest. A workplace can report excellent figures for years and still sit one failed control away from a serious event.
Why a low injury rate can hide a high risk
The clearest evidence sits in the history of the BP Texas City refinery. On 23 March 2005 an explosion there killed 15 workers and injured around 180. In the years before the explosion, the refinery’s personal injury record was better than the industry average.
The independent review that followed, known as the Baker Panel, reported in 2007. It found that BP had read its low personal injury rates as a sign of sound process safety, and that the inference was wrong. Personal injury rates did not predict process safety performance. By trusting them, the company had built a false sense of confidence that its major risks were under control.
The same gap appears in Australian data. In 2024, 188 workers were fatally injured at work, while 146,700 serious workers’ compensation claims were lodged in 2023–24, more than 400 a day (Safe Work Australia, Key Work Health and Safety Statistics Australia 2025). These two measures do not describe the same risk. Most worker deaths cluster in a small set of high-energy events, with vehicle incidents, falls from a height, and being hit by moving objects accounting for close to two-thirds of fatalities in 2024. A falling lost-time injury rate, driven mostly by the frequency of lower-severity harm, says little about whether the controls against those fatal events are sound.
The Australian researcher Andrew Hopkins made the point directly in his study of Texas City, Failure to Learn. A strong lost-time injury record can produce exactly the complacency that precedes a disaster.
How the watermelon grows
Several forces push the skin green while the flesh turns red. The first is over-reliance on lagging indicators. Injury counts, incident totals, and audit results describe what has already happened, or what the paperwork records. They are straightforward to collect and easy to present, so they dominate most dashboards.
The second is convenient indicator selection. Measures are often chosen for being simple to collect and already well controlled. Whether they test a known weakness rarely enters the decision. An indicator that has shown green for three years may be tracking something that was never genuinely at risk.
The third is a good-news culture. Where reporting bad news carries a career or commercial cost, dashboards turn green for the wrong reason. Near misses go unrecorded, and the absence of bad news gets read as safety when it may only be the absence of reporting. The Baker Panel found this pattern at BP, noting that incidents and near misses were not investigated effectively across the company.
The fourth is the gap between the system on paper and the work as it is actually done. A procedure can exist, pass an audit, and score well, while the task on the floor runs differently. An auditor who samples documents rather than watching the work will miss it. Safetysure has examined this paper-to-practice gap at length. The sociologist Diane Vaughan described the slow, unremarked acceptance of this kind of drift as the normalisation of deviance.
Safety performance indicators that tell the truth
The remedy is to measure the health of the controls that prevent serious harm. Counting the harm after it happens is not enough. Three bodies of guidance, drawn from different industries, point the same way.
The HSE guidance HSG254 sets out a model it calls dual assurance. For each critical risk control system, an organisation pairs a lagging indicator, which shows whether the control has failed, with a leading indicator, which shows whether the activities that keep the control working are actually being carried out. Each checks the other. HSG254 was written for major hazard industries, yet its authors state that the same model suits any organisation seeking similar assurance.
API RP 754, developed in response to the Baker Panel, arranges indicators along a continuum from lagging to leading across four tiers. The tiers run from an actual loss of containment at the top down to the everyday management activities that keep barriers sound. Organisations that map their existing measures against the four tiers often find the same imbalance: plenty of data at the lagging end, very little at the leading end.
The most demanding current practice goes a step further. The International Council on Mining and Metals sets out Critical Control Management, which asks an organisation to identify the few controls that genuinely prevent a fatality or catastrophe, define what good performance looks like for each, give each an owner, and then verify in the field that the control is present and effective. The stated aim is justified confidence, and the method deliberately looks for the gap between work as imagined and work as done.
For senior leaders the implication is sharp. As Safetysure has set out in its work on officer due diligence, the safety information reaching a board should show whether the critical controls are working, not merely how many incidents occurred. One idea runs through all of this. Confidence in safety should rest on verified evidence that controls work. The absence of incidents is not that evidence.
Three questions for the boardroom
At the conclusion of the Buncefield prosecution in 2010, the United Kingdom Health and Safety Executive issued a challenge to the boards of major-hazard companies. Gordon MacDonald, then Director of its Hazardous Installations Directorate, put three questions to senior leaders on behalf of the UK regulators. They are often called the Buncefield questions, although they came from the regulator at the close of the case rather than from the 2005 incident itself:
- Do we understand what can go wrong?
- Do we know what systems are in place to prevent this happening?
- Do we have assurance that these systems are working effectively?
The third question is the one the watermelon effect cannot survive. A green dashboard answers a question nobody asked. It reports that nothing has gone wrong lately. It does not show that the critical controls have been checked and found effective, and that is the assurance senior leaders actually need.
Chronic unease: the cultural antidote
Indicators and frameworks only work inside a culture that wants to hear bad news. The organisations that manage major hazards most reliably, such as air traffic control and nuclear operators, share a trait that researchers call chronic unease. Weick and Sutcliffe describe it as a preoccupation with failure and a refusal to treat a quiet stretch as proof of safety.
Research led by Dr Laura Fruhen at the University of Western Australia, with Professor Rhona Flin, defines chronic unease as a healthy discomfort about the management of risk and a scepticism about one’s own decisions. In day-to-day terms it means treating the absence of surprises as a reason to look harder, thanking the person who reports a near miss, and asking the third Buncefield question even when the board is green.
One limit matters here. The same researchers note that unease follows an inverted U. A little sharpens judgement; too much tips into anxiety that helps no one. The goal is calibrated wariness, not alarm.
A Safetysure view: choose a few indicators that can be verified
Most organisations do not need more indicators. They need a few that can be trusted. In Safetysure’s WHS audit and advisory work across Australian industries, the safety performance indicators that earn their place tend to share three traits.
They attach to a control that prevents a fatality or serious injury. General activity counts for little here. Confirming that the energy isolation on a specific high-risk task was in place and verified tells a board something that matters, where a count of toolbox talks does not.
They can be checked independently, by observation or test, rather than self-reported on a form. An indicator that relies on the same people reporting on their own compliance will drift toward green over time.
They are few. A board that watches three or four well-chosen leading indicators, each tied to a critical control, sees more than a board reading a forty-line dashboard. The same logic applies to what a safety committee reviews each month.
For Australian conditions this usually means starting from the hazards most likely to kill or maim in the particular operation, often vehicle and mobile plant interaction, work at height, and uncontrolled energy, then choosing one verifiable leading indicator for each. The test is simple. Could the organisation show a regulator, today, the evidence behind each green light? We recommend building the indicator set backwards from that question.
What senior leaders and WHS practitioners can do
Measurement is essential. The task is to measure the right things, and to treat good news with the same scrutiny as bad.
A handful of steps help close the distance between reported and real safety. The indicator set itself can be audited, with each measure tested against a single question: what weakness does this detect? Indicators that have stayed green for years without ever being at risk can be retired. Each critical control can be paired with both a leading and a lagging indicator, on the dual assurance model. The small number of controls that prevent a fatality can be verified in the field, through observation and testing of the actual work. The reporting culture can be examined, so that a scarcity of near misses is read as a reason to look harder. The three questions can be taken to the leadership table, with the third kept firmly in view. Indicators can then be refreshed on a regular cycle, so the measures keep pace with where the risk has moved.
The organisations that avoid the next serious incident tend to be the ones that distrust their own green dashboards enough to look underneath. A green light is best understood as a question, not an answer.
Safetysure helps organisations review their safety performance indicators and test whether the controls that matter most are working in practice, through its WHS audit and advisory work. We welcome a conversation with any senior team that wants to know whether its dashboard is telling the whole truth.
Frequently asked questions
What are work health and safety performance indicators?
They are the measures an organisation uses to judge whether its safety risks are under control. They fall into two broad types: lagging indicators, which record harm or failures that have already happened, and leading indicators, which show whether the activities that keep controls effective are being carried out.
What is the difference between leading and lagging safety indicators?
A lagging indicator looks backward, counting outcomes such as injuries, incidents, or claims. A leading indicator looks forward, checking that a control is in place and working before anything fails. Strong programs pair the two for each critical control, so that one confirms what the other suggests.
What is the watermelon effect in safety?
It describes safety indicators that appear green on the surface while serious risk builds underneath, like a watermelon that is green outside and red inside. It usually happens when an organisation relies on convenient lagging measures, or on indicators chosen because they already look good, rather than measures that test the controls against fatal events.
How many safety performance indicators should an organisation track?
Fewer than most organisations expect. A small set of verifiable leading indicators, each tied to a control that prevents a fatality or serious injury, tells a board more than a large dashboard of convenient measures. The practical test is whether the organisation could show a regulator the evidence behind each one.
References
- ABB 2017, Avoiding the ‘watermelon’ effect, ABB white paper (origin of the watermelon framing in safety).
- Safe Work Australia 2025, Key Work Health and Safety Statistics Australia 2025, Safe Work Australia, Canberra.
- The BP US Refineries Independent Safety Review Panel 2007, The Report of the BP US Refineries Independent Safety Review Panel (the Baker Panel report).
- US Chemical Safety and Hazard Investigation Board 2007, Investigation Report: Refinery Explosion and Fire, BP Texas City, Texas.
- Hopkins, A 2008, Failure to Learn: The BP Texas City Refinery Disaster, CCH Australia, Sydney.
- Health and Safety Executive and Chemical Industries Association 2006, Developing Process Safety Indicators: A Step-by-step Guide for Chemical and Major Hazard Industries (HSG254), HSE Books, Sudbury.
- Health and Safety Executive 2010, three questions posed to major-hazard company boards by G MacDonald, Director of the Hazardous Installations Directorate, at the conclusion of the Buncefield prosecution; recorded in IChemE Hazards 27 symposium proceedings (2017).
- American Petroleum Institute 2021, Process Safety Performance Indicators for the Refining and Petrochemical Industries (ANSI/API Recommended Practice 754), 3rd edn (1st edn 2010), API, Washington DC.
- International Council on Mining and Metals 2015 (updated 2026), Critical Control Management: Good Practice Guide, ICMM, London.
- Reason, J 1997, Managing the Risks of Organizational Accidents, Ashgate, Aldershot.
- Weick, KE and Sutcliffe, KM, Managing the Unexpected: Sustained Performance in a Complex World, John Wiley and Sons, San Francisco.
- Fruhen, LS, Flin, R and McLeod, R 2014, ‘Chronic unease for safety in managers: a conceptualisation’, Journal of Risk Research, vol. 17, no. 8, pp. 969–979.
- Vaughan, D 1996, The Challenger Launch Decision: Risky Technology, Culture, and Deviance at NASA, University of Chicago Press, Chicago.
