Research · 30 Jun 2026 · 31 min read

The Numbers Were Always Green – Safety Metrics Boards Misread

John Ninness
John Ninness John is Safetysure's Principal Consultant

Over the decade to 2023-24, the rate at which Australian workers died on the job fell by roughly a quarter. Over the same decade, serious workers’ compensation claims for psychological injury rose from approximately 6,700 in 2013-14 to 17,600 in 2023-24, an increase of 161 per cent, the largest of any category of harm, until mental health conditions made up 12 per cent of all serious claims, the highest share ever recorded. One set of numbers was moving in the direction that’s boards desire, and the other was moving the opposite direction, faster than anything else, and for most of that decade it was wasn’t on many boards agendas at all.

Part of the rise in psychological claims reflects changes that have nothing to do with worsening conditions. Stigma around submitting stress related claims has fallen, workers are more willing to lodge a claim, and the psychosocial regulations introduced across the jurisdictions during this period gave those claims a clearer legal position.. Some of the 161 per cent was likely recognition catching up with harm that probably already existed but wasn’t captured. The leading causes of these claims are likely not artefacts of reporting such as harassment and bullying at a third of claims, work pressure at a quarter, and exposure to violence, which are features of how work is designed and led rather than features of how willing people are to complain about it. Whatever share is recognition and whatever share is deterioration, the conditions producing the claims sit upstream of the injury, in the same place the conventional work safety report has never looked.

If that were a novel problem, it would be likely forgivable. Unfortunately it is not. Three decades earlier, in the Queensland coal industry, a disease that had been largely declared eradicated was quietly accumulating in the lungs of working miners while every formal indicator implied the system was sound.

A disease the data had already recorded

Coal workers’ pneumoconiosis or black lung, was thought to have been eliminated from Queensland mining by the regulatory and health surveillance systems built through the 1980s and 1990s. In 2015 it was confirmed in a working miner for the first time in three decades. By the time the Queensland Parliament’s select committee reported in 2017, twenty-one cases had been confirmed, and the committee’s language was not the careful hedging usual to such reports. It found a catastrophic failure, at almost every level, of the regulatory system meant to protect coal workers’ health.

The detail that matters for any board, in any industry is not that the dust controls failed, it is that the data which would have revealed the harm already existed and the systems built by government and industry potentially obliged no competent person to examine it for early disease rather than fitness for work. The Monash review that preceded the inquiry confirmed that design choice, and found that something in the order of 100,000 miners’ chest X-rays had never been examined by a qualified medical practitioner. Dust exposure was monitored, and there was evidence that a majority of mines had at times operated above the exposure limit with no regulatory consequence. The chest X-rays existed,  the respirable dust readings existed and the disease was legible in the data the system already held by organisations that were created to prevent that disease. What was ultimately missing was an organisation or person required to read it as a warning rather than as routine work as done.

The two cases are not identical, and the difference is likely instructive for readers. Black lung was a failure to read data the system had already collected ie the chest X rays and dust respirable dust readings existed and went unexamined without a thought. The psychosocial injury surge is closer to a failure to collect, since most organisations gathered little systematic data on the exposures that produced psychological harm. One is blindness to information at hand, the other is blindness to information that was never gathered. What ultimately unites them is the deeper failure beneath both systems, harm that was knowable, through data held or data obtainable, while the formal work safety indicators reported that things were under control and potentially the people closest to the work were the last to accept what the evidence suggested. The median psychological claim now costs 35.7 weeks of lost time and a median payout of $67,400, which is the price of not gathering, or not reading, what could have been known. This essay is about why that pattern recurs (in the author’s experience), what the research says about the metrics most organisations have adopted to prevent it, and what boards should actually be asking to see in work safety reporting.

Measuring downstream of the harm

The dominant safety metrics of the last fourty years are typically counts of injuries that have already occurred. The lost time injury frequency rate and the total recordable injury frequency rate measure how many people were hurt badly enough to meet a reporting threshold, per million hours worked. They are simple, comparable between organisations, and typically easy to present for most organisations. They follow a standard methodology that has been ingrained in industry for many years. Those properties explain what they have continued to be used. They do not explain their ultimate value, because by the time an injury is counted, the conditions that produced it have potentially been present in the organisation for some period of time. The metric records the outcome of a failure but it says nothing about the failure, and nothing about the failures that have not yet produced an injury.

This is not a fringe critique of safety metrics, it has been the consistent finding of the most serious scholarship in the field. Andrew Hopkins, whose forensic analyses of the Longford gas explosion, the Moura mine disaster, and the Texas City refinery fire at the Australian National University did more than any other Australian work to establish how major accidents happen, argued repeatedly that a low injury rate tells an organisation almost nothing about whether it is exposed to catastrophe. An organisation can record falling personal injury rates while its defences against a low-frequency, high-consequence event quietly degrade, because the two have different causes and the injury rate measures only the first. James Reason gave this its theoretical form in that catastrophic failures arise from latent conditions, the resident pathogens introduced by decisions about resourcing, design, maintenance, and workload, which lie dormant until they align with a local trigger. Latent conditions are, by definition, present and consequential before any incident occurs and that is why a measurement system built on counting incidents cannot see them.

Neil Gunningham’s work on safety regulation supplies the behavioural mechanism, and it is the safety-specific form of a regularity that economists and social scientists had already named. Goodhart’s law holds that any statistical regularity used as a target tends to collapse under the pressure placed on it for control purposes (Goodhart 1975), sharpened by Strathern, that when a measure becomes a target it ceases to be a good measure (Strathern 1997). Campbell’s law is its social-policy twin, that the more a quantitative indicator is used to drive decisions, the more it will be gamed in the orhganisation and the more it will corrupt the process it was meant to monitor (Campbell 1979). When a regime measures and rewards a particular indicator, organisations manage to that indicator. A compliance system built around recordable injury rates produces organisations that become skilled at managing recordable injury rates, through classification, through return-to-work timing, through the chasing of ambulances and at the margin through suppression, none of which requires preventing a single injury. The metric and the harm come apart, and the organisation can improve the former while the latter is unchanged or worsening. The lost time injury rate is not merely a weak signal in that where it is the dominant signal, it actively shapes behaviour toward managing the number.

A fair objection to the argument so far is that it is built from disasters, and disasters are precisely the cases in which the indicators were, in hindsight, misleading. The catastrophes that fill the long time inquiry records are selected on their outcome and no commission is convened for the refinery that ran for forty years on a falling injury rate and was, in fact, safe, so unfortunately that population data never enters the account. Read against it, a low injury rate is ultimately not meaningless and for the high-frequency, lower-consequence harm that makes up the bulk of any organisation’s injury count, the sprains, the lacerations, the falls from the same level, the rate measures something real, and an organisation whose count is climbing there has a problem worth knowing about. The claim of this essay is therefore narrower than “the numbers lie”, and survives the objection in its narrower form. The lagging rate is reasonably informative about the class of harm it counts, and close to uninformative about the classes it does not, the low-frequency catastrophe, the latent disease, and the psychological injury. Its failure is not that it is untrue but that it is asked to reassure across a range of harm it was never able to visualise. The cases that open this essay are remembered precisely because they are the ones in which that misplaced reassurance proved fatal to workers; the discipline is to keep the measure for what it can do and to stop accepting it as evidence about everything it cannot.

The comfortable answer and why the evidence does not support it

The safety profession is aware of all this, and has likely has been for many years. The standard response, in 2026, is that the field has already moved on, that mature organisations now run leading indicators, the proactive measures of activity and condition that are meant to predict and prevent rather than merely record. Safety audits completed, training delivered, hazard inspections conducted, behavioural observations logged, critical control checks performed. The premise is that measuring these forward-looking activities closes the gap that lagging indicators leave open. It is a reasonable assumption, and leading indicators are a genuine advance in intent over counting injuries after the fact. The difficulty is not that they have been shown to fail, it is that the evidence they succeed is considerably thinner than their near-universal adoption likely implies, and an organisation that has adopted them in good faith may have less assurance than it believes.

In 2023 a research team conducted a scoping review of the evidence base for safety leading indicators, searching eight databases and identifying forty-eight primary studies that evaluated whether leading indicators actually improve safety outcomes. The review was commissioned by the Lloyd’s Register Foundation, a charity whose stated mission includes the promotion of leading indicators, so any bias in its framing would run toward a favourable conclusion. Its conclusion was not favourable and while most of the individual studies reported a positive effect of leading indicators on lagging outcomes, the review found the overall evidence base to be weak on three independent grounds. The study designs were not capable of establishing causation, the internal validity of the studies was moderate to low, and the findings were poorly generalisable. Most strikingly, no two of the forty-eight studies addressed the same research question, which meant their results could not be combined into any coherent synthesis. The body of evidence was substantial in volume and largely incoherent in substance.

No two of the forty-eight studies addressed the same research question. The body of evidence for leading indicators was substantial in volume and largely incoherent in substance.

The construction safety literature, where leading indicators have been studied most intensively, reaches compatible conclusions. Systematic reviews note that the connection between the leading indicators organisations actually use and the situations that actually cause incidents remains poorly established, that safety practitioners are given little evidence-based guidance on which indicators to select, and that alternative metrics are typically studied in isolation, with their benefits asserted and their weaknesses rarely examined. The recurring methodological recommendation from this literature is highly revealing, that metrics should be evaluated against explicit criteria of whether they are predictive, objective, and valid, and that no single metric should be relied upon, because each has weaknesses that only a deliberate combination can offset.

The uncomfortable implication is that metric adoption has typically run ahead of validation. Many organisations measure leading indicators, but the indicators they have chosen tend to be the ones that are easy to count rather than the ones with a demonstrated causal connection to harm. Counting safety meetings held, or audits completed, or observation cards submitted, measures the volume of safety activity. It ultimately does not establish that the activity reduces harm, and a leading indicator that measures activity for its own sake is vulnerable to exactly the decoupling that Gunningham identified in the lagging measures. The number improves, the activity is performed, and the underlying risk is largely untouched. A fair objection is that most leading indicators were aimed at acute physical hazards, so their adoption tells us little about the psychosocial rise, which they never targeted. That is substantially correct, and it is the point. The leading-indicator movement concentrated on the hazards that were already being managed, and the harm that was growing fastest sat in a domain those indicators were never designed to watch. Weak evidence for the indicators in use, and near-absence of indicators for the harm that was rising, are two faces of the same problem: a measurement effort directed by what is countable rather than by where the harm is.

Why the signal is not enough

There is a deeper reason that better metrics, on their own, do not solve the problem, and it is the reason the black lung X-rays could sit unread and the psychosocial pressures could build unremarked. Diane Vaughan named it in her study of the Challenger launch decision, published in 1996. Investigating why NASA launched a shuttle its own engineers had worried about, Vaughan found no villains and no single bad decision. She found instead a gradual process she called the normalisation of deviance, in which evidence that something was departing from expected performance was repeatedly reinterpreted, each time, as falling within the bounds of acceptable risk. The O-ring erosion was observed, discussed, analysed, and absorbed. Each flight that returned safely with damage made the next instance of damage more normal. The deviation did not slip past unnoticed. It was noticed, and renormalised, until catastrophic risk had been redefined as routine.

Vaughan also identified what drove the renormalising, and it is the thread that ties the argument together: production pressure, combined with scarcity of time and money. The pressure to keep flying was what made each reinterpretation feel reasonable. The launch-schedule pressure on a space programme is not the same thing as the work pressure that Australian workers cite in a quarter of psychological injury claims, and it would be too neat to call them one phenomenon. What they share is a mechanism. In each, the demand to keep producing competes with the signals that producing has become unsafe, and over time the demand reframes the signal rather than heeding it. Hopkins, applying compatible analysis to the Longford gas explosion and the Texas City refinery fire, reached the same conclusion in organisational settings much closer to those most boards govern: the failures arose not from ignorance of the hazard but from a culture in which signals of deterioration were repeatedly absorbed and reclassified as acceptable variation, under the same competitive pressure to keep producing (Hopkins 2000, 2008). The same force runs through the major-accident record at Deepwater Horizon. Production pressure does not only cause harm directly. It degrades the organisation’s capacity to read the warning that harm is coming, by making each erosion of a margin look like a reasonable adjustment rather than a deviation.

This is the finding that should trouble any board, because it means that supplying an organisation with better data is necessary but not sufficient on its own. A workforce and a management layer operating under sustained production pressure will typically normalise a rising leading indicator as readily as they normalised the falling one, explaining each adverse movement as the new baseline. The black lung health scheme had the radiographs and NASA had the erosion data. In both cases the people closest to the work possessed the signal and had culturally reorganised themselves around not acting on it in a meaningful way. No metric is self-executing and the question is never only what is measured, but who reads it, and whether they are positioned to read it honestly.

A candid account has to concede an asymmetry at the centre of its own argument. Production pressure is named here as the force that drives the renormalising, the deepest link in the chain, and yet it is the one thing the recommended core does not measure directly, because it is the one thing organisations have least learned to instrument. This is not an oversight to be tidied away before publication; it is the frontier the argument runs up against. Part of its footprint is already visible in measures the combination does include: the concentration of excessive hours, the crew on its third consecutive night, the after-hours demand clustered in particular teams are production pressure leaving its trace in the psychosocial and control-integrity panels, and they should be read with that in mind rather than as isolated wellbeing figures. What none of them supplies is the direct article, a measure of the distance between what the organisation has committed to produce and the resources it has actually provided to produce it safely, tracked over time. Until that measure exists, the honest position is threefold, to name production pressure as a driver the dashboard does not yet capture, to read its proxies deliberately rather than mistake them for the thing itself, and, consistent with this essay’s argument about the reader, to make “are we operating beyond our safe capacity” a question a board puts directly to management rather than one it waits for a green cell to answer. The root cause being the least instrumented is not a reason to leave it unspoken. It is the strongest reason to govern it by question until it can be governed by measure.

What a better indicator has to do

If neither lagging nor leading indicators resolve the problem on their own, the answer is not a third category to be adopted with the same misplaced confidence. It is a set of design requirements that any indicator should be tested against, and then a deliberate combination of measures that together satisfy them. Three requirements follow from the evidence above.

The first is that a meaningful indicator should sit upstream of the harm, measuring the conditions that produce it rather than the count of those already hurt. The second is that it should resist normalisation, which in practice means it should be read against an external or original standard rather than against the organisation’s own drifting baseline, and ideally read by someone outside the operational culture doing the drifting. The third, which the psychosocial and black lung cases make non-negotiable, is that the set of indicators must span the full range of consequential harm, the acute and traumatic, the chronic and latent, and the psychological, because an organisation that instruments only the first will be blind to the other two in exactly the way the last decade has demonstrated at national scale.

No single number satisfies all three, which is the whole point. What follows is a combination, and the logic of the combination matters more than any individual item in it. It retains the conventional outcome metrics for the narrow purpose they genuinely serve, adds the better-validated forward measures where a real causal pathway exists, and adds upstream and cross-domain signals presented as matters for interpretation rather than as dashboard cells that turn green. A supplementary figure available with the full published version maps the fuller landscape these measures occupy, across the three harm domains and along the line from upstream condition to recorded outcome, and it includes drivers such as production pressure that the evidence identifies as causes of harm but that few organisations yet measure as standard. The prose below recommends the tighter core a board could adopt now; the figure shows the territory that core sits within, including the ground still left to instrument.

A combination of safety metrics worth putting in front of a board

The measures that follow are subject to the same evidential caution this essay has applied to leading indicators generally. They are not proven in the sense of having a demonstrated causal connection to reduced harm under controlled conditions, because the literature reviewed above shows no indicator currently meets that test. The argument for each is comparative rather than absolute: that it is anchored to a plausible causal pathway, that it spans the harm domains the conventional report ignores, and that it resists normalisation in ways the lost time injury rate does not. That is a reason to prefer this combination over the alternatives currently in use. It is not a warrant to adopt it with the same misplaced confidence this essay has criticised elsewhere.

Keep the serious-outcome metrics, but demote them to a position reflective of their value in preventing incidents. Fatalities, permanent-disability injuries, and confirmed occupational disease should still be reported, in full and without softening, because at the serious end these counts are unambiguous, hard to manipulate, and carry a moral weight no proxy can replace. What should change is their status in the reporting framework. Ultimately they belong at the back of the report as the record of what the system failed to prevent, not at the front as the headline that signals all is well. The lost time injury frequency rate can remain for external benchmarking and continuity, clearly labelled as a number that describes the past and predicts very little. Stripped of its false role as the primary signal of safety, it does no harm. Left in that role, it manufactures the false comfort this entire essay is arguing  against.

Add control-integrity measures for the acute risks that can kill or seriously injure workers. For the number of hazards in any organisation capable of causing a fatality, the meaningful question is not how often someone was hurt but whether the specific controls that prevent the fatal event were verified to be in place and working. This is not a novel proposal, and it should not be presented as one, it is the logic of critical control management, the barrier-based discipline set out in the ICMM’s Health and Safety Critical Control Management good-practice and implementation guides (ICMM 2015, updated 2026) and now embedded in major-hazard practice across mining and the process industries. The contribution here is not the measure but its placement and its reading: that verification of the controls standing between an energy source and a person belongs at the front of what a board sees, where the injury rate has sat for forty years, and that the verification figure has to be read for its quality rather than its completeness. A critical control verification rate, reported by hazard and by site, measures the barriers between a worker and a catastrophe directly. This is itself a leading indicator, and the scepticism applied above applies here too in that it carries no stronger evidence base than any other, and it must earn its place on the same terms. The argument for it is not that it is proven but that it is anchored to a defined causal pathway, the barrier between an energy source and a person, rather than to a count of activity whose link to harm is assumed. That anchoring is a reason to prefer it, not a warrant to exempt it. Its weakness must be stated plainly: it can be gamed by narrowing the register of what counts as critical, or by shallow verification that records a ticked box without formal testing whether the control would actually hold. Reported as a single completion percentage it invites exactly that gaming, so it should be read as a small composite of three things rather than one number. The first is coverage, the proportion of critical controls verified in the period. The second is independence, how much of that verification was carried out by someone outside the line that operates and maintains the control, together with the time elapsed since each control was last examined independently, because a control checked only by the crew that runs it drifts toward the crew’s own normalised view of what is acceptable, and the longer since an outside look, the more that drift has set in. The third is the findings in as much as what the verification actually discovered, since a process that never finds a degraded control is not assuring safety but performing it.

Add exposure-against-limit measures for the chronic-health risks, with a mandatory competent review. Black lung is the standing proof that latent disease is governed by exposure data and health surveillance, not by injury counts, and that the data is worthless if no competent person is required to read it. For the relevant hazards, respirable dust, silica, noise, diesel particulate, the indicator is the proportion of monitoring results within the exposure standard, reported as a trend, paired with confirmation that health surveillance has been conducted and, critically, examined by a qualified practitioner. The governance failure in Queensland was not unmeasured exposure, it was measured exposure and unexamined chest X-rays from a correct disease standpoint. The indicator has to close both halves of that gap or it repeats the failure.

Hearing loss also exposes the limits where an organisation may hold noise surveys and periodic audiometry, but neither captures what determines the disease. The surveys are episodic and often measure the area rather than the person, so they do not record the cumulative dose a particular worker absorbed across years of changing tasks. The audiometry, the one genuine piece of health data, is lagging rather than leading, because by the time a threshold shift appears on the audiogram the damage to the cochlea is permanent and done and the dose that caused it was delivered across a working life and across employers, most of it before the worker arrived and outside any record the current organisation holds. This is a different failure from black lung. There the data existed and went unread; here, for an individual’s hearing, the determining data was largely never the organisation’s to hold. A whole class of latent harm has this shape, much occupational cancer, noise-induced hearing loss, slow musculoskeletal degeneration: cumulative, long in latency, distributed across a mobile working life. For this class the honest position is narrow but very real. An organisation can measure and control the exposure it is creating now, ideally as a personal dose recorded while the worker is still being exposed and the harm is not yet locked in, and it cannot reconstruct the lifetime dose that actually governs the disease. The lifetime record would have to follow the worker rather than the employer, and it does not yet exist. A board should at least know that this is a thing it does not have, rather than mistake a folder of area noise surveys for a measure of whether its people are going deaf.

Add a psychosocial exposure measure built on the established hazard categories. The national data identifies the leading causes of psychological injury with precision: harassment and bullying, work pressure, and exposure to violence. These are measurable through validated survey instruments and through operational proxies, the distribution of excessive working hours, the concentration of after-hours demand, turnover and absence clustered in particular teams, and the volume and pattern of complaints. None is definitive alone, and each is open to the objection that it is indirect. Read together and over time, they indicate whether psychological harm is accumulating in identifiable parts of the organisation, which is precisely the information that was missing while claims rose 161 per cent. They have to be presented as a panel of individually inconclusive signals whose value lies in their convergence, not as a single index pretending to a precision it does not have.

Add the reporting-culture measures that tell you whether the rest can be trusted at all. Every metric above depends on people being willing to report honestly, so the health of the reporting system is a meta-indicator governing the reliability of all the others. Hazard and near-miss reporting rates, read as a trend, and the time taken to close a reported hazard, indicate whether the workforce believes that reporting is worthwhile and whether that belief is justified. The interpretation is counter-intuitive and must be taught once and then held: a falling volume of reports, absent independent evidence that conditions have improved, is a warning rather than a success, because it usually means people have stopped believing reporting changes anything. This is the indicator most vulnerable to the naive reading, and the one a board most needs to read correctly.

A falling volume of hazard reports, absent evidence that conditions have improved, is a warning, not a success. It usually means people have stopped believing that reporting changes anything.

Why the easy safety metrics survived

It is worth asking why the lost time injury rate has held its place for forty years, because the answer is not that it is a poor measure. It is a good measure of the wrong thing. It counts acute injury that has already happened, accurately and comparably, and then it is asked to reassure a board about the safety of everything that has not happened yet, across every kind of harm. The rate did not survive despite its limits. It survived because of what its strengths offered the people reporting it. A single, standardised, externally comparable number is safe to table, easy to benchmark, and easy to be seen to be managing, in a way that a set of indirect signals requiring interpretation never is. A falling rate may also let an organisation feel its duty was discharged, which is a powerful thing for a number to provide but its almost unrelated to whether anyone was safer at the workplace. A precise answer to the easy question will always hold an institutional advantage over an uncertain answer to the hard one. The rate was never the problem., allowing its narrow accuracy to stand in for an assurance it could not give was the fundamental problem, and the people that substitution served least, were the ones whose safety it claimed to describe.

The reader matters more than the safety dashboard

A combination of measures along these lines satisfies the three design requirements where no single metric can. It sits upstream, it spans acute, chronic, and psychological harm, and it retains the outcome counts for the honest record they provide. But the deepest finding of this essay is that even the right measures will be normalised away by the people closest to the work, because that is what sustained operational pressure does to human judgement. The x-rays were read, or not read, by a system built to find fitness for work. The erosion data was read by engineers who had learned to expect erosion. Improving the data without changing who reads it, and from what vantage point, would solve the smaller half of the problem and leave the larger half intact.

This is where the board’s role becomes structural rather than ceremonial. The value of a board here is not that it is more expert than management, because it is clearly not, and not that it is more diligent because it need not be. Its value is that it sits outside the operational culture that is doing the normalising, and can read the signals against an external standard rather than against the comfortable internal baseline that has drifted by degrees. That distance cuts both ways, and the essay should not pretend otherwise. The same separation that protects a board from the operational culture also leaves it dependent on what management chooses to surface, meeting too infrequently and too far from the work to see much for itself. Distance is an asset only if the integrated picture actually reaches the boardroom, which is the hard part and is not yet solved in most organisations. The upstream signals likely sit in separate operational, financial, and human-resource systems, owned by different functions, rarely assembled into one view. The board’s contribution is therefore conditional, in that it can be the reader the operational culture cannot be for itself, but only if it insists on receiving the integrated upstream picture, and treats the absence of that picture as a finding in its own right rather than as a gap to be excused.

Reading from outside also means reading at a higher level. The measures set out above are management’s instruments, fine-grained and, increasingly, real-time. They are the live state of a control, the dust accumulating across a shift, the crew on its third consecutive night. A board has no business in that detail, and a board that finds itself watching live operational readings has stopped governing and started supervising. What the board needs is not the feed but the answer to a small number of higher-order questions that the feed informs. Four are likely enough to create concern. Are the controls for our fatal and chronic risks verified, and is that verification independent of the people who operate them? Across our key signals, is the direction of travel improving or deteriorating? When a signal deteriorates, does the organisation act, and how quickly, because a signal can worsen for months while nothing happens, and the distance between deterioration and response is itself a measure. And can we trust what we are being shown, which means is the reporting culture healthy and are we seeing the disaggregated picture rather than an average that hides the one site or team where harm is gathering? The fourth question is the one boards most rarely ask and the one that matters most, because it is the question that interrogates the instrument rather than the reading, and an instrument that cannot be interrogated is the green safety dashboard in another form.

The numbers were always green. They were green at Longford and green before Texas City, green in the Queensland coal fields while the disease accumulated, and green across a decade in which psychological injury became the fastest-growing harm in the Australian workforce. The colour was never the problem, the problem was a measurement system that reported on the wrong things, read by the people least able to see past the reassurance, governed by boards that accepted the colour as the answer.

Changing the metrics is the necessary first step and changing who reads them, and what they are prepared to see, is the one that ultimately matters.

You might like to read an article by the author “What is Safe?” or Work Safety Indicators – The Watermelon Effect or Industrial Manslaughter A Guide for Boards

References

Campbell, D.T. (1979). Assessing the impact of planned social change. Evaluation and Program Planning, 2(1), 67–90.

Coal Workers’ Pneumoconiosis Select Committee (2017). Black Lung White Lies: Inquiry into the Re-identification of Coal Workers’ Pneumoconiosis in Queensland. Report No. 2, 55th Parliament, Queensland Parliament, Brisbane.

Goodhart, C.A.E. (1975). Problems of monetary management: the U.K. experience. In Papers in Monetary Economics, Vol. 1. Reserve Bank of Australia, Sydney.

Gunningham, N. and Johnstone, R. (1999). Regulating Workplace Safety: Systems and Sanctions. Oxford University Press, Oxford.

Hopkins, A. (2000). Lessons from Longford: The Esso Gas Plant Explosion. CCH Australia, Sydney.

Hopkins, A. (2008). Failure to Learn: The BP Texas City Refinery Disaster. CCH Australia, Sydney.

Hopkins, A. (2009). Thinking About Process Safety Indicators. Safety Science, 47(4), 460–465.

International Council on Mining and Metals (2015). Health and Safety Critical Control Management: Good Practice Guide (and companion Implementation Guide). ICMM, London. [Updated 2026.]

Monash Centre for Occupational and Environmental Health (2016). Review of the Respiratory Component of the Coal Mine Workers’ Health Scheme for the Queensland Department of Natural Resources and Mines. Monash University, Melbourne.

Reason, J. (1997). Managing the Risks of Organizational Accidents. Ashgate, Aldershot.

Safe Work Australia (2025). Key Work Health and Safety Statistics Australia 2025. Safe Work Australia, Canberra.

Strathern, M. (1997). ‘Improving ratings’: audit in the British university system. European Review, 5(3), 305–321.

Vaughan, D. (1996). The Challenger Launch Decision: Risky Technology, Culture, and Deviance at NASA. University of Chicago Press, Chicago.

Watkins, D., Bishop, E., Naylor, S., Frankish, R., Staves, M., Dony, J., Miller, P., McCool, R. and Ferrante di Ruffano, L. (2025). A scoping review of the evidence base for the performance of leading indicators for improving safety outcomes: available evidence, implications for practice and future directions. Journal of Safety Research, 95, 530–544. DOI 10.1016/j.jsr.2025.10.024.

Xu, J., Cheung, C., Manu, P. and Ejohwomu, O. (2021). Safety leading indicators in construction: a systematic review. Safety Science, 139, article 105250. DOI 10.1016/j.ssci.2021.105250.