Something happens on a plant at half past nine in the morning. A record of it appears in the safety system at twenty past four that afternoon. In between, a man carried it around in his head through six hours of the work he was actually being paid to do.

Very little of what mattered about it survives that journey. Nobody is careless. Nothing in the process was built to carry it.

The distinction I want to make in this piece is between a record that is complete and a record that is a good account of what happened. Those are two different tests. Almost every safety system on the market applies the first one rigorously, at the point of submission, with validation rules and mandatory fields and a refusal to accept anything with a gap in it. Almost none of them apply the second, because there is no way to write a validation rule for whether the report tells you anything.

So you can run a reporting programme at ninety-odd per cent completeness and hold a body of data that would not help you predict anything.

Why do safety reports lose detail?

Safety reports lose detail because the record gets created at a different time, in a different place, by a different person, and often in a different language to the event it describes. Each of those four steps strips something out, and none of them are visible in the finished record.

Worth taking them separately, because they are not the same problem and they do not have the same fix.

Six hours of forgetting

Human memory for detail degrades fast, and it degrades unevenly. What survives is the gist. What goes first is precisely the specific, situational, sensory material that makes a safety report useful to somebody reading it later: where exactly he was standing, what he heard before he looked, what he assumed was about to happen next.

By the time the card gets filled in at the end of the shift, he is not describing the event. He is describing his memory of having noticed the event, which is a thinner thing. He will still tell you the guard was loose. He is unlikely to tell you it was the third time that month, or that he only looked because the noise had changed.

Ask a crew where the observation cards actually get written and the answer is usually the vehicle, the crib room or the gate, at the end of shift. Calling that a discipline problem misses what is going on. It is the only slot in the day where nobody is asking him to do something else.

The category that nearly fits

Every safety taxonomy is a compression algorithm. Somebody sat down, probably years ago, and decided that everything that will ever happen on your sites can be sorted into somewhere between eight and forty categories. There is nothing wrong with the taxonomy. It is always going to be slightly smaller than reality.

So the man in front of the dropdown has something in his hand that is obviously a housekeeping issue and obviously an equipment guarding issue and slightly a training issue, and he has to pick one. He picks the one that generates the least follow-up work for him, which is a rational thing to do and which nobody has ever told him not to do.

That choice is invisible afterwards. Nothing in your data records that a judgement call was made, or which options were rejected, or that a different supervisor on a different shift would have coded it somewhere else entirely.

The box that runs out

Free-text fields are where the useful material lives, and they are the least completed part of most safety records.

Partly this is physical. He is filling it in with gloved hands, or on a phone, or on a paper card with a box the size of a stamp. Partly it is learned. Long free-text answers get queried, generate follow-up questions and create work, so the people who report the most learn to write the least. A four-word entry closes cleanly. A forty-word entry gets a phone call.

If you want an uncomfortable half hour, pull the free-text field from your last twelve months and sort by character count. In most operations, the median is short enough to fit in this sentence.

Reporting in your second language

On a multilingual site the report gets written in the language of the form, which for a large share of your workforce is their second or third. What comes back is thin, and the reason it is thin is that he had a great deal to say and none of it available in that language, at speed, at the end of a shift, with somebody waiting on him.

The version he would give you in his own language is longer, more specific and more useful. There has never been a mechanism for collecting it, so most operations have quietly concluded that their non-English-speaking crews are low reporters. Some are. Mostly what you are measuring is the cost of the form.

What makes a good safety observation?

A good safety observation lets somebody who was not there understand what the worker saw and why it concerned him. Completeness is a separate test, and it is the only one most systems apply.

The practical difference shows up when a safety manager reads the record three weeks later, or when an investigator goes looking for precursors after something has gone wrong. A complete record tells you an event of type X occurred at location Y on date Z. A good account tells you what was actually going on, which is the only version with any predictive value in it.

Most reporting programmes are optimised for the first and then judged on the second. That is where a lot of the frustration in this job comes from.

Who is actually writing your safety records?

In most industrial operations the record gets created by somebody who was not there. Usually a supervisor, working from a card, a conversation or a memory belonging to someone else.

This job does not appear in anyone's job description and rarely gets costed. It runs somewhere between forty minutes and two hours a day. He reads handwriting. He interprets what somebody meant. He picks the category. Where the account runs out he completes it himself, sensibly, from what he knows about that area and that crew, because a record that fails validation comes straight back to him and he has a shift to run.

That last step deserves more attention than it gets. A reasonable inference, made at half four on a Thursday by a competent person who was not there, is now safety data. It gets counted. It enters the trend. Some months later it is in a slide in front of a board.

He is doing the only thing available to him. Nothing about that is fixed by better training or a more strongly worded reporting policy, and both get tried repeatedly.

Why are near miss numbers unreliable?

Near miss numbers are unreliable because reporting one costs the worker time and returns him nothing, so the figure that reaches your board reflects how many were worth the paperwork rather than how many occurred. The arithmetic gets done in about a second and it comes out the same way most times.

Then a target gets set, and this is where it becomes actively misleading rather than merely soft. Sites start reporting more. The number climbs. In the room, that reads as a strengthening reporting culture, and sometimes it is exactly that. Sometimes it is people writing up trivia to hit a count, which leaves you worse off than under-reporting, because the two reports that mattered are now sitting underneath ninety that did not.

Nobody sets a target on the fifteen minutes, and the fifteen minutes is the only variable in this that anyone actually controls.

What this does to a group safety picture

For a single site, all of the above is a quality problem. For a multi-site operator it becomes a comparison problem, and that is the expensive version.

The same event at four sites produces four different records. Housekeeping at one. Slips and trips at the next. A near miss at the third, because that supervisor takes the wider view of most things. Nothing at all at the fourth, because that crew worked out some time ago that reporting generates work and rarely generates anything back.

Group safety then builds a trend on top of that, and the site that reports least comes out looking safest.

The taxonomy is already standardised, so standardising it again does nothing. What varies is the human step between the event and the box, and that step is a different person at every site, reading the same taxonomy his own way, with his own view of what the number is for. Most group safety leaders I have spoken to know their site comparisons are shaky. Most have stopped saying so out loud, because there has been nothing obvious to do about it.

Meanwhile the leading indicators sitting on top of this data are the ones you are using to decide where to send money and attention.

Can conversational reporting produce a structured record?

Yes, provided the structuring happens after the worker has spoken rather than being demanded of him while he speaks. The fields your safety system requires still have to be populated, and the only question is who populates them.

This is the part people get wrong when they first hear about it, so it is worth being exact. Severity, location, category, time, injury type, whatever else your system insists on before it will accept a record: none of that goes away. Auditors want it. Your own trending depends on it. What we claim is narrower than it sometimes gets repeated as. There is no form for the worker. The form itself is still there.

What that looks like in practice. He describes what happened, in his own language, at the job, while he can still see the thing he is describing. Whatever is missing for that record type gets asked for, one question at a time, in the way a supervisor would ask if he were standing there. The completed account is then read back to him, and nothing reaches your safety system until the worker confirms the record is right. If it is wrong he says so and it gets corrected before anything is committed under his name.

Voice matters here only because of the environment. Gloves, noise, movement, height, and a workforce for whom typing in a second language at the end of a shift is a real barrier. The mechanism is not the interesting part and every serious vendor in this category will have it before long. What changes the data is where the record starts and who does the translation work.

Two things this does not solve, since we would rather say so up front. It does not improve records for events nobody chooses to report at all, though lowering the cost of reporting moves that number. And it does not retrospectively fix the twelve months of data you already have.

Six things to check in your own data

None of this needs a new system to investigate. Most of it is sitting in an export you can pull this week.

The gap between event date and record creation date, as a distribution rather than an average. The share of your total volume coming from your five most frequent reporters. The proportion of free-text fields running under fifteen words. Whether your multilingual sites report at the same rate as your English-speaking ones once you normalise for headcount. Whether the same event type is coded consistently across sites. And the one most operations have never measured at all: what share of records were created by the person who was actually present.

If you run the authorship check and the number surprises you, that is the finding. Everything above is an explanation of it.