Dark data, written once and never read

Where operational data ends up

More than half of what an operation records is never looked at by anyone, ever. Not because it is worthless, but because nobody can find it, nobody trusts it, or nobody knows it exists. That share has a name: dark data.

Recorded Read Never read Stored, retained, secured, unread

What it means

What dark data is, and what it is not

Dark data is data an organisation collects and stores in the normal course of running, and then never uses for anything. Sensor histories nobody queries, log files kept for a retention rule, inspection photos in a folder with one reader, exports that outlived the person who made them.

It is not broken, and it is not noise. Most of it was recorded because somebody once needed it. The failure is not in the data. It is that nothing connects it to the people and questions that would use it.

What it costs

Why plants produce so much of it

Nobody sets out to collect data they will never read. Dark data is what the defaults produce when recording is easy and finding is hard.

Recording is cheap, context is not

A historian tag or a log line costs nothing to write. Knowing which pump, which process and which question it belongs to takes modelling work, and that work is what gets skipped on a busy week.

The finder leaves

Data is findable as long as the person who recorded it remembers it. When they change roles or leave, the data stays and the map of it goes with them.

Tools reward writing, not reading

Most systems make it far easier to store a value than to discover one. Ten years of ingestion pipelines and no catalogue is the standard shape of an industrial estate.

Our approach

How data stops being dark

Data goes dark when it has no address. The fix is not another storage system. It is a model that gives every value a place someone can navigate to.

Give every value a place in the model

When time-series, events and files attach to assets and processes in one connected model, what do we know about this pump becomes a query instead of an archaeology project.

Make finding cheaper than measuring again

A searchable model with lineage means the first move is to look, not to install another sensor. Most dark data is rediscovered the first time someone can actually browse it.

Let agents read the archive

AI agents are tireless readers. Pointed at a connected model, they can sweep the history nobody had time for and surface what deserves a human look.

Common questions

Asked and answered

What is dark data?

Dark data is data an organisation collects and stores as part of normal operations and then never uses for analysis or decisions. Industry surveys estimate that roughly half of all stored operational data is dark, kept but never read.

Why is dark data a problem?

It costs storage and carries retention, privacy and security obligations while returning nothing, and its invisibility hides answers that teams then measure again. The data a decision needed often already exists, unread.

How do you reduce dark data?

Give the data an address. When values attach to assets and processes in a connected model with search and lineage, finding becomes cheaper than collecting again, and the dark share shrinks every time somebody looks.

The rest of the split

Where the other data points go

The drain has five outlets, and each one fails an operation in its own way. This page covers one of them. The other four are worth ten minutes each.

Want to see this against your own data?