Dark data, written once and never read
Where operational data ends up
More than half of what an operation records is never looked at by anyone, ever. Not because it is worthless, but because nobody can find it, nobody trusts it, or nobody knows it exists. That share has a name: dark data.
What it means
What dark data is, and what it is not
Dark data is data an organisation collects and stores in the normal course of running, and then never uses for anything. Sensor histories nobody queries, log files kept for a retention rule, inspection photos in a folder with one reader, exports that outlived the person who made them.
It is not broken, and it is not noise. Most of it was recorded because somebody once needed it. The failure is not in the data. It is that nothing connects it to the people and questions that would use it.
- It sits in normal storage, costs normal money, and answers to nobody.
- It is invisible to the teams who would use it, so they measure again what already exists.
- It carries real obligations: retention, privacy and security apply whether or not anyone reads it.
- Surveys put it at roughly half of all stored data, and the operations we meet rarely disagree.
What it costs
Why plants produce so much of it
Nobody sets out to collect data they will never read. Dark data is what the defaults produce when recording is easy and finding is hard.
Recording is cheap, context is not
A historian tag or a log line costs nothing to write. Knowing which pump, which process and which question it belongs to takes modelling work, and that work is what gets skipped on a busy week.
The finder leaves
Data is findable as long as the person who recorded it remembers it. When they change roles or leave, the data stays and the map of it goes with them.
Tools reward writing, not reading
Most systems make it far easier to store a value than to discover one. Ten years of ingestion pipelines and no catalogue is the standard shape of an industrial estate.
Our approach
How data stops being dark
Data goes dark when it has no address. The fix is not another storage system. It is a model that gives every value a place someone can navigate to.
Give every value a place in the model
When time-series, events and files attach to assets and processes in one connected model, what do we know about this pump becomes a query instead of an archaeology project.
Make finding cheaper than measuring again
A searchable model with lineage means the first move is to look, not to install another sensor. Most dark data is rediscovered the first time someone can actually browse it.
Let agents read the archive
AI agents are tireless readers. Pointed at a connected model, they can sweep the history nobody had time for and surface what deserves a human look.
Common questions
Asked and answered
What is dark data?
Dark data is data an organisation collects and stores as part of normal operations and then never uses for analysis or decisions. Industry surveys estimate that roughly half of all stored operational data is dark, kept but never read.
Why is dark data a problem?
It costs storage and carries retention, privacy and security obligations while returning nothing, and its invisibility hides answers that teams then measure again. The data a decision needed often already exists, unread.
How do you reduce dark data?
Give the data an address. When values attach to assets and processes in a connected model with search and lineage, finding becomes cheaper than collecting again, and the dark share shrinks every time somebody looks.
The rest of the split
Where the other data points go
The drain has five outlets, and each one fails an operation in its own way. This page covers one of them. The other four are worth ten minutes each.