Industry examples
DataHub has no vertical editions. The same four building blocks and the same platform serve every industry, the difference between deployments is the model you load into it, and how much of one you need on day one.
Five deployments that share nothing but the platform. The first is worked through in detail, because it shows the whole arc, from a number nobody can defend, to a measured one, to a process that costs less and discharges less.
Oil and gas: cutting chemical use at a processing facility
The clearest example of what a model buys you, because the same work delivers an environmental result and a financial one from a single change, and because it starts from a number almost everybody reports and almost nobody can substantiate.
The situation
An oil and gas processing facility runs on chemicals. They are injected continuously, all over the process, each doing a specific job:
| Chemical | Injected into | Doing what |
|---|---|---|
| Demulsifier | The separator | Breaking the oil–water emulsion so the two separate |
| Scale inhibitor | The gas train | Stopping mineral scale building up in pipework |
| Biocide | Produced water | Keeping bacteria out of the system |
| Glycol | Subsea lines | Preventing hydrate ice forming and blocking flow |
| Oxygen scavenger | Water systems | Removing dissolved oxygen that would corrode steel |
Every one of these ends up somewhere, discharged to sea, emitted to air, or retained in the product. All of it must be reported.
Why today's reported number is not really a measurement
Here is the part that surprises people outside the industry: discharge figures are usually built from what was purchased, not from what was actually used.
Tank deliveries are invoiced, so procurement records exist and are accurate. What happens between the tank and the sea is estimated, dose rates assumed from design values, run hours assumed from schedules, losses assumed from convention. The resulting figure is defensible as an accounting exercise and is not a measurement of anything.
Two consequences follow, and they point in opposite directions:
- You cannot prove the number is right, which is a compliance exposure that grows as reporting obligations tighten.
- You cannot find the waste, because a figure derived from purchase records is structurally incapable of showing that a particular pump has been over-dosing for a year.
The environmental classification makes it sharper
Offshore chemicals are not interchangeable from an environmental standpoint. They are classified by hazard, conventionally by colour:
So "how much chemical did we discharge" is really four questions, and the red and black ones carry far more weight than the total. An aggregate figure built from purchase records cannot answer them separately, which means the number that matters most is the one you have least evidence for.
What gets modelled
| Layer | Modelled as |
|---|---|
| Assets | Chemical tanks, dosing pumps, injection valves, the separator, the gas train, produced-water system, metering skid |
| Functions | Separation, gas treatment, produced-water treatment, reinjection, each chemical's injection duty |
| Business knowledge | Discharge permits and their per-chemical limits, the environmental classification, substitution obligations, the reporting cycle |
| Time series | Tank levels, pump strokes, valve positions, flow rates, separator performance, water quality |
| Events | Dose changes, pump faults, tank deliveries, permit periods, sampling |
The critical relationships are the ones connecting a dosing pump to the part of the process it injects into, and that process to the discharge route and the permit it falls under. Those relationships are exactly what nobody writes down today, and they are what turns a pump stroke into a reportable discharge figure.
Measuring what cannot be measured directly
One obstacle: some of the values you most want are impractical to measure continuously. Oil content in water sent for reinjection is the classic case: it is a slow laboratory test, so the process is corrected after the fact rather than kept in specification.
A soft sensor closes that gap. It is a virtual meter: it watches the cheap signals you already have, pressure, flow, temperature, and a model trained against historical lab samples infers the value you actually want, second by second. You get a live reading where before there was a delayed lab result. Today that computation runs beside the platform, against the same APIs, fed by subscription; running it in-platform is exactly what functions are being built for.
That only works if the model knows which signals belong to which vessel, which lab samples correspond to which stream, and what the process configuration was at the time. In other words, the soft sensor is downstream of the knowledge graph, which is why this is not simply a sensor purchase.
What changes
The metering, the model and the reproducible figure work today. The traceability that makes a submission self-evidencing, here and in the grid and water examples below, depends on lineage, which is planned.
Tank levels, pump strokes, valve positions and the separator's response are fused into a continuous account of how much of each chemical entered each part of the process, and how much left, to air or water.
Discharge per chemical and per classification, traceable through lineage to metered injection and measured flows, with data-quality flags where a reading was gap-filled. The submission carries its own evidence instead of resting on purchase records.
Once you can see dose against process response, the over-dosing becomes visible. The target is the smallest dose that still does the job, which is almost never the design dose that has been running unchanged for years.
With per-chemical usage measured, a proposal to move a duty from a red chemical to a yellow one can be argued from data, how much is actually used, where, and what the process response has been.
Why the return is unusually good
Most efficiency projects trade one benefit against another. This one compounds in three directions at once:
The third one is easy to overlook and often the largest. Chemical treatment exists to protect steel; getting it right adds years to pipes and equipment, and that shows up as deferred capital rather than as a line in an operating budget.
Why it needs the model
Every step above depends on knowing which pump feeds which process, which process discharges by which route, and which permit that route falls under. That is not data any single source system holds, the tank levels are in one place, the pump telemetry in another, the permits in a document, and the connection between them in somebody's head.
That connection is the model. Build it, and the environmental report, the cost reduction and the equipment-life benefit all fall out of the same work. Building your model →
This sector is also where the model stops being only an analytical asset. Offshore, installations designed to run with no permanent crew are operated through the model itself, so the description built for a discharge report becomes, eventually, the operating surface. Where this is heading →
Starting play: audit-ready reporting on the discharge submission, it already exists, already costs effort, and already cannot be substantiated, which makes it an unusually easy case to fund.
The rest of the sector
The same modelling approach carries across oil and gas generally:
| Layer | Modelled as |
|---|---|
| Assets | Wells, separators, compressors, pipelines, valves, instruments |
| Functions | Separation, compression, export, flaring |
| Business knowledge | Production targets, emissions obligations, integrity policies |
| Time series | Pressure, temperature, flow, vibration |
| Events | Alarms, work permits, isolations, interventions |
This sector has the strongest head start, because the naming is already standardised, ISA-5.1 instrument tags, NORSOK Z-DP-002 coding, IEC/ISO 81346 reference designations, CFIHOS handover classes. A large part of the model is already specified and in daily use; DataHub mirrors it rather than replacing it. Standards →
If a capital project is in flight, capital project handover is the cheapest complete model you will ever get.
A grid operator
| Layer | Modelled as |
|---|---|
| Assets | Generating units, substations, battery storage, wind farms |
| Functions | Balancing, delivery, load management |
| Business knowledge | Tariffs, power purchase agreements, regulated reporting obligations |
| Time series | SCADA telemetry |
| Events | Grid events, trips, switching operations |
What it makes possible. "Which assets contributed to yesterday's capacity shortfall, and what events were raised against them?" becomes one query. The hourly emissions figure submitted to the regulator is reproducible today, and once lineage ships it traces back to individual generating units with data-quality flags attached, an auditable artefact rather than a claim.
Starting play: audit-ready reporting, because the regulated submission already exists and already costs real effort.
A municipal water utility
| Layer | Modelled as |
|---|---|
| Assets | Treatment plants, pumping stations, network segments |
| Functions | Treatment stages, delivery zones |
| Business knowledge | The regulatory framework each site reports under |
| Time series | Flow, pressure, quality parameters |
| Events | Excursions, maintenance, sampling |
What it makes possible. When an effluent reading goes out of spec, the team traverses from the excursion to the contributing treatment stages to the upstream events and maintenance history. Compliance submissions stop being spreadsheets assembled by hand and become queries against the model, each number reproducible, and carrying its full audit trail once lineage lands.
Starting play: faster investigations, then audit-ready reporting once the model covers the treatment train.
A manufacturer
| Layer | Modelled as |
|---|---|
| Assets | Lines, stations, machines, quality instruments |
| Functions | Production steps, quality gates, changeovers |
| Business knowledge | OEE targets, product specifications, customer commitments |
| Time series | Throughput, cycle time, quality signals, energy per unit |
| Events | Stoppages, defects, changeovers, maintenance |
What it makes possible. Traceability from a defect back through the stations, settings and material that produced it. Energy per unit compared across lines, with the differences attributable to specific equipment rather than assumed. Downtime attributed to cause rather than to a category somebody picked from a dropdown.
Starting play: data liberation across the MES, maintenance system and energy meters, the fastest route to a visible win on a shop floor.
An IT operations team
| Layer | Modelled as |
|---|---|
| Assets | Servers, storage arrays, switches, virtual machines |
| Functions | The services and pipelines they host |
| Business knowledge | Service levels, business criticality, ownership |
| Time series | Latency, utilisation, error rates |
| Events | Deploys, alerts, incidents |
What it makes possible. When a service degrades, the investigation traverses the graph: service, to host, to the array showing elevated error rates, to the switch that dropped packets, to the deploy that went out twenty minutes earlier. "What is the blast radius of taking down this switch?" is a relationship query rather than tribal knowledge.
This one is not hypothetical, IntelliStream runs DataHub on its own infrastructure this way.
Starting play: faster investigations. The data is already accessible and the relationships are already known, so the model comes together unusually fast.
A plant that just wants its own data back
Not every deployment starts with an ontology. Sometimes the value is plain data liberation.
The data sits in a historian, a maintenance system, a work-permit system and an ERP, and each guards it behind a data model that takes vendor training to query. DataHub liberates it by simplifying it: whatever arrives from a silo becomes one of a few primitives. A work permit opening or closing, a new purchase order, a sensor alarm, a system state change, all become events. Signals become time series. The things they describe become resources.
You never need to understand a source system's schema to use its data, and there is one interface for querying and reading all of it. Lineage, the knowledge graph and everything above can be layered on later. Having every silo readable through one simple model is worth doing by itself, and because the platform is open source, the liberated data has not simply moved into a newer cage.
The common thread
Same platform, same building blocks, no vertical editions. What changes between these five is the vocabulary in the model, and that vocabulary comes from your own people rather than from a taxonomy anybody imposed on you. Each of the five, grown to whatever size its first question needed, is a digital twin of a different operation.
The starting plays differ, but the sequence rarely does: liberate the data, model one area, answer one question, then widen.
- Value paths: the six plays referenced above, in detail
- The three layers: the structure every example follows
- Where to start: turning an example into a 90-day plan
- Building your model: producing your own version of these tables