Skip to main content

Industry examples

LeadershipDomain expertsEngineers
In one minute

DataHub has no vertical editions. The same four building blocks and the same platform serve every industry, the difference between deployments is the model you load into it, and how much of one you need on day one.

Five deployments that share nothing but the platform. The first is worked through in detail, because it shows the whole arc, from a number nobody can defend, to a measured one, to a process that costs less and discharges less.

Oil and gas: cutting chemical use at a processing facility

The clearest example of what a model buys you, because the same work delivers an environmental result and a financial one from a single change, and because it starts from a number almost everybody reports and almost nobody can substantiate.

The situation

An oil and gas processing facility runs on chemicals. They are injected continuously, all over the process, each doing a specific job:

ChemicalInjected intoDoing what
DemulsifierThe separatorBreaking the oil–water emulsion so the two separate
Scale inhibitorThe gas trainStopping mineral scale building up in pipework
BiocideProduced waterKeeping bacteria out of the system
GlycolSubsea linesPreventing hydrate ice forming and blocking flow
Oxygen scavengerWater systemsRemoving dissolved oxygen that would corrode steel

Every one of these ends up somewhere, discharged to sea, emitted to air, or retained in the product. All of it must be reported.

Why today's reported number is not really a measurement

Here is the part that surprises people outside the industry: discharge figures are usually built from what was purchased, not from what was actually used.

Tank deliveries are invoiced, so procurement records exist and are accurate. What happens between the tank and the sea is estimated, dose rates assumed from design values, run hours assumed from schedules, losses assumed from convention. The resulting figure is defensible as an accounting exercise and is not a measurement of anything.

Two consequences follow, and they point in opposite directions:

  • You cannot prove the number is right, which is a compliance exposure that grows as reporting obligations tighten.
  • You cannot find the waste, because a figure derived from purchase records is structurally incapable of showing that a particular pump has been over-dosing for a year.

The environmental classification makes it sharper

Offshore chemicals are not interchangeable from an environmental standpoint. They are classified by hazard, conventionally by colour:

Green
Considered to pose little or no risk, glycol, for example
Yellow
In use and permitted, but tracked, scale inhibitors
Red
Substitution actively expected, biocides
Black
Discharge prohibited outright; use only under specific permit

So "how much chemical did we discharge" is really four questions, and the red and black ones carry far more weight than the total. An aggregate figure built from purchase records cannot answer them separately, which means the number that matters most is the one you have least evidence for.

What gets modelled

LayerModelled as
AssetsChemical tanks, dosing pumps, injection valves, the separator, the gas train, produced-water system, metering skid
FunctionsSeparation, gas treatment, produced-water treatment, reinjection, each chemical's injection duty
Business knowledgeDischarge permits and their per-chemical limits, the environmental classification, substitution obligations, the reporting cycle
Time seriesTank levels, pump strokes, valve positions, flow rates, separator performance, water quality
EventsDose changes, pump faults, tank deliveries, permit periods, sampling

The critical relationships are the ones connecting a dosing pump to the part of the process it injects into, and that process to the discharge route and the permit it falls under. Those relationships are exactly what nobody writes down today, and they are what turns a pump stroke into a reportable discharge figure.

Measuring what cannot be measured directly

One obstacle: some of the values you most want are impractical to measure continuously. Oil content in water sent for reinjection is the classic case: it is a slow laboratory test, so the process is corrected after the fact rather than kept in specification.

A soft sensor closes that gap. It is a virtual meter: it watches the cheap signals you already have, pressure, flow, temperature, and a model trained against historical lab samples infers the value you actually want, second by second. You get a live reading where before there was a delayed lab result. Today that computation runs beside the platform, against the same APIs, fed by subscription; running it in-platform is exactly what functions are being built for.

That only works if the model knows which signals belong to which vessel, which lab samples correspond to which stream, and what the process configuration was at the time. In other words, the soft sensor is downstream of the knowledge graph, which is why this is not simply a sensor purchase.

What changes

On the roadmap

The metering, the model and the reproducible figure work today. The traceability that makes a submission self-evidencing, here and in the grid and water examples below, depends on lineage, which is planned.

Injection becomes metered rather than assumed

Tank levels, pump strokes, valve positions and the separator's response are fused into a continuous account of how much of each chemical entered each part of the process, and how much left, to air or water.

The reported figure becomes a measurement

Discharge per chemical and per classification, traceable through lineage to metered injection and measured flows, with data-quality flags where a reading was gap-filled. The submission carries its own evidence instead of resting on purchase records.

Doses get trimmed to what actually works

Once you can see dose against process response, the over-dosing becomes visible. The target is the smallest dose that still does the job, which is almost never the design dose that has been running unchanged for years.

Substitution becomes evidence-based

With per-chemical usage measured, a proposal to move a duty from a red chemical to a yellow one can be argued from data, how much is actually used, where, and what the process response has been.

Why the return is unusually good

Most efficiency projects trade one benefit against another. This one compounds in three directions at once:

Less
Discharged to sea and air
A real environmental reduction, weighted toward the chemicals that matter most
Less
Chemical purchased
Dosing to what works rather than to a design assumption cuts consumption directly
Longer
Asset life
Correct dosing protects pipework and equipment rather than over- or under-treating it
Defensible
Reported figures
A measurement with an audit trail, instead of an estimate from invoices

The third one is easy to overlook and often the largest. Chemical treatment exists to protect steel; getting it right adds years to pipes and equipment, and that shows up as deferred capital rather than as a line in an operating budget.

Why it needs the model

Every step above depends on knowing which pump feeds which process, which process discharges by which route, and which permit that route falls under. That is not data any single source system holds, the tank levels are in one place, the pump telemetry in another, the permits in a document, and the connection between them in somebody's head.

That connection is the model. Build it, and the environmental report, the cost reduction and the equipment-life benefit all fall out of the same work. Building your model →

This sector is also where the model stops being only an analytical asset. Offshore, installations designed to run with no permanent crew are operated through the model itself, so the description built for a discharge report becomes, eventually, the operating surface. Where this is heading →

Starting play: audit-ready reporting on the discharge submission, it already exists, already costs effort, and already cannot be substantiated, which makes it an unusually easy case to fund.

The rest of the sector

The same modelling approach carries across oil and gas generally:

LayerModelled as
AssetsWells, separators, compressors, pipelines, valves, instruments
FunctionsSeparation, compression, export, flaring
Business knowledgeProduction targets, emissions obligations, integrity policies
Time seriesPressure, temperature, flow, vibration
EventsAlarms, work permits, isolations, interventions

This sector has the strongest head start, because the naming is already standardised, ISA-5.1 instrument tags, NORSOK Z-DP-002 coding, IEC/ISO 81346 reference designations, CFIHOS handover classes. A large part of the model is already specified and in daily use; DataHub mirrors it rather than replacing it. Standards →

If a capital project is in flight, capital project handover is the cheapest complete model you will ever get.


A grid operator

LayerModelled as
AssetsGenerating units, substations, battery storage, wind farms
FunctionsBalancing, delivery, load management
Business knowledgeTariffs, power purchase agreements, regulated reporting obligations
Time seriesSCADA telemetry
EventsGrid events, trips, switching operations

What it makes possible. "Which assets contributed to yesterday's capacity shortfall, and what events were raised against them?" becomes one query. The hourly emissions figure submitted to the regulator is reproducible today, and once lineage ships it traces back to individual generating units with data-quality flags attached, an auditable artefact rather than a claim.

Starting play: audit-ready reporting, because the regulated submission already exists and already costs real effort.


A municipal water utility

LayerModelled as
AssetsTreatment plants, pumping stations, network segments
FunctionsTreatment stages, delivery zones
Business knowledgeThe regulatory framework each site reports under
Time seriesFlow, pressure, quality parameters
EventsExcursions, maintenance, sampling

What it makes possible. When an effluent reading goes out of spec, the team traverses from the excursion to the contributing treatment stages to the upstream events and maintenance history. Compliance submissions stop being spreadsheets assembled by hand and become queries against the model, each number reproducible, and carrying its full audit trail once lineage lands.

Starting play: faster investigations, then audit-ready reporting once the model covers the treatment train.


A manufacturer

LayerModelled as
AssetsLines, stations, machines, quality instruments
FunctionsProduction steps, quality gates, changeovers
Business knowledgeOEE targets, product specifications, customer commitments
Time seriesThroughput, cycle time, quality signals, energy per unit
EventsStoppages, defects, changeovers, maintenance

What it makes possible. Traceability from a defect back through the stations, settings and material that produced it. Energy per unit compared across lines, with the differences attributable to specific equipment rather than assumed. Downtime attributed to cause rather than to a category somebody picked from a dropdown.

Starting play: data liberation across the MES, maintenance system and energy meters, the fastest route to a visible win on a shop floor.


An IT operations team

LayerModelled as
AssetsServers, storage arrays, switches, virtual machines
FunctionsThe services and pipelines they host
Business knowledgeService levels, business criticality, ownership
Time seriesLatency, utilisation, error rates
EventsDeploys, alerts, incidents

What it makes possible. When a service degrades, the investigation traverses the graph: service, to host, to the array showing elevated error rates, to the switch that dropped packets, to the deploy that went out twenty minutes earlier. "What is the blast radius of taking down this switch?" is a relationship query rather than tribal knowledge.

This one is not hypothetical, IntelliStream runs DataHub on its own infrastructure this way.

Starting play: faster investigations. The data is already accessible and the relationships are already known, so the model comes together unusually fast.


A plant that just wants its own data back

Not every deployment starts with an ontology. Sometimes the value is plain data liberation.

The data sits in a historian, a maintenance system, a work-permit system and an ERP, and each guards it behind a data model that takes vendor training to query. DataHub liberates it by simplifying it: whatever arrives from a silo becomes one of a few primitives. A work permit opening or closing, a new purchase order, a sensor alarm, a system state change, all become events. Signals become time series. The things they describe become resources.

You never need to understand a source system's schema to use its data, and there is one interface for querying and reading all of it. Lineage, the knowledge graph and everything above can be layered on later. Having every silo readable through one simple model is worth doing by itself, and because the platform is open source, the liberated data has not simply moved into a newer cage.


The common thread

Same platform, same building blocks, no vertical editions. What changes between these five is the vocabulary in the model, and that vocabulary comes from your own people rather than from a taxonomy anybody imposed on you. Each of the five, grown to whatever size its first question needed, is a digital twin of a different operation.

The starting plays differ, but the sequence rarely does: liberate the data, model one area, answer one question, then widen.

Go deeper