Implementation guide
How an enterprise data hub implementation actually runs
Most data platform programmes fail on sequencing, not technology. This is the order the work goes in, what each stage produces, what typically goes wrong, and what a realistic 60–90 day timeline looks like — including the parts that depend on you rather than on the platform.
The principle that decides the timeline
Traditional data programmes are sequenced consolidation-first: choose a target platform, migrate sources into it, model the data, then finally show the business something. Value arrives at the end, which is why so many of these programmes are cancelled before it does.
A hub implementation inverts that. Because the hub governs data where it already lives, there is no migration milestone standing between kick-off and a working view. The business sees a governed 360 view in roughly the first 60 days — the thing the traditional model delivers only at the end — and the estate expands from there.
That inversion is the whole reason the timeline is measured in weeks. Everything below follows from it.
Stage 1 — Scope one domain, not the estate (week 1)
The single largest predictor of failure is starting with “all of it”. Pick one domain where the pain is measurable and the owner is identifiable: a Customer 360 where marketing and billing disagree, a Finance 360 where the close is slow, a Factory 360 where OEE is not comparable between plants.
What this stage produces: a named business owner, three to five metrics that will be considered proof, and the list of source systems those metrics genuinely depend on.
What goes wrong: scoping by system rather than by question. “Connect SAP” is not a scope; “landed cost, calculated the same way on every lane” is. The first produces an integration project with no finish line, the second produces a test you can pass.
Stage 2 — Connect the sources (weeks 1–3)
Connection is no-code: point the platform at the source and authenticate. That covers on-premise databases, cloud warehouses, SaaS applications, APIs and flat files, across 187+ connectors, without replication and without a pipeline team.
SAP is the exception worth planning for — not because it is hard here, but because it is where most estates have historically stalled. Extraction from S/4HANA, BW/BW4HANA and ECC is non-invasive and requires no ABAP development or specialist consultants, which removes the dependency that usually sets the timeline. With SAP ECC mainstream maintenance ending in 2027, this is also the stage where most organisations discover how much of their operational truth is locked in a system they are about to have to think about anyway.
What this stage produces: live connections to every in-scope source, with metadata harvested and catalogued automatically.
What goes wrong: credentials and network access. This is genuinely the most common cause of slippage, and it is entirely on the customer’s side. Start the access requests in week one; a firewall change request that takes three weeks will cost you three weeks.
Stage 3 — Resolve entities (weeks 3–6)
This is the stage that produces the value, and the stage people underestimate. The same customer exists five times under four spellings. The same supplier is a different legal entity in each ERP. Matching and deduplication run across sources to produce one golden record per entity, with lineage showing which system contributed each attribute, so any match can be inspected and overruled.
What this stage produces: golden records for the domain’s core entities, and the first honest count of how badly the systems actually disagreed.
What goes wrong: nobody is empowered to arbitrate. When two systems disagree about a customer’s legal name, that is a business decision, not a technical one. If there is no steward with the authority to decide, resolution stalls at 95% and stays there.
Stage 4 — Govern it (weeks 4–8, overlapping)
Governance runs alongside resolution rather than after it, because the definitions are being argued about anyway. Six things get established: a business glossary, policy-as-code, end-to-end lineage, quality rules with trust scores, privacy and compliance guardrails mapped to the frameworks you are subject to, and stewardship with role-based access — including least-privilege credentials for service accounts and AI agents, not only people.
What this stage produces: metric definitions the business has signed off, and audit-ready evidence produced continuously rather than reconstructed at audit time.
What goes wrong: treating governance as a documentation exercise. A glossary nobody queries through is a spreadsheet with better formatting. It matters because the copilot and agents resolve questions against these definitions — governance is load-bearing, not paperwork.
Stage 5 — Prove it against what the team already trusts (weeks 6–9)
Before anyone builds on the governed layer, run the three to five metrics from stage 1 and compare them to the numbers the business currently uses. Expect differences. Every difference is either a bug in the hub or a bug that was already in the old number — and in our experience it is more often the second, which is uncomfortable and is exactly the point.
What this stage produces: reconciled numbers, and the credibility that everything afterwards depends on.
What goes wrong: skipping it to save two weeks. A governed layer the business has not personally checked will be quietly ignored, and no amount of subsequent capability recovers that.
Stage 6 — Activate, then hand over (weeks 8–12)
Only now does activation make sense: dashboards, data products, the plain-language copilot, and agents that act within defined limits with a human in the loop for anything irreversible and every action logged.
Handover is the part vendors skip and the part that determines whether phase two happens. The objective is that your team runs the next domain without us.
What goes wrong: a successful phase one that is entirely dependent on the implementation partner. If your stewards cannot add a source and define a metric unaided at the end of phase one, the deployment has not actually landed.
What this looked like at EFS
EFS, a facilities management group, ran this sequence across an estate of 11 disconnected systems. The hub unified 6.5 million records spanning more than 6,100 attributes, with over 110,000 transactions flowing daily through the governed layer.
The detail that matters most for anyone reading an implementation guide is the last one: phase two was run in-house by EFS’s own team. That is what a completed handover looks like.
Their CIO summarised the change as “from siloed to connected, from reactive to proactive”.
For a larger estate, ECU Worldwide unified 30+ systems across 300+ offices, 2,400 trade routes and roughly 40,000 port pairs into one governed view — work their CIO described as what “three consulting firms couldn’t solve in two years, SCIKIQ delivered in three months”.
What you need to provide
An honest implementation guide has to include the customer’s side of the work, because that is where timelines actually slip.
- A named business owner for the domain, with the authority to sign off metric definitions.
- Source system access — credentials, network routes and any firewall changes. Start in week one.
- One or two data stewards who know why the systems disagree. This knowledge usually lives with long-serving staff, and capturing it in the glossary is part of the value.
- A reconciliation session at stage 5 with the people who own the current numbers.
- Roughly a day a week of that group’s time. Not a full-time programme team.
Realistic timeline expectations
A live 360 view for one domain in about 60 days. A full hub live in 60–90 days. Each additional domain after the first is faster, because the connections, glossary and stewardship model already exist — the model is additive rather than restarted.
What will not happen in 60 days: every system in a large enterprise connected, or an organisation-wide change in how people make decisions. The first is a sequencing choice, the second is a cultural one, and neither is a platform capability.
Explore further
- Enterprise data hub — the platform this guide implements
- Data hub vs data fabric vs data lakehouse — choosing the right pattern first
- SAP data integration — no-code, non-invasive extraction
- Data Governance — the six foundations established in stage 4
- Impact studies — eight deployments, and what changed
Scope it against your own estate
A 30-minute session on your sources and your first domain — not a slide deck.
Frequently asked questions
How long does a data hub implementation take?
A live 360 view for one domain in about 60 days, and a full hub live in 60–90 days. Each domain after the first is faster, because the connections, glossary and stewardship model already exist — the model is additive rather than restarted.
What are the stages of a data hub implementation?
Six: scope one domain rather than the estate, connect the sources no-code, resolve entities into golden records, govern them with glossary, lineage, quality and policy, prove the metrics against numbers the business already trusts, then activate and hand over so your own team can run the next domain.
What causes data hub projects to slip?
Most often credentials and network access, which sit on the customer side — start those requests in week one. After that: scoping by system instead of by question, having nobody empowered to arbitrate when two systems disagree about an entity, and skipping the reconciliation stage to save two weeks.
What do we need to provide?
A named business owner with authority to sign off metric definitions, source system access, one or two data stewards who know why the systems disagree, and a reconciliation session with the people who own the current numbers — roughly a day a week of that group’s time, not a full-time programme team.