Virtual Data Fabric
Connect, federate and query data wherever it exists.
Lesson 10 of 17 · Find it on the platform map
01 What it is
What this layer does
A virtual data fabric connects to data where it already lives — ERP, CRM, warehouses, files, APIs and streams — and lets people and systems query it through one access layer instead of copying everything into a single store first. In SCIKIQ it combines connectors, query and API federation, pushdown optimisation and caching with incremental ingest, data quality rules and master data, so that what comes out the other side is a set of certified data products rather than raw extracts.
02 Concepts
Four ideas to hold on to
Federation
Querying several sources through one interface and combining the results, rather than moving all the data into one place first. SCIKIQ offers query and API federation across its connectors to SAP, Oracle, Salesforce, Snowflake, SQL databases, files, BI tools, IoT and streaming sources.
Pushdown and materialisation
Pushdown sends filtering and aggregation to the source system so only the needed rows travel; materialisation stores a result when repeated live queries would be too slow or costly. SCIKIQ pairs pushdown optimisation with caching and on-demand materialisation, guided by a query planner and tier advisor.
Change data capture
Capturing inserts, updates and deletes as they happen in a source, so downstream copies stay current without full reloads. SCIKIQ supports CDC and streaming alongside incremental ingest using upsert and merge.
Entity resolution and data products
Entity resolution decides when records in different systems describe the same customer, supplier or product; a data product is a curated, owned and documented dataset offered for reuse. In SCIKIQ, master data and entity resolution feed certified data products that can be shared through the data marketplace and data exchange.
03 In SCIKIQ
What the layer contains
04 What good looks like
Signs it is working
- A new source system can be connected and queried without building a bespoke pipeline for each report.
- Every certified data product has a named owner, documented quality rules and a stated refresh schedule.
- The same customer or supplier carries one resolved identity across ERP, CRM and finance data.
- For any heavy query the team can explain whether it runs live, from cache or from a materialised copy, and why.
05 Diagnostics
Questions to ask your team
- 1
Which of our data copies exist only because we had no way to query the source directly?
- 2
Who owns each data product we rely on, and what quality rules does it have to pass before it is certified?
- 3
How do we decide what to query live, what to cache and what to materialise?
- 4
When a customer appears in three systems with three different IDs, how do we know it is the same customer?
06 Keep going
Related reading
— Questions
Frequently asked
Does a virtual data fabric replace the data warehouse?
Not necessarily. It can query the warehouse as one source among many, and materialise data where performance needs it; the point is that copying becomes a deliberate choice rather than the default.
Is federation slower than querying a copy?
It can be for heavy, repeated workloads, which is why pushdown, caching and on-demand materialisation exist. The query planner and tier advisor help decide which approach suits each workload.
What makes a data product certified?
In practice it means the dataset has an owner, passes its agreed data quality rules, has resolved entities where relevant and is refreshed on a known schedule. Certification tells consumers they can build on it without re-checking it themselves.