Data Academy · Layer 4 · Discover & ground

Enterprise Discovery Engine

The top-down logical model meets bottom-up metadata discovery.

Lesson 6 of 17 · Find it on the platform map

01 What it is

What this layer does

The discovery layer builds an inventory of what the enterprise holds and what it means, working from two directions at once. Top-down, it starts from the business model — industry, processes, domains, personas, concepts, metrics, taxonomy and glossary. Bottom-up, it scans applications, databases, APIs, BI tools, documents and IoT sources, and SCIKIQ reconciles the two into one ontology with technical metadata, business metadata, classification, lineage, usage, ownership and allowed actions.

02 Concepts

Four ideas to hold on to

1

Top-down and bottom-up discovery

Top-down discovery describes the business in its own terms; bottom-up discovery profiles what physically exists in systems. Neither is enough alone, so mature practice runs both and reconciles them.

2

Ontology reconciliation

The work of matching business concepts to discovered assets and resolving conflicts, such as two systems that both claim to hold the customer. SCIKIQ treats this as an explicit step, not an afterthought.

3

Technical and business metadata

Technical metadata records types, keys, nulls and value ranges; business metadata records meaning, owner and criticality. Good catalogues keep both against the same asset.

4

Lineage

A record of where data comes from and how it is transformed on the way. SCIKIQ includes SQL lineage parsed from views and stored procedures, where much undocumented logic lives.

03 In SCIKIQ

What the layer contains

Top-down: industry, process, domain, persona, concept, metric, taxonomy, glossaryOntology reconciliationBottom-up: applications, databases, APIs, BI, documents, IoTLogical systems and flows (landscape)Technical metadata (types, keys, nulls, ranges)Business metadataAutomatic classificationLineage, including SQL lineage from views and proceduresUsage patternsOwner, criticality, AI suitability and allowed actionsSchema-change rediscovery, process mining and domain induction

04 What good looks like

Signs it is working

  • Every critical data asset has a named owner, a criticality rating and a stated AI suitability.
  • Lineage for a headline report can be traced back to source tables, including logic inside views and procedures.
  • A schema change in a source system triggers rediscovery rather than a silent break.
  • Each business concept in the glossary is linked to at least one discovered physical asset, and unlinked concepts are listed as gaps.

05 Diagnostics

Questions to ask your team

  1. 1

    Which of our systems have never been profiled, and what do we assume they contain?

  2. 2

    When a source schema changes, how and when do we find out?

  3. 3

    Who owns each critical data asset, and which actions is an AI agent allowed to take on it?

  4. 4

    Where do our business glossary and our physical systems disagree, and who resolves it?

06 Keep going

Related reading

— Questions

Frequently asked

Is discovery a one-off project?

No. Systems change constantly, so discovery has to repeat; SCIKIQ includes schema-change rediscovery, process mining and domain induction for that reason.

Why classify data automatically?

Manual classification does not keep pace with the number of tables and files in a large estate. Automatic classification gives a first pass that owners then confirm or correct.

What does usage data add?

Usage patterns show which assets people and reports actually rely on. That tells you where to focus quality and governance effort first.