Data Academy · Sources · Connect

Enterprise Sources

Data can remain anywhere: SCIKIQ connects in place and federates on demand.

Lesson 1 of 17 · Find it on the platform map

01 What it is

What this layer does

Enterprise Sources is the layer that reaches the systems where the business already keeps its data: applications, databases, files, BI tools, APIs, streams and messages. In SCIKIQ the data can stay where it is; the platform connects in place and federates queries on demand rather than copying everything into a new store first.

02 Concepts

Four ideas to hold on to

1

Connect in place

Reading data from the system that owns it instead of first moving it into a central store. It reduces duplicate copies and keeps the owning system as the point of truth.

2

Federation

Answering a query by reaching into several sources at the time of asking and combining the results. It trades some query cost for freshness and less data movement.

3

Unstructured sources

Documents, PDFs, email, chat and web pages that hold policy, contracts and know-how. Extraction and OCR turn them into content a platform can index and reason over.

4

Shadow spreadsheets

Business logic that lives in spreadsheets feeding BI models rather than in governed systems. Recovering them from BI models brings hidden calculations and reference lists into view.

03 In SCIKIQ

What the layer contains

Enterprise applications (SAP, Oracle, Salesforce, Workday, ServiceNow, Dynamics)Databases and data platforms (Snowflake, Databricks, BigQuery, AWS, Azure, SQL Server)Data-lake catalogs (Snowflake, Databricks)Files and documents (SharePoint, Drive, OneDrive, Box, Office, PDF)Document extraction (PDF layout, Office, email, OCR)Cloud document stores (S3, Azure Blob, GCS)BI and analytics (Power BI, Tableau, Looker, Qlik, MicroStrategy, ThoughtSpot)Shadow spreadsheets recovered from BI modelsAPIs and services (REST, GraphQL, OpenAPI)Streaming and IoT (Kafka, MQTT, PLM, MES, WMS, QMS)Other sources (email, Teams, Slack, web)Document crawlers (SharePoint, OneDrive, Drive, web)

04 What good looks like

Signs it is working

  • Every connected source has a named owner and a recorded connection method.
  • A new question can be answered from an existing source without building a new copy of its data.
  • Documents and file stores are crawled on a schedule, and the date of the last crawl is visible.
  • Spreadsheets feeding BI reports are listed alongside the reports that depend on them.

05 Diagnostics

Questions to ask your team

  1. 1

    Which of our systems of record are connected today, and which are still reached through manual extracts?

  2. 2

    Where are we copying data that could be read in place?

  3. 3

    Which important decisions depend on spreadsheets or documents that no platform can see?

  4. 4

    Who owns each source connection, and what happens when that system changes?

06 Keep going

Related reading

— Questions

Frequently asked

Does SCIKIQ require us to move our data into a new warehouse?

No. The layer connects in place and federates on demand, so data can remain in the applications, databases and stores where it already lives.

What kinds of sources can be connected?

Enterprise applications, databases and data platforms, data-lake catalogs, files and cloud document stores, BI tools, APIs, streaming and IoT feeds, and other sources such as email, Teams, Slack and the web.

How are PDFs and scanned documents handled?

Document extraction reads PDF layout, Office files and email, and applies OCR to scanned content, while document crawlers cover SharePoint, OneDrive, Drive and the web.