Evaluating a data platform
Most data platform evaluations are decided by the demo, then justified by the scorecard. This is the other way round: what to ask, what proof to insist on, how to score answers so they can be compared, and the handful of things that only become a problem in year two. Written to be used against us as much as anyone else — a fair evaluation is the only kind worth winning.
01 Before you write an RFP
Decide what you are actually buying
The word "data platform" covers at least five different products. Buying the wrong category is far more expensive than picking the wrong vendor inside the right one, and it is the mistake nobody catches until the second year.
Storage and compute
Warehouses and lakehouses. They hold data and run queries fast. They do not, on their own, tell you what anything means or who owns it.
Right answer when: query cost and scale are the problemMovement
Ingestion, ELT and CDC tools. They get data from A to B reliably. Meaning, ownership and quality remain entirely yours.
Right answer when: pipelines are the bottleneck · vs FivetranCatalogue and governance
Metadata, glossaries, policies, lineage. They describe your estate. Whether anything changes depends on whether the descriptions are enforced anywhere.
Right answer when: audit and compliance drive the case · vs CollibraSemantic and activation layer
One agreed meaning per number, served to every consumer — dashboards, applications and agents — over data that stays where it is.
Right answer when: two teams get two numbers · Enterprise Data HubWrite down which of these your problem actually is, in one sentence, before anyone books a demo. If the sentence is "our reports disagree and nobody trusts them", more storage and faster pipelines will not fix it. If it is "our queries cost too much", a semantic layer will not fix that either.
02 The question set
Thirty questions, in six categories
Send these in writing, before the demo. The answers you get in writing are the ones you can hold people to, and the pattern of what a vendor answers precisely versus vaguely is itself information.
1. Connectivity — can it reach what you have?
- Which of our named systems do you connect to today, in production, at a customer of our size? Not "we support SAP" — which SAP modules, which extraction method, and who has it live?
- What does connecting a new source require: your engineers, our engineers, or a business user? How long, measured from access granted to first governed dataset?
- Does data have to move for the platform to work? If yes, where does it land, who owns that copy, and what happens to it if we leave?
- How do you handle on-premise and legacy systems that will not be replaced in this decade?
- What happens when an upstream schema changes without warning? Show us the alert, not the architecture diagram.
2. Meaning — will two teams get the same number?
- Where does a metric definition live, and which consumers read it? If it lives in each dashboard, you are buying storage, not agreement.
- Can a definition be version controlled and reviewed like code, with an owner and a change history?
- How do you resolve the same customer appearing in five systems with four spellings and three identifiers?
- Can the business see the definition next to the number, in the tool they already use?
- What happens to historical reports when a definition changes — are they restated, versioned, or silently different?
3. Governance — who owns it and can you prove it?
- Show us lineage for one figure, end to end, from a report back to the source entry — using our data, in the trial.
- Is lineage captured automatically from the pipelines, or drawn by hand? Hand-drawn lineage is wrong within a month.
- How are quality rules defined, where do they run, and what happens when one fails at 2am?
- How is access control applied — per table, per row, per column? Does it extend to service accounts and AI agents?
- Which frameworks do you produce evidence for, and can we see a sample export rather than a logo slide?
4. Time to value — how long until something is real?
- What is live at 30, 60 and 90 days — named deliverables, not phases?
- What must we do before you can start, and who on our side is needed, for how long?
- Can we start with one domain and extend, or does the model require the whole estate first?
- Show a reference of similar size and shape — and let us talk to them without you in the room.
- What did the last three implementations get wrong, and what changed as a result?
5. AI — grounded, or guessing?
- When someone asks a question in natural language, does the answer come from governed definitions or from generated SQL against raw tables?
- Can the answer cite its sources, hop by hop, in a way an auditor would accept?
- Do agents inherit our access control, or do they run with broad service credentials?
- Where do our prompts and data go, which model providers are involved, and can we choose or self-host?
- What happens when the model is confidently wrong — what stops it reaching a decision-maker?
6. Commercials and exit — the year-two questions
- What drives the price — users, sources, volume, compute? Model our year three, not our year one.
- What is not included: connectors, environments, support tiers, professional services days?
- If we double our data volume, what happens to the bill?
- If we leave, what do we keep — the data, the models, the definitions, the lineage? In what format?
- Which parts are open standards and which are proprietary? Where exactly is the lock-in, since there always is some?
03 Scoring
A sheet that survives a disagreement
Weight the categories before you see any answers, and agree the weights with the people who will live with the decision. Doing it afterwards produces a sheet that confirms whatever the room already decided.
| Category | Typical weight | Score 1 | Score 3 | Score 5 |
|---|---|---|---|---|
| Connectivity | 20% | Our systems need custom work | Most supported, some engineering | All named systems live elsewhere, no-code |
| Meaning | 25% | Definitions live in dashboards | Catalogue, loosely enforced | One definition, read by every consumer |
| Governance | 20% | Documentation only | Rules and owners on key data | Automated lineage and quality, with evidence |
| Time to value | 15% | Phase one is discovery | Something live in a quarter | Named deliverable inside 60 days |
| AI grounding | 10% | Chat over raw tables | Grounded on some curated data | Answers from governed metrics, with citations |
| Commercials & exit | 10% | Opaque, usage-driven | Predictable with caveats | Predictable, and we keep our assets |
Score each category from the written answers first, then adjust only for what the trial actually demonstrated. A demo is a rehearsed performance; a trial on your own data is evidence.
04 Proof
Four things to insist on seeing
05 Red flags
Things that only hurt later
"Migrate everything first"
A plan whose first value arrives after a full migration is a plan whose value arrives after the sponsor has changed jobs. Insist on a first domain that stands alone.
Ask: what is live in 60 days without a migration?Governance that is a document
A glossary nothing enforces changes nobody's behaviour. If the definition is not read by the tools that produce numbers, it is a wiki page with a licence fee.
Ask: which consumers read this definition at query time?AI demos on clean sample data
Natural-language querying looks magical on a tidy demo schema and falls apart on a real estate with four customer tables. Ask for it on yours, including the messy source.
Ask: run that question against our data, nowPricing that moves with success
Per-query and per-row pricing means the better it works, the more it costs — and teams start avoiding the platform to control the bill.
Ask: model our year three at triple the volumeNo named owner on their side
If the implementation lead is unnamed until after signature, the people in the room are not the people you get.
Ask: who exactly, and what else are they on?Every answer is yes
Every platform is bad at something. A vendor who cannot name their own weak spot either does not know the product or is not being straight with you.
Ask: where are you the wrong choice?06 Next
Where SCIKIQ fits, and where it does not
SCIKIQ is a governed semantic and activation layer over sources that stay where they are. It is the right answer when your problem is that numbers disagree, nothing is traceable, and AI has nothing trustworthy to stand on. It is not a replacement for a warehouse when your problem is query cost at scale, and it is not a pipeline tool competing on connector count alone — those comparisons are on the pages below, written the same way as this one.
Bring the hardest question on this page
Trace one of your own board-pack figures back to source, live, on your data.