Data Academy · Buyer's guide

Evaluating a data platform

Most data platform evaluations are decided by the demo, then justified by the scorecard. This is the other way round: what to ask, what proof to insist on, how to score answers so they can be compared, and the handful of things that only become a problem in year two. Written to be used against us as much as anyone else — a fair evaluation is the only kind worth winning.

Read time ~14 min Includes 30 questions & a scoring sheet For CIOs, architects, procurement

01 Before you write an RFP

Decide what you are actually buying

The word "data platform" covers at least five different products. Buying the wrong category is far more expensive than picking the wrong vendor inside the right one, and it is the mistake nobody catches until the second year.

Category

Storage and compute

Warehouses and lakehouses. They hold data and run queries fast. They do not, on their own, tell you what anything means or who owns it.

Right answer when: query cost and scale are the problem
Category

Movement

Ingestion, ELT and CDC tools. They get data from A to B reliably. Meaning, ownership and quality remain entirely yours.

Right answer when: pipelines are the bottleneck · vs Fivetran
Category

Catalogue and governance

Metadata, glossaries, policies, lineage. They describe your estate. Whether anything changes depends on whether the descriptions are enforced anywhere.

Right answer when: audit and compliance drive the case · vs Collibra
Category

Semantic and activation layer

One agreed meaning per number, served to every consumer — dashboards, applications and agents — over data that stays where it is.

Right answer when: two teams get two numbers · Enterprise Data Hub

Write down which of these your problem actually is, in one sentence, before anyone books a demo. If the sentence is "our reports disagree and nobody trusts them", more storage and faster pipelines will not fix it. If it is "our queries cost too much", a semantic layer will not fix that either.

02 The question set

Thirty questions, in six categories

Send these in writing, before the demo. The answers you get in writing are the ones you can hold people to, and the pattern of what a vendor answers precisely versus vaguely is itself information.

1. Connectivity — can it reach what you have?

  • Which of our named systems do you connect to today, in production, at a customer of our size? Not "we support SAP" — which SAP modules, which extraction method, and who has it live?
  • What does connecting a new source require: your engineers, our engineers, or a business user? How long, measured from access granted to first governed dataset?
  • Does data have to move for the platform to work? If yes, where does it land, who owns that copy, and what happens to it if we leave?
  • How do you handle on-premise and legacy systems that will not be replaced in this decade?
  • What happens when an upstream schema changes without warning? Show us the alert, not the architecture diagram.

2. Meaning — will two teams get the same number?

  • Where does a metric definition live, and which consumers read it? If it lives in each dashboard, you are buying storage, not agreement.
  • Can a definition be version controlled and reviewed like code, with an owner and a change history?
  • How do you resolve the same customer appearing in five systems with four spellings and three identifiers?
  • Can the business see the definition next to the number, in the tool they already use?
  • What happens to historical reports when a definition changes — are they restated, versioned, or silently different?

3. Governance — who owns it and can you prove it?

  • Show us lineage for one figure, end to end, from a report back to the source entry — using our data, in the trial.
  • Is lineage captured automatically from the pipelines, or drawn by hand? Hand-drawn lineage is wrong within a month.
  • How are quality rules defined, where do they run, and what happens when one fails at 2am?
  • How is access control applied — per table, per row, per column? Does it extend to service accounts and AI agents?
  • Which frameworks do you produce evidence for, and can we see a sample export rather than a logo slide?

4. Time to value — how long until something is real?

  • What is live at 30, 60 and 90 days — named deliverables, not phases?
  • What must we do before you can start, and who on our side is needed, for how long?
  • Can we start with one domain and extend, or does the model require the whole estate first?
  • Show a reference of similar size and shape — and let us talk to them without you in the room.
  • What did the last three implementations get wrong, and what changed as a result?

5. AI — grounded, or guessing?

  • When someone asks a question in natural language, does the answer come from governed definitions or from generated SQL against raw tables?
  • Can the answer cite its sources, hop by hop, in a way an auditor would accept?
  • Do agents inherit our access control, or do they run with broad service credentials?
  • Where do our prompts and data go, which model providers are involved, and can we choose or self-host?
  • What happens when the model is confidently wrong — what stops it reaching a decision-maker?

6. Commercials and exit — the year-two questions

  • What drives the price — users, sources, volume, compute? Model our year three, not our year one.
  • What is not included: connectors, environments, support tiers, professional services days?
  • If we double our data volume, what happens to the bill?
  • If we leave, what do we keep — the data, the models, the definitions, the lineage? In what format?
  • Which parts are open standards and which are proprietary? Where exactly is the lock-in, since there always is some?

03 Scoring

A sheet that survives a disagreement

Weight the categories before you see any answers, and agree the weights with the people who will live with the decision. Doing it afterwards produces a sheet that confirms whatever the room already decided.

CategoryTypical weightScore 1Score 3Score 5
Connectivity20%Our systems need custom workMost supported, some engineeringAll named systems live elsewhere, no-code
Meaning25%Definitions live in dashboardsCatalogue, loosely enforcedOne definition, read by every consumer
Governance20%Documentation onlyRules and owners on key dataAutomated lineage and quality, with evidence
Time to value15%Phase one is discoverySomething live in a quarterNamed deliverable inside 60 days
AI grounding10%Chat over raw tablesGrounded on some curated dataAnswers from governed metrics, with citations
Commercials & exit10%Opaque, usage-drivenPredictable with caveatsPredictable, and we keep our assets

Score each category from the written answers first, then adjust only for what the trial actually demonstrated. A demo is a rehearsed performance; a trial on your own data is evidence.

04 Proof

Four things to insist on seeing

Your data, your worst sourceNot the sample dataset. Pick the system everyone complains about and see it connected in the trial.
One number, traced to sourcePick a figure from your board pack and follow it back, hop by hop, in front of you.
A definition changed liveChange a metric definition and watch every consumer pick it up — or watch them not.
A reference call without the vendorAsk what went wrong and how long it really took. Everyone has a difficult month; the useful question is what happened next.

05 Red flags

Things that only hurt later

Flag

"Migrate everything first"

A plan whose first value arrives after a full migration is a plan whose value arrives after the sponsor has changed jobs. Insist on a first domain that stands alone.

Ask: what is live in 60 days without a migration?
Flag

Governance that is a document

A glossary nothing enforces changes nobody's behaviour. If the definition is not read by the tools that produce numbers, it is a wiki page with a licence fee.

Ask: which consumers read this definition at query time?
Flag

AI demos on clean sample data

Natural-language querying looks magical on a tidy demo schema and falls apart on a real estate with four customer tables. Ask for it on yours, including the messy source.

Ask: run that question against our data, now
Flag

Pricing that moves with success

Per-query and per-row pricing means the better it works, the more it costs — and teams start avoiding the platform to control the bill.

Ask: model our year three at triple the volume
Flag

No named owner on their side

If the implementation lead is unnamed until after signature, the people in the room are not the people you get.

Ask: who exactly, and what else are they on?
Flag

Every answer is yes

Every platform is bad at something. A vendor who cannot name their own weak spot either does not know the product or is not being straight with you.

Ask: where are you the wrong choice?