Master data decides the outcome
Consumer goods AI programmes fail on product, location and calendar resolution far more often than on modelling, because every downstream number depends on them.
A brand owner sells to a retailer and the shopper buys from that retailer. Everything difficult about this industry follows from that gap: your demand signal is someone else’s data, on their calendar, in their hierarchy, and the money you spend with them comes back as a deduction you have to argue about months later.
Part 1 maps a chain that ends before the purchase does, and names the people who plan against a signal they do not own.
Six stages from innovation to deduction. Select one.
Concept, formulation, packaging, claims and launch readiness.
Forecasting, S&OP, production planning and inventory across the network.
Production, quality, warehousing and distribution to customer.
Customer planning, trade terms, promotional calendars and joint business plans.
Shopper and category insight, distribution, share of shelf and digital presence.
Invoicing, deductions, claims validation and dispute recovery.
None of these people can see the shopper. All of them are judged on what the shopper did. Select one to light the stages they own.
An industry rich in data it does not own, and poor in the master data that would join it.
Consumer goods AI programmes fail on product, location and calendar resolution far more often than on modelling, because every downstream number depends on them.
Shipments measure what the retailer ordered; consumption measures what the shopper bought. Forecasting on the first is the sector’s most persistent error.
Every promotional claim depends on a counterfactual, and agreeing how the baseline is computed matters more than the modelling that follows it.
Trade deductions arrive with backup in every format, and validating them by hand costs more than many organisations recover.
Content compliance and share of search on retailer ecommerce now behave like physical distribution did, and are measurable in the same way.
Pricing, promotion, mix and trade terms are increasingly managed as one discipline, which requires them to be measured on one set of definitions.
Consulted when this program was written, in July 2026. Each entry names the claim on this page that it backs, so the pairing stays checkable if the copy is later edited.
The defining architectural problem is that the most important data belongs to somebody else, arrives late, and uses hierarchies that do not match yours. Reconciliation is not preparation here; it is the product.
Six families, and only three of them are yours.
Items, packs, hierarchies, attributes and their lifecycle.
PIM, SAP material masterCustomer orders, shipments and invoices — the sell-in signal.
ERP, order managementRetailer scan data, panel and market measurement — sell-out.
Market measurement providers, retailer portalsAgreements, promotional plans, accruals and funding.
TPM systems, trade agreementsCustomer deductions, claims, backup documentation and disputes.
Deduction management, ARProduction plans, capacity, inventory and distribution.
APS, WMS, ERPConsumer businesses fail at AI more often on master data than on modelling. If product, location and calendar do not resolve consistently, every number the agent quotes is a different number from the one the business reports.
Six subject areas, each with the entities it holds, its critical data elements and the function accountable for it.
What is sold, at every level from case to consumer unit.
Where things are made, held and sold.
Retail customers who buy from you, and the consumers who buy the product.
Forecasts, orders, shipments and point-of-sale movement.
Price lists, promotions, trade spend and the agreements behind them.
How the product appears and performs on retailer ecommerce.
The entities and the typed relationships between them — what the agent traverses instead of guessing joins. Select any entity.
The lowest-level sellable unit.
The marketed identity an item belongs to.
The classification the item competes within.
The retailer or account that buys from you.
A selling location.
A grouping of stores behaving alike.
Expected demand at an item-location-period grain.
A customer request for goods.
Goods dispatched against an order.
Units actually sold to the consumer.
A time-bound price or display mechanic.
The item as presented on a retailer site.
The classification hierarchies that make records comparable across systems.
Beverages → Carbonates → Cola → BrandX → 330ml can 6-pack
RetailCo → RetailCo Express → South → Store 4412
Price → Temporary price reduction → 20% off
The definitions an agent must use rather than invent. Most wrong answers in this industry are a term used loosely.
The layers a deployment needs, what each one holds, and what the engineer owns there.
Seven layers, with hierarchy reconciliation treated as a governed capability, not a script.
Your master data and shipments, plus syndicated, POS and retailer portal data you do not control.
Nothing. This is the upstream edge.
Sell-in and sell-out feeds mapped to one hierarchy, with mapping coverage stated on each.
Items, packs, hierarchies and attributes. Every downstream number is computed across this, which is why its defects propagate everywhere.
Customer orders, shipments and invoices — the sell-in signal, and the one you actually own.
Retailer scan data and market measurement: the sell-out signal, arriving late, restated, and in somebody else’s hierarchy.
Promotional plans, mechanics, funding and accruals against customer agreements.
Customer deductions with their backup documentation, in every format a customer has ever used.
Production plans, capacity, inventory and distribution across the network.
Getting external data on a reliable cadence, and knowing exactly how stale each feed is.
Passing a syndicated number downstream without its mapping coverage attached. A share figure computed across 70% item coverage is a 70% figure however precisely it is later reported.
Ingestion plus the mapping between your hierarchies and every external one.
Read-mostly access paths, change feeds and document streams from the systems of record.
Replayable, resolved, quality-checked records with their lineage back to the source row.
Reads the source’s own change log rather than polling it, so the platform sees every state a record passed through instead of only where it ended up.
Every extract kept as it arrived, immutable and timestamped. The thing you replay from when a downstream definition turns out to have been wrong.
Layout-aware extraction over the unstructured half of the estate: sectioning, tables, signatures, and the page reference every later citation depends on.
Decides that two records are the same real-world thing, with a survivorship rule and a confidence, so downstream joins are a decision rather than an assumption.
Schema, freshness and volume expectations asserted at the boundary, so a bad load fails loudly here rather than quietly two layers later.
Maps internal items and locations to every retailer and syndicated equivalent, and reports the coverage of that mapping as a first-class output.
Mapping your item to the retailer SKU and the syndicated code, with coverage measured.
Reporting a market number without its mapping coverage. Every figure downstream inherits the gap silently, and the reader has no way to know it is there.
Product, location, customer and calendar as governed dimensions with one definition each.
Replayable, resolved records with lineage from ingestion.
Typed, classified, defined data an agent can be grounded on without inferring meaning.
One agreed definition per term, owned by a named person, so the number an agent quotes means what the business means by it.
The typed entities and relationships the domain actually has, so an agent can traverse "which customers are exposed to this" rather than guess from adjacent text.
Metrics defined once, in one place, with their filters and grain. Removes the class of error where the agent computed something plausible and wrong.
Sensitivity labels and the purpose each classification permits, applied at the field level and inherited by everything downstream.
Where a value came from and what it touched on the way. The layer that makes an answer defensible rather than merely correct.
One agreed baseline definition, because every promotional claim is computed from it.
Letting the model infer meaning from column names. It will, it will be plausible, and nobody will notice until the number reaches a regulator or a board pack.
Retrieval over trade agreements, deduction backup, claims substantiation and specifications.
Typed, classified, defined data from the governance layer.
Ranked, permission-trimmed, citable evidence scoped to the caller and the moment.
Vector and lexical retrieval together, because exact identifiers, codes and clause numbers are the thing semantic search is worst at.
Splits on the document’s own structure and carries effective dates, version and source into every chunk, so a retrieved passage knows when it was true.
Applies the caller’s permissions inside the query rather than filtering results afterwards, so the model never sees what the user may not.
Traverses the knowledge graph for questions that are joins rather than similarity — exposure, genealogy, ownership, causation.
Reorders candidates on relevance and returns the passage identifier behind every sentence, so the answer can be checked rather than trusted.
Answering "what did we agree" from the agreement rather than from the account manager.
Filtering results after retrieval instead of constraining the query. The model has already seen the rows you removed, and it will use them.
Agents for deduction validation, distribution gap detection and negotiation preparation.
Ranked, permission-trimmed, citable evidence from retrieval.
Actions taken or proposed, each with the evidence, the identity and the trace behind it.
Plans, calls tools and holds the loop. Where the step budget, the timeout and the stopping condition are enforced rather than hoped for.
Typed, permissioned tools with declared schemas. An agent’s real capability surface is this list, which is why the list is a security artefact.
Durable task state with checkpoints, so a run that dies mid-way resumes instead of restarting and re-doing side effects.
Approval steps on the actions that need one, carrying enough context for the approver to actually decide rather than rubber-stamp.
Writes back into the systems people already work in, under a service identity with its own audit trail.
Deciding which deductions an agent may clear and which need a human decision.
An agent that acts under a service account rather than on behalf of the user. It will eventually do something the requesting user had no right to do, and the log will not show it.
Claims substantiation, competition-law limits on data sharing, and audit.
Proposed actions and drafted responses from the agent layer.
Permitted, grounded, logged output — or a refusal with a stated reason.
The rules that decide whether an action is permitted at all, evaluated before the action and independently of the model that proposed it.
Injection detection on the way in, and on the way out the checks for leakage, unsupported claims and content the domain forbids.
Verifies each assertion resolves to retrieved evidence, and fails the response rather than shipping the sentence that does not.
Model inventory, intended use, validation evidence and the sign-off that lets a model be used for a purpose. Not optional in a regulated estate.
Immutable record of what was asked, retrieved, decided and done — the artefact a reviewer reads when they do not take your word for it.
Restricts what may be asserted about a product to the claims regulatory has approved, which is narrower than what the evidence might support.
Preventing an agent from asserting a product claim that regulatory has not approved.
An agent composing a plausible product claim. That is a regulatory exposure created by your system, not a marketing asset.
Forecast accuracy, recovery rate, mapping coverage and cost per validation.
Permitted, grounded, logged output from the guardrail layer.
Measured quality, cost and latency — and the evidence to change any of the three.
Held-out sets and graded runs on every change, so a prompt edit is a measured change rather than a hopeful one.
Blocks a release when quality drops, in CI, on the same evidence for everyone. The difference between a system and a demo.
End-to-end spans across retrieval, model and tool calls, so a bad answer can be opened and read rather than argued about.
Cost per task and per tenant, against throughput and quota. The number that decides whether the pilot can become the rollout.
Watches quality against production traffic rather than the test set, and routes real corrections back into the eval suite.
Reporting mapping coverage honestly, since every metric above it inherits the gaps.
Evaluating once, before launch. Quality moves with the data, the model and the traffic, and a system with no live measurement has no idea which of the three moved.
Syndicated and POS data arrive on someone else’s calendar, in their hierarchy, with their revisions. Design for lateness and restatement rather than treating them as exceptions.
Every number is computed across a mapping between your items and theirs. Report its coverage, because a metric built on 70% mapping is a 70% metric however precisely it is stated.
A promotional lift is a claim about a counterfactual. If commercial and finance have not agreed how the baseline is derived, the model is arguing rather than measuring.
What may be said about a product is controlled by regulatory approval. An agent that composes a plausible claim has created a compliance problem, not a marketing asset.
These use cases divide into two groups: those that fix the data nobody owns, and those that recover money the organisation has already spent.
Filter by stage or by earned autonomy. Selecting a use case jumps the value chain and the architecture to the stage and layer it depends on.
No use cases match that combination.
How a consumer goods deployment actually runs.
Establish item and location mapping coverage against every external source. Publish it. Every metric you deliver later is capped by it and must carry it.
Commercial and finance must sign one baseline definition before anyone computes a lift. Doing this after the model exists turns measurement into negotiation.
Deduction validation recovers cash within a quarter, needs no forecasting maturity, and forces the agreement corpus into shape for everything that follows.
Distribution gaps and consumption forecasting only mean anything at adequate mapping coverage. Do not start them before the number justifies it.
Prove the agent cannot assert an unapproved product claim, then hand over runbook, mapping coverage report and the agreed baseline definition.
8 modules, 80 taught hours, 40 hands-on labs and 8 assessments — every lab provisioned and graded by the SCIKIQ Agentic AI Playground. Open a module to see its labs.
End-to-end agent design, development, deployment and testing, 10 hours each. Reviewed by a Senior SCDAI Engineer against a published rubric.
Build an agent that validates customer deductions against the operative trade agreement and the backup supplied, citing agreement, clause and date, escalating what it cannot substantiate, and measured on cash recovered against a pre-agent baseline.
Build an agent that computes a promotional baseline to an agreed, documented definition and reports incremental lift with mapping coverage attached to every figure, so a commercial reader can see exactly how much of the category the number covers.
Every lab in this program follows the shape below. This is lab 07 in full — the brief you are given, the environment that is provisioned for you, the code you start from and the assertions that decide whether you passed.
Measure incremental lift rather than promoted sales.
You are given two years of POS across 900 stores and a promotion calendar. Design and run a matched-store holdout, estimate the baseline, and report lift with a confidence interval. The lab includes one promotion that genuinely did not work — your method must say so.
def measure_lift(pos, promo, store_master):
"""Return {promo_id: (lift_pct, ci_low, ci_high, n_test, n_control)}.
Matching happens BEFORE the promotion window, on pre-period behaviour."""
raise NotImplementedError
Every module ends with a timed, randomised assessment delivered through the SCIKIQ Agentic AI Playground. The certificate requires a pass on all of them plus two reviewed capstones.
What each test covers, how long it runs, and how many items are currently in the versioned bank behind it.
| # | Assessment & coverage | Items | Time | In bank |
|---|---|---|---|---|
| 01 | Operating model & scopingDelivery mandateSell-in vs sell-outValue sizingThin slicing | 25 | 35 min | 4 |
| 02 | Consumer goods data landscapeMaster dataSyndicated dataCalendarsItem matching | 25 | 35 min | 4 |
| 03 | Context engineeringHierarchy resolutionUnit conversionRefusalStructured outputs | 25 | 35 min | 4 |
| 04 | Retrieval & groundingAgreement retrievalAmendmentsBackup matchingRetrieval metrics | 25 | 40 min | 4 |
| 05 | Agent design for commercial workValidationBaselinesAmbiguityCoverage reporting | 25 | 40 min | 4 |
| 06 | Systems integrationMCP designPortal retrievalWrite-back safetyLate data | 25 | 35 min | 4 |
| 07 | Claims, competition & securityClaims controlSubstantiationCompetition lawLLM security | 30 | 45 min | 4 |
| 08 | Production & outcomesEval gatingCoverageRecovery measurementEconomics | 30 | 45 min | 4 |
Real items, drawn from 32 in this program's bank — weighted toward the scenario and diagnosis types, because those are the ones that predict field performance. Instant feedback, nothing saved.
Four sample items — one attempt each, then the reasoning is shown.
Q1Before promising a forecasting improvement, what must you check?
Master data determines feasibility and timeline. A forecast built on unreconciled hierarchies produces confident nonsense.
Q2POS and shipment data disagree for the same week. What is the correct handling?
They measure different things at different points. Hiding the discrepancy destroys the planner’s ability to reason.
Q3A planner overrides the agent’s forecast every week. What is the most likely cause?
Explanation is the adoption feature in planning. An unexplained forecast is overridden regardless of its quality.
Q4What makes analogue selection the critical method in new-product forecasting?
Superficially similar products behave differently. Analogue selection is where the method lives or dies.
The same ladder whichever specialisation you enter through — what changes is the domain you go deep in. Below: how the program is delivered, the skills it moves, the roles it leads to, and the specialisations closest to this one.
The same labs, assessments and capstones, delivered to an enterprise cohort or to individual professionals.
Cohorts of 20 to 2,000+ on your own tenancy, with your data patterns and your cloud. Skill-gap baselining up front, per-team mastery reporting throughout, and capstones scoped against your real backlog so the output is deployable work.
The same labs, assessments and capstones for individual engineers and analysts, run on shared infrastructure with a fixed cohort calendar. You leave with a graded portfolio, not a certificate of attendance.
Find your row and aim one column right. The Playground scores you against this after every module.
| Skill | Beginner | Intermediate | Advanced |
|---|---|---|---|
| Consumer goods fluency | Knows the trade terms. | Separates sell-in from sell-out in every metric. | Sizes trade spend and recovery value credibly. |
| Master data engineering | Reads a product master. | Maps items across internal, retailer and syndicated codes. | Runs mapping with coverage published on every metric. |
| Commercial measurement | Computes a lift. | Computes it to an agreed baseline definition. | Governs baseline definitions across commercial and finance. |
| Context engineering | Writes clear prompts. | Structures retrieval, tools and state deliberately. | Designs context strategy for reliability and cost at scale. |
| Retrieval & grounding | Builds basic vector search. | Tunes chunking, hybrid search and reranking. | Designs graph + vector grounding with measured recall. |
| Agent orchestration | Runs a single tool-calling agent. | Builds supervised multi-step and multi-agent flows. | Designs autonomy boundaries and failure containment. |
| Tool & system integration | Calls a documented API. | Writes an MCP server over a system of record. | Designs a least-privilege tool estate across systems. |
| Evaluation | Eyeballs outputs. | Builds labelled eval sets and regression gates. | Runs online evals with drift and judge calibration. |
| Observability & cost | Reads logs. | Traces runs, tracks tokens and latency. | Owns cost per task and capacity planning in production. |
| Security & guardrails | Adds output filters. | Mitigates the OWASP LLM Top 10 in a build. | Threat-models an agent estate and proves controls. |
| Client delivery | Takes notes in a workshop. | Runs discovery and scopes a thin slice. | Owns the account technically, from scope to handover. |
The SCDAI ladder is the same whichever specialisation you enter through — what changes is the domain you go deep in.
Skill a team, or join a cohort
B2B cohorts run on your tenancy with capstones scoped to your backlog. B2C cohorts run on a fixed calendar.
Stated plainly enough to rule yourself in or out without a sales call: the prerequisites, how the program runs, exactly what the credential is worth, and the questions everyone asks.
Stated plainly so you can rule yourself in or out without a sales call. Nothing here is a formal qualification — it is what the first lab assumes you can already do.
You should already be able to do these
What we assume, and what we teach
What the program asks of your week
Cohort dates and pricing are confirmed on enquiry rather than printed here, because both move with the intake.
The credential is awarded per specialisation, so it names the domain or stack you were assessed in rather than claiming general competence. On this program the badge reads SCDAI — FMCG & CPG.
A certificate that cannot be checked is decoration. Every award resolves to a record showing the specialisation, the award date and the assessments passed.
This is the credential the program awards, shown exactly as it is issued — with the specialisation named, the assessment record attached and a verification link anyone can check without an account.
This is to certify that
Your name
has been assessed and certified as
SCIKIQ Certified Data and AI Engineer
FMCG & CPG
Not “Data and AI Engineer” but the domain or stack you were actually assessed in. A general claim would be a weaker one.
Modules passed, labs graded and both capstones reviewed — so the credential states what was measured rather than that you attended.
The credential ID resolves to a public record showing the specialisation, the award date and the assessments passed.
Add it to your LinkedIn profile in one step. The link pre-fills the certification fields from the credential record, so the entry on your profile matches the record a reader can check.
Add to LinkedIn profile The button is live on your real certificate; here it opens LinkedIn pre-filled with this specialisation so you can see exactly what the profile entry will say.A certificate that cannot be checked is decoration. Every award resolves to a record showing the specialisation, the award date and the assessments passed.
The objections that come up in every conversation about this program, answered without the brochure voice.
Both, and the second is the point. Eight timed assessments and two reviewed capstones stand between you and the credential, so a pass means someone measured the skill rather than recorded your attendance.
Every lab is provisioned, graded and unblocked by the Playground rather than by an instructor. That is what lets a cohort of 2,000 cost the same faculty time as a cohort of 20 — and why you are never waiting on someone to mark your work.
Two retakes are included per assessment, each drawing a fresh item set from the bank, so a retake is a genuinely new paper rather than the same questions again.
For the domain specialisations, no — labs run in provisioned sandboxes. For the four tech-stack programs you will want access to that platform, since deploying into a real subscription is much of the point.
Yes. B2B cohorts run on your own tenancy with your data patterns, a skills baseline before kick-off, per-team mastery reporting, and capstones scoped against your actual backlog so the output is deployable work rather than an exercise.
Take the domain you deploy into. If you move across industries, take a tech-stack program instead and pick up domain context on the engagement. The chooser on the programs page will narrow it.
The trends, platform capabilities and regulatory positions are reviewed each quarter, and every external claim on these pages links to its source so you can check the date yourself.
A graded portfolio: forty machine-graded labs, two reviewed end-to-end agent builds with measured evaluation and cost per task, and a verifiable credential naming your specialisation.
Still deciding?
Tell us the systems you deploy into and we will say plainly whether this specialisation is the right one — or which of the 21 is.