Phantom inventory is the quiet tax
Stock the system believes is on shelf and is not suppresses replenishment silently, and the lost sale is never recorded because it never happened.
Retail is the one consumer business that owns the shelf, the stock on it and the person standing next to it. That makes the data unusually direct — you see the actual purchase — and the constraints unusually physical. A recommendation the store cannot execute, because the stock is not there or nobody is rostered to do it, is not a recommendation.
Part 1 follows a product from the buying decision to the shelf and, sometimes, back again. It names the people who make each of those calls.
Six stages from range to returns. Select one.
Assortment planning, buying, supplier terms and range reviews.
Everyday pricing, promotions, markdown strategy and competitor response.
Allocation to stores, replenishment, safety stock and inbound flow.
Labour scheduling, task execution, shelf compliance and store standards.
Service, availability at the moment of purchase, loyalty and substitution.
Shrink, waste, returns, refunds and the fraud around them.
The store manager is the person every central decision eventually lands on. Select one to light the stages they own.
A business with excellent transaction data and a persistent gap between system and shelf.
Stock the system believes is on shelf and is not suppresses replenishment silently, and the lost sale is never recorded because it never happened.
Most margin is decided by when the markdown starts rather than by how deep it goes, which makes timing the higher-value modelling problem.
Central task generation that ignores rostered hours produces plans stores cannot execute, and the tasks that get dropped are chosen by whoever is on shift.
Removing an item does not remove its demand, and treating delisting as pure loss systematically under-ranges.
Measuring what the system holds is easy and increasingly seen as insufficient; what the shopper can actually pick up is the number that matters.
Shrink and return analytics touch individual staff and customers, so explainability and proportionality are design requirements rather than afterthoughts.
The defining problem is the gap between the inventory record and the shelf. Almost every retail use case is either exploiting good transaction data or compensating for that gap.
Six families. The second one is the one that lies to you.
Line-level sales, baskets, tenders, returns and refunds.
POS, transaction logStock by item and location, movements, counts and adjustments.
Inventory systems, RFID, countsItems, hierarchy, attributes, assortment and store clustering.
Merchandising systems, PIMPrices, markdowns, promotional mechanics and competitor prices.
Pricing systems, competitor feedsRosters, tasks, completion, compliance and store attributes.
Workforce management, task systemsIdentified customers, loyalty, consent and interaction history.
Loyalty platform, CRMA retail agent that treats the stock figure as a measurement rather than as a record will recommend confidently and wrongly, every day, on exactly the items that are selling. These are the domains, the graph and the definitions it must be grounded on.
Six subject areas, each with the entities it holds, its critical data elements and the function accountable for it.
What is sold, and where it is ranged.
Where it is sold, and which locations behave alike.
What the system believes is where, and how it got there.
What it costs and what it is being sold for.
What the shopper actually bought or brought back.
Who is available and what they were asked to do.
The entities and the typed relationships between them — what the agent traverses instead of guessing joins. Select any entity.
The lowest-level sellable unit.
The classification the item competes within.
A selling location.
A grouping of stores that behave alike.
What the system believes is on hand.
A recorded change to stock.
A purchase at the till.
Goods brought back by a shopper.
The retail price in force.
A time-bound price or display mechanic.
Work issued to a store.
The hours actually available in a store.
The classification hierarchies that make records comparable across systems.
Grocery → Dairy → Yoghurt → 4-pack strawberry
South → Convenience → High-footfall commuter
Process → Mis-scan at till → Addressable by training
The definitions an agent must use rather than invent. Most wrong answers in this industry are a term used loosely.
The layers a deployment needs, what each one holds, and what the engineer owns there.
Seven layers, with the inventory-truth problem addressed explicitly rather than assumed away.
Transactions, the inventory record, range and pricing, plus rosters and task completion.
Nothing. This is the upstream edge.
Trustworthy transactions, stock figures carrying a confidence, and the roster capacity bound.
Line-level transactions, baskets, tenders and returns. High volume, and the one retail dataset that is genuinely trustworthy.
Stock by item and location with movements, counts and adjustments. Drifts from physical reality continuously, which is the central problem of the domain.
Range, assortment, product hierarchy, store attributes and clustering.
Everyday prices, markdowns, promotional mechanics and competitor price feeds.
Rosters, available hours, task issue and completion. The constraint that decides whether any central plan is executable at all.
Identified customers, loyalty state, consent and interaction history.
Understanding how the inventory record is maintained, and therefore how it fails.
Treating the stock figure as truth. It is a record that drifts through theft, damage, mis-scans and misplacement, and a replenishment system that believes it will silently stop ordering the items that are selling fastest.
High-volume transaction ingestion plus inference of where the stock record is untrustworthy.
Read-mostly access paths, change feeds and document streams from the systems of record.
Replayable, resolved, quality-checked records with their lineage back to the source row.
Reads the source’s own change log rather than polling it, so the platform sees every state a record passed through instead of only where it ended up.
Every extract kept as it arrived, immutable and timestamped. The thing you replay from when a downstream definition turns out to have been wrong.
Layout-aware extraction over the unstructured half of the estate: sectioning, tables, signatures, and the page reference every later citation depends on.
Decides that two records are the same real-world thing, with a survivorship rule and a confidence, so downstream joins are a decision rather than an assumption.
Schema, freshness and volume expectations asserted at the boundary, so a bad load fails loudly here rather than quietly two layers later.
Infers where the stock record has diverged from the shelf, and attaches a confidence to the figure rather than passing it through unchanged.
Producing a confidence on the stock figure, rather than passing it through unchanged.
Passing the stock figure downstream as though it were a measurement. It is a record, it drifts, and replenishment built on it stops ordering exactly what is selling.
Product, store, cluster and calendar as governed dimensions with agreed definitions.
Replayable, resolved records with lineage from ingestion.
Typed, classified, defined data an agent can be grounded on without inferring meaning.
One agreed definition per term, owned by a named person, so the number an agent quotes means what the business means by it.
The typed entities and relationships the domain actually has, so an agent can traverse "which customers are exposed to this" rather than guess from adjacent text.
Metrics defined once, in one place, with their filters and grain. Removes the class of error where the agent computed something plausible and wrong.
Sensitivity labels and the purpose each classification permits, applied at the field level and inherited by everything downstream.
Where a value came from and what it touched on the way. The layer that makes an answer defensible rather than merely correct.
One definition of availability, agreed between supply chain and stores.
Letting the model infer meaning from column names. It will, it will be plausible, and nobody will notice until the number reaches a regulator or a board pack.
Retrieval over supplier agreements, range review packs, policies and store procedures.
Typed, classified, defined data from the governance layer.
Ranked, permission-trimmed, citable evidence scoped to the caller and the moment.
Vector and lexical retrieval together, because exact identifiers, codes and clause numbers are the thing semantic search is worst at.
Splits on the document’s own structure and carries effective dates, version and source into every chunk, so a retrieved passage knows when it was true.
Applies the caller’s permissions inside the query rather than filtering results afterwards, so the model never sees what the user may not.
Traverses the knowledge graph for questions that are joins rather than similarity — exposure, genealogy, ownership, causation.
Reorders candidates on relevance and returns the passage identifier behind every sentence, so the answer can be checked rather than trusted.
Grounding a recommendation in the supplier term or store policy that governs it.
Filtering results after retrieval instead of constraining the query. The model has already seen the rows you removed, and it will use them.
Agents for phantom detection, markdown timing, task generation and loss review.
Ranked, permission-trimmed, citable evidence from retrieval.
Actions taken or proposed, each with the evidence, the identity and the trace behind it.
Plans, calls tools and holds the loop. Where the step budget, the timeout and the stopping condition are enforced rather than hoped for.
Typed, permissioned tools with declared schemas. An agent’s real capability surface is this list, which is why the list is a security artefact.
Durable task state with checkpoints, so a run that dies mid-way resumes instead of restarting and re-doing side effects.
Approval steps on the actions that need one, carrying enough context for the approver to actually decide rather than rubber-stamp.
Writes back into the systems people already work in, under a service identity with its own audit trail.
Bounds generated store work by rostered hours, so the plan issued is the plan that can actually be executed.
Generating tasks that fit the roster, because an unexecutable task is worse than none.
Issuing more tasks than the roster supports. The list is then triaged by whoever is on shift, and you will never know which half was dropped.
Loss analytics fairness, customer consent, staff monitoring limits and audit.
Proposed actions and drafted responses from the agent layer.
Permitted, grounded, logged output — or a refusal with a stated reason.
The rules that decide whether an action is permitted at all, evaluated before the action and independently of the model that proposed it.
Injection detection on the way in, and on the way out the checks for leakage, unsupported claims and content the domain forbids.
Verifies each assertion resolves to retrieved evidence, and fails the response rather than shipping the sentence that does not.
Model inventory, intended use, validation evidence and the sign-off that lets a model be used for a purpose. Not optional in a regulated estate.
Immutable record of what was asked, retrieved, decided and done — the artefact a reviewer reads when they do not take your word for it.
Tests shrink and return models for systematic disparity, and requires findings to point at process before they point at people.
Ensuring shrink analytics point at process before people, and are explainable when they do not.
Shrink analytics that name individuals first. It is the fastest way to turn a data programme into an employee-relations incident.
Availability lift, phantom precision, task completion and cost per store.
Permitted, grounded, logged output from the guardrail layer.
Measured quality, cost and latency — and the evidence to change any of the three.
Held-out sets and graded runs on every change, so a prompt edit is a measured change rather than a hopeful one.
Blocks a release when quality drops, in CI, on the same evidence for everyone. The difference between a system and a demo.
End-to-end spans across retrieval, model and tool calls, so a bad answer can be opened and read rather than argued about.
Cost per task and per tenant, against throughput and quota. The number that decides whether the pilot can become the rollout.
Watches quality against production traffic rather than the test set, and routes real corrections back into the eval suite.
Proving on-shelf availability moved, not that the dashboard did.
Evaluating once, before launch. Quality moves with the data, the model and the traffic, and a system with no live measurement has no idea which of the three moved.
Stock figures drift from reality through theft, damage, mis-scans and misplacement. Any system that treats the record as truth will silently suppress replenishment on exactly the items that are selling.
A store has a fixed number of rostered hours. Task generation that ignores them produces a plan that is silently triaged by whoever is on shift, and you will not know which parts survived.
Shrink and return models implicate staff and customers. Proportionality, explainability and pointing at process before individuals are design requirements.
Per-transaction inference is not affordable at retail scale. Reserve reasoning for decisions, not for records.
The most valuable use cases here are the ones that stop trusting a number the whole business has been trusting for decades.
Filter by stage or by earned autonomy. Selecting a use case jumps the value chain and the architecture to the stage and layer it depends on.
No use cases match that combination.
How a retail deployment actually runs.
Measure inventory record accuracy against physical reality on a sample. Every replenishment and availability use case is capped by this number.
Supply chain and stores must agree one definition, including whether it is measured in the system or on the shelf. They usually mean different things by it.
It needs only transaction data, recovers sales nobody knew were lost, and produces the inventory confidence every later use case depends on.
Do not generate store tasks until you have rostered hours. Without them you are adding to a queue that is already being triaged without you.
Test loss analytics for disparity, prove on-shelf availability actually moved, then hand over runbook, thresholds and the definitions agreed in week two.
8 modules, 80 taught hours, 40 hands-on labs and 8 assessments — every lab provisioned and graded by the SCIKIQ Agentic AI Playground. Open a module to see its labs.
End-to-end agent design, development, deployment and testing, 10 hours each. Reviewed by a Senior SCDAI Engineer against a published rubric.
Build an agent that detects items the inventory record shows on shelf but which have stopped selling abnormally, distinguishes them from genuine slow sellers, propagates a confidence into the replenishment output, and is measured on precision against a physically counted set.
Build an agent that generates a store’s daily task list bounded by actual rostered hours, prioritised by measured sales impact, citing the store policy governing each task — and demonstrably never issuing more work than the roster supports.
Every lab in this program follows the shape below. This is lab 05 in full — the brief you are given, the environment that is provisioned for you, the code you start from and the assertions that decide whether you passed.
Detect phantom inventory from sales behaviour, without flagging every slow seller.
You are given 12 weeks of line-level sales, the stock record and the movement history for 40,000 item-store combinations, plus a physically counted set. Some items stopped selling because they are not physically on the shelf; most stopped because nobody wanted them. Return the phantoms, with a confidence, and propagate that confidence into the replenishment signal.
def detect_phantoms(sales, stock, movements):
"""Return [(item, store, confidence, evidence)] for suspected phantoms.
A slow seller and a phantom look identical in a stock figure. The
difference is in the SHAPE of the sales stop, against that item's own
prior rate of sale at that store.
"""
# 1. baseline rate of sale per item-store, not per item
# 2. a phantom is an abrupt stop while stock is still recorded positive
# 3. never return a bare flag -- return the evidence that produced it
raise NotImplementedError
Every module ends with a timed, randomised assessment delivered through the SCIKIQ Agentic AI Playground. The certificate requires a pass on all of them plus two reviewed capstones.
What each test covers, how long it runs, and how many items are currently in the versioned bank behind it.
| # | Assessment & coverage | Items | Time | In bank |
|---|---|---|---|---|
| 01 | Operating model & scopingDelivery mandateSystem vs shelfValue sizingThin slicing | 25 | 35 min | 3 |
| 02 | Retail data landscapeTransactionsInventory movementsClusteringLabour data | 25 | 35 min | 2 |
| 03 | Context engineeringIdentifier resolutionPack arithmeticConfidenceRefusal | 25 | 35 min | 2 |
| 04 | Retrieval & groundingAgreement retrievalPolicy retrievalRange historyRetrieval metrics | 25 | 40 min | 2 |
| 05 | Agent design for retailPhantom detectionConfidenceCapacity boundingFairness | 25 | 40 min | 2 |
| 06 | Systems integrationMCP designWrite-back safetyBatch windowsIdentity | 25 | 35 min | 2 |
| 07 | Fairness, privacy & securityFairnessProportionalityConsentLLM security | 30 | 45 min | 2 |
| 08 | Production & outcomesEval gatingAvailabilityPhantom precisionEconomics | 30 | 45 min | 2 |
Real items, drawn from 17 in this program's bank — weighted toward the scenario and diagnosis types, because those are the ones that predict field performance. Instant feedback, nothing saved.
Four sample items — one attempt each, then the reasoning is shown.
Q1A retailer asks for AI to improve availability. What do you measure in week one?
Every availability and replenishment use case is capped by inventory record accuracy. Measuring it first sets an honest ceiling instead of a later surprise.
Q2Should a rate-of-sale baseline be computed per item or per item-store?
A chain-level baseline flags every slow store as anomalous and misses genuine stops in fast ones. The baseline has to be local to be meaningful.
Q3Why should an agent refuse to attribute shrink to a named individual?
Shrink patterns have many innocent explanations. Attribution is an HR process with its own standards, and an agent that pre-empts it causes real harm.
Q4A phantom model achieves high recall by flagging 15% of item-store combinations. Is this useful?
The constraint is store labour. A model tuned for recall against a workforce that can check twenty items a day produces nothing but noise.
The same ladder whichever specialisation you enter through — what changes is the domain you go deep in. Below: how the program is delivered, the skills it moves, the roles it leads to, and the specialisations closest to this one.
The same labs, assessments and capstones, delivered to an enterprise cohort or to individual professionals.
Cohorts of 20 to 2,000+ on your own tenancy, with your data patterns and your cloud. Skill-gap baselining up front, per-team mastery reporting throughout, and capstones scoped against your real backlog so the output is deployable work.
The same labs, assessments and capstones for individual engineers and analysts, run on shared infrastructure with a fixed cohort calendar. You leave with a graded portfolio, not a certificate of attendance.
Find your row and aim one column right. The Playground scores you against this after every module.
| Skill | Beginner | Intermediate | Advanced |
|---|---|---|---|
| Retail domain fluency | Knows the functions. | Maps range, price and stock decisions to margin. | Sizes availability, markdown and shrink value credibly. |
| Inventory truth | Reads a stock figure. | Detects where the record and the shelf diverge. | Propagates stock confidence through every downstream decision. |
| Executable operations | Generates a task. | Bounds task generation by rostered capacity. | Proves plans were executed rather than silently triaged. |
| Context engineering | Writes clear prompts. | Structures retrieval, tools and state deliberately. | Designs context strategy for reliability and cost at scale. |
| Retrieval & grounding | Builds basic vector search. | Tunes chunking, hybrid search and reranking. | Designs graph + vector grounding with measured recall. |
| Agent orchestration | Runs a single tool-calling agent. | Builds supervised multi-step and multi-agent flows. | Designs autonomy boundaries and failure containment. |
| Tool & system integration | Calls a documented API. | Writes an MCP server over a system of record. | Designs a least-privilege tool estate across systems. |
| Evaluation | Eyeballs outputs. | Builds labelled eval sets and regression gates. | Runs online evals with drift and judge calibration. |
| Observability & cost | Reads logs. | Traces runs, tracks tokens and latency. | Owns cost per task and capacity planning in production. |
| Security & guardrails | Adds output filters. | Mitigates the OWASP LLM Top 10 in a build. | Threat-models an agent estate and proves controls. |
| Client delivery | Takes notes in a workshop. | Runs discovery and scopes a thin slice. | Owns the account technically, from scope to handover. |
The SCDAI ladder is the same whichever specialisation you enter through — what changes is the domain you go deep in.
Skill a team, or join a cohort
B2B cohorts run on your tenancy with capstones scoped to your backlog. B2C cohorts run on a fixed calendar.
Stated plainly enough to rule yourself in or out without a sales call: the prerequisites, how the program runs, exactly what the credential is worth, and the questions everyone asks.
Stated plainly so you can rule yourself in or out without a sales call. Nothing here is a formal qualification — it is what the first lab assumes you can already do.
You should already be able to do these
What we assume, and what we teach
What the program asks of your week
Cohort dates and pricing are confirmed on enquiry rather than printed here, because both move with the intake.
The credential is awarded per specialisation, so it names the domain or stack you were assessed in rather than claiming general competence. On this program the badge reads SCDAI — Retail.
A certificate that cannot be checked is decoration. Every award resolves to a record showing the specialisation, the award date and the assessments passed.
This is the credential the program awards, shown exactly as it is issued — with the specialisation named, the assessment record attached and a verification link anyone can check without an account.
This is to certify that
Your name
has been assessed and certified as
SCIKIQ Certified Data and AI Engineer
Retail
Not “Data and AI Engineer” but the domain or stack you were actually assessed in. A general claim would be a weaker one.
Modules passed, labs graded and both capstones reviewed — so the credential states what was measured rather than that you attended.
The credential ID resolves to a public record showing the specialisation, the award date and the assessments passed.
Add it to your LinkedIn profile in one step. The link pre-fills the certification fields from the credential record, so the entry on your profile matches the record a reader can check.
Add to LinkedIn profile The button is live on your real certificate; here it opens LinkedIn pre-filled with this specialisation so you can see exactly what the profile entry will say.A certificate that cannot be checked is decoration. Every award resolves to a record showing the specialisation, the award date and the assessments passed.
The objections that come up in every conversation about this program, answered without the brochure voice.
Both, and the second is the point. Eight timed assessments and two reviewed capstones stand between you and the credential, so a pass means someone measured the skill rather than recorded your attendance.
Every lab is provisioned, graded and unblocked by the Playground rather than by an instructor. That is what lets a cohort of 2,000 cost the same faculty time as a cohort of 20 — and why you are never waiting on someone to mark your work.
Two retakes are included per assessment, each drawing a fresh item set from the bank, so a retake is a genuinely new paper rather than the same questions again.
For the domain specialisations, no — labs run in provisioned sandboxes. For the four tech-stack programs you will want access to that platform, since deploying into a real subscription is much of the point.
Yes. B2B cohorts run on your own tenancy with your data patterns, a skills baseline before kick-off, per-team mastery reporting, and capstones scoped against your actual backlog so the output is deployable work rather than an exercise.
Take the domain you deploy into. If you move across industries, take a tech-stack program instead and pick up domain context on the engagement. The chooser on the programs page will narrow it.
The trends, platform capabilities and regulatory positions are reviewed each quarter, and every external claim on these pages links to its source so you can check the date yourself.
A graded portfolio: forty machine-graded labs, two reviewed end-to-end agent builds with measured evaluation and cost per task, and a verifiable credential naming your specialisation.
Still deciding?
Tell us the systems you deploy into and we will say plainly whether this specialisation is the right one — or which of the 21 is.