Catalogue quality caps everything
Search, filtering, recommendation and comparison all read the same attributes, so attribute coverage sets a ceiling no ranking model can lift.
Online you observe everything — every search, every hesitation, every abandoned basket — and control almost nothing. On a marketplace you may not even own the inventory or the seller. The engineering problem is turning an enormous behavioural signal into decisions fast enough to matter while the shopper is still on the page.
Part 1 follows the shopper from the click that brought them to the return that may follow, and names the people who own each step.
Six stages from traffic to retention. Select one.
Paid and organic acquisition, landing experience and channel economics.
Catalogue, search, browse, filters, ranking and recommendations.
Product page, basket, checkout, payment and abandonment.
Sourcing, picking, carrier selection, delivery promise and exception handling.
Returns, refunds, disposition, return fraud and the margin they consume.
Lifecycle marketing, loyalty, seller quality and marketplace trust.
Everyone here is optimising a funnel where the shopper leaves if you make them wait. Select one to light the stages they own.
Total observability of behaviour, and a catalogue that limits what you can do with it.
Search, filtering, recommendation and comparison all read the same attributes, so attribute coverage sets a ceiling no ranking model can lift.
Zero-result and low-result queries are a directly attributable loss that most catalogues could fix, and most organisations do not measure.
A sale counted at checkout can lose money after return handling and disposition, which makes return probability a merchandising input rather than a back-office metric.
A platform that does not own its inventory competes on seller quality and dispute handling, which are data problems before they are policy ones.
Anything inserted into the shopping path is measured against its effect on conversion, so inference budgets are commercial decisions.
Personalisation now has to respect consent where the profile is read rather than where the message is sent, which changes the architecture.
Two constraints shape everything: the catalogue is the ceiling on quality, and anything on the shopping path is on a latency budget the shopper can feel. Most design decisions here are about which work happens before the visit rather than during it.
Six families. The first limits the others, and it is usually the worst maintained.
Products, attributes, media, descriptions and seller-supplied data.
PIM, seller feeds, content storesSessions, page views, searches, clicks and basket events at very high volume.
Clickstream, event pipelinesOrders, payments, fraud outcomes and order state.
Order management, payment platformsSourcing, picking, shipments, carrier events and exceptions.
WMS, carrier integrationsReturn requests, reasons, inspection outcomes and disposition.
Returns platforms, warehouse systemsCustomers, consent, loyalty, and on marketplaces the sellers themselves.
CRM, consent platform, seller systemsOnline you observe the shopper completely and know the product barely at all — especially on a marketplace, where the catalogue was written by someone else. These are the domains, the graph and the definitions an agent must be grounded on before it answers a shopper.
Six subject areas, each with the entities it holds, its critical data elements and the function accountable for it.
What is listed, and how well it is described.
On a marketplace, who is actually selling and how well.
What the shopper did, at a resolution that identifies them.
What was bought, paid for and promised.
How it was sourced, shipped and whether it arrived.
What came back, why, and what it was worth afterwards.
The entities and the typed relationships between them — what the agent traverses instead of guessing joins. Select any entity.
The identified shopper.
One visit, and everything done during it.
What the shopper asked for.
The item as catalogued.
A specific size, colour or configuration.
A structured property search can filter on.
The marketplace party offering the listing.
A completed purchase.
Goods dispatched against an order.
The date shown to the shopper.
Goods sent back after purchase.
Customer-written content about a product.
The classification hierarchies that make records comparable across systems.
Home → Lighting → Table lamp → Shade material, bulb fitting
Zero result → Attribute missing from catalogue → Recoverable by enrichment
Fit → Smaller than expected → Description
The definitions an agent must use rather than invent. Most wrong answers in this industry are a term used loosely.
The layers a deployment needs, what each one holds, and what the engineer owns there.
Seven layers, with the shopping-path latency budget drawn explicitly.
Seller and internal catalogue feeds, clickstream, orders, fulfilment and returns.
Nothing. This is the upstream edge.
A normalised catalogue with attribute provenance, sessionised behaviour, and consent state.
Products, variants, attributes, media and descriptions. On a marketplace much of it is seller-supplied and unverified, and it caps everything discovery can do.
Sessions, page views, searches, clicks and basket events at very high volume. Personal data, however anonymous the identifiers look.
Orders, order state, payments, fraud outcomes and cancellations.
Sourcing decisions, picking, shipments and carrier tracking events including exceptions.
Return requests, stated reasons, inspection outcomes and disposition decisions.
Customer identity and consent state, and on a marketplace the sellers and their quality history.
Establishing catalogue attribute coverage, because it caps every downstream capability.
Republishing a seller-supplied claim as a platform statement. On a marketplace the content is not yours and neither is its accuracy, but the liability becomes yours the moment you assert it.
Event ingestion at volume, plus attribute extraction and normalisation across sellers.
Read-mostly access paths, change feeds and document streams from the systems of record.
Replayable, resolved, quality-checked records with their lineage back to the source row.
Reads the source’s own change log rather than polling it, so the platform sees every state a record passed through instead of only where it ended up.
Every extract kept as it arrived, immutable and timestamped. The thing you replay from when a downstream definition turns out to have been wrong.
Layout-aware extraction over the unstructured half of the estate: sectioning, tables, signatures, and the page reference every later citation depends on.
Decides that two records are the same real-world thing, with a survivorship rule and a confidence, so downstream joins are a decision rather than an assumption.
Schema, freshness and volume expectations asserted at the boundary, so a bad load fails loudly here rather than quietly two layers later.
Derives structured attributes from seller descriptions and media, marking every inferred value distinctly from a seller-stated one.
Turning inconsistent seller-supplied content into attributes search can actually filter on.
Merging inferred attributes into seller-stated ones. You have then published a claim nobody made and cannot tell which is which.
Taxonomy, attributes, variants and customer identity as governed dimensions.
Replayable, resolved records with lineage from ingestion.
Typed, classified, defined data an agent can be grounded on without inferring meaning.
One agreed definition per term, owned by a named person, so the number an agent quotes means what the business means by it.
The typed entities and relationships the domain actually has, so an agent can traverse "which customers are exposed to this" rather than guess from adjacent text.
Metrics defined once, in one place, with their filters and grain. Removes the class of error where the agent computed something plausible and wrong.
Sensitivity labels and the purpose each classification permits, applied at the field level and inherited by everything downstream.
Where a value came from and what it touched on the way. The layer that makes an answer defensible rather than merely correct.
One taxonomy and attribute schema that sellers, search and merchandising all obey.
Letting the model infer meaning from column names. It will, it will be plausible, and nobody will notice until the number reaches a regulator or a board pack.
Retrieval over catalogue content, reviews, policies and seller documentation.
Typed, classified, defined data from the governance layer.
Ranked, permission-trimmed, citable evidence scoped to the caller and the moment.
Vector and lexical retrieval together, because exact identifiers, codes and clause numbers are the thing semantic search is worst at.
Splits on the document’s own structure and carries effective dates, version and source into every chunk, so a retrieved passage knows when it was true.
Applies the caller’s permissions inside the query rather than filtering results afterwards, so the model never sees what the user may not.
Traverses the knowledge graph for questions that are joins rather than similarity — exposure, genealogy, ownership, causation.
Reorders candidates on relevance and returns the passage identifier behind every sentence, so the answer can be checked rather than trusted.
Answering a shopper question from the catalogue and reviews rather than from the model.
Filtering results after retrieval instead of constraining the query. The model has already seen the rows you removed, and it will use them.
Agents for attribute enrichment, zero-result recovery, exception handling and returns triage.
Ranked, permission-trimmed, citable evidence from retrieval.
Actions taken or proposed, each with the evidence, the identity and the trace behind it.
Plans, calls tools and holds the loop. Where the step budget, the timeout and the stopping condition are enforced rather than hoped for.
Typed, permissioned tools with declared schemas. An agent’s real capability surface is this list, which is why the list is a security artefact.
Durable task state with checkpoints, so a run that dies mid-way resumes instead of restarting and re-doing side effects.
Approval steps on the actions that need one, carrying enough context for the approver to actually decide rather than rubber-stamp.
Writes back into the systems people already work in, under a service identity with its own audit trail.
Pre-computes expensive decisions offline and serves them from a store the page can read inside its latency budget.
Deciding what is pre-computed offline and what may run in the shopper path at all.
Calling a model from the shopping path. The median looks acceptable and the tail, which is where abandonment happens, does not.
Consent enforcement, content safety, seller trust, counterfeit signals and audit.
Proposed actions and drafted responses from the agent layer.
Permitted, grounded, logged output — or a refusal with a stated reason.
The rules that decide whether an action is permitted at all, evaluated before the action and independently of the model that proposed it.
Injection detection on the way in, and on the way out the checks for leakage, unsupported claims and content the domain forbids.
Verifies each assertion resolves to retrieved evidence, and fails the response rather than shipping the sentence that does not.
Model inventory, intended use, validation evidence and the sign-off that lets a model be used for a purpose. Not optional in a regulated estate.
Immutable record of what was asked, retrieved, decided and done — the artefact a reviewer reads when they do not take your word for it.
Prevents the platform restating an unverified seller claim as its own, which is how a seller’s liability becomes the platform’s.
Preventing an agent from presenting seller-supplied claims as though the platform made them.
Answering a shopper question by paraphrasing a seller description. The paraphrase is now a platform statement about a product you never inspected.
Search conversion, attribute coverage, latency at percentile and cost per session.
Permitted, grounded, logged output from the guardrail layer.
Measured quality, cost and latency — and the evidence to change any of the three.
Held-out sets and graded runs on every change, so a prompt edit is a measured change rather than a hopeful one.
Blocks a release when quality drops, in CI, on the same evidence for everyone. The difference between a system and a demo.
End-to-end spans across retrieval, model and tool calls, so a bad answer can be opened and read rather than argued about.
Cost per task and per tenant, against throughput and quota. The number that decides whether the pilot can become the rollout.
Watches quality against production traffic rather than the test set, and routes real corrections back into the eval suite.
Measuring latency at the tail, because the slow sessions are the ones that abandon.
Evaluating once, before launch. Quality moves with the data, the model and the traffic, and a system with no live measurement has no idea which of the three moved.
Anything the shopper waits for is measured against conversion. Most model work belongs offline, materialised into something the page can read in milliseconds.
Search, filters, comparison and recommendation all read the same attributes. No amount of ranking sophistication compensates for attributes that are missing or wrong.
Seller-supplied descriptions and images may be inaccurate or infringing. An agent that repeats a seller claim as platform fact transfers the seller’s liability to you.
Sessions, searches and baskets identify people. Consent and purpose have to be enforced where the profile is retrieved, not where a message is sent.
Note how many of these run offline. Online you observe everything, but the shopper will not wait while you think about it.
Filter by stage or by earned autonomy. Selecting a use case jumps the value chain and the architecture to the stage and layer it depends on.
No use cases match that combination.
How an ecommerce deployment actually runs.
Establish what proportion of the catalogue has the attributes search and filters need. This number is the ceiling on everything discovery-related.
Agree what may run in the shopping path and what must be pre-computed. Deciding this late means rebuilding the serving design.
Both run offline, both lift the ceiling, and zero-result recovery produces a revenue number within weeks that funds everything after it.
Anything the shopper waits for must be materialised. Measure latency at the tail percentiles, because the slow sessions are the ones that leave.
Prove consent is enforced at retrieval and that inferred attributes are never presented as seller-stated, then hand over runbook, coverage report and latency budgets.
8 modules, 80 taught hours, 40 hands-on labs and 8 assessments — every lab provisioned and graded by the SCIKIQ Agentic AI Playground. Open a module to see its labs.
End-to-end agent design, development, deployment and testing, 10 hours each. Reviewed by a Senior SCDAI Engineer against a published rubric.
Build an agent that extracts structured attributes from seller descriptions, media and specifications at catalogue scale, marks every inferred attribute distinctly from seller-stated ones, and measurably lifts attribute coverage and filter recall in a target category.
Build an agent that diagnoses failed searches, distinguishes "we do not stock it" from "we could not find it", routes recoverable queries to intent-matching results and reports genuine range gaps — serving from pre-computed output inside a strict tail-latency budget.
Every lab in this program follows the shape below. This is lab 05 in full — the brief you are given, the environment that is provisioned for you, the code you start from and the assertions that decide whether you passed.
Separate "we could not surface it" from "we do not sell it", and fix only the first.
You are given 30,000 searches that returned nothing, the catalogue with its patchy attributes, and a labelled set stating for each query whether a matching product actually existed. Recover the findable ones by understanding the query and the catalogue; report the rest as genuine range gaps. Returning near-matches for products you do not stock is the failure this lab catches.
def recover(query, catalogue):
"""Return (results, verdict) where verdict is 'recovered' or 'range_gap'.
A near-match for something you do not stock is worse than an honest
empty result: the shopper leaves annoyed instead of merely disappointed.
"""
intent = parse_intent(query) # attributes, not just tokens
# only recover when the intent genuinely matches a stocked product
raise NotImplementedError
Every module ends with a timed, randomised assessment delivered through the SCIKIQ Agentic AI Playground. The certificate requires a pass on all of them plus two reviewed capstones.
What each test covers, how long it runs, and how many items are currently in the versioned bank behind it.
| # | Assessment & coverage | Items | Time | In bank |
|---|---|---|---|---|
| 01 | Operating model & scopingDelivery mandateLatency constraintValue sizingThin slicing | 25 | 35 min | 3 |
| 02 | Ecommerce data landscapeCatalogue feedsClickstreamOrder stateConsent | 25 | 35 min | 2 |
| 03 | Context engineeringQuery understandingNormalisationProvenanceRefusal | 25 | 35 min | 2 |
| 04 | Retrieval & groundingCatalogue retrievalReview groundingVariantsRetrieval metrics | 25 | 40 min | 2 |
| 05 | Agent design & latencyOffline vs in-pathMaterialisationZero-resultFairness | 25 | 40 min | 2 |
| 06 | Systems integrationMCP designCatalogue write-backServing layerIdentity | 25 | 35 min | 2 |
| 07 | Consent, liability & securityConsentSeller liabilityCounterfeitLLM security | 30 | 45 min | 2 |
| 08 | Production & latencyEval gatingLatency percentilesRelevanceEconomics | 30 | 45 min | 2 |
Real items, drawn from 17 in this program's bank — weighted toward the scenario and diagnosis types, because those are the ones that predict field performance. Instant feedback, nothing saved.
Four sample items — one attempt each, then the reasoning is shown.
Q1An online retailer asks for AI to lift conversion. What do you measure first?
Discovery quality is bounded by the attributes it can read. No ranking work compensates for attributes that are absent or wrong.
Q2On a marketplace, who is responsible for the accuracy of a product description?
Hosting and asserting are different. An agent that paraphrases a seller claim into a platform answer has transferred the liability.
Q3An agent infers a product attribute the seller never stated. How should it be stored?
Merging publishes a claim nobody made and makes it impossible to tell later which values you can stand behind.
Q4A shopper asks a question the description does not answer but reviews do. What should the agent do?
Reviews are genuinely useful and genuinely not specifications. Attribution is what lets you use them without asserting them.
The same ladder whichever specialisation you enter through — what changes is the domain you go deep in. Below: how the program is delivered, the skills it moves, the roles it leads to, and the specialisations closest to this one.
The same labs, assessments and capstones, delivered to an enterprise cohort or to individual professionals.
Cohorts of 20 to 2,000+ on your own tenancy, with your data patterns and your cloud. Skill-gap baselining up front, per-team mastery reporting throughout, and capstones scoped against your real backlog so the output is deployable work.
The same labs, assessments and capstones for individual engineers and analysts, run on shared infrastructure with a fixed cohort calendar. You leave with a graded portfolio, not a certificate of attendance.
Find your row and aim one column right. The Playground scores you against this after every module.
| Skill | Beginner | Intermediate | Advanced |
|---|---|---|---|
| Ecommerce domain fluency | Knows the funnel. | Maps discovery and fulfilment decisions to margin. | Sizes zero-result, promise and return value credibly. |
| Catalogue engineering | Reads a product feed. | Normalises seller attributes into one schema. | Runs catalogue-scale enrichment with provenance and measured coverage. |
| Latency engineering | Aware of the budget. | Moves expensive work offline and materialises it. | Holds a tail-percentile budget under peak traffic. |
| Context engineering | Writes clear prompts. | Structures retrieval, tools and state deliberately. | Designs context strategy for reliability and cost at scale. |
| Retrieval & grounding | Builds basic vector search. | Tunes chunking, hybrid search and reranking. | Designs graph + vector grounding with measured recall. |
| Agent orchestration | Runs a single tool-calling agent. | Builds supervised multi-step and multi-agent flows. | Designs autonomy boundaries and failure containment. |
| Tool & system integration | Calls a documented API. | Writes an MCP server over a system of record. | Designs a least-privilege tool estate across systems. |
| Evaluation | Eyeballs outputs. | Builds labelled eval sets and regression gates. | Runs online evals with drift and judge calibration. |
| Observability & cost | Reads logs. | Traces runs, tracks tokens and latency. | Owns cost per task and capacity planning in production. |
| Security & guardrails | Adds output filters. | Mitigates the OWASP LLM Top 10 in a build. | Threat-models an agent estate and proves controls. |
| Client delivery | Takes notes in a workshop. | Runs discovery and scopes a thin slice. | Owns the account technically, from scope to handover. |
The SCDAI ladder is the same whichever specialisation you enter through — what changes is the domain you go deep in.
Skill a team, or join a cohort
B2B cohorts run on your tenancy with capstones scoped to your backlog. B2C cohorts run on a fixed calendar.
Stated plainly enough to rule yourself in or out without a sales call: the prerequisites, how the program runs, exactly what the credential is worth, and the questions everyone asks.
Stated plainly so you can rule yourself in or out without a sales call. Nothing here is a formal qualification — it is what the first lab assumes you can already do.
You should already be able to do these
What we assume, and what we teach
What the program asks of your week
Cohort dates and pricing are confirmed on enquiry rather than printed here, because both move with the intake.
The credential is awarded per specialisation, so it names the domain or stack you were assessed in rather than claiming general competence. On this program the badge reads SCDAI — Ecommerce & Marketplaces.
A certificate that cannot be checked is decoration. Every award resolves to a record showing the specialisation, the award date and the assessments passed.
This is the credential the program awards, shown exactly as it is issued — with the specialisation named, the assessment record attached and a verification link anyone can check without an account.
This is to certify that
Your name
has been assessed and certified as
SCIKIQ Certified Data and AI Engineer
Ecommerce & Marketplaces
Not “Data and AI Engineer” but the domain or stack you were actually assessed in. A general claim would be a weaker one.
Modules passed, labs graded and both capstones reviewed — so the credential states what was measured rather than that you attended.
The credential ID resolves to a public record showing the specialisation, the award date and the assessments passed.
Add it to your LinkedIn profile in one step. The link pre-fills the certification fields from the credential record, so the entry on your profile matches the record a reader can check.
Add to LinkedIn profile The button is live on your real certificate; here it opens LinkedIn pre-filled with this specialisation so you can see exactly what the profile entry will say.A certificate that cannot be checked is decoration. Every award resolves to a record showing the specialisation, the award date and the assessments passed.
The objections that come up in every conversation about this program, answered without the brochure voice.
Both, and the second is the point. Eight timed assessments and two reviewed capstones stand between you and the credential, so a pass means someone measured the skill rather than recorded your attendance.
Every lab is provisioned, graded and unblocked by the Playground rather than by an instructor. That is what lets a cohort of 2,000 cost the same faculty time as a cohort of 20 — and why you are never waiting on someone to mark your work.
Two retakes are included per assessment, each drawing a fresh item set from the bank, so a retake is a genuinely new paper rather than the same questions again.
For the domain specialisations, no — labs run in provisioned sandboxes. For the four tech-stack programs you will want access to that platform, since deploying into a real subscription is much of the point.
Yes. B2B cohorts run on your own tenancy with your data patterns, a skills baseline before kick-off, per-team mastery reporting, and capstones scoped against your actual backlog so the output is deployable work rather than an exercise.
Take the domain you deploy into. If you move across industries, take a tech-stack program instead and pick up domain context on the engagement. The chooser on the programs page will narrow it.
The trends, platform capabilities and regulatory positions are reviewed each quarter, and every external claim on these pages links to its source so you can check the date yourself.
A graded portfolio: forty machine-graded labs, two reviewed end-to-end agent builds with measured evaluation and cost per task, and a verifiable credential naming your specialisation.
Still deciding?
Tell us the systems you deploy into and we will say plainly whether this specialisation is the right one — or which of the 21 is.