AI-first carriers are separating from the pack
Carriers redesigning core processes around AI report roughly 20% cost reduction and 3-5% GWP growth from share gains and productivity.
Insurance is the industry where agentic AI has moved fastest from pilot to production — and where the failure modes are most expensive. Quote in three minutes instead of three days is real; so is an unfair-discrimination finding. This specialisation covers both.
Insurance is a chain of judgements about risk, made under time pressure with incomplete information. Part 1 maps where those judgements happen, who makes them, and what the strategy houses see changing.
From product design to reserving. Select a stage to see what breaks and who owns it.
Product design, rating factors, filings and portfolio strategy.
Submission intake, appetite matching, triage and quotation.
Risk assessment, pricing decision, terms, referral and binding.
Policy administration, endorsements, renewals and customer service.
First notice, triage, coverage check, investigation, reserve and settlement.
Reserving, reinsurance cessions, regulatory reporting and portfolio review.
Five roles whose decisions the agent will either sharpen or undermine. Select one to light the stages they own.
Insurance is further ahead than most industries — and further exposed.
Carriers redesigning core processes around AI report roughly 20% cost reduction and 3-5% GWP growth from share gains and productivity.
Agentic effort concentrates in claims, underwriting and servicing; distribution and customer service follow rather than lead.
Work shifts from human coordination to agent-led execution, with human judgement reserved for complex risk and accountability.
STP on simple claims has gone from roughly 10-15% to 70-90% at leading carriers, and time-to-quote on standard SME risk from days to minutes.
More than 20 US jurisdictions have adopted the NAIC model bulletin on insurer use of AI, which expects governance, testing and documentation.
Proxy discrimination is the reputational risk that ends programmes. It has to be tested for during the build, not discovered in audit.
Consulted when this program was written, in July 2026. Each entry names the claim on this page that it backs, so the pairing stays checkable if the copy is later edited.
Insurance data is unusually document-heavy and unusually regulated. The architecture that works is one where every coverage determination can be traced to a clause and every rating input to a source.
Six families, and the wording corpus is the one that decides your accuracy.
Policies, endorsements, coverage terms and the transaction history.
Guidewire PolicyCenter, Duck CreekFNOL, claim notes, payments, reserves and adjuster narrative.
Guidewire ClaimCenter, SapiensPolicy wordings, endorsements, ISO forms. Where coverage truth lives.
Forms library, ECMBroker email, ACORD forms, schedules and loss runs — the messiest input in insurance.
ACORD, email, spreadsheetsProperty, geospatial, catastrophe, credit and telematics enrichment.
Cat models, geospatial, MVRExperience data, triangles, reserves and cession records.
Actuarial marts, GLCoverage is contractual, so meaning is contractual too. If the agent’s idea of "policy", "claim" and "exposure" does not match the carrier’s, every determination it makes is unsafe.
Six subject areas, each with the entities it holds, its critical data elements and the function accountable for it.
Everyone attached to a policy and the role each one plays on it.
The contract, its coverages, limits, exclusions and endorsements.
What is insured, where it sits, and what it is worth.
Everything from first notice to settlement, including reserves and payments.
Premium, reserving, reinsurance cessions and experience.
What is sold, the rating factors, and the filings behind them.
The entities and the typed relationships between them — what the agent traverses instead of guessing joins. Select any entity.
Any person or organisation attached to a policy or claim.
The party that contracts with the carrier.
The insurance contract for a period of cover.
A specific protection granted within a policy.
The exact contract text edition attached to the policy.
A change to the contract mid-term.
The property, vehicle, life or liability covered.
A cause of loss the coverage responds to or excludes.
A notified loss made against a policy.
The estimated ultimate cost of a claim.
Money paid on a claim.
The rated offering the policy instantiates.
The classification hierarchies that make records comparable across systems.
Commercial → Property → Package → Business interruption
Natural → Windstorm → Named storm
Manufacturing → Food processing → Bakery
The definitions an agent must use rather than invent. Most wrong answers in this industry are a term used loosely.
The layers a deployment needs, what each one holds, and what the engineer owns there.
Seven layers. The clause-citation path is the one that gets audited.
Policy admin, claims, forms library, broker submissions and third-party enrichment.
Nothing. This is the upstream edge.
Policy, claim and wording feeds, each carrying the edition and effective dates it was valid under.
Issues and services the contract. Holds the wording edition attached to each policy, which is the fact most coverage answers turn on.
FNOL through to settlement, including reserves movements. The reserve history, not the final figure, is what tells you how a claim actually developed.
Risk capture, rating engines and referral rules. Where a decision was made is recorded separately from why, and joining the two is a routine piece of work.
Premium, instalments and lapse. Drives the retention questions, and disagrees with policy admin more often than anyone expects.
Treaty and facultative cessions, and the bordereaux exchanged with partners. Spreadsheet-shaped in practice, however the contract describes it.
Wordings across editions, schedules, adjuster reports and medical evidence. Multiple editions of one wording live here simultaneously.
Turning unstructured broker submissions into something a pipeline can trust.
Treating the wording corpus as one document set. Four editions of one wording differ in a single exclusion, and an agent that answers from the newest one is confidently wrong.
Document AI over ACORD forms, schedules and loss runs, with confidence-based routing.
Read-mostly access paths, change feeds and document streams from the systems of record.
Replayable, resolved, quality-checked records with their lineage back to the source row.
Reads the source’s own change log rather than polling it, so the platform sees every state a record passed through instead of only where it ended up.
Every extract kept as it arrived, immutable and timestamped. The thing you replay from when a downstream definition turns out to have been wrong.
Layout-aware extraction over the unstructured half of the estate: sectioning, tables, signatures, and the page reference every later citation depends on.
Decides that two records are the same real-world thing, with a survivorship rule and a confidence, so downstream joins are a decision rather than an assumption.
Schema, freshness and volume expectations asserted at the boundary, so a bad load fails loudly here rather than quietly two layers later.
The confidence threshold that decides what a human still sees.
Snapshot loads instead of change capture. It looks identical in a demo and it silently loses every intermediate state, which is exactly what an audit later asks you for.
Coverage ontology, governed metrics, entitlements and the wording-to-clause map.
Replayable, resolved records with lineage from ingestion.
Typed, classified, defined data an agent can be grounded on without inferring meaning.
One agreed definition per term, owned by a named person, so the number an agent quotes means what the business means by it.
The typed entities and relationships the domain actually has, so an agent can traverse "which customers are exposed to this" rather than guess from adjacent text.
Metrics defined once, in one place, with their filters and grain. Removes the class of error where the agent computed something plausible and wrong.
Sensitivity labels and the purpose each classification permits, applied at the field level and inherited by everything downstream.
Where a value came from and what it touched on the way. The layer that makes an answer defensible rather than merely correct.
Making “covered” a resolvable reference rather than a model opinion.
Letting the model infer meaning from column names. It will, it will be plausible, and nobody will notice until the number reaches a regulator or a board pack.
Clause-level retrieval over wordings, precedent claims and guideline documents.
Typed, classified, defined data from the governance layer.
Ranked, permission-trimmed, citable evidence scoped to the caller and the moment.
Vector and lexical retrieval together, because exact identifiers, codes and clause numbers are the thing semantic search is worst at.
Splits on the document’s own structure and carries effective dates, version and source into every chunk, so a retrieved passage knows when it was true.
Applies the caller’s permissions inside the query rather than filtering results afterwards, so the model never sees what the user may not.
Traverses the knowledge graph for questions that are joins rather than similarity — exposure, genealogy, ownership, causation.
Reorders candidates on relevance and returns the passage identifier behind every sentence, so the answer can be checked rather than trusted.
Constrains the corpus to the wording edition attached to the policy before ranking, so the newest edition can never win by relevance.
Clause-level, not document-level, citation — the difference an auditor cares about.
Ranking across editions and trusting the top hit. One cross-edition leak is a coverage answer that is wrong in a way the customer can enforce.
Triage, coverage determination, fraud signal and settlement recommendation under approval gates.
Ranked, permission-trimmed, citable evidence from retrieval.
Actions taken or proposed, each with the evidence, the identity and the trace behind it.
Plans, calls tools and holds the loop. Where the step budget, the timeout and the stopping condition are enforced rather than hoped for.
Typed, permissioned tools with declared schemas. An agent’s real capability surface is this list, which is why the list is a security artefact.
Durable task state with checkpoints, so a run that dies mid-way resumes instead of restarting and re-doing side effects.
Approval steps on the actions that need one, carrying enough context for the approver to actually decide rather than rubber-stamp.
Writes back into the systems people already work in, under a service identity with its own audit trail.
The fast-track boundary: what settles automatically and what always sees an adjuster.
An agent that acts under a service account rather than on behalf of the user. It will eventually do something the requesting user had no right to do, and the log will not show it.
Proxy-discrimination testing, NAIC-aligned governance, documentation and action allow-lists.
Proposed actions and drafted responses from the agent layer.
Permitted, grounded, logged output — or a refusal with a stated reason.
The rules that decide whether an action is permitted at all, evaluated before the action and independently of the model that proposed it.
Injection detection on the way in, and on the way out the checks for leakage, unsupported claims and content the domain forbids.
Verifies each assertion resolves to retrieved evidence, and fails the response rather than shipping the sentence that does not.
Model inventory, intended use, validation evidence and the sign-off that lets a model be used for a purpose. Not optional in a regulated estate.
Immutable record of what was asked, retrieved, decided and done — the artefact a reviewer reads when they do not take your word for it.
Evidence that the model was tested for unfair discrimination before it priced anything.
Guardrails that only run on the prompt. Most real failures are on the output side: the unsupported sentence, the leaked field, the claim nobody approved.
Coverage-accuracy evals, leakage tracking, drift monitoring and cost per claim.
Permitted, grounded, logged output from the guardrail layer.
Measured quality, cost and latency — and the evidence to change any of the three.
Held-out sets and graded runs on every change, so a prompt edit is a measured change rather than a hopeful one.
Blocks a release when quality drops, in CI, on the same evidence for everyone. The difference between a system and a demo.
End-to-end spans across retrieval, model and tool calls, so a bad answer can be opened and read rather than argued about.
Cost per task and per tenant, against throughput and quota. The number that decides whether the pilot can become the rollout.
Watches quality against production traffic rather than the test set, and routes real corrections back into the eval suite.
Proving the agent reduced leakage rather than merely reduced touch time.
Evaluating once, before launch. Quality moves with the data, the model and the traffic, and a system with no live measurement has no idea which of the three moved.
Rating and claims outcomes must be testable for proxy discrimination. Design the test set with the actuary, early.
Adopted in 20+ jurisdictions: expects an AI governance programme, testing and documentation you can produce on request.
A coverage answer without a clause citation is worthless. Retrieval must resolve to clause level.
Catastrophe events multiply volume overnight. Capacity and cost per claim must hold under a 10× spike.
Twelve use cases, each tied to a stage of the chain and the architecture layer that carries it. The labs build the hard ones; the capstones prove you can run one end to end.
Filter by stage or by earned autonomy. Selecting a use case jumps the value chain and the architecture to the stage and layer it depends on.
No use cases match that combination.
How a carrier deployment actually runs.
Choose one product, one claim type, one jurisdiction. Breadth is what kills insurance pilots.
Version the wording corpus and build clause-level retrieval before anything else. Everything downstream depends on it.
Build the coverage eval set with the people who make the call today. Their disagreements are the signal.
Run proxy-discrimination testing and set the fast-track boundary using measured accuracy, not appetite.
Prove it holds at 10× volume, then hand over the runbook, evals and the NAIC-aligned documentation pack.
8 modules, 80 taught hours, 40 hands-on labs and 8 assessments — every lab provisioned and graded by the SCIKIQ Agentic AI Playground. Open a module to see its labs.
End-to-end agent design, development, deployment and testing, 10 hours each. Reviewed by a Senior SCDAI Engineer against a published rubric.
Build an agent that determines coverage for a claim against the correct wording edition, cites the operative clauses and exclusions, and escalates genuine ambiguity. Prove accuracy on a labelled set.
Build a pipeline that ingests a broker submission, pre-fills the risk, checks appetite and guidelines, and produces a referral rationale — with proxy-discrimination testing evidence.
Every lab in this program follows the shape below. This is lab 04 in full — the brief you are given, the environment that is provisioned for you, the code you start from and the assertions that decide whether you passed.
Guarantee a coverage answer resolves to the wording edition attached to that policy.
The corpus holds four editions of one wording, differing in a single exclusion. Given a claim and its policy, retrieve the operative clause from the correct edition and answer the coverage question. Answering from the newest edition is the failure this lab is built to catch.
def coverage_answer(policy_id, question, index, registry):
"""Answer grounded in the wording edition bound to this policy."""
edition = registry.edition_for(policy_id) # never assume 'latest'
# constrain retrieval to `edition` BEFORE ranking, not after
raise NotImplementedError
Every module ends with a timed, randomised assessment delivered through the SCIKIQ Agentic AI Playground. The certificate requires a pass on all of them plus two reviewed capstones.
What each test covers, how long it runs, and how many items are currently in the versioned bank behind it.
| # | Assessment & coverage | Items | Time | In bank |
|---|---|---|---|---|
| 01 | Operating model & scopingDelivery mandateSegment selectionValue sizingRegulatory posture | 25 | 35 min | 4 |
| 02 | Insurance data landscapePolicy and claims dataACORD and submissionsWording versioningEnrichment | 25 | 35 min | 4 |
| 03 | Context engineeringClause-level groundingStructured outputsAmbiguity and escalationCost | 25 | 35 min | 4 |
| 04 | Retrieval & groundingClause chunkingVersion fidelityPrecedent retrievalRetrieval metrics | 25 | 40 min | 4 |
| 05 | Agent design & autonomyTriage and determinationFast-track boundariesMulti-agent flowContainment | 25 | 40 min | 4 |
| 06 | Systems integrationMCP tool designWrite-back safetyIdentitySurge and resilience | 25 | 35 min | 4 |
| 07 | Fairness, compliance & securityProxy discriminationNAIC expectationsLLM securityRecourse | 30 | 45 min | 4 |
| 08 | Production & operationsEval gatingLeakage measurementDriftUnit economics and handover | 30 | 45 min | 4 |
Real items, drawn from 32 in this program's bank — weighted toward the scenario and diagnosis types, because those are the ones that predict field performance. Instant feedback, nothing saved.
Four sample items — one attempt each, then the reasoning is shown.
Q1Why is segment selection the first decision in a carrier deployment?
One product, one claim type, one jurisdiction. Every additional dimension multiplies the wording, data and regulatory surface.
Q2Why is wording versioning critical in insurance retrieval?
Every policy attaches to a specific wording edition. Retrieval must resolve to that edition, not to the newest or most similar one.
Q3A wording is genuinely ambiguous. What should a well-designed agent do?
Resolving ambiguity is a human judgement with legal consequences. Surfacing it is the agent’s job; deciding it is not.
Q4Why chunk wordings at clause boundaries rather than fixed token windows?
Chunking should follow document structure. A clause split across chunks retrieves partially and cites incorrectly.
The same ladder whichever specialisation you enter through — what changes is the domain you go deep in. Below: how the program is delivered, the skills it moves, the roles it leads to, and the specialisations closest to this one.
The same labs, assessments and capstones, delivered to an enterprise cohort or to individual professionals.
Cohorts of 20 to 2,000+ on your own tenancy, with your data patterns and your cloud. Skill-gap baselining up front, per-team mastery reporting throughout, and capstones scoped against your real backlog so the output is deployable work.
The same labs, assessments and capstones for individual engineers and analysts, run on shared infrastructure with a fixed cohort calendar. You leave with a graded portfolio, not a certificate of attendance.
Find your row and aim one column right. The Playground scores you against this after every module.
| Skill | Beginner | Intermediate | Advanced |
|---|---|---|---|
| Insurance domain fluency | Knows the products. | Maps the chain to decisions and owners. | Sizes leakage and designs the operating change. |
| Coverage grounding | Retrieves a document. | Retrieves the right clause and edition. | Designs version-safe, auditable citation. |
| Fairness engineering | Aware of bias risk. | Runs proxy-discrimination tests. | Designs testable fairness into the build. |
| Context engineering | Writes clear prompts. | Structures retrieval, tools and state deliberately. | Designs context strategy for reliability and cost at scale. |
| Retrieval & grounding | Builds basic vector search. | Tunes chunking, hybrid search and reranking. | Designs graph + vector grounding with measured recall. |
| Agent orchestration | Runs a single tool-calling agent. | Builds supervised multi-step and multi-agent flows. | Designs autonomy boundaries and failure containment. |
| Tool & system integration | Calls a documented API. | Writes an MCP server over a system of record. | Designs a least-privilege tool estate across systems. |
| Evaluation | Eyeballs outputs. | Builds labelled eval sets and regression gates. | Runs online evals with drift and judge calibration. |
| Observability & cost | Reads logs. | Traces runs, tracks tokens and latency. | Owns cost per task and capacity planning in production. |
| Security & guardrails | Adds output filters. | Mitigates the OWASP LLM Top 10 in a build. | Threat-models an agent estate and proves controls. |
| Client delivery | Takes notes in a workshop. | Runs discovery and scopes a thin slice. | Owns the account technically, from scope to handover. |
The SCDAI ladder is the same whichever specialisation you enter through — what changes is the domain you go deep in.
Skill a team, or join a cohort
B2B cohorts run on your tenancy with capstones scoped to your backlog. B2C cohorts run on a fixed calendar.
Stated plainly enough to rule yourself in or out without a sales call: the prerequisites, how the program runs, exactly what the credential is worth, and the questions everyone asks.
Stated plainly so you can rule yourself in or out without a sales call. Nothing here is a formal qualification — it is what the first lab assumes you can already do.
You should already be able to do these
What we assume, and what we teach
What the program asks of your week
Cohort dates and pricing are confirmed on enquiry rather than printed here, because both move with the intake.
The credential is awarded per specialisation, so it names the domain or stack you were assessed in rather than claiming general competence. On this program the badge reads SCDAI — Insurance.
A certificate that cannot be checked is decoration. Every award resolves to a record showing the specialisation, the award date and the assessments passed.
This is the credential the program awards, shown exactly as it is issued — with the specialisation named, the assessment record attached and a verification link anyone can check without an account.
This is to certify that
Your name
has been assessed and certified as
SCIKIQ Certified Data and AI Engineer
Insurance
Not “Data and AI Engineer” but the domain or stack you were actually assessed in. A general claim would be a weaker one.
Modules passed, labs graded and both capstones reviewed — so the credential states what was measured rather than that you attended.
The credential ID resolves to a public record showing the specialisation, the award date and the assessments passed.
Add it to your LinkedIn profile in one step. The link pre-fills the certification fields from the credential record, so the entry on your profile matches the record a reader can check.
Add to LinkedIn profile The button is live on your real certificate; here it opens LinkedIn pre-filled with this specialisation so you can see exactly what the profile entry will say.A certificate that cannot be checked is decoration. Every award resolves to a record showing the specialisation, the award date and the assessments passed.
The objections that come up in every conversation about this program, answered without the brochure voice.
Both, and the second is the point. Eight timed assessments and two reviewed capstones stand between you and the credential, so a pass means someone measured the skill rather than recorded your attendance.
Every lab is provisioned, graded and unblocked by the Playground rather than by an instructor. That is what lets a cohort of 2,000 cost the same faculty time as a cohort of 20 — and why you are never waiting on someone to mark your work.
Two retakes are included per assessment, each drawing a fresh item set from the bank, so a retake is a genuinely new paper rather than the same questions again.
For the domain specialisations, no — labs run in provisioned sandboxes. For the four tech-stack programs you will want access to that platform, since deploying into a real subscription is much of the point.
Yes. B2B cohorts run on your own tenancy with your data patterns, a skills baseline before kick-off, per-team mastery reporting, and capstones scoped against your actual backlog so the output is deployable work rather than an exercise.
Take the domain you deploy into. If you move across industries, take a tech-stack program instead and pick up domain context on the engagement. The chooser on the programs page will narrow it.
The trends, platform capabilities and regulatory positions are reviewed each quarter, and every external claim on these pages links to its source so you can check the date yourself.
A graded portfolio: forty machine-graded labs, two reviewed end-to-end agent builds with measured evaluation and cost per task, and a verifiable credential naming your specialisation.
Still deciding?
Tell us the systems you deploy into and we will say plainly whether this specialisation is the right one — or which of the 21 is.