Agentic AI orchestrates the close
BCG’s 2026 finance perspective is that agentic AI moves from task assistance to orchestrating multi-step workflows such as the accounting close.
Finance is the function where an agent is a control, not a tool. If it touches anything in scope for internal controls over financial reporting, your auditor will ask how it works, who reviews it and what evidence exists. Build for that question and finance becomes the fastest domain to deploy in.
Finance is a chain of controls as much as a chain of activities. Part 1 maps record-to-report, order-to-cash, procure-to-pay and the assurance layer over all of them.
Six stages from transaction to assurance. Select one.
Procure-to-pay, order-to-cash and the capture of every accounting event.
Account reconciliation, intercompany matching and balance substantiation.
Period-end close, journals, accruals, consolidation and flux analysis.
Statutory reporting, management reporting, disclosures and narrative.
FP&A, budgeting, forecasting, scenario analysis and business partnering.
Internal controls, SOX testing, internal audit and external audit support.
Five roles, and the auditor is the one who decides whether it stays. Select one to light the stages they own.
Finance is adopting quickly — and being audited on it just as quickly.
BCG’s 2026 finance perspective is that agentic AI moves from task assistance to orchestrating multi-step workflows such as the accounting close.
Finance functions are using AI for forecasting, real-time working capital monitoring and faster reporting cycles — with human checks retained on judgement.
Agentic workflows can automate more than half of close tasks and cut days-to-close by around 30%, with anomaly detection running continuously.
Big Four firms have trained staff specifically to scrutinise AI-generated evidence, AI-touched controls and AI-driven exception resolution.
Continuous evidence across 100% of transactions lets auditors shift from sampling to reviewing exception reports — a structural change in how assurance works.
The prudent 2026 posture is agentic assistance under human review for documentation, evidence collection and analysis — not autonomous conclusion-forming.
Consulted when this program was written, in July 2026. Each entry names the claim on this page that it backs, so the pairing stays checkable if the copy is later edited.
In finance the architecture question is not “can it compute the number” but “can we evidence how the number was produced”. Lineage and immutability are the load-bearing layers.
Six families, all of which an auditor may ask you to trace end to end.
Journals, balances, master data and the sub-ledgers behind them.
SAP S/4HANA, Oracle, NetSuiteAP, AR, payroll, treasury, expenses and procurement.
Ariba, Coupa, ConcurConsolidation, disclosure management and management reporting.
EPM, consolidation toolsControl matrices, test results, findings and evidence repositories.
GRC platforms, audit toolsContracts, accounting policy, standards and internal guidance.
CLM, policy repositoriesThe non-financial data that explains a variance — volume, headcount, activity.
Operational systems, HRFinance already has a semantic layer — the chart of accounts and the close calendar. The work is making it machine-usable, so an agent quotes the number the business reports rather than one it derived itself.
Six subject areas, each with the entities it holds, its critical data elements and the function accountable for it.
The legal and management structures the numbers roll up through.
The account structure giving every posting its meaning.
Everything that posts, and the documents behind it.
The period-end process and what it produces.
The control environment and the evidence it generates.
Budgets, forecasts and the drivers behind them.
The entities and the typed relationships between them — what the agent traverses instead of guessing joins. Select any entity.
A reporting unit within the group.
A chart-of-accounts node giving a posting meaning.
The management unit a cost belongs to.
A posting to the general ledger.
A billed or received demand for payment.
A commitment to buy.
Settlement of an obligation.
The accounting period a posting falls in.
Substantiation of a balance at period end.
A procedure mitigating a financial reporting risk.
A control test result outside tolerance.
A reported statement or note.
The classification hierarchies that make records comparable across systems.
P&L → Operating expense → Travel → Airfare
Order to cash → Revenue cut-off → Manual review of period-end shipments → Sample 25
Group → Industrial → EU Manufacturing Ltd → Plant maintenance
The definitions an agent must use rather than invent. Most wrong answers in this industry are a term used loosely.
The layers a deployment needs, what each one holds, and what the engineer owns there.
Seven layers, built so that every figure resolves to a source.
ERP, sub-ledgers, consolidation and evidence repositories — read through controlled paths.
Nothing. This is the upstream edge.
Period-locked balances, sub-ledger detail and the policy corpus, each pinned to a close period.
The book of record for the entity. Balances by account, period and dimension, closed and locked on a calendar that does not move for you.
Receivables, payables, fixed assets and revenue. Detail that must reconcile to the ledger, and the place reconciliation breaks are actually caused.
Group structure, intercompany elimination, currency translation and the close checklist that governs the whole cycle.
Requisitions, purchase orders, receipts and expense claims. The highest-volume transaction population and the usual target for testing.
Control descriptions, test plans, evidence and findings. Where an assertion has to be supported by something a reviewer can open.
Accounting policy, group manuals and the standards themselves — the corpus a treatment question is grounded on.
Establishing read access that IT and audit both accept.
Reading a period that is still open. A number that moves after you have cited it is worse than no number, so pin the period and record which close you answered from.
Extracts landed immutably with timestamps, so evidence cannot be quietly rewritten.
Read-mostly access paths, change feeds and document streams from the systems of record.
Replayable, resolved, quality-checked records with their lineage back to the source row.
Reads the source’s own change log rather than polling it, so the platform sees every state a record passed through instead of only where it ended up.
Every extract kept as it arrived, immutable and timestamped. The thing you replay from when a downstream definition turns out to have been wrong.
Layout-aware extraction over the unstructured half of the estate: sectioning, tables, signatures, and the page reference every later citation depends on.
Decides that two records are the same real-world thing, with a survivorship rule and a confidence, so downstream joins are a decision rather than an assumption.
Schema, freshness and volume expectations asserted at the boundary, so a bad load fails loudly here rather than quietly two layers later.
Proving the data the agent saw is the data that existed at that moment.
Snapshot loads instead of change capture. It looks identical in a demo and it silently loses every intermediate state, which is exactly what an audit later asks you for.
Chart of accounts semantics, governed metrics and entitlements by entity and role.
Replayable, resolved records with lineage from ingestion.
Typed, classified, defined data an agent can be grounded on without inferring meaning.
One agreed definition per term, owned by a named person, so the number an agent quotes means what the business means by it.
The typed entities and relationships the domain actually has, so an agent can traverse "which customers are exposed to this" rather than guess from adjacent text.
Metrics defined once, in one place, with their filters and grain. Removes the class of error where the agent computed something plausible and wrong.
Sensitivity labels and the purpose each classification permits, applied at the field level and inherited by everything downstream.
Where a value came from and what it touched on the way. The layer that makes an answer defensible rather than merely correct.
One definition of revenue, margin and headcount that an agent must use.
Letting the model infer meaning from column names. It will, it will be plausible, and nobody will notice until the number reaches a regulator or a board pack.
Retrieval over accounting policy, contracts, prior filings and control documentation.
Typed, classified, defined data from the governance layer.
Ranked, permission-trimmed, citable evidence scoped to the caller and the moment.
Vector and lexical retrieval together, because exact identifiers, codes and clause numbers are the thing semantic search is worst at.
Splits on the document’s own structure and carries effective dates, version and source into every chunk, so a retrieved passage knows when it was true.
Applies the caller’s permissions inside the query rather than filtering results afterwards, so the model never sees what the user may not.
Traverses the knowledge graph for questions that are joins rather than similarity — exposure, genealogy, ownership, causation.
Reorders candidates on relevance and returns the passage identifier behind every sentence, so the answer can be checked rather than trusted.
Grounding treatment questions in the standard and the policy, not the model.
Filtering results after retrieval instead of constraining the query. The model has already seen the rows you removed, and it will use them.
Close orchestration, reconciliation, exception triage, commentary drafting and control testing.
Ranked, permission-trimmed, citable evidence from retrieval.
Actions taken or proposed, each with the evidence, the identity and the trace behind it.
Plans, calls tools and holds the loop. Where the step budget, the timeout and the stopping condition are enforced rather than hoped for.
Typed, permissioned tools with declared schemas. An agent’s real capability surface is this list, which is why the list is a security artefact.
Durable task state with checkpoints, so a run that dies mid-way resumes instead of restarting and re-doing side effects.
Approval steps on the actions that need one, carrying enough context for the approver to actually decide rather than rubber-stamp.
Writes back into the systems people already work in, under a service identity with its own audit trail.
Separate preparer and reviewer agents with disjoint tool sets, so no single identity can both propose and approve a journal.
Segregation of duties: the agent that prepares must not be the one that approves.
One agent with both propose and approve tools. It is a control failure the moment it exists, whether or not it is ever used.
Control design for AI-touched processes, review evidence, access control and audit trail.
Proposed actions and drafted responses from the agent layer.
Permitted, grounded, logged output — or a refusal with a stated reason.
The rules that decide whether an action is permitted at all, evaluated before the action and independently of the model that proposed it.
Injection detection on the way in, and on the way out the checks for leakage, unsupported claims and content the domain forbids.
Verifies each assertion resolves to retrieved evidence, and fails the response rather than shipping the sentence that does not.
Model inventory, intended use, validation evidence and the sign-off that lets a model be used for a purpose. Not optional in a regulated estate.
Immutable record of what was asked, retrieved, decided and done — the artefact a reviewer reads when they do not take your word for it.
Designing the compensating control that makes the agent auditable.
Answering from an open period. The number moves after you cite it, and every downstream working paper is now wrong.
Accuracy evals, exception-rate monitoring, drift and cost per transaction.
Permitted, grounded, logged output from the guardrail layer.
Measured quality, cost and latency — and the evidence to change any of the three.
Held-out sets and graded runs on every change, so a prompt edit is a measured change rather than a hopeful one.
Blocks a release when quality drops, in CI, on the same evidence for everyone. The difference between a system and a demo.
End-to-end spans across retrieval, model and tool calls, so a bad answer can be opened and read rather than argued about.
Cost per task and per tenant, against throughput and quota. The number that decides whether the pilot can become the rollout.
Watches quality against production traffic rather than the test set, and routes real corrections back into the eval suite.
Evidence that the control operated effectively throughout the period, not just today.
Evaluating once, before launch. Quality moves with the data, the model and the traffic, and a system with no live measurement has no idea which of the three moved.
If the agent touches a process in scope for internal controls over financial reporting, it is part of the control environment and will be tested.
An agent that both prepares and approves breaks a fundamental control. Separation must be architectural.
Calculations must be deterministic and reproducible. Use the model for language and judgement support, not arithmetic.
Evidence that can be regenerated differently later is not evidence. Stage immutably and timestamp everything.
These use cases sit deliberately on the assist-and-approve side of the line, because that is what an auditor will accept in 2026 — and because it is where the value actually is.
Filter by stage or by earned autonomy. Selecting a use case jumps the value chain and the architecture to the stage and layer it depends on.
No use cases match that combination.
How a finance deployment actually runs.
Determine whether the process is in scope for internal controls. That answer sets the documentation and review burden for everything that follows.
Decide who reviews what and how evidence is produced before writing the agent. Retrofitting a control around a working agent never goes well.
Build the evidence path first. If you cannot show the auditor what data the agent saw, nothing else you build will survive.
Run the agent alongside the existing process for a full cycle. Compare outputs and document every difference. This is your operating effectiveness evidence.
Do the walkthrough before go-live, not after. Hand over the control documentation, runbook and the continuous evidence pipeline.
8 modules, 80 taught hours, 40 hands-on labs and 8 assessments — every lab provisioned and graded by the SCIKIQ Agentic AI Playground. Open a module to see its labs.
End-to-end agent design, development, deployment and testing, 10 hours each. Reviewed by a Senior SCDAI Engineer against a published rubric.
Build an agent that reconciles an account, explains residual differences with attached evidence and reports close status — with preparer/reviewer separation enforced architecturally.
Build continuous testing of a control across the full transaction population, producing an exception report and an evidence trail an external auditor could rely on.
Every lab in this program follows the shape below. This is lab 05 in full — the brief you are given, the environment that is provisioned for you, the code you start from and the assertions that decide whether you passed.
Make it architecturally impossible for the preparing agent to approve its own work.
Build a reconciliation agent that matches transactions, explains residual differences and proposes a journal. A separate reviewer path approves. Then demonstrate to a mock auditor that the preparer cannot approve, by capability rather than by configuration.
# The preparer holds propose_journal. It does not hold post_journal.
PREPARER_TOOLS = [read_gl, read_subledger, propose_journal]
REVIEWER_TOOLS = [read_proposal, approve_journal]
def reconcile(account, period):
raise NotImplementedError
Every module ends with a timed, randomised assessment delivered through the SCIKIQ Agentic AI Playground. The certificate requires a pass on all of them plus two reviewed capstones.
What each test covers, how long it runs, and how many items are currently in the versioned bank behind it.
| # | Assessment & coverage | Items | Time | In bank |
|---|---|---|---|---|
| 01 | Operating model & control scopeDelivery mandateICFR scopeControl designValue sizing | 25 | 35 min | 4 |
| 02 | Finance data landscapeGL and sub-ledgersChart of accountsImmutable stagingEntitlements | 25 | 35 min | 4 |
| 03 | Context engineeringMetric groundingDeterminismPolicy groundingRefusal | 25 | 35 min | 4 |
| 04 | Retrieval & groundingPolicy retrievalContract retrievalPrior-period comparisonRetrieval metrics | 25 | 40 min | 4 |
| 05 | Agent design & segregation of dutiesPreparer/reviewer separationClose orchestrationException triageContainment | 25 | 40 min | 4 |
| 06 | Systems integrationMCP over ERPPosting safetyGRC integrationResilience | 25 | 35 min | 4 |
| 07 | SOX, audit & securityAI in the control environmentEvidenceWalkthroughsLLM security | 30 | 45 min | 4 |
| 08 | Production & operationsEval gatingContinuous evidenceException managementEconomics and handover | 30 | 45 min | 4 |
Real items, drawn from 32 in this program's bank — weighted toward the scenario and diagnosis types, because those are the ones that predict field performance. Instant feedback, nothing saved.
Four sample items — one attempt each, then the reasoning is shown.
Q1Why is ICFR scope the first question on a finance engagement?
In-scope means documentation, review evidence and auditor walkthroughs. Discovering that late is expensive.
Q2Why must staging be immutable and timestamped?
Evidence that can be regenerated differently later is not evidence. Immutability is what makes it auditable.
Q3Where should a figure in generated commentary come from?
Model arithmetic is unnecessary risk. Separating calculation from narration is the core discipline here.
Q4Why must accounting policy retrieval be version-aware?
Standards change and periods differ. Version-aware retrieval is the same discipline as wording versioning in insurance.
The same ladder whichever specialisation you enter through — what changes is the domain you go deep in. Below: how the program is delivered, the skills it moves, the roles it leads to, and the specialisations closest to this one.
The same labs, assessments and capstones, delivered to an enterprise cohort or to individual professionals.
Cohorts of 20 to 2,000+ on your own tenancy, with your data patterns and your cloud. Skill-gap baselining up front, per-team mastery reporting throughout, and capstones scoped against your real backlog so the output is deployable work.
The same labs, assessments and capstones for individual engineers and analysts, run on shared infrastructure with a fixed cohort calendar. You leave with a graded portfolio, not a certificate of attendance.
Find your row and aim one column right. The Playground scores you against this after every module.
| Skill | Beginner | Intermediate | Advanced |
|---|---|---|---|
| Finance fluency | Knows the cycles. | Maps the chain to controls and owners. | Sizes value and designs the control change. |
| Control engineering | Aware of SOX. | Designs review controls around agents. | Produces audit-ready operating effectiveness evidence. |
| Evidence & lineage | Logs outputs. | Stages immutably with lineage. | Designs an evidence path that survives external audit. |
| Context engineering | Writes clear prompts. | Structures retrieval, tools and state deliberately. | Designs context strategy for reliability and cost at scale. |
| Retrieval & grounding | Builds basic vector search. | Tunes chunking, hybrid search and reranking. | Designs graph + vector grounding with measured recall. |
| Agent orchestration | Runs a single tool-calling agent. | Builds supervised multi-step and multi-agent flows. | Designs autonomy boundaries and failure containment. |
| Tool & system integration | Calls a documented API. | Writes an MCP server over a system of record. | Designs a least-privilege tool estate across systems. |
| Evaluation | Eyeballs outputs. | Builds labelled eval sets and regression gates. | Runs online evals with drift and judge calibration. |
| Observability & cost | Reads logs. | Traces runs, tracks tokens and latency. | Owns cost per task and capacity planning in production. |
| Security & guardrails | Adds output filters. | Mitigates the OWASP LLM Top 10 in a build. | Threat-models an agent estate and proves controls. |
| Client delivery | Takes notes in a workshop. | Runs discovery and scopes a thin slice. | Owns the account technically, from scope to handover. |
The SCDAI ladder is the same whichever specialisation you enter through — what changes is the domain you go deep in.
Skill a team, or join a cohort
B2B cohorts run on your tenancy with capstones scoped to your backlog. B2C cohorts run on a fixed calendar.
Stated plainly enough to rule yourself in or out without a sales call: the prerequisites, how the program runs, exactly what the credential is worth, and the questions everyone asks.
Stated plainly so you can rule yourself in or out without a sales call. Nothing here is a formal qualification — it is what the first lab assumes you can already do.
You should already be able to do these
What we assume, and what we teach
What the program asks of your week
Cohort dates and pricing are confirmed on enquiry rather than printed here, because both move with the intake.
The credential is awarded per specialisation, so it names the domain or stack you were assessed in rather than claiming general competence. On this program the badge reads SCDAI — Finance, Accounting & Auditing.
A certificate that cannot be checked is decoration. Every award resolves to a record showing the specialisation, the award date and the assessments passed.
This is the credential the program awards, shown exactly as it is issued — with the specialisation named, the assessment record attached and a verification link anyone can check without an account.
This is to certify that
Your name
has been assessed and certified as
SCIKIQ Certified Data and AI Engineer
Finance, Accounting & Auditing
Not “Data and AI Engineer” but the domain or stack you were actually assessed in. A general claim would be a weaker one.
Modules passed, labs graded and both capstones reviewed — so the credential states what was measured rather than that you attended.
The credential ID resolves to a public record showing the specialisation, the award date and the assessments passed.
Add it to your LinkedIn profile in one step. The link pre-fills the certification fields from the credential record, so the entry on your profile matches the record a reader can check.
Add to LinkedIn profile The button is live on your real certificate; here it opens LinkedIn pre-filled with this specialisation so you can see exactly what the profile entry will say.A certificate that cannot be checked is decoration. Every award resolves to a record showing the specialisation, the award date and the assessments passed.
The objections that come up in every conversation about this program, answered without the brochure voice.
Both, and the second is the point. Eight timed assessments and two reviewed capstones stand between you and the credential, so a pass means someone measured the skill rather than recorded your attendance.
Every lab is provisioned, graded and unblocked by the Playground rather than by an instructor. That is what lets a cohort of 2,000 cost the same faculty time as a cohort of 20 — and why you are never waiting on someone to mark your work.
Two retakes are included per assessment, each drawing a fresh item set from the bank, so a retake is a genuinely new paper rather than the same questions again.
For the domain specialisations, no — labs run in provisioned sandboxes. For the four tech-stack programs you will want access to that platform, since deploying into a real subscription is much of the point.
Yes. B2B cohorts run on your own tenancy with your data patterns, a skills baseline before kick-off, per-team mastery reporting, and capstones scoped against your actual backlog so the output is deployable work rather than an exercise.
Take the domain you deploy into. If you move across industries, take a tech-stack program instead and pick up domain context on the engagement. The chooser on the programs page will narrow it.
The trends, platform capabilities and regulatory positions are reviewed each quarter, and every external claim on these pages links to its source so you can check the date yourself.
A graded portfolio: forty machine-graded labs, two reviewed end-to-end agent builds with measured evaluation and cost per task, and a verifiable credential naming your specialisation.
Still deciding?
Tell us the systems you deploy into and we will say plainly whether this specialisation is the right one — or which of the 21 is.