When AI pilots stall in banks, the post-mortem rarely blames the model. It blames the data: definitions that differ between finance and risk, customer records that cannot be joined, documents with no metadata, lineage that stops at the data warehouse. Agents make the problem sharper. A copilot that misreads a table produces a bad answer that a person may catch; an agent that misreads it takes a bad action.
The fix is not a bigger data lake. AI-ready data is data whose meaning, provenance and quality are known at the moment an agent uses it, which means moving metadata from passive documentation to an active, continuously updated asset.
Banks already have the mandate — but not the finish
Banks have been under a supervisory obligation to fix their data for over a decade. The Basel Committee's principles for effective risk data aggregation and risk reporting, BCBS 239, require accurate, complete and timely risk data with clear governance and architecture. Yet the Committee's November 2023 progress report found that only two of the 31 G-SIBs assessed were fully compliant with all principles, and that no single principle was fully implemented across all banks1. The European Central Bank's 2024 guide raised the bar further, expecting complete and up-to-date data lineage at data-attribute level — from capture through extraction, transformation and loading — for the risk indicators in scope2.
BCBS 239: a decade on, full compliance is rare
Global systemically important banks assessed by the Basel Committee, 2023 (banks)
Note: 31 G-SIBs assessed; 'not yet fully compliant' is the remainder (31 − 2).
The overlap with AI readiness is large. The lineage, data dictionaries, quality controls and ownership that BCBS 239 requires are exactly what an agent needs to know where a number came from, what it means and whether it can be trusted. Banks that treat BCBS 239 as a reporting exercise miss the chance to make it their AI data foundation.
From tables to meaning: metadata, ontology and knowledge graphs
Large language models are good at language and weak at institutional meaning. They do not know that ‘exposure’ in the credit-risk mart is post-mitigation while in finance it is gross, or that two customer identifiers refer to the same legal entity. Three layers supply that meaning:
- Active metadata. Technical, business and operational metadata — definitions, owners, quality scores, usage — captured continuously rather than documented once and forgotten.
- An ontology or semantic layer. A shared model of the bank's core concepts (customer, account, facility, exposure, product) and how they relate, so that finance, risk and operations query the same meaning.
- Knowledge graphs. Entities and relationships instantiated from the ontology — ownership chains, guarantor links, product hierarchies — that agents can traverse.
Knowledge graphs also improve retrieval. Microsoft Research found that GraphRAG, which builds a knowledge graph from source documents, substantially outperformed baseline vector-search RAG on questions that require connecting disparate information or summarising themes across a whole dataset3. In banking, those are precisely the questions that matter: who ultimately owns this counterparty, and what is our total exposure to the group?
Context engineering is the new data discipline
Anthropic describes context engineering as the set of strategies for curating and maintaining the optimal set of tokens an LLM sees during inference4. For a bank, that translates into data-management questions: which policies, records and definitions should an agent retrieve for this task; which fields must be masked; how fresh must the data be; and how is each retrieved item traced back to source? Poor context engineering is the root cause of many agent errors that get blamed on the model.
What AI-ready data adds to BCBS 239 foundations
Mapping supervisory data requirements to agent needs
| Foundation | Supervisory anchor | What agents need from it |
|---|---|---|
| Attribute-level lineage | ECB RDARR guide2 | Provenance to cite in every answer and action |
| Data quality controls | BCBS 239 principles1 | Confidence signals to decide when to escalate to a human |
| Active metadata | Data dictionaries and ownership | Definitions, owners and freshness at retrieval time |
| Semantic layer / ontology | Consistent risk and finance definitions | Shared meaning across finance, risk and operations |
| Knowledge graph + RAG | Entity and ownership resolution | Multi-hop reasoning over relationships3 |
Note: SCIKIQ synthesis; citations indicate the anchoring source for each row.
An agent is only as good as the context it is given. In a bank, context is lineage, definitions and entitlements — which is to say, data management. (SCIKIQ view)
Two further foundations are often overlooked. The first is entitlements: an agent retrieving context on behalf of a user must see only what that user — and the agent itself — is permitted to see, which means access policies must be expressed in the data layer, not only in applications. The second is unstructured content. Much of the knowledge agents need sits in credit memos, policies, contracts and emails. Bringing those documents under the same metadata, classification and lineage disciplines as structured data is what allows an agent to combine a covenant clause with a live exposure figure and cite both.
Data quality, finally, needs to become machine-readable. Quality scores and freshness indicators that an agent can read at retrieval time allow it to decide when to proceed, when to caveat an answer and when to hand the task to a human. That is a small change in data engineering with a large effect on the reliability of agentic workflows.
A pragmatic sequence
Banks do not need to finish an enterprise ontology before starting. The sequence that works is use-case led: pick a domain where agents will act — reconciliation, KYC, regulatory reporting — model its core concepts, connect lineage and quality scores for the data it uses, and expose that through a governed retrieval layer. Each domain adds to the shared semantic layer, and the second domain is faster than the first.