Outcome & Observability
Measure outcomes and feed continuous learning back into every layer.
Lesson 17 of 17 · Find it on the platform map
01 What it is
What this layer does
The outcome and observability layer measures what happened after an action and whether it helped. It tracks both the immediate result of the action and the change in the business KPI it was meant to move, alongside the health, cost and accuracy of the agents and models involved. In SCIKIQ it covers action and business outcomes, agent and model performance, an evaluation harness, OpenTelemetry tracing, adoption, exceptions and human feedback, and feeds what is learnt back into ontology, metadata, context and agents.
02 Concepts
Four ideas to hold on to
Action outcome versus business outcome
An action outcome says whether the step succeeded, such as an order being created. A business outcome says whether the KPI it targeted moved against a baseline, which is what matters to the business.
Baselines
A reference value for a KPI before an intervention, or for a comparable group without it. Without a baseline a change cannot be attributed to the action rather than to normal variation.
Evaluation harness
A repeatable test set-up for AI behaviour using golden sets of known good answers, judges that score responses, and regression runs that catch quality falling after a change.
Distributed tracing
Recording each step of a request across components as linked spans, commonly with the OpenTelemetry standard, so a slow, costly or wrong answer can be traced to the step that caused it.
03 In SCIKIQ
What the layer contains
04 What good looks like
Signs it is working
- Each automated action is linked to the KPI it was meant to affect and a baseline to compare against.
- Token use, latency and cost are visible per model and per agent.
- Changes to agents or models run against a golden set before release, and regressions block the change.
- Exceptions and human feedback are captured and reviewed, and resulting fixes are traceable to the layer they changed.
05 Diagnostics
Questions to ask your team
- 1
For our automated decisions, which KPI does each one move and how do we know it moved?
- 2
How would we notice if an agent’s answers got worse after a model change?
- 3
What does each agent cost to run, and is that cost justified by its outcome?
- 4
When users override or correct an agent, where does that feedback go and who acts on it?
06 Keep going
Related reading
— Questions
Frequently asked
Is this the same as application monitoring?
It includes technical monitoring such as tracing, latency and cost, but also measures business outcomes, prediction accuracy, adoption and human feedback, which ordinary monitoring does not.
How does learning flow back?
Outcomes, exceptions and feedback are used to improve the ontology, metadata, context and agents, so later decisions start from better information.
What is a golden set?
A curated collection of questions or cases with agreed correct answers, used by the evaluation harness to check that agent quality holds or improves between releases.