Telemetry in AI: why it matters and how it improves testing
Boidra Expert3 min read
Traditional software is deterministic: same input, same output, so a pass/fail assertion is enough. AI systems are not. The same prompt can yield different answers, an agent can take a different path each run, and "correct" is often a matter of degree. That is exactly why telemetry - instrumenting every step so you can see what happened - is not a nice-to-have for AI. It is the foundation testing stands on.
What "telemetry" means for an AI agent
Telemetry is the practice of emitting a span for every step an agent takes. A span captures:
- What ran - the tool called, the model invoked, the prompt used.
- Inputs and outputs - what went in, what came back.
- Cost and latency - tokens consumed, milliseconds elapsed.
- A quality signal - a score, ideally from an LLM-as-judge against criteria you define.
String those spans together and you get a trace: the full, inspectable story of one agent run.
from dataclasses import dataclass
@dataclass
class Span:
name: str # e.g. "retrieve_context"
inputs: dict
outputs: dict | None = None
latency_ms: float = 0.0
tokens: int = 0
score: float | None = None # 0..1, from an LLM-as-judgeWhy telemetry matters
1. You cannot debug a black box
When an agent does something wrong - or surprisingly right - logs at the HTTP boundary tell you nothing about the reasoning step that caused it. A span does. "The AI did something weird" becomes an inspectable, comparable trace.
2. Non-determinism needs measurement, not assertions
You cannot assert output === expected when output legitimately varies. What you can do is measure a distribution - score many runs and track whether quality holds. Telemetry is what makes that measurement possible.
3. Cost and latency are correctness too
An agent that produces the right answer for ten dollars and forty seconds is failing in a way a boolean test never catches. Per-span token and latency data surfaces it.
How telemetry improves testing
This is where it pays off. Once every step emits a scored span, several testing practices that are impossible on a black box become routine:
Evaluation sets ("evals")
Build a fixed set of representative inputs. Run the agent over them and let an LLM-as-judge score each output against a rubric. Now you have a quantitative quality baseline - a number, not a vibe.
Regression testing across versions
Changed a prompt, swapped a model, tweaked orchestration? Re-run the eval set and compare score distributions between versions. A drop from 0.91 to 0.78 on plan_actions is a regression you can see - before it ships.
| Step | Latency (ms) | Tokens | Score (v1) | Score (v2) |
|---|---|---|---|---|
| retrieve_context | 120 | 450 | 0.95 | 0.95 |
| plan_actions | 340 | 1 200 | 0.91 | 0.78 |
| call_sap_api | 80 | 0 | 1.00 | 1.00 |
Pinpointing where it broke
Because the trace is step-by-step, a failing run tells you which span degraded - not just that the final answer was wrong. That turns a vague "the agent is worse now" into "plan_actions regressed after the model swap."
Threshold gates in CI
Set a minimum score per step. Any run - or any new version - that dips below the threshold routes into human review or fails the build. Telemetry makes AI quality a gate, like test coverage.
Observability isn't a dashboard you bolt on afterwards. It's a span you emit on every step - and it's the only thing that makes an AI system genuinely testable.
The loop
- Instrument - a span on every agent step.
- Score - LLM-as-judge against your rubric.
- Evaluate - run the eval set, get a baseline.
- Gate - regression-test every change against that baseline.
Each stage depends on the telemetry beneath it. No spans, no scores; no scores, no evals; no evals, no regression testing.
Where Boidra fits
This is exactly what Boidra Platform does - a span for every agent step, LLM-as-judge scoring, and the traces surfaced inside S/4HANA so your team can test and trust agents without leaving SAP. For a hands-on, code-level walkthrough, see Building observable AI agents.
Keep reading
New SAP AI Core Calculator 2.0
What does an SAP AI agent really cost? Explore SAP AI Core Cost Calculator 2.0, understand Capacity Units, and compare model consumption. A worked example shows how orchestration and prompt optimization can outweigh foundation-model costs and why budgeting for the complete workflow matters.
Joule Studio Classic vs the new Joule Studio: what changed
Joule Studio went from low-code skill building (2024) to autonomous agent building (2025) to intent-based development in Joule Studio 2.0 (2026). Here is the progression and what it means for you.
How SAP Joule works: from copilot to a network of agents
Joule is SAP’s AI copilot - natively embedded in SAP apps and grounded in business data. In 2025–2026 it grew from a single assistant into an orchestrated network of specialized agents.