AI

Telemetry in AI: why it matters and how it improves testing

Boidra Expert3 min read

Traditional software is deterministic: same input, same output, so a pass/fail assertion is enough. AI systems are not. The same prompt can yield different answers, an agent can take a different path each run, and "correct" is often a matter of degree. That is exactly why telemetry - instrumenting every step so you can see what happened - is not a nice-to-have for AI. It is the foundation testing stands on.

What "telemetry" means for an AI agent

Telemetry is the practice of emitting a span for every step an agent takes. A span captures:

  • What ran - the tool called, the model invoked, the prompt used.
  • Inputs and outputs - what went in, what came back.
  • Cost and latency - tokens consumed, milliseconds elapsed.
  • A quality signal - a score, ideally from an LLM-as-judge against criteria you define.

String those spans together and you get a trace: the full, inspectable story of one agent run.

from dataclasses import dataclass
 
 
@dataclass
class Span:
    name: str          # e.g. "retrieve_context"
    inputs: dict
    outputs: dict | None = None
    latency_ms: float = 0.0
    tokens: int = 0
    score: float | None = None   # 0..1, from an LLM-as-judge

Why telemetry matters

1. You cannot debug a black box

When an agent does something wrong - or surprisingly right - logs at the HTTP boundary tell you nothing about the reasoning step that caused it. A span does. "The AI did something weird" becomes an inspectable, comparable trace.

2. Non-determinism needs measurement, not assertions

You cannot assert output === expected when output legitimately varies. What you can do is measure a distribution - score many runs and track whether quality holds. Telemetry is what makes that measurement possible.

3. Cost and latency are correctness too

An agent that produces the right answer for ten dollars and forty seconds is failing in a way a boolean test never catches. Per-span token and latency data surfaces it.

How telemetry improves testing

This is where it pays off. Once every step emits a scored span, several testing practices that are impossible on a black box become routine:

Evaluation sets ("evals")

Build a fixed set of representative inputs. Run the agent over them and let an LLM-as-judge score each output against a rubric. Now you have a quantitative quality baseline - a number, not a vibe.

Regression testing across versions

Changed a prompt, swapped a model, tweaked orchestration? Re-run the eval set and compare score distributions between versions. A drop from 0.91 to 0.78 on plan_actions is a regression you can see - before it ships.

Step Latency (ms) Tokens Score (v1) Score (v2)
retrieve_context 120 450 0.95 0.95
plan_actions 340 1 200 0.91 0.78
call_sap_api 80 0 1.00 1.00

Pinpointing where it broke

Because the trace is step-by-step, a failing run tells you which span degraded - not just that the final answer was wrong. That turns a vague "the agent is worse now" into "plan_actions regressed after the model swap."

Threshold gates in CI

Set a minimum score per step. Any run - or any new version - that dips below the threshold routes into human review or fails the build. Telemetry makes AI quality a gate, like test coverage.

Observability isn't a dashboard you bolt on afterwards. It's a span you emit on every step - and it's the only thing that makes an AI system genuinely testable.

The loop

  1. Instrument - a span on every agent step.
  2. Score - LLM-as-judge against your rubric.
  3. Evaluate - run the eval set, get a baseline.
  4. Gate - regression-test every change against that baseline.

Each stage depends on the telemetry beneath it. No spans, no scores; no scores, no evals; no evals, no regression testing.

Where Boidra fits

This is exactly what Boidra Platform does - a span for every agent step, LLM-as-judge scoring, and the traces surfaced inside S/4HANA so your team can test and trust agents without leaving SAP. For a hands-on, code-level walkthrough, see Building observable AI agents.

ShareXLinkedIn

Keep reading