Back to Services

GenAI applications for SAP

Generative AI that reaches production - built on SAP BTP and AI Core, grounded in your business data, evaluated before release, and running with cost and quality in plain sight.

What we build

Why most GenAI pilots never ship

A demo needs one good answer; production needs a distribution of acceptable ones. The gap is almost never the model - it is grounding, evaluation, cost and the operational question of what happens when the answer is wrong. Those are engineering problems, and they are the ones we spend our time on.

The other common failure is scope. A tool that drafts one document well beats a general assistant that does everything approximately, because only the first one can be measured and improved.

How we work

We start from a use case with a real owner and a real definition of “good”, build a working prototype quickly, then put evaluation and guardrails around it before it reaches users. Applications ship instrumented with Boidra Platform - traces, token cost and judge scores in one console.

Where the work is conversational or process-driven, it usually belongs with agents; where it needs live SAP facts, it sits on MCP servers. Most real projects use all three.

Frequently asked questions

What counts as a GenAI application here?

Anything where a model does real work inside a business flow - summarising a case, drafting a document from SAP data, classifying incoming requests, extracting structure from a PDF and posting the result. It is an application with a model in it, not a chat window bolted to the side.

Which models, and where do they run?

Usually through SAP AI Core and its generative-AI hub, which gives you a choice of models behind one contract and keeps traffic inside your BTP account. Where a use case needs a model the hub does not carry, we integrate it directly - with the same telemetry and the same guardrails.

How do you know the output is good enough to ship?

We build an evaluation set from real examples before release, and score against it with LLM-as-judge plus whatever deterministic checks the domain allows. That turns "it seems better" into a number you can compare across versions.

What does it cost to run?

That depends on the use case, which is exactly why we instrument it. Token spend and Capacity Units are attributed per feature, so the running cost is a line you can see and manage rather than a surprise at renewal.

Talk to us about a GenAI applicationSee how we deliver

Have a use case in mind?

Tell us what you’re trying to automate in SAP. We’ll tell you - honestly - whether an agent is the right tool, and how we’d start.

Give us a use case →