AI

New SAP AI Core Calculator 2.0

Boidra Admin9 min read

SAP AI Core Cost Calculator 2.0: Understanding AI Agent Costs

SAP AI Core costs include more than model calls. This article explains how tokens become Capacity Units, breaks down a monthly estimate, and compares model consumption. It also explains prompt optimization and why translation, testing, and other services belong in the budget.

The SAP AI Core Cost Calculator estimates monthly use in Capacity Units, or CUs. Version 2.0 adds nine starting templates, an expanded model catalog, and PDF export.

The calculator showed 68 catalog entries when checked on September 28, 2026. These include different versions, context limits, and model types. They are not 68 interchangeable models for agents.

All calculator figures below are estimates from that date.

How tokens become CUs

SAP separates three units:

Unit Meaning
Model tokens Pieces of input or output processed by a model
GenAI tokens SAP's common consumption unit across models
Capacity Units (CUs) The unit used for SAP AI Core billing

Input and output have separate conversion rates. The rates depend on the model. SAP points to SAP Note 3437766 for model details and token conversion rates.

The calculation has two steps:

GenAI tokens per request =
    (input tokens / 1,000 × input conversion rate)
  + (output tokens / 1,000 × output conversion rate)
 
CUs per request = GenAI tokens per request × CU conversion factor

SAP's Metering and Pricing for Generative AI uses 1.90385 as the CU conversion factor in its example. This is a consumption factor, not a euro or dollar price.

The page labels the example values as fictitious. Use it to understand the calculation, and check the current rates for your selected model before budgeting.

A worked conversion example

Using the values in SAP's example:

Input Value
Input tokens per request 3,500
Output tokens per request 300
Input conversion rate per 1,000 tokens 0.00112
Output conversion rate per 1,000 tokens 0.00320
CU conversion factor 1.90385
Requests per month 25,000
Input:  3,500 / 1,000 × 0.00112 = 0.00392 GenAI tokens
Output:   300 / 1,000 × 0.00320 = 0.00096 GenAI tokens
 
Total per request = 0.00488 GenAI tokens
CUs per request   = 0.00488 × 1.90385 = 0.009290788
Monthly CUs       = 25,000 × 0.009290788 = 232.2697

The result is about 232.27 CU per month for model calls, before supporting services.

SAP shows 232.25 because it rounds each request to 0.00929 CU first. Its calculation also prints a dollar sign at that step, although the formula calculates CUs. A currency price must be applied separately. Source: SAP's worked example.

How much does one CU cost?

The calculator states that your SAP contract defines the currency rate.

Monthly cost = monthly CUs × agreed price per CU

For example, at an assumed price of €1 per CU, 2,112.36 CU would cost €2,112.36. This is an illustration, not a confirmed price for your account.

Check your agreement and the SAP Discovery Center Estimator for the applicable price. Do not apply 1.90385 again to totals that are already shown in CUs.

SAP AI Core CUs are also different from the AI Units used for SAP Business AI offerings. See SAP Business AI pricing.

A monthly cost breakdown

The example below shows why model pricing alone is not enough.

Component CU/month Share of known subtotal
Foundation Models 331.41 15.69%
Orchestration 817.34 38.69%
Prompt Optimization 958.15 45.36%
Evaluations 5.46 0.26%
Compute Instances 0.00 0.00%
Object Storage Not supplied Not included
Known subtotal 2,112.36 100%

If Object Storage is zero and there are no other charges, the total is 2,112.36 CU per month.

Model calls account for 15.7% of this estimate. Orchestration and prompt optimization account for 84.1%. This split comes from the selected settings. It is not a standard split for every SAP AI application.

What is included in orchestration?

Orchestration provides a common API for model access and supporting services. These services can retrieve business data, filter content, translate text, mask personal data, and record requests.

The following settings reproduced the example's 817.34 CU orchestration total in the calculator. They are one matching configuration, not proof of the original settings.

Service Tested monthly settings Displayed CU/month
Content filtering 10,000 text blocks 5.71
Grounding 10,000 queries; storage field set to 1 GB/day 51.28
Translation 10,000 requests 740.22
Data masking 5,000 requests 11.23
Inference observability 10,000 requests 8.90
Total 817.34

Translation makes up 90.6% of this orchestration estimate. The calculator assumes about 24 text blocks per translation request. Set translation volume to match how often your application will actually use it.

The conversion factor helps explain several displayed totals:

Filtering:
10,000 × 0.00030 × 1.90385 = 5.71155 CU
 
Translation:
10,000 × 24 × 0.00162 × 1.90385 = 740.21688 CU
 
Data masking:
5,000 × 0.00118 × 1.90385 = 11.232715 CU

These calculations match the rounded calculator results. The calculator's helper text labels the starting values as CU rates, but the displayed totals behave as though a further 1.90385 conversion is applied. This is an inference from the tested numbers, not a confirmed definition of those labels.

Do not apply that multiplier to every service. Storage and observability have their own measures and conversion rules.

The grounding storage amount still needs care. The tested total is consistent with about 51.07 CU for retrieval plus 0.21 CU for storage. It does not appear to multiply the storage field by 30 days. SAP's documentation treats grounding storage as gigabyte-days. Check the storage period when preparing a monthly estimate. See SAP's metering guide.

What is prompt optimization in SAP AI Core?

Prompt optimization is a service that improves the instructions you send to a model. You provide a starting prompt, example data, a target model, and a way to score the answers. SAP AI Core runs an optimization job that tests and refines the prompt against that score.

It changes the prompt rather than retraining the model.

A typical process is:

  1. Prepare a prompt template and representative examples with expected answers.
  2. Choose the model you want to use.
  3. Choose a metric that measures whether the answers are useful.
  4. Run the optimization job.
  5. Review the resulting prompt and test it on separate examples before using it.

The optimized prompt can be saved and reused. SAP provides the workflow through SAP AI Core and SAP AI Launchpad. See SAP's prompt optimization tutorial.

For example, an agent might read a support request and return its category, urgency, and responsible team. Prompt optimization can help improve those outputs against a set of correctly labelled requests. It can also be useful when moving an existing prompt to another model.

For agents that call tools, SAP provides a tool-calling optimization tutorial. It shows how to compare a starting prompt with an optimized prompt using test data and a metric for structured output.

An improved test score does not guarantee lower costs or better results on every request. Check answer quality, prompt length, and running cost before adopting the new prompt.

How prompt optimization and evaluations affect the estimate

The calculator reproduced the example values with these inputs:

Component Monthly input Displayed CU/month Rate derived from this test
Prompt Optimization 25 optimizations 958.15 38.326 CU per optimization
Evaluations 10 node-hours 5.46 0.546 CU per node-hour

These rates are derived from the calculator settings. They are not a separate confirmation of contractual billing terms. Source: SAP AI Core Cost Calculator.

The 25 optimizations are optimization jobs, not 25 user questions. Budget for how often you plan to improve prompts. A month spent developing or migrating prompts may need more optimization work than a month of normal operation.

Evaluations measure how well a prompt and model perform. Prompt optimization uses scoring to improve the prompt; evaluations let you check and compare results. SAP supports separate training and test datasets for optimization. See SAP AI Core release notes.

The calculator estimates evaluation use in node-hours. That is why Evaluations can have a cost even when the separate Compute Instances line is zero.

For managed foundation models, zero custom compute does not mean free model calls. Token charges still apply. Custom model workloads can also have compute, storage, and baseline charges. See SAP AI Core metering.

Comparing models for agents

The catalog offers a popularity sort, but it does not provide enough evidence here to rank models by actual agent usage. The table below is a practical comparison of selected models.

Each model was tested in the calculator with:

  • 1,000 requests per month
  • 1,000 input tokens per request
  • 1,000 output tokens per request
  • Standard pricing mode

This is one million input tokens plus one million output tokens per month. The figures include only the model line item.

Model CU/month Possible use to test
GPT-5.4 Nano 1.75 Classification, routing, and simple extraction
GPT-5 Mini 2.95 Routine agent steps and structured tasks
Gemini 3.5 Flash Lite 6.04 High-volume tasks with text or other inputs
Claude 4.5 Haiku 8.49 Short assistant tasks
Gemini 3.5 Flash 21.89 Document and multimodal tasks
Claude 4.6 Sonnet 24.94 General agent tasks and coding
Claude 4.8 Opus 41.37 Difficult tasks that need more reasoning
GPT 5.6 Sol, up to 272k context tier 44.87 Complex planning and agent tasks

Source: SAP calculator, checked September 28, 2026. The suggested uses are starting points for testing, not performance results.

These are not prices for one million input tokens alone. Different input and output volumes, caching, and context tiers can change the comparison. Model availability also depends on the region.

Other catalog entries do different jobs. Embedding models help find relevant content. Cohere Reranker helps order search results. SAP RPT and Prior Labs TabPFN handle tabular tasks. They can support an agent without replacing its main conversational model.

Budget for the whole task

An agent may call a model several times to finish one task. It may plan an action, call a tool, read the result, retry, and write an answer. Each model call adds tokens.

Monthly model calls = tasks per month × average model calls per task

For example, 10,000 tasks with five model calls each produce 50,000 calls before extra retries. Each step can use a different model and a different number of tokens.

Start by testing a smaller model for routine steps. Use a more capable model where it improves results enough to justify the cost. Compare cost per completed task as well as accuracy and response time.

In this example, translation and prompt optimization together account for 1,698.37 CU per month, or 80.4% of the known subtotal. Review those two settings first. They have more effect on this estimate than a small change in model price.

ShareXLinkedIn

Keep reading