SpendLens AI

Workload economics

What does one useful AI result cost?

See what each workload costs, find wasted attempts, and measure whether a cheaper model actually saves money. A workload is simply a job your application does, such as answering support questions or reading invoices.

A simple support example

Imagine your support assistant spends $10 answering 100 questions, but only 80 answers are usable. Each useful answer costs $0.125: $10 divided by 80.

Your provider bill shows the $10 total. SpendLens AI helps explain the cost per useful answer, including retries, refusals and cache savings. Your application supplies whether an answer was usable; token counts alone cannot tell us that.

A cheaper model might cost half as much per call but need three attempts to get the answer right. Workload economics helps you see whether that change really saves money.

What appears in the dashboard

Open Analyze → Workload economics to see cache-hit rate, actual cache savings, estimated avoidable uncached spend, refusal and retry costs, cost per usable outcome, and gross versus net routing savings.

Missing facts remain unavailable. SpendLens AI does not interpret a missing overhead as zero or treat an HTTP 200 response as proof of a usable business result.

Enable Anthropic cache diagnostics

Diagnostics are enabled in your own Anthropic request. SpendLens AI observes the returned reason and estimated missed-token count; it does not add request parameters or retain diagnostic hashes.

response = anthropic.beta.messages.create(
    model="claude-opus-5-5",
    diagnostics={"previous_message_id": previous_id},
    messages=messages,
    max_tokens=1024,
)

# The SDK captures only scalar diagnostic facts such as:
# cache_miss_reason, cache_missed_input_tokens,
# previous_message_id_available and cache_diagnostic_status.

cache_missed_input_tokens is an Anthropic estimate and may differ from billed input tokens. Avoidable uncached spend is therefore an estimate, not an invoice adjustment.

Already choosing between models? Measure the savings

Some applications choose a cheaper model for simple requests and a stronger model for difficult ones. SpendLens AI compares the selected model with the model the application would otherwise have used, then subtracts retry and routing costs.

Example: a password-reset question goes to a cheaper model, while a difficult contract review goes to a stronger model.

A direct provider client normally chooses models from the same provider: OpenAI Sol → Luna, or Claude Opus/Sonnet → Haiku. An OpenAI client cannot call Claude, and an Anthropic client cannot call an OpenAI model. Cross-provider selection requires a separate customer-managed gateway or router with credentials for both providers.

If your application always calls the same model, skip this section. SpendLens AI does not choose or route models.

Illustrative example: the same observed token usage would cost $0.01 on Sol. The chosen Luna call costs $0.002, and deciding which model to use costs $0.00002. With no other overhead or retries, estimated net savings are $0.00798.

The baseline cost is an estimate, not a second Sol call. This comparison measures cost; it does not prove that Luna produces equally good answers. Validate quality before changing production models.

classification_cost means the cost of deciding which model to use for this request. It is unrelated to SpendLens AI’s internal workload classification. All cost fields are in US dollars. Use zero only when you know the cost is zero; leave unknown costs unspecified and net savings will remain unavailable.

Show the advanced SDK metadata
# Your application chooses between models from the same provider.

# Sol is the model you would normally use.
baseline_model = "gpt-6.1-sol"

# Placeholder for YOUR application logic: easy question -> Luna.
# Implement this function yourself; it is not a SpendLens AI function.
selected_model = choose_openai_model(request)

with client.tag(workload="customer-support", metadata={
    # A label for your application's model chooser.
    "router_name": "support-router",

    # The usual model, used to estimate what this call would have cost.
    "requested_model": baseline_model,

    # The model your application actually chose.
    "selected_model": selected_model,

    # USD spent deciding which model to use, for THIS request.
    # Example: a separate AI call checks whether the question is easy.
    # For a simple if/else rule with no extra billed cost, use 0.
    "classification_cost": 0.00002,

    # Extra USD spent running the chooser, excluding the AI answer.
    "routing_overhead_cost": 0,

    # Extra USD spent retrying this request. Use 0 if there were no retries.
    # If you report linked retry calls separately, do not count them here too.
    "incremental_retry_cost": 0,

    # Your application's identifier connecting attempts for this request.
    "trace_id": trace_id,

}):
    # OpenAI answers using the chosen model; SpendLens AI records usage.
    response = client.responses.create(model=selected_model, ...)

Usable outcomes and refusals

The SDK automatically recognizes Anthropic refusals, including refusals returned with HTTP 200. Set usable_result only when your application knows whether the request completed its business task. This keeps provider transport success separate from workload success.

Upgrade and save recommendations

Candidates use the workload’s observed token, cache and processing-mode mix, company provider restrictions and available evaluation evidence. Every candidate remains marked Quality validation required; a lower price is never treated as proof of equivalent quality.

Privacy and compatibility

New fields are optional. Older ingest clients continue to work. Current Python and Node.js SDKs accept legacy prompt-sampling settings for compatibility but always emit metadata-only telemetry: no prompts, responses, tool definitions, message history, provider hashes or diagnostic objects.