recommend-products
Strong evidenceshopping/catalog.py::recommend_products
- Current
- gpt-4o
- Starting candidate
- Gemini Flash
- Spend
- $5,210
- Savings
- $2,140/mo
Compact product-ranking prompts and stable catalog fields make this a strong replay candidate.
Sample savings report
SpendLens AI shows which workload to test first, ranks three policy-approved models, and measures cost, speed and reliability before you change production.
First Compare Mode test
Starting candidates: Gemini 2.5 Flash, DeepSeek V4 Flash and Mistral Small 4. Potential savings: $2,140/month before validation.
Projected monthly spend
$12,654
Based on 18,291 calls yesterday
Potential savings
$4,812/mo
Directional estimate before testing
Ready to evaluate
$3,284/mo
Top candidates available for staging
Payback signal
1 day
Pro plan covered by first accepted switch
Opportunity queue
shopping/catalog.py::recommend_products
Compact product-ranking prompts and stable catalog fields make this a strong replay candidate.
shopping/reviews.py::summarize_reviews
Summaries are under 500 tokens and quality risk is low after replay.
shopping/search.py::extract_filters
Structured filter extraction looks safe, but needs validation on ambiguous shopping queries.
shopping/compare.py::compare_products
Open-ended answers need human review before any rollout.
| Endpoint | Current | Starting candidate | Spend | Savings | Evidence strength |
|---|---|---|---|---|---|
recommend-products shopping/catalog.py::recommend_products | gpt-4o | Gemini Flash | $5,210 | $2,140/mo | Strong evidence Compact product-ranking prompts and stable catalog fields make this a strong replay candidate. |
summarize-product-reviews shopping/reviews.py::summarize_reviews | claude-sonnet-4 | Qwen Turbo | $3,420 | $1,144/mo | Strong evidence Summaries are under 500 tokens and quality risk is low after replay. |
extract-shopping-filters shopping/search.py::extract_filters | gpt-4o | DeepSeek Chat | $1,980 | $792/mo | Moderate evidence Structured filter extraction looks safe, but needs validation on ambiguous shopping queries. |
compare-products shopping/compare.py::compare_products | claude-sonnet-4 | Llama 3.1 70B | $2,044 | $736/mo | Early evidence Open-ended answers need human review before any rollout. |
Savings graph
$4,812/month total
The first two strong-evidence opportunities cover most of the upside without asking the team to touch risky customer-facing QA yet.
Prompt Waste Signals
These are advisory, rule-based signals from existing metadata. SpendLens does not rewrite prompts automatically or send full user prompts to external services.
Workload: rag-answer
Avg input
18,400
Avg output
640
Review whether all retrieved chunks are needed. Try reducing RAG chunks, filtering context more aggressively, or using prompt caching if available.
Workload: write-product-description
Avg input
4,900
Avg output
2,300
Add concise-output instructions or max token limits where appropriate.
Workload: recommend-products
Avg input
3,200
Avg output
80
Shorten static instructions or move reusable context outside the repeated prompt.
Evidence strength
High
Move to replay testing now.
Medium
Test carefully on edge cases.
Low
Keep as manual review only.
Suggested next actions