claude-opus-5 → claude-haiku-4-5: ~$103 /mo on "Ticket #41200 … Before you answer, check: all four keys pre…"
$103Monthly saving
Evidence
- Model
- claude-opus-5
- Recommended model
- claude-haiku-4-5
- Calls in window
- 1,400
- Measured window
- 30 days
- Cost today
- $128.99/mo
- Cost after change
- $25.80/mo
- Replay samples attached
- Yes
- Replay verified
- Yes
- Mean judge score
- 95 / 100
- Samples replayed
- 10
- Samples attempted
- 10
- Call count
- 1,400
- Target tier
- small
- Record count
- 1,400
- Fallback used
- 0
- Candidate tier
- small
- Matched simple
- category,label,summary,one sentence
- Output token cv
- 0.04
- Avg input tokens
- 17,509.58
- Avg prompt chars
- 6,058.88
- Avg output tokens
- 127.85
- Complexity score
- -4
- Reasoning marker
- 0
- Prompt sample size
- 8
- Simple task marker
- 1
- Matched structured
- json,return only
- Aggregate dominated
- 0
- Output token cv scored
- -1
- Structured output marker
- 1
Replay verification
Sample 1 · replayed on claude-haiku-4-5 · 93/100
“Same category, same priority, same owning rota. The candidate's summary is a few words shorter but names the same request.”
Sample 2 · replayed on claude-haiku-4-5 · 96/100
“Both outputs route the ticket identically. Wording of the summary differs; the meaning a human would act on does not.”
Sample 3 · replayed on claude-haiku-4-5 · 92/100
“Identical labels across all three enumerated fields. The summary paraphrases rather than copies, which the contract allows.”
Sample 4 · replayed on claude-haiku-4-5 · 95/100
“The candidate matched the reference on category and team, and chose the same priority band. No material difference.”
Sample 5 · replayed on claude-haiku-4-5 · 98/100
“Equivalent. The candidate dropped a redundant clause from the summary and was otherwise a character-for-character match on the labels.”
Sample 6 · replayed on claude-haiku-4-5 · 94/100
“Same routing decision. The candidate's summary leads with the customer's request rather than the symptom, which reads slightly better.”
Sample 7 · replayed on claude-haiku-4-5 · 97/100
“All enumerated fields agree. The summary is within the word limit in both, and neither leaks account identifiers.”
Sample 8 · replayed on claude-haiku-4-5 · 93/100
“No difference that would change what the agent does next. Labels identical, summary reworded.”
Sample 9 · replayed on claude-haiku-4-5 · 96/100
“Same category, same priority, same owning rota. The candidate's summary is a few words shorter but names the same request.”
Sample 10 · replayed on claude-haiku-4-5 · 92/100
“Both outputs route the ticket identically. Wording of the summary differs; the meaning a human would act on does not.”
How to fix it
- 1. This call site averages 128 output tokens per call (spread 0.04) across 1400 calls.
- 2. Change the model parameter at this call site from claude-opus-5 to claude-haiku-4-5 and re-run its evaluation set.
- 3. Do not ship it yet: replay verification over the 50 attached sample calls will confirm output quality holds on claude-haiku-4-5 before you act on this.