DocsQuick StartAI News
AI NewsClaude Haiku 5.5 Costs Plummet by 75%
Product Update

Claude Haiku 5.5 Costs Plummet by 75%

2026-10-08T00:04:32.189Z
Claude Haiku 5.5 Costs Plummet by 75%

Anthropic released Claude Haiku 5.5 yesterday, cutting short-context input and output prices to $0.10 and $0.50 per million tokens, respectively. What it truly changes is not benchmark scores, but the cost structure of high-frequency agentic, classification, and coding tasks.

Anthropic Drives Down Haiku Pricing

Anthropic released Claude Haiku 5.5 on October 7, positioning it as the fastest, lowest-priced, and most capable model in the Haiku family to date.

The most noteworthy part of this update is not the “5.5” version number, but the pricing: for prompts of no more than 100,000 tokens, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens. The previous-generation Haiku 4.5 costs $1 and $5, respectively.

Looking only at standard token pricing for short contexts, the new model is effectively 90% cheaper. Anthropic’s broader estimate is that Haiku 5.5 costs approximately 75% less to run on average than Haiku 4.5.

The two figures are not contradictory. Haiku 5.5 uses a different pricing tier for long prompts exceeding 100,000 tokens, with input and output prices rising to $0.50 and $2.50 per million tokens, respectively—only about half the price of Haiku 4.5. Since context length, input-to-output ratios, and cache hit rates vary across workloads, final bills will not uniformly fall by 90%.

Comparison of input, output, and cache pricing for Claude Haiku 5.5, Haiku 4.5, and Sonnet 5.5

The 100K Threshold Is What Really Matters in the Pricing Table

Haiku 5.5’s official pricing can be divided into two tiers:

| Item (per million tokens) | Haiku 5.5, prompt ≤100K | Haiku 5.5, prompt >100K | Haiku 4.5 | Sonnet 5.5 | |---|---:|---:|---:|---:| | Input | $0.10 | $0.50 | $1.00 | $2.00 | | Output | $0.50 | $2.50 | $5.00 | $10.00 | | Cache writes | $0.125 | $0.625 | $1.25 | $2.50 | | Cache reads | $0.01 | $0.05 | $0.10 | $0.10 |

For developers, 100K is not an insignificant billing footnote—it is a cost boundary that must be incorporated into architecture design. Once a request enters the long-context tier, Haiku 5.5’s input, output, and cache prices all become five times those of the short-context tier.

Of course, even in the higher-priced tier, it still costs only half as much as Haiku 4.5. The problem is that if a team sees only “$0.10 per million input tokens” and uses that figure to estimate its annual budget, it may significantly underestimate production costs.

This is especially relevant for agent applications. They repeatedly feed system prompts, tool definitions, message histories, retrieval results, and intermediate steps back into the context. The user may see only the word “continue,” while the model receives a complete working context containing hundreds of thousands of tokens.

More sensible practices include:

  • Use rolling summaries for conversation history instead of appending the original text indefinitely;
  • Restrict retrieval results to genuinely relevant excerpts rather than sending entire documents to the model;
  • Keep system prompts and tool definitions in fixed positions to improve cache reuse;
  • Set budget alerts as the context approaches 100K instead of investigating only after the bill arrives;
  • Assign long-document preprocessing, classification, and retrieval to Haiku, while passing a small number of high-value excerpts to Sonnet or a more capable model for evaluation.

Haiku 5.5’s low pricing comes with conditions, but those conditions are not particularly restrictive. Most customer-service classification, structured extraction, search-query rewriting, content moderation, and short-chain tool-calling tasks do not require a 100K context in the first place.

A Simple Calculation: The Same Traffic Could Drop from $200 to $20

Suppose a workload processes 100 million input tokens and 20 million output tokens per month, with each request remaining below 100K and caching temporarily excluded from the calculation:

  • Haiku 4.5: $100 for input and $100 for output, totaling $200;
  • Haiku 5.5: $10 for input and $10 for output, totaling $20.

For this type of workload, the cost does indeed fall by 90%.

If the same tokens all fall into the tier above 100K, Haiku 5.5 would cost $50 for input and $50 for output, totaling $100—a 50% reduction compared with Haiku 4.5. Anthropic’s claim of an “average reduction of approximately 75%” appears to be an overall estimate based on a typical mix of workloads, rather than a fixed discount that applies to every API request.

This is the most important point to clarify about this release: Haiku 5.5 does not simply reduce all Haiku 4.5 pricing by 75%. Instead, it uses extremely low short-context pricing to encourage developers to migrate high-frequency, well-defined, batchable tasks.

Cache Pricing Matters More Than Per-Token Model Pricing for Agent Teams

In the short-context tier, Haiku 5.5 charges $0.125 per million tokens for cache writes and only $0.01 for cache reads. Compared with the standard input price of $0.10, the initial write is slightly more expensive, but subsequent reads cost only one-tenth as much.

If a fixed prompt segment is reused, the second call will generally have an opportunity to offset the additional cost of the initial write. For applications containing extensive tool schemas, rule descriptions, product catalogs, or codebase context, this is more effective than merely compressing user messages.

Prompt Cache can be understood as a “read-only mirror” of the model context: creating the mirror incurs a write fee the first time, but subsequent requests do not need to pay the full input price to read the same content again.

Anthropic has also reduced Sonnet 5.5’s cache-read price from $0.20 to $0.10 per million tokens, saying that this will lower Sonnet 5.5’s operating costs by approximately 20% for most agentic tasks. This adjustment shows that competition among model providers is shifting from the “list price per million tokens” toward the cost of the complete agent loop.

A typical Q&A interaction usually calls the model only once. An agent may need more than a dozen calls to complete a task, with each step carrying similar tool definitions and execution history. In such cases, cache hit rates, tool-response lengths, and the number of failed retries often have a greater impact on the bill than the base input price.

What Haiku 5.5 Is Best Suited For

Anthropic describes the new model as the most capable Haiku yet, but at this stage, it should not be treated as a drop-in replacement for Sonnet 5.5 based solely on the vendor’s claims. The core value of a small model remains low latency and high throughput—not the ability to solve every complex problem.

Haiku 5.5 is better suited to the following tasks:

  1. Classification, routing, and moderation: Determining intent, risk level, or which backend model should be called. These tasks produce short outputs and are invoked frequently, allowing them to benefit most from the lower pricing.
  2. Structured information extraction: Extracting JSON fields from emails, support tickets, or contract excerpts. The task boundaries are clear, and the results are easy to validate programmatically.
  3. Lightweight agent steps: Generating search terms, selecting tools, and organizing tool outputs, rather than independently handling an entire open-ended task.
  4. Code-assistance pipelines: Explaining errors, generating test cases, and performing preliminary code reviews before escalating difficult issues to Sonnet.
  5. High-concurrency interactions: Scenarios sensitive to time to first token, such as real-time completion, game-character dialogue, and customer-service assistance.

Conversely, if a task requires sustained planning under ambiguous objectives, self-correction across dozens of steps, or involves high business costs when errors occur, Sonnet 5.5 remains the more natural default choice. Haiku can handle screening and material preparation, but it should not be forced into the final decision-making role merely because it is cheaper.

A more practical production architecture uses “tiered routing”: first let Haiku 5.5 assess the task’s difficulty and answer simple requests directly, then escalate complex coding tasks, long-horizon agent workflows, and high-risk decisions to Sonnet. As long as routing errors remain under control, this combination is generally more economical than sending all traffic to Sonnet and more reliable than forcing every request through Haiku.

How Developers Can Integrate It

Haiku 5.5 is a newly released proprietary Claude model. If a project already uses the OpenAI SDK, an OpenAI-compatible gateway can reduce migration effort. OpenAI Hub can also provide unified access to models such as Claude, GPT, Gemini, and DeepSeek, eliminating the need to maintain multiple API protocols separately in mainland China.

Below is a Python example. Model identifiers may change depending on the platform’s listing strategy, so the model list in the console should be treated as authoritative before production deployment:

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_OPENAI_HUB_KEY",
    base_url="https://api.openai-hub.com/v1"
)

response = client.chat.completions.create(
    model="claude-haiku-5-5",
    messages=[
        {
            "role": "system",
            "content": "You are a support ticket classifier. Return only billing, bug, feature, or other."
        },
        {
            "role": "user",
            "content": "My credit card was charged, but the credit has not appeared in my account."
        }
    ],
    temperature=0
)

print(response.choices[0].message.content)

During an actual migration, do not stop at a smoke test that merely checks whether the model can return text. At a minimum, also verify:

  • The reliability of JSON or structured output;
  • Whether tool-call parameters are fully compatible with existing schemas;
  • Actual billing after the context exceeds 100K;
  • Whether Prompt Cache hits occur and how long cached content remains valid;
  • Concurrency limits, time to first token, and overall throughput;
  • Whether changes to content-safety policies cause false refusals;
  • Whether tasks can be automatically escalated to Sonnet when Haiku cannot complete them.

What This Update Means: Small Models Are Beginning to Take Over Mainstream Requests

Over the past two years, vendors have tended to emphasize flagship benchmark scores when releasing new models, but the bulk of the real API market does not consist entirely of complex reasoning. Many requests simply rewrite a query, extract a few fields, determine whether a tool should be called, or turn machine output into human-readable text.

For these tasks, once model capabilities clear the required threshold, price, latency, and reliability become the deciding procurement factors. By reducing short-context input pricing to $0.10 per million tokens, Haiku 5.5 is directly competing for this high-frequency, foundational traffic.

In our view, Haiku 5.5’s greatest value is not “chatting with a cheaper Claude,” but making pipeline stages viable that previously did not justify using a large model because they were called too frequently. Log attribution, batch tagging, RAG query rewriting, agent routing, and tool-output cleanup may all shift from rule-based systems to model-based systems as a result.

However, developers should also be alert to a counterintuitive outcome: once each call becomes cheaper, teams often increase the number of calls, meaning the total bill may not decline in proportion to the advertised price. This is especially true for agents, where lower token prices can easily conceal unproductive loops, repeated tool calls, and excessively long contexts.

Haiku 5.5 is therefore a highly competitive update, but it is not a substitute for cost governance. Teams that successfully realize the promised 75% cost reduction will also implement effective model routing, context compression, cache reuse, and call-chain observability.

References

Related Articles

View All

Contact Us

We usually reply quickly during business hours

Scan WeChat

Support: Hub Assistant

WeChat ID: