Analysis · Tools

Haiku 5.5 puts a fivefold price step inside a million-token window.

Anthropic's new small model is much cheaper, but prompts above 100,000 tokens face a different tariff. A tokenizer change means the old prompt budget needs recounting.

A cheap model can become a more expensive component without changing models at all. Anthropic's October 7 release of Claude Haiku 5.5 puts a fivefold rate increase at a 100,000-token prompt boundary, inside a model that can accept a million tokens. Reading the context-window headline as a budget promise would miss the decision that matters: how much context each job actually needs.

The discount has two prices

On the Claude Platform's standard rate card, prompts up to 100,000 tokens cost $0.10 per million input tokens and $0.50 per million output tokens. Longer prompts cost $0.50 and $2.50 respectively. Both rates rise fivefold. These are API prices, not Claude subscription prices. Anthropic pricing documentation.

The reduction against Haiku 4.5 is still substantial. Anthropic describes a roughly 75% average saving, based partly on its previous model's request mix. That estimate is a vendor-wide description, not a forecast for one company's traffic. Launch announcement.

Consider two illustrative, uncached requests, each producing 1,000 output tokens. At the published standard rates, 90,000 input tokens cost $0.0095 including the output; 110,000 cost $0.0575. About 22% more input produces a bill about six times as large because the second request falls into the higher tariff. This is arithmetic, not a measured workload, and excludes tools, regional premiums and other modifiers.

The old token budget will not carry over

Migration adds a second boundary. Anthropic's model documentation says Haiku 5.5 uses a newer tokenizer—the system that splits text into billable units—and the same text counts as approximately 30% more tokens than on Haiku 4.5. The model also supports a one-million-token context window. Model specifications.

An old 80,000-token payload could therefore land around 104,000 tokens under that approximate change. The exact result depends on its contents. Recount the actual request before moving it; a character limit copied from the old integration cannot establish which tariff applies.

Small jobs need a smaller payload

For a support product, the routing question may need one message and a short policy, rather than the customer's entire history. A document extractor may need a section rather than every attachment.

That makes prompt construction a product decision. Sending less irrelevant context can keep a narrow job within the lower tariff. Splitting a job can also introduce extra calls, missing context and reconciliation work, so a smaller payload must still contain enough information to answer correctly.

A sensible comparison keeps the task fixed, then measures the result after trimming. If quality falls, the cheaper request has not delivered the same job.

The rate card cannot establish cost per successful job

We have not benchmarked Haiku 5.5 or measured a production migration. Published prices cannot tell us how many retries a particular application will need, whether its answers will pass review, or how much human correction remains.

Those unknowns prevent a blanket recommendation to replace a larger model. A lower token price gives a candidate for testing; the acceptance rate decides whether the saving survives.

Watch the share of requests that cross 100K

Before changing the default model, count representative payloads with the new tokenizer. Track how often they exceed 100,000 tokens, then compare total cost for accepted results, including retries and review.

The operational trigger is concrete: when ordinary conversation growth pushes a request into the higher band, inspect what context it is carrying before increasing the budget.

Reporting note: Source-based analysis of Anthropic's October 7 announcement and live Claude Platform documentation, checked October 8, 2026. Savings estimates and tokenizer behavior are Anthropic's claims. Cost examples are our calculations using published standard API rates; they are not invoices or benchmark results. No hands-on evaluation, vendor contact or sponsorship informed this article.

Primary source: review the source