today at 10:39 PM
Ran our DataAnalyticsBench benchmark on it: https://plotly.com/blog/claude-haiku-5-5-plotly-data-analyti...
9x cheaper than Haiku 4.5 and 2 letter grades better. It's also now the fastest model (using the default speeds, not trying any of the other models "Fast" mode) to complete the exam.
Similar ballpark to Luna in price, cost, and accuracy. These are very cheap models: $0.38 to answer 40 in-depth data analytics questions (compared to $15 for Opus 5.5 or $20 for Astra).
Overall very good at data analysis - handling all of the straightforward data analytics questions correctly. It fell short answering some of the questions that required some deeper statistical analysis like looking into other variables. In other words, it's not as persistent as other models in its analysis, which I think we'd expect from how they're positioning the model.
Compared to OpenAI: GPT-6 Luna did a bit better and was about 30% the cost of Haiku 5.5. GPT-6.1 Sol got all answers correct, but was 10x more expensive.
today at 7:35 PM
Pelicans riding bicycles for Haiku at the different thinking levels: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Low messes up the bicycle frame, but medium/high/xhigh/max all get the bicycle frame right.
The max one took 5 minutes 9 seconds and cost 3.3826 cents. The cheapest one (low) cost 0.0936 cents and took 7 seconds.
The most recent release of my llm-anthropic plugin queries the Anthropic model listing API directly, so I didn't have to upgrade the plugin to add support for this model:
llm install llm-anthropic -U
llm anthropic refresh
llm -m claude-haiku-5.5 'prompt goes here'
EDIT: Here's the Haiku 4.5 pelican from a year ago for comparison, it was terrible: https://simonwillison.net/2025/Oct/15/claude-haiku-45/today at 8:23 PM
Of course, the sun again. Everyone knows that a pelican can't ride a bicycle without a sun in the frame and can only go right.
today at 8:52 PM
I wonder if we'll start to see pelicans like a mascot of sorts. You could have a pelican pin on your backpack.
> "What's up with the pelican?"
Well you see in the early days of LLMs we wanted a fun way to test new models, and there was this blog, ...
today at 9:16 PM
Will Smith eating spaghetti is the OG benchmark
today at 10:29 PM
The medium thinking effort one doesn't have a sun at all?
today at 9:20 PM
I always find the time/token differences between the xhigh and the max effort levels for Claude models absolutely insane.
Even more so, because in a lot of their benchmarks they use the max models. I honestly think I'd rather these labs use their xhigh models as the default for benchmarking instead since I don't think the average person is even using max.
today at 10:17 PM
Benchmarks are the entire reason why max exists
today at 10:06 PM
I use max all the time, a bit annoyed that they keep trying to silently switch me off it. (Claude Code will refuse to remember a setting of max and will continually reset it to xhigh - I have an objection to these patterns in general)
I'm definitely not the average person though.
today at 10:27 PM
I say half facetiously - have you tried writing a skill or rule to remember your setting as a workaround?
I actually don't like that it sometimes remembers the last model/effort i used. I should be able to set a default model/effort that is separate from the one off fable runs I use.
today at 10:28 PM
I thought the thinking effort was specified out of band from that, though maybe it's not. Not sure if the model was trained to listen in other areas. The biggest issue is, it's difficult to tell if it works because you can no longer see the thinking! Though I guess if you can't tell a difference in the output, was there any point to max in the first place?
today at 10:09 PM
I'm still getting network errors. Seems to be CORS-related.
today at 8:39 PM
I find it helpful when you post your link that compares the model to other models in the same class or family, or shows progression over time.
The pelicans all start to look the same after a while.
But seeing the comparison to other models by class, family, or historical progression gives an excellent frame of reference.
today at 9:18 PM
Good call, I've edited my comment.
Here's the Haiku 4.5 pelican from a year ago - it sucked in comparison to Haiku 5.5: https://simonwillison.net/2025/Oct/15/claude-haiku-45/
today at 10:35 PM
How does Haiku 5.5 compare with modern alts in its class, like Luna-6 or OSS models of similar speed/cost?
today at 7:49 PM
I thought Anthropic models didn’t generate images.
today at 7:55 PM
This is SVG, but recent Anthropic models have got extremely good at other forms of visual data.
Here's a Blender model I had Claude Opus 5.5 create: https://tools.simonwillison.net/blender-viewer?url=https%3A%...
And here's some animated pixel art by Opus 5.5: https://tools.simonwillison.net/kakapo-party
And some Monkey Island style music (Opus can compose music too): https://tools.simonwillison.net/scrimshaw-jukebox
Anthropic's models do all of this by outputting code. GPT-6 Astra has similar capabilities - I got this Blender model using that: https://tools.simonwillison.net/blender-viewer?url=https%3A%...
today at 10:04 PM
Pardon, I have a lot of questions about that Scrimshaw music text format. It's clever. Did you invent it, and is it specifically intended to be written to by LLMs? Is the editor/player LLM-coded as well, and was this its recommendation for a format that would be easy for LLMs to write? I'm wondering why this instead of say, asking it to write a .MOD file.
today at 9:07 PM
This is great! Love the pixel art and tunes.
today at 8:37 PM
Tried it with GPT-6 Astra with Ultra but the outcome was underwhelming with Blender. Maybe it was my prompting ¯\_(ツ)_/¯
today at 8:05 PM
They are really good at generating artifacts, which are windows within the replies containing all kind of visualization, often interactive.
They are still not great at SVG. I just asked Opus and Fable to add a background to an SVG and the results were, well, not great.
today at 8:29 PM
SVG is hard.
today at 7:51 PM
They generate svg. You can paste in pngs and they'll convert them to svg with varying degrees of success.
today at 8:23 PM
I’ve been playing around with Opus 5.5 which has made a big leap over previous generations in its ability to use a simple drawing-instruction prompt to generate images.
This creates Sierra AGI-style adventure game scenes painted live from simple Turtle-esque drawing instructions so you can basically provide it an empty canvas and then position text labels on the canvas where you want certain things (tavern, oak tree, etc) and it will generate a custom script for rendering them in a EGA graphics style.
today at 8:22 PM
They don’t do raster images.
today at 7:51 PM
Those are SVGs not images.
today at 8:28 PM
I've created multiple videos using Claude Code, including music and speech. It generates python which in turn generates frame PNGs that it runs through ffmpeg.
Please don't judge me too harshly for this particular poop video. But here is an example of something 100% generated with claude prompts only.
https://www.youtube.com/watch?v=2EqMplbt0gU
today at 8:56 PM
To clarify the ”100%” part - the Python script generated the video output, and you did nothing? No video edit at all? Then I think it is impressive! Are you able to share the prompts you used?
today at 9:20 PM
Source code is linked from the video! Scan the QR code. I will try to /resume tonight and give you some prompts.
today at 10:11 PM
I mean Claude code sessions are all jsonl files that it can interrogate on its own. Get a new agent to capture how it was made and what the prompts were. No need for tedious /resume’ing and prompting.
today at 9:17 PM
The Purple Screen of Death at the end :)
today at 7:52 PM
SVG is code
today at 7:52 PM
They're SVGs
today at 7:52 PM
Bruh. Svg. It is like drawing something with geometric shapes which are represented using equations.
today at 9:51 PM
[dead]
today at 6:05 PM
Pricing is...a bit weird.
Input
$0.10 / MTok for prompts up to 100,000 tokens
$0.50 / MTok for prompts over 100,000 tokens
Output
$0.50 / MTok for prompts up to 100,000 tokens
$2.50 / MTok for prompts over 100,000 tokens
100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents; for typical generation or Jev-like classifiers, it's a good value and as noted in this article, that is apparently the vast majority of Haiku use.In both cases, still much cheaper than Haiku 4.5's $1 input / $5 output and these prices better compete with GPT-6 Luna. ($0.10 input / $0.50 output, but with no token threshold [EDIT: the threshold for Luna is apparently 272k])
today at 6:16 PM
Haiku 5.5 is noticeably smarter than GPT-6 Luna, so I can see their pricing strategy here.
For a while Anthropic has lacked a cost effective “cheap” LLM for summarisation, compacting, RAG helpers, etc.
These ‘ephemeral’ workloads are often under 100k tokens, or can be structured to be under 100k.
In some coding benchmarks, Haiku 5.5 beats Sonnet 5! (Especially implementation; do a well defined Jira ticket; etc), it’s really impressive how much intelligence per dollar has grown in just a few short months.
today at 9:59 PM
I said this in another comment, but Artificial Analysis has the cost per task of Haiku on max roughly equal to that of Sol on medium, and the latter is significantly more intelligent. (And I'd wager that Sol probably finishes tasks more quickly, even with Haiku inference being faster.) So Haiku really only makes sense on lower reasoning levels, and only if you care about intelligence and speed more than you do about cost effectiveness (where Luna currently dominates). And that's without even bringing Chinese models into the mix.
today at 7:13 PM
noticeably smarter remains to be seen in practice. For now, Haiku is a bit more expensive than Luna on < 100k token, but I just don't have any agentic work below 100k, so this is going to be 5x more expensive than shown on these charts. It's hardly competitive ...
today at 9:14 PM
> it’s really impressive how much intelligence per dollar has grown in just a few short months.
Open weights models giving a distant salute from afar
today at 9:59 PM
Yes, it was weird to see MiMo and DeepSeek missing in the article's comparison...
today at 10:00 PM
It's not that weird. Most companies considering paying Anthropic are probably not considering Chinese models as alternatives. Many don't even realize they exist.
today at 10:26 PM
The thing is: availability of near-SOTA cheap Chinese models is forcing OAI and Anthropic to bring prices down and offer more efficient models, instead of simply focusing on super expensive SOTA LLMs.
today at 10:31 PM
"companies" is a meaningless metric.
If you want to make it about 99% of real world companies, they are all on Gemini or Copilot anyway, nobody is going through legal and procurement to get models from dubious silicon valley startups when you have relations with Microsoft or Google or Amazon from ages because some benchmark is showing some minor digit benefit when vibe coding GTA 6.
today at 9:03 PM
The benchmarks are very long form logic, knowledge, and coding tasks though. I'm very interested in Haiku 5.5's performance on ObviousBench where Luna 6 is currently SotA.
today at 10:09 PM
If you can't get any coding done with 100K context that is either a broken model, a broken harness or a skill issue. I would mostly use Haiku in task or explorer subagents. I'm not saying I stay under that on every task, but I do have quite a few sessions that cap out well below that, so that price difference would be very meaningful.
I use Luna for this day in and out and its excellent - if Haiku is that much better I will be changing things up.
today at 6:18 PM
There's also a tokenizer efficiency difference: modern Claude's 100K tokens are about ~60-65K modern GPT tokens, so in reality the Luna cutoff is much further away than the Haiku one.
You can test with Anthropic's count_tokens endpoint or with https://crates.io/crates/tokwc
today at 6:38 PM
> ...this tokenizer, the same input text produces approximately 30% more tokens on Claude Haiku 5.5 than on Claude Haiku 4.5.
So, it is might be even worse.
today at 6:41 PM
No, it's just Haiku 4.5 is so old that it predates the new Claude tokenizer change in Claude 4.7+
today at 6:14 PM
It's actually existing flat per-token pricing that is weird.
Neither encode nor decode are linear in compute, so providers need to price for average expected length.
This is just getting closer to the true cost of generating tokens.
today at 8:51 PM
Flat per-token pricing is likely just logistically easier, particularly if these closed models are also picking up the kv cache efficiency improvements seen in recent open weight models.
today at 7:21 PM
Flat pricing is weird too but jumping up 5x at one cutoff is surprising in the other direction IMO
today at 7:37 PM
My theory here is that providers cover the non-constant costs of output tokens as context length caries using the cache input fees.
today at 6:24 PM
Notable that one suggested use case for Haiku is "classification requests", i.e. Jev competitor, and the pricing matches GPT-6 Luna which is behind OpenAI's "Decisions API" Jev competitor.
For this application 100K token input is plenty.
Of course Anthropic and OpenAI, both at $0.10/M, are still 2.5x the cost of Jev's $0.04/M.
today at 6:27 PM
I think the 2.5 times cost but actually pays off in terms of intelligence compared to jev and the general capability of using it beyond classification
today at 6:41 PM
The classification performance remains to be seen, but presumably we'll soon start to see classification benchmarks.
For other tasks like summaries (another suggested usage) it's good to see Haiku and Luna now competing against each other on cost.
I'd love to know how the business automation market breaks down by volume of call type though - hard to imagine that decision making (e.g. branching, triage) isn't a very large part of it, greater than these other suggested Haiku use cases.
today at 6:10 PM
Luna does as well, but just at a higher limit.
From OpenAI's website: Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request.
today at 6:15 PM
Huh, that disclaimer is on the model page (https://developers.openai.com/api/docs/models/gpt-6-luna) but not the pricing page. Annoying.
Fixed.
today at 7:06 PM
So even at the 1.5x/2x rate luna is still half the price of this. Weird pricing strategy from Anthropic. I'm sticking with Luna if I don't need a super smart model
today at 8:23 PM
You're judging purely by token cost I assume, not cost per completed task?
The benchmark in the article showed it as lower per completed task than luna, but I guess we'll find out how representative that is. Anthropic has generally been fairly honest in their benchmarking though.
today at 9:42 PM
The cost per task from Artificial Analysis is roughly 3x higher at every reasoning effort level for Haiku than Luna. Sol 6.1 on medium has the same cost per task as Haiku with significantly higher intelligence. According to those numbers (which you should take with a grain of salt), from a pure cost vs intelligence standpoint, you're better off using Luna for economics and Sol for intelligence.
With that said, the real reason to use Haiku is that it's faster than all of these models. OpenRouter is showing an average so far of 93 tokens/sec, and AA got at least 137 in each of their benchmarks. So it might be valuable for speed at lower thinking levels. (At higher thinking levels, it's likely going to take longer to produce results than Sol on low/medium.)
https://artificialanalysis.ai/models/releases/comparisons/cl...
today at 8:28 PM
yes, that's true. I should be looking at the $/completed task
today at 7:27 PM
Isn't it less than a year since Claude models went from 100k token limit to 1M limit? Don't get me wrong - my main agent normally gets to 25% or so before I clear it these days, but as a subagent, doing research or summarisation, I don't think 100k is "absurdly low".
today at 10:33 PM
This is pretty good tbh
today at 6:34 PM
They are targeting businesses/API use for fast decision making and agent integration. Plus they now need to be competitive with Jev-type models in that space.
today at 8:45 PM
I think they are also trying to make sure Deepseek and other chinese models don't eat their lunch. They need something price competitive.
today at 8:41 PM
If you look at how different reasoning levels can easily exceed task cost of sonnet 5.5 you will see that you will basically never fall into that under 100,000 token threshold. I mean maybe you can choose low and do a basic summary task, but then you could choose something much cheaper instead. I don't know what Anthropic is thinking with its dumber models.
today at 7:07 PM
You could also use it as a subagent prompted eg by Sonnet/Opus orchestrator agent and for many agentic workflows significant part of the dispatched tasks might be under 100k budget.
today at 6:08 PM
I with they'd give Haiku like 400k tokens roughly, I think between 400k or even 600k tokens is a sweet spot, but Haiku is basically designed to be for small edits is my understanding, but it sucks because any time I ask Opus to "try" letting Haiku do the work, it just falls apart and Opus comes back and tells me it switched to Sonnet (even before Sonnet finally jumped up to 5.x).
I will try the new Haiku, but it would be worthwhile if Haiku could take sane instructions and do all file editing for Opus / Sonnet / Fable then it would be worth using.
today at 6:16 PM
I mostly use Haiku for really, really basic stuff, never for actual engaging work. I've used it for first-pass analysis to triage bugs, for example - all it does is related N bugs together to see if any potentially relate. Then I have Sonnet investigate further.
today at 6:18 PM
>> 100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents
Your vibes don't appear to be supported by facts. From the announcement:
>> Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens. On Haiku 4.5, 90% of requests fell into the former category.
today at 6:21 PM
People weren't using Haiku 4.5 for agents before. 5.5 is good enough that it might be.
today at 9:54 PM
[dead]
today at 6:56 PM
Haiku 4.5 users were using it for Kleenex requests because that was the best it could do.
today at 7:54 PM
Not really. We use Haiku 4.5 to turn users' natural language queries and requests into fairly complex structured specs for interior design and construction. It has near perfect accuracy.
today at 9:09 PM
How many examples are in your prompt? How large is that prompt? Or do you have some other way of tuning the output?
I'm asking to learn for a similar project, not to discount anything you're saying.
today at 8:23 PM
Chatbot could be < 100k tokens.
today at 6:17 PM
encode and decode tok/s which is ($/s) when it comes to pricing drops heavily above 100k tokens.
There are plenty of workflows like translations where you'd easily be under the cap.
today at 6:17 PM
Who in their right mind would use haiku while Mimo or GLM cost 10% of what they are charging with much smarter models?
today at 6:58 PM
That's not what any benchmarks that look at cost per task or similar says in terms of cost. The Chinese models, generally speaking, might be cheaper per token but need a lot more tokens to get there.
today at 9:46 PM
Except for the new MiMo V2.6 models, which appear to give some of the best value right now, at least on paper. (I haven't tried them so I can't speak from experience.)
today at 6:45 PM
Some people/organizations are ideologically opposed to using Chinese models. Not me, I use GLM-5.3-Flash for almost everything (the subscription-subsidized pricing on a legacy Z.ai plan makes it the best value model by a wide margin), along with some MiMo and DeepSeek. Still, I use Luna for certain tasks where speed is more valuable than performance; I can see this new Haiku displacing Luna for those. If you mean Haiku 4.5 though I agree, that model was a waste of time and money.
today at 9:30 PM
Luna is not really the fastest. You need to use it in high/max to get the good output for what it is good for: summarizing. And that is already close to two minutes per task...
today at 9:40 PM
I’m on the Legacy v2 plan and same: nothing comes close to 5.3 Flash’s value on it. It’s crazy, no wonder they discontinued them!
today at 10:05 PM
Isn't the point of this release that it's comparable?
AAI Index // Input // Output
Haiku 5.5: 43 // $0.10 // $0.50
Mimo 2.6 Pro: 46 // $0.43 // $0.87
Mimo 2.6 Flash: 38 // $0.10 // $0.28
Seems competitive to me? Plus then I don't have to manage multiple providers
today at 6:21 PM
Presumably everyone who doesn't bother integrating a third party API key into their harness, which would probably be most of the Claude Code users.
today at 8:28 PM
On subscription pricing a $20 Anthropic subscription gives >$500 equivalent tokens, which is not so different, and you get smarter models. API pricing has decent margins.
And Opus 5.5 is really good.
today at 8:28 PM
Where do you get this 10% number? Checking providers I know/respect, and GLM 5.3 flash is $0.15/m. Haiku is $0.10/m.
today at 7:11 PM
Well, unless you're using OpenCode Go, it's per-token costs (even if already super low), while Haiku falls under the Claude sub. It's just more straight forward and you aren't feeling a "loss" with the sub.
today at 7:05 PM
There really aren't any models at 10% of the price of Luna or Haiku.
today at 8:31 PM
People who are stuck using Bedrock in-geo due to their company policy (me).
today at 6:17 PM
It's their creative way of 'matching' Luna's prices.
today at 6:12 PM
It could be to incentivize people to not be lazy users of tokens.
today at 6:21 PM
> Second, this week, we’ll roll out a new monthly API credit to all Max and Team subscribers for use on the Claude Platform. Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users
This is a very big benefit for me. I can now ship actual ai enhanced features behind my subscription without paying extra or fully relying on on-device models. I do worry that this is to soften the blow for user-unfriendly changes
today at 6:47 PM
This is them sneaking in taking the Claude Agent SDK (claude -p) off of subscription plans through the back door along with a model release. They previously wanted to do this in June, but backpedaled after huge backlash:
https://support.claude.com/en/articles/15036540-use-the-clau...
today at 10:36 PM
Sorry that that help center article was misleading; we’ve updated it to clarify that `claude -p` has not been removed from subscriptions as part of this change!
today at 7:35 PM
I hope people notice again that this is happening this time around.
Being forced through the non-OSS Claude Code with all of its quirks and issues is... such an exhausting use of force by Anthropic.
To the extent that you _can_ choose to disable telemetry and training on your traces in CC, it's not all that obvious what they gain by crippling your ability to use the subscription with other – better – tools.
It's also remarkable that it's coincident with OpenAI adding "Sign in with OpenAI", so that you can use your tokens with other tools.
today at 8:05 PM
The page does not say anything about changing the way Agent SDK bills. I just tested Agent SDK and nothing has changed (yet).
You might be right and they will change this in the future, but that's speculative
today at 8:07 PM
> Claude Max and Team plans now include monthly API credits, which cover the Claude Agent SDK, the Claude API, and Claude Managed Agents.
This text has replaced the entirety of the page called "Use the Claude Agent SDK with your Claude plan."
today at 8:16 PM
Yes. Previously the page was all about how they were going to start charging for Agent SDK use with a banner on the top saying that, actually, they weren't going to do that.
today at 8:19 PM
...yes, a banner which has also now disappeared and been replaced, with the explicit mention that API credits "cover the Claude Agent SDK"?
What more do you need?
today at 8:27 PM
Well, it doesn't currently work that way on the latest SDK. If there's a change coming, it hasn't happened yet.
today at 8:57 PM
The monthly credit allocation hasn't rolled out yet either, so as of right now, we're effectively at the status quo. I'd expect the billing change to land once you can actually collect your Claude Console account.
today at 10:16 PM
How credits work with `claude -p` is a common question. we're updating the faq now to make sure it's more clear
today at 7:08 PM
Those mfers. I'm using this for work! I use my work teams plan with pi so I can do all kinds of custom workflows that I can't in Claude Code. Time to convince management I need OpenAI instead.
today at 7:32 PM
Time to convince management (and yourself) to build some skills. :)
today at 7:52 PM
Spent many years building skills, I'm just working on a different level now.
today at 7:02 PM
Even after the June changes there was some allowance to use agent SDK on the pro plan. This will move me to codex tomorrow if agent SDK is really blocked on pro
today at 8:29 PM
These are api tokens, you can build a business with them using any harness.
today at 8:32 PM
Yes, but they're wildly lower in value than the corresponding subscription usage.
today at 9:47 PM
Nothing is being removed as part of this!
today at 9:57 PM
*For now. If a company were to degrade something, it shouldn't be so obvious that the "goodwill" was just a reallocation. Just a good strategy. For example, it allows them to claim that they're "just going from 150% to 125% usage allowance, which is still more than 100%".
today at 7:10 PM
Do we know if claude -p is now drawing from this API usage?
today at 10:17 PM
How credits work with `claude -p` is a common question we're seeing. we're updating the faq now to make sure it's more clear. Will share updated docs soon
today at 8:02 PM
Not yet, as of version 2.1.293. But I suspect this is coming next.
today at 7:29 PM
That's what the Agent SDK is, according to this help page:
https://code.claude.com/docs/en/headless
So, as written, yes.
today at 9:45 PM
today at 6:28 PM
This is literally for you to get tangled in their api and when they stop giving you the allowance they hope you will just continue to pay
today at 6:50 PM
Nah, it's pretty trivial to switch providers (especially with Claude's help, ha).
This is more to encourage people to try out adding AI into their product, which is a totally different flow and experience from using AI to build the product.
today at 8:30 PM
Or to discourage people from using cheap subscription tokens as part of automated workflows
today at 6:38 PM
What does "tangled in their api" mean? Switching is pretty easy.
today at 6:50 PM
Not really, you have to fiddle with generating api keys and setting environment variables. Meanwhile with Anthropic it will just start charging you API prices for the tokens you are generating without even a single warning.
today at 6:57 PM
>> Not really, you have to fiddle with generating api keys and setting environment variables.
That's 5-15 minutes of work at most. Not exactly the type of lock-in the parent is implying.
today at 7:06 PM
The user could have always done that regardless of if the user has the option to be charged API rates on or off.
today at 6:45 PM
That's an old tactic for an old world. You only need, what, half an hour with your agent of choice to write you out of that?
today at 7:05 PM
This is massive. So on top of the regular usage, we now have USD 200,- to freely use via the API however we please, even resell? That is a statement, even knowing that inference does not cost them nearly as much as they charge, this is very developer-friendly. Does some minor de-risking for testing concepts. Terms seem to be reasonable [0].
Of course, they don't do this out of pure kindness, but I really struggle to see a negative for subscribers already using a Claude Max subscription, especially given changing to another model is essentially frictionless via OpenRouter.
Compared with "Sign in via OpenAI" which they just announced, this is far less lock-in for anyone hosting services but less interesting for users of said services. With Anthropics approach, you can just use the allowance on your users however you see fit along with any other models and once it's used up, you can still just decide not to use their models for the remainder. With users bringing their tokens meanwhile, there is less flexibility in terms of switching for you, though might be cheaper for users.
Both interesting, each approaching this from a very different direction, each having their own trade-offs. On the OpenAI front, will be interesting whether developers can set specific temp, reasoning budgets, etc. for such "provided tokens" or whether OpenAI exposes that only via the actual API.
[0] https://www.anthropic.com/legal/credit-terms
today at 6:31 PM
OpenAI will probably add this to their plans within a week
today at 6:43 PM
With OpenAI you can just use Oauth and get a token to use your subscription.
Anthropic isn't even close to being this useful.
today at 6:52 PM
Biggest reason for an OAI subscription instead of Ant imo.
Biggest loss is that Ant models look like they are genuinely better.
today at 7:06 PM
> Biggest loss is that Ant models look like they are genuinely better.
This changes on a weekly basis, I ended up with subscriptions to most of the providers (except for X.ai).
today at 7:33 PM
My complaint about Haiku 4.5 was that it was 10x the price of GPT-6 Luna.
> Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens
Haiku and Luna now have the exact same price up to 100,000 tokens. Luna is now cheaper for anything after 100,000 tokens, even after Luna's own price increases at 270,000 it's still less than Haiku.
So it sounds like they've directly addressed that problem. Their self-reported benchmarks are all higher than Luna too.
today at 7:45 PM
Yep. I was looking at the prices of lower tier models a few weeks ago for zero/few shot tasks (pre Jev) and Haiku rates just didn't make sense at all. I ended up using 5.6-luna.
Good to know that is going back to being an actual option from perf/price perspective.
today at 9:56 PM
[dead]
today at 6:30 PM
Ran image -> html tests for this. I was curious if this smaller model was good enough for complex UI. It was not.
Haiku 5.5: https://html.non.io/lcars-haiku-5.5/
Opus 5.5 for comparison: https://html.non.io/lcars-opus-5.5
Designs it was building from: https://diffui.ai/app/canvas/5093e689-1e74-4f26-b632-2a4500f...
One interesting thing is it took a look at the job at hand, and immediately delegated it to Opus 5.5. It at least knows what it isn't good at. Very fast though, and likely best used for small subagent tasks / tightly scoped work.
today at 7:07 PM
Pac-Man Bench:
Considering the price, no model comes close to being as good as this. However, it did take an extremely long time.
TIME 19m COST $0.16 https://jonclegg.github.io/pacman-bakeoff/#claude-haiku-5-5
All results: https://jonclegg.github.io/pacman-bakeoff/
today at 7:31 PM
Interesting that you have gpt-6-luna at $0.01 vs. claude-haiku-5-5 at $0.16 for this task. I see the score disparity though and I played them briefly. My takeaway from this is that the choice between Luna and Haiku 5.5 may remain nuanced. Luna may be a lot cheaper still and good enough for some jobs. Is that your read of the results?
today at 8:40 PM
Actually, I misspoke. At least as far as Pac-Man Bench, Luna does about as good of a job. The ghost logic's not quite as good, but it also makes a map that doesn't have nonsensical sections in it. So maybe call it a wash.
today at 9:00 PM
Yeah, I was mainly thinking about how much cheaper Luna appeared to be in this case.
today at 7:43 PM
Something is not right there. DSv4.1 flash shows $1.89 for tens of thousands of tokens? What am I missing?
today at 7:14 PM
How have you avoided being sued by Namco?
today at 7:28 PM
I'm pretty sure they'll never see this. It's pretty much impossible for anything you do you build nowadays to get noticed anyways.
today at 7:37 PM
> likely best used for small subagent tasks / tightly scoped work.
Hasn't this always been the case with Haiku?
today at 7:08 PM
To be fair, you're making it compete with the best public LLM right now that's 2 size/price tiers above it.
today at 7:21 PM
Sure, but presumably Haiku was distilled from the same training data. Part of this is seeing how much the capabilities degrade as their model size goes down.
today at 6:34 PM
Neither of these look "good" to me. There is so much visual noise on the page, like someone turned the "AI Slop" dial to 11. In fact I prefer the simpler design Haiku made.
today at 6:46 PM
It's not really about whether the design looks good. It's about if the model can take the design given to it and replicate it in code. Opus 5.5 matches the designs almost to the pixel. Haiku built something else entirely.
today at 7:00 PM
I guess I'm giving GP feedback about their product diffui.ai, not really about Opus' performance.
today at 6:47 PM
Totally fair, but I'd encourage you not to look at the design so much as the task. This was a design that's part of a benchmark test suite specifically for image->html conversion. The dense visual noise / complexity / flowing svg shapes are things that most LLMs have trouble with.
It's meant to be a good test, not a good design.
today at 6:43 PM
It’s not really AI slop, it’s how most modern SAAS websites look like.
today at 6:12 PM
The monthly API credits for Max plan seems fantastic, especially considering Haiku pricing. Being able to actually use my Claude plan for other harnesses and use-cases on top of regular CC usage is everything I wanted.
Anthropic has really been doing all the right things in the past few weeks, while OpenAI continues to fumble the bag.
today at 6:33 PM
Note that this is Anthropic Trojan-Horsing the previously announced June change in with a model release, where the Claude Agent SDK can no longer be used with Claude subscriptions and is now billed with API credits only.
https://support.claude.com/en/articles/15036540-use-the-clau...
today at 10:37 PM
Sorry for the misleading wording on this page; we’ve updated it to clarify that the Claude Agent SDK can still be used with subscriptions.
today at 9:45 PM
Yep they're definitely getting ready to yank using your subscription with the agent SDK - this page has just been pulled: https://support.claude.com/en/articles/15036540-use-the-clau...
That's not a good sign for Conductor...
today at 10:15 PM
this is a common question. we're updating the faq now to make sure it's more clear
today at 6:40 PM
Ah, that sucks. I'm using Paseo to run Claude Code; I guess that just got a whole lot more complicated.
today at 9:41 PM
Thanks for sharing this! These vendor lock-in attempts are very annoying.
today at 6:18 PM
yeah totally agree. esp how efficient it can be to have a subscription quota-paid orch spin up a bunch of API agents, this is kind of like free money to encourage what was already an easy way to save money (via batch pricing)
today at 7:18 PM
> Being able to actually use my Claude plan for other harnesses
Wait what? This has gotten their blessing?
today at 8:34 PM
You could always use Claude models on other harnesses via API... just not via subscription. Now they give you $100 worth of API tokens to use on opencode or Pi. Which is better, but still not the same as OpenAI were you can use the subscription on Pi without problems.
today at 8:38 PM
Absolutely not.
today at 6:29 PM
About time Anthropic released a competitive cheap model. Haiku 4.5 has been too expensive compared to its performance for months now (in fact I don't remember being too impressed even when it was released). This one actually looks worth using in some scenarios. If it's really as much of a step up from Luna as the benchmarks they've shown indicate, it'll probably replace Luna in my workflows. 100k tokens is a pretty low threshold before the price goes up, but I tend to use these smaller models for smaller tasks anyway.
today at 10:16 PM
It's around Qwen-3.8, and Sonnet 5.5 level, but a lot cheaper. It is also really fast.
My tests for Haiku 5.5: https://aibenchy.com/compare/anthropic-claude-haiku-5-5-xhig...
Twice as expensive as Luna, but also considerably smarter too:
https://aibenchy.com/compare/anthropic-claude-haiku-5-5-xhig...
today at 10:37 PM
A really big difference can be seen in this basic CSS animation generation of a solar system of Luna vs Haiku:
https://aibenchy.com/compare/anthropic-claude-haiku-5-5-xhig...
today at 6:35 PM
This is great! Been using GPT 6 Luna for decompiling my childhood favorite game (Age of Mythology) and this means I can throw Haiku into the mix as well. 17352/21965 functions matched so far...
today at 9:25 PM
How do you validate the functions are correct? I did something similar, letting it (mostly deepseek 4.1) translate from assembly to C but it commonly made mistakes, some really hard to discover and fix.
today at 10:07 PM
You should be able to validate it by creating tests against the assembler.
today at 6:42 PM
can you share more details? was this very involved or asking codex/claude/open code with a 1 shot like approach?
today at 6:56 PM
I'll write a blogpost when I actually have it working, but basically I gave the game .msi installer to claude opus 5.5 and said to read these blogs:
- https://blog.chrislewis.au/using-coding-agents-to-decompile-nintendo-64-games/
- https://blog.chrislewis.au/the-long-tail-of-llm-assisted-decompilation/
And to setup a harness that will decompile the game and start doing a matching decompilation of every function. It set up a bunch of tooling and started a service in the background to do this actual decompilation campaign. I put some instructions into the main opus chat now and then to e.g. add automatic git pushing including a nice svg chart of progress and to switch model strategies here and there i.e. to do a first pass with a cheap model and then switch to opus/sol if the small model can't solve it.I could now one-shot a new game, yeah.
today at 7:16 PM
Maybe a Show HN? I would be quite interested in seeing the results of this project
today at 9:42 PM
decompilation doesn't trigger any safeguard refusals? I would have assumed it would but glad it doesn't. Very cool and would also love to hear more.
today at 10:25 PM
Surprisingly, no. I've been using Fable and Astra both to orchestrate decompilation of a relatively modern game (delivered via Steam) and they have no qualms about it.
today at 7:07 PM
i love Age of Mythology, but why did you feel the need to decompile it? Its got a great world editor if you were trying to "mod" it.
today at 7:19 PM
I want to get the original (Age of Mythology Gold Edition) running natively on macOS and then port it to WASM to run it on the web so I can easily play it with friends
today at 7:59 PM
Intriguing, I wonder how far you could go with turning games into websites.
Like could total war become a browser game?
today at 8:33 PM
Probably. There have been dozens of examples of taking old games (Crazy Taxi, Super Monkey Ball, Quake, etc) and making WASM browser equivalents using AI to decompile them just on "Show HN" alone.
They often ship the original assets in a somewhat brazen disregard for basic copyright law even when the games are still for sale on places like GOG though.
today at 8:55 PM
I saw recently that someone had ported Halo CE to the web and had 1024 players in Blood Gulch.
today at 8:09 PM
Last time I gave that a try (without LLM assistance though) it was really hard as games DirectX calls cannot simply be glued to WebGL so performance was bad.
today at 8:50 PM
My weekly limit __on a Pro sub__ has not gone over 50% since before the pre-Fable promos, but usage has been pretty much the same from my point of view. Maybe I am holding it right? Anyone else getting this?
As such, I do not need to even reach for Haiku, and 4.5 was so inaccurate that it often cost more to do so in the past. Sonnet 5.5/low has been good for this kind of thing, and i didn't even touch thinking tokens or any of that. Opus 5.5 low for questions/repros, medium for implementation, basically never reaching for anything above that anymore. 5.5 has been great, so I'll try Haiku, but don't see myself going out of my way to integrate it.
today at 8:58 PM
I'm excited for API use. I run some agents, mostly on Luna 6 right now. It's just tool use, web browsing, etc, so something dirt cheap, but also not super dumb, is much appreciated. Having a Luna competitor is nice.
today at 8:52 PM
there is no way in hell that im gonna use Haiku too, tho weekly limits become a problem for me in a last couple of month tbh
today at 7:08 PM
At work we use haiku 4.5 for a handful of latency sensitive tasks that are fairly simple. It performs well. Just started testing 5.5 as I’ve been anticipating a nice improvement since it was teased. Results so far are trash. Prompt leakage even. And it’s slower. I guess it’s cheap but I think they got the balance wrong on this.
today at 9:41 PM
Exact same thing here.
Both evals and Human pairwise tests for our use case are giving Haiku 4.5 first place in pretty much all tests.
No we'll try understand if we need to change our prompts to match performance ...
edit: maybe this will help: https://platform.claude.com/docs/en/build-with-claude/prompt...
today at 7:18 PM
Curious as to why. Haiku 4.5 has been far away from pareto frontier for a long time. Maybe you need to update your prompt for the newer model in your workflow.
today at 6:04 PM
I wonder if we have an AI LLM equivalent to Moore's Law. Like how often do we expect improvement in this technology and with what timing?
today at 6:18 PM
Yes -> every 18 months they've gotten 90% more efficient for the same level of quality for about 5 years. There's little sign that trend is slowing. If anything, there's reason to believe that System 1 models (plus potentially 1-2-3 workflows) may increase that over the next 3-5 years.
You'll know when the trend stops -> when the intelligence differential between smaller models like 7B starts to grow instead of shrink from 32B models -> that means 7B is getting about as smart as it can get. Then, 32B will follow next, then 70B, etc etc.
We haven't yet seen that at any size AFAIK.
today at 6:22 PM
Andrej Karpathy said once that he expects superintelligence could fit in 1 billion parameters.
today at 6:39 PM
Super intelligence that doesn't have to deal with the real world, maybe.
I wouldn't be surprised if less than 1B param equivalent of our brain deals with solving math and writing computer programs and physics and all the things we tend to associate with "intelligence" - especially if you ultra optimized for that, I doubt our brain works like that.
Dealing with the real world, I highly highly doubt it.
today at 9:34 PM
It would be interesting if running ends up being a more complex task than advanced math. And our brains are just 95% allocated to dealing with the real world.
today at 6:47 PM
How about if we get away from written text as the input, to something more fundamental, that then also is able to produce text (among other things)?
Given that humans learn to talk while having encountered a measly number of word instances, and, given enough time, we should always be able to improve on the lottery that is biology, it does seems fairly likely.
today at 6:25 PM
According to Epoch AI:
> The cost of achieving a given level of AI performance has fallen about 47% per quarter since 2023, or 13× per year. [0]
0. https://epoch.ai/publications/the-plunging-price-of-thought
today at 6:46 PM
Then why are AI plans still so super expensive, and AI spending going through the roof, while all the subsidies are ending?
today at 8:58 PM
https://en.wikipedia.org/wiki/Jevons_paradox
AI gets cheaper, people use it everywhere. Google searches, for example. Now we want to crack math problems and spend weeks with unreleased models.
If you used GPT-2, it'd be incredibly cheap. You basically can't use it for anything and it's simple to serve.
today at 6:57 PM
The cost per fixed level of intelligence is dropping, but we're also getting dramatically more intelligent models.
today at 9:44 PM
Reddit is full of people complaining how they burn their 200$ sub in half an hour by starting ten Max sub agents. That’s to say, many people just don’t know what they’re doing.
today at 7:22 PM
Because models are only getting better at a rate of 10% per year, people always want the best quality possible. You can get SotA performance from a year ago for a fraction of the cost, but why would you use Opus 4.5 when you can use Opus 5.5?
today at 6:48 PM
Because it's increasingly useful and the thing you are substituting (human time) is much more expensive.
today at 7:57 PM
Apart from what others said about using more intelligent models instead of cheaper ones, token usage is also increasing a lot. Classic Jevons paradox
today at 6:56 PM
At least for me the Claude plans seem like an incredible deal and I never hit my limit.
today at 7:07 PM
today at 6:08 PM
reminds me of this blog post: https://campedersen.com/singularity
today at 9:28 PM
Whew, at least I won't have to hand-code solutions to the 2K38 problem!
today at 6:07 PM
I've heard tell about 100% of certain types of work being ended in batches of six months. For years. Truthfully, I'm skeptical, but accuracy wasn't prioritized.
today at 6:05 PM
double the information density every 2 days?
serious bit: if you think about how these smaller models work, at the end of the day it seems that they are now capable of forgetting useless information because they're able to derive it in reasoning allowing models to become smaller at the cost of requiring more reasoning tokens to solve a task.
today at 6:09 PM
Knowledge will be shifted to systems like n-gram augmentation which are relatively cheap and will not compete with reasoning capabilities for weight saturation.
today at 6:06 PM
Hopefully enough runway for an existing model to train the next to be better than itself with absolutely no human intervention.
today at 6:23 PM
> Claude Haiku 5.5 is our fastest model to date at each model’s standard speed, although it runs less quickly than our Opus models in Fast Mode.
Opus 5.5 runs 117 tps average on Openrouter, so it must be at least 10-20 tps slower for them to mention. IDK why they mention this as it does not help for marketing though. https://openrouter.ai/anthropic/claude-opus-5.5
today at 6:25 PM
Maybe they think it's of interest.
today at 7:12 PM
Is that page showing Opus TPS stats in fast mode? IIRC fast mode is 2.5x speed, so that would be 293 TPS, no?
today at 7:28 PM
117 tps is the fast one, regular speed is 69 tps.
today at 10:19 PM
Has Anthropic released a decision model ala Jev? I wonder if they’ll launch something soon
today at 6:13 PM
Good to see Anthropic back alternative OSes.
today at 8:01 PM
https://www.anthropic.com/claude-haiku-5-5#further-updates
This section makes the reader think: why would I not pick Sonnet 5.5 instead of Haiku 5.5?
today at 6:52 PM
131tok/s P50 according to OpenRouter currently, though might move up or down over the coming days. If it sticks at that speed, roughly twice the throughput of Luna and far lower latency (up to 2sec depending on provider) is impressive, though the 5x price increase beyond 100k is painful.
Was a big fan of Haiku 4.5, though understand why for most Sonnet was the far better option back then.
today at 7:30 PM
> we’ll roll out a new monthly API credit to all Max and Team subscribers for use on the Claude Platform. Max 5x users will get $100 in credits per month, Max 20x users will get $200
They’re definitely planning to make the subscriptions API based so they can charge you full price.
today at 6:10 PM
From these selected benchmarks, it looks like it smokes Luna capability-wise. Excited to put it through its paces
today at 8:13 PM
Alright its still little early since there is not enough independent testing but this looks very promising and I wasn't expecting anthropic to beat GPT-6 Luna especially at the same price. Haiku 5.5 beats Luna on every shared benchmark Anthropic published, particularly computer use and agentic coding.
today at 6:19 PM
Top of the page in 17 minutes? Now I know what y'all do while your agents are working.
today at 7:30 PM
It's a brave new world... of idleness!
today at 7:15 PM
> Second, this week, we’ll roll out a new monthly API credit to all Max and Team subscribers for use on the Claude Platform. Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users. These credits are designed to allow our users to experiment with building tools, apps, and agents that call our API.
This is kind of nuts
today at 8:00 PM
Is it realistic or cynical for me to assume this is to wean developers off the heavily subsidized subscriptions? Presumably it's using similar compute.
today at 8:42 PM
Anthropic was ignoring the usage of third party harnesses. Not anymore.
today at 6:37 PM
The forgotten model is back on the map. I actually got OK mileage when I tried it for coding months ago. Maybe I'll try it again, with Opus guiding it, and see how it goes.
today at 6:40 PM
Apparently quite a bit smarter than Luna, I wonder what use cases it can cover. I actually honestly don't need a Haiku level AI to be that smart, and looks like you pay for it in the per token cost, I need speed mainly. I might even rather have a dumber but much faster model for things like web searching and parsing to retrieve results for the app or other LLM to do things with.
today at 6:12 PM
How are y'all using Haiku though? I rarely select it.
today at 6:18 PM
I have a zsh functions that calls claude code with haiku to suggest commit messages, is faster and the instructions are two lines.
I also have an "ask" script that I use daily to ask simple stuff, it can access websearch and webfetch, it's more than enough to parse logs, ask for commands, quick research on the internet, small stuff. https://github.com/mariocesar/dotfiles/blob/main/common/.loc...
I use haiku for things that needs to be quick, have really clear instructions.
today at 8:44 PM
with claude -p seemingly now using api credits I guess this approach will have to change unfortunately :/ I wonder what will be the best cmdline way to do things like these
today at 6:21 PM
I've been using GPT-6 Luna in some capacity for nearly all my agent workflows. It's just a really good model, and the pricing is cheap. If Haiku 5.5 is better, and the same price (under 100k context... which is a big caveat) i'd probably swap it.
today at 6:25 PM
It’s absolutely better than Luna. It feels closer to a “sonnet 5.2” if that makes sense.
Of course it’s not as big, and hence falls-off quicker. I’d consider the 100k a “promotional price” to match Luna’s token pricing while delivering noticeably more intelligence.
today at 6:15 PM
Opus often picks it when it's doing a "find me something" subagent. But largely it's been held back by being fully a year old at this point, and priced at a much higher price than models that are far more capable.
today at 6:49 PM
not haiku, but luna - last week i used it for things like "read this historical dump of 15k support tickets and break them into categories that make sense, then propose help docs that i could write to handle the most frequent queries in each category"
used <10% of my 5hr limit on a $100 codex plan.
today at 6:21 PM
I was waiting for this.
Planning on doing flash analyses of PRs that impact evals in some way, and then post comments on GitHub whenever there’s flaws in them
today at 6:16 PM
It's great at parsing documents inexpensively. For the few skills/plugins I've made, I usually instruct Claude to use Haiku for low-reasoning grunt work.
today at 7:14 PM
My software application uses Haiku in production more or less as a Jev. I do not use it for coding or development.
today at 8:33 PM
Why not have a CPU-first decision model for free? check out gutsy
today at 6:33 PM
I'm building a game that incorporates LLMs as a game mechanic.
I've been using Luna, but I'll probably switch to Haiku.
today at 7:21 PM
[dead]
today at 7:02 PM
[flagged]
today at 8:15 PM
Finally! I understand Haiku is the less intelligent model, but the gap between Sonnet and Opus has been far too wide for about a year now.
today at 7:44 PM
I am perpetually confused about every name and version combination from both OpenAI and Anthropic. Especially in conjunction with the effort levels.
today at 6:09 PM
It’s about time Haiku got an update!
today at 6:52 PM
It fails the "How many r's in <word>?" test.
I ask:
> how many r's in diminished
It answers:
> Diminished has 1 r.
today at 7:19 PM
According to AA benchmarks, it uses 162k output tokens per task (with max reasoning) - over double GLM-5.3 Flash for similar level of Intelligence
today at 6:14 PM
The important question though...how does it do making a pelican on a bicycle?
today at 7:27 PM
no AA benchmarks yet and the last chart in the announcement makes Haiku look useless vs new Sonnet pricing, interesting to see what 3rd party benchmarks show because i think Anthropic are costpertaskmaxxing here and it's going to look more like that bottom chart than the top ones.
today at 7:10 PM
Seeing people talking about the Agents SDK -> credits change makes me wonder does it impact Zed or likes.
today at 6:05 PM
Happy about the Sonnet cache read price cut.
today at 6:11 PM
That was effectively required to match GPT-6.1 Sol (costs and caching prices are now equal). Sonnet 5.5 made zero sense to use over Opus 5.5 under the old cache prices.
today at 7:02 PM
[flagged]
today at 6:11 PM
How does the price compare to Luna? At least looking at the numbers it is noticeably better at most tasks.
today at 6:12 PM
IMO, this is better. Luna is super cheap, but it's not that capable. At higher levels of reasoning, it's not that fast.
This is more expensive, but it also looks like it's better enough that it's far more useful.
I also won't be surprised if you look at cost per completed task + wall clock time that it comes out ahead for the majority of what you'd want to actually use it for.
Luna will still be a great option for doing non-engineering tasks super cheaply.
today at 6:13 PM
For prompts under 100k tokens, it's priced the same as Luna - $0.10 in, $0.50 out.
For prompts over 100k tokens it's 5 times more expensive - $0.50 in, $2.50 out.
today at 9:44 PM
Really excited to use this
today at 7:00 PM
It's finally here ! Need to take a look at some benchmark now
today at 8:59 PM
request to anthropic team release haiku as os model
today at 6:07 PM
Wow, the rate of improvements in the AI era is staggering.
GDPval-AA v2.1 as of now: 1620
GDPval-AA v2.1 for Haiku 4.5: 735
The 100k tokens pricing makes sense, looks to be a hedge against OpenAI's decisions API and Jev or its open source alternatives that are springing up.
Nice release, congrats to Anthropic.
today at 6:10 PM
today at 7:34 PM
Finally it has arrived.
today at 7:14 PM
Begging Anthropic to let us use Claude subs with harnesses other than Claude Code at this point.
today at 6:23 PM
The important question is, does it talk in incomprehensible Claude-ese like the other Claude 5.x models?
today at 6:17 PM
> Second, this week, we’ll roll out a new monthly API credit to all Max and Team subscribers for use on the Claude Platform. Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users. These credits are designed to allow our users to experiment with building tools, apps, and agents that call our API. They can be used on any of our models. For more information, see our Help Center article.
Did anyone read this? We get free API credits on some plans now
today at 7:07 PM
will use it to replace luna in production !
today at 6:19 PM
Probably the same scam as the last Haiku update I guess. Uses more tokens to compensate for the lower price.
today at 7:41 PM
To correct myself. The price was not lower. It was up about 20% for the tokens. But, the big price hike was that it used a lot more tokens for the same tasks.
today at 6:09 PM
Where is Pelican? ehehhe
today at 6:14 PM
[flagged]
today at 6:23 PM
I think it's a joke at this point, but also the visual benchmark is a remarkably dense method for demonstrating how good a model is.
today at 6:42 PM
Yeah, people like to poop on the pelican. But pelican quality still correlated with overall model capabilities reasonably well, and you can immediately see and interpret it. It's a running gag, but it also does have some actual value.
today at 6:19 PM
So do you have the pelican or no?
today at 7:08 PM
[flagged]
today at 10:39 PM
What is i want to code svg files though?
today at 7:42 PM
Can you please not be so sneery/grouchy? That's far worse for HN than suboptimal benchmarks. The guidelines specifically ask us to avoid being curmudgeonly.
today at 8:25 PM
Ok, so I looked at all of your links, but nary a pelican to be found. How am I supposed to know what all these numbers mean if there is no pelican?
today at 6:12 PM
I remember a friend asking me why LLMs suck so bad. She was using Haiku 4.5 and that poor model couldn't keep track of the context within 3 messages.
She said she was using Haiku 4.5 because she was advised to be careful with the spending.
I hate that model so much lol.
today at 6:48 PM
> but they still block penetration testing and other techniques more likely to be used by attackers.
>
> Haiku 5.5’s biology safeguards are the same as for Sonnet 5, Sonnet 5.5, and Opus 5. They allow research biology questions but restrict access to requests that we judge as likely to cause harm. Organizations working on wider-ranging biology and cyber activities can apply to our Life Sciences Verification Program and Cyber Verification Program.
I would like to take a moment of your time to tell you about some of the "bioweapons" Anthropic has blocked that involved Haiku!These are the examples from "Detecting and countering misuse of AI: September 2026" - https://news.ycombinator.com/item?id=49647300
> Importantly, because our biological safety classifiers robustly block content involving high-risk biological research (in this case, the construction of enhanced pandemic potential pathogens), all of these exchanges occurred on models in our weakest class of models (specifically, the models were Claude Sonnet 4 and Haiku 4.5, the latter of which the user began using after Sonnet 4 was deprecated).
>
> Upon a detailed examination of the exchanges, we estimate that the uplift provided by Claude was primarily clerical assistance in data analysis, study ideation and design. This is consistent with our understanding of the capabilities of Sonnet 4 and Haiku 4.5, which are not able to perform expert-level biology research tasks; we estimate that the uplift provided to the researcher was limited and substantially lower than it would have been from one of our more capable models.
Anthropic then says for the above, "we estimate that the uplift provided by Claude was primarily clerical assistance in data analysis, study ideation and design"While doing my best to avoid comment, please note, they're talking about a domain expert in a state research institution using Claude to do paperwork.
What did they save us from? What bioweapons did these filters prevent? From the front matter report,
> The above LLM platform is not the only route via which researchers engaged in viral gain-of-function research have used our platform. In May 2026, we discovered a researcher outside the US using Claude in their research on highly-pathogenic avian influenza (“bird flu”). The research focused on viruses’ adaptation to mammals, and the mechanism by which it causes severe disease beyond the respiratory tract.
OK. Sounds serious. "Gain of function research..." but who and why? > The researcher pursued this work in a credible institutional context, and interacted with Claude over the course of several weeks, exchanging thousands of messages. In these exchanges, the researcher leveraged Claude’s knowledge of the scientific literature to assist the researcher in study planning and design, data analysis, and the interpretation and prioritization of experiments. The researcher also used Claude for editorial assistance in writing up the research.
So this was a researcher inside of some country's national lab ("credible institutional context") doing research on dangerous viruses using Claude for "for editorial assistance in writing up the research."What "uplift" are you providing to scientists working at specialized global BSL-4 labs that already have – and I quote their report - "physical access to such isolates." (as in samples of viruses)? Are we uplifting their grammar?
These "safeguards" are being expanded. The scientists I know can't use Claude for grammar checks or anything serious. You can try it for yourself.
today at 6:16 PM
> Haiku 5.5’s cybersecurity safeguards are more restrictive than Haiku 4.5’s, but somewhat less restrictive than those we’ve applied to other recent models. In cybersecurity, they permit a wider range of defensive tasks than our safeguards for Sonnet 5.5, but they still block penetration testing and other techniques more likely to be used by attackers.
If you block pentest or "other techniques more likely to be used by attackers", then what does "permit a wider range of defensive tasks" even mean?
Any defensive task that's meaningful is almost indistinguishable from legitimate red-teaming that then falls under 'likely to be used by attackers". If only they would just stop nerfing these models, that'd be great. No APT is waiting around for Anthropic's permission, so might as well let us have some cool stuff.
today at 7:01 PM
Yeah it's too little too late, cat's out of the bag as people know that GLM 5.3 exists and is great at defensive and offensive cybersec.
(sadly Mistral Large 4 isn't up to par - but Mistral serves GLM at 130 tps!)
today at 10:04 PM
[flagged]
today at 9:57 PM
today at 8:14 PM
[flagged]
today at 8:16 PM
[dead]
today at 6:57 PM
[flagged]
today at 6:22 PM
[flagged]
today at 6:22 PM
Does anyone still use Haiku model?
today at 7:04 PM
I use it for title generation basically. Will have to see where this one fits in
today at 6:23 PM
Opus 5.5 does :^)