SpaceXAI's Grok 4.6 Scores 61 on the Artificial Analysis Intelligence Index
87 points - today at 4:54 PM
Sourcesatvikpendem
today at 5:08 PM
Cursor, since Grok 4.5, has had an incredible deal for frontier level models, their subscription now goes way further than OpenAI or Anthropic. Even on their lower tier plans you can use a lot tokens on their of their first party models (Grok and Composer) and not really run out comparatively. Combine them with an orchestrator and implementor type setup and it goes even further.
How does Grok 4.5 compare to Opus >= 4.8 though?
I'm willing to pay 2x for a 10% smarter model. Intelligence matters that much (because 10% smarter probably saves, on average, several hours of human time).
hmokiguess
today at 5:44 PM
I believe they are the only western provider that has Kimi K3 on a subscription plan today as well. I would love to ditch Anthropic and be on Kimi if there were a subsidized plan like that with ZDR
Can you explain what you mean? These days courtesy of an addictive reset game OpenAI is playing, I can't find anything with frontier intelligence that's more cost efficient...
jesse_dot_id
today at 5:14 PM
Goes even further to exfiltrate your data, yeah.
greenavocado
today at 5:34 PM
That would be Muse Spark Contributor Tier. 12-21x price reduction at the expense of your digital existence.
CuriouslyC
today at 5:59 PM
I'd be the first model I'd reach for if I was providing a free service to AI gooners though. Serves them both right.
Grok is not the best model around, but it's decent. It gets the basic job done at a low price. I don't think it can advance frontier Math, yet.
Probably can't advance frontier math yet, yeah. But please let us know other places you want to see Grok improve for future models!
nylonstrung
today at 5:11 PM
I have never met a single human being who uses Grok for coding
I have. He was using it due to philosophical reasons the same way many people have philosophical reasons for avoiding it. I don't know how many people are like that, but it's not exactly where you want to position your product if you're a business.
Personally - and I know I'm not alone with this sentiment based on comments I see on this site - I wouldn't touch Grok no matter how good or cheap it is. I don't trust Elon and I don't want to give another dollar to the world's richest person who turns around and uses the money to interfere with elections. The guy I know uses it for essentially the same reason I won't use it.
I had a security incident the other day and Grok was the only model that would help. Claude and GPT refused on ethical grounds and only gave general advice. In an emergency, I'd only trust Grok. However, that's the only time I used Grok for coding (since Opus 4.8 it would take a lot to get me to switch away from Anthropic)
A bunch of SWEs at my work use it as their primary model.
We have Claude, ChatGPT, and Cursor with essentially no cap on spend (top guy is spending over 10K a month on AI at API prices), and he hasn't had his hand slapped.
So it's not like they are using it purely because it's cheaper.
I think people like to use it for its speaking style, pretty solid performance, and its speed.
I use. I used to be a Claude user. Since trying Grok 4.5 and especially Grok 4.6, I don't want to go back to Claude any more (I have early access to 4.6).
Grok is 3x+ faster than Claude and I can't tell the diff in engineering work quality. As an engineer, speed is important to me.
For $30/month, I'd expect it to have higher usage limits than Claude Code and Codex.
It really does. I felt like I could have spent $1000+ api token on claude for the amount of work on my $30 grok subscription.
An hour in, I've been running four terminals full bore on my $20/mo Grok sub and I'm at 9% for the week. Codex or Claude would easily have hit 5-hour or weekly limits.
I refuse to use that product because of the parent company.
I can respect if you say you hate their guts. Everyone has their worldview. But over moral or ethical stand? You don't have any if you're using Chinese models, or fly Middle East airlines, or countless of other products. Don't delude yourself.
Recurecur
today at 5:38 PM
I wonder why folks feel the need to virtue signal like this…?
As an aside, do you use Apple equipment despite Apple’s use of Chinese slave labor?
forestrywat
today at 5:59 PM
It's not virtue signaling. People are allowed to have morals. And they are allowed to act in ways that align with those moral beliefs.
> do you use Apple equipment despite Apple’s use of Chinese slave labor?
If someone said they didn't buy Apple because of this, I'd support them in that too.
People. Are. Complicated. We're allowed to have beliefs, we're allowed to have beliefs that aren't internally consistent. You gotta get over the idea that people aren't genuine believers in their beliefs.
rootusrootus
today at 5:57 PM
Do you not ever 'vote with your wallet'?
100%, it's an easy pass given that it is always playing catch up.
Does it matter? Why turn it into a popularity contest?
the new models are quite good, give it a shot
Folks working in US govt tend to, based on convos I've had with one such person.
supriyo-biswas
today at 5:13 PM
I'm only being forced to use it at $WORK since some people overran their Cursor bill, so everyone gets Cursor Auto enabled by default which routes to Grok 4.5.
Recurecur
today at 5:34 PM
Hi! Grok’s worked quite well for my use cases…
It’s also a great deal!
locknitpicker
today at 5:42 PM
> I have never met a single human being who uses Grok for coding
Me too. The only people I ever saw using grok were using it by accident as they used copilot in auto mode and noticed some prompts were thrown it's way.
I saw far more people using Mistral than grok.
DetroitThrow
today at 5:15 PM
I've tried it on my "let's run every model in parallel and see which finds more edge cases" type of tasks, and Grok 4.5 was really behind Opus/ChatGPT but ahead of Gemini - despite having a strong showing on benchmarks.
That makes me really skeptical of it being GPT5.6-tier, much less Fable-tier, based on some of these benchmarks alone. But I'll test here shortly.
Seems the cache read pricing almost doubled from $0.30 in Grok 4.5 to $0.50 in Grok 4.6.
In my experience in heavy coding sessions most pricing is just cache read and cache write like 80% of my token bill.
petesergeant
today at 5:43 PM
Interesting. Grok 4.5 is a capable model, although not quite at Fable/Sol levels. Will be interesting to see how this holds up. Musk appears to have made a savvy choice buying Cursor's data.
thiago_fm
today at 5:08 PM
I often wonder if there's a chance, even if minimal... that they stole the weights of the Anthropic models they run on their datacenter... or are actively destillating it.
I think the more likely explanation is that the Cursor data they effectively acquired for $10B was extremely valuable for their training when combined with the insane number of GB300s xAI has for training.
Cursor was 60B. The 10B number was the breakup fee if the deal fell through.
> ... or are actively destillating it.
I just assumed every model manufacturer is distilling from the frontier models. If they aren't they are definitely trying to do it.