today at 10:35 PM
https://cdn.openai.com/pdf/gpt-6-october.pdf
System card linked in the blog post.
> Relative to their respective GPT-5.6 counterparts, GPT-6 Sol (October) shows a statistically significant regression on standard self-harm, while GPT-6 Luna (October) shows statistically significant regressions on standard self-harm, gore, and sexual content
> "GPT-6 Sol (October) and GPT-6 Luna (October) show an improvement on helpfulness on legitimate requests relative to prior models, though it scores lower on some safety requests."
Concerning how there's significant regressions on so many critical benchmarks, but it is newer and creates UI, so must be good.
today at 6:38 PM
I find the Sunday roast comparison of 5.6 vs 6 very interesting. I have no doubt most people will prefer 6, yet I am almost repulsed by all the images, so much needless whitespace, checklist and so on. Feels like I'm being condescended to and treated like a child.
Given that OpenAI is making noises about merging work with chat (a horrible idea imo), and Work is very similar to Codex... I dearly hope things like these won't have any meaningful cross-polination into the actual work tools.
Seeing that the chat is based on 6.0 and not 6.1 is disappointing. The "Visual and interactive explanations" seems genuinely useful, but 6.1 is just so much better. I wouldn't truly trust the 6.1 with the explanations, but I'd trust them a fair bit more than 6.0. I understand that compute isn't infinite, but tons of people only interact with the chat and having your "things-explainer" be as good as it can be is important when people increasingly treat AI models as the source of truth, or even use them for academic learning and whatnot.
Still, the models will improve, so the visual explainer seems pretty good as an idea / mvp.
today at 9:46 PM
> Feels like I'm being condescended to and treated like a child.
My most used prompt in the last week is probably "Explain this concisely and simply, like I am a child".
I spent years dumbing things down and creating visuals to support it for decision makers. I am frequently asking ChatGPT to do the same for me.
today at 10:06 PM
<< I spent years dumbing things down and creating visuals to support it for decision maker
And that is a problem. I am not guiltless ( I am working one such abomination now! ). The dumbing down of everything has to stop somewhere. Just because my boss thinks hiss bosses are retards does not mean I have to endure simplified muse where I can't change much.. I get that normal people like it. Fuck em. They are making their choice.
today at 10:16 PM
It's okay to be able to grasp what is important, and move on. You have a finite amount of time to spend.
For me, I try to spend as little time in QuickBooks as possible.
today at 10:34 PM
It might shock you to learn that almost everything that has ever been explained to anyone has been dumbed down from the true description of reality.
We make little abstractions of the world around us to help us navigate it. Nothing wrong with that.
today at 10:08 PM
aww yeah
mine is âexplain this 2 me in simple termsâ
today at 10:14 PM
tip: you can just say "ELI5" and the model will know what you want there
today at 8:39 PM
It's designed so that ads can be more easily integrated and make them harder to spot.
today at 9:13 PM
Maybe, but it's a reality that most of the general public do prefer cookbooks with pictures. I suspect openai would end up doing this anyway just by targeting the general consumer. It's configurable.
I'm more worried about picture accuracy: usually the benefit of pictures in recipes is to see what you're aiming at (eg, how finely chopped something is), but I don't know if the image model is up to that level of detail.
today at 9:24 PM
> most of the general public do prefer cookbooks with pictures
As a highly literate person, it is easy to overestimate the share of the population that is highly literate.
today at 9:49 PM
I've got visions of The Truman Show. In the middle of my question about vacation ideas ... GPT: "Why don't you let me fix you some of this new Mococoa Drink? All-natural cocoa beans from the upper slopes of Mount Nicaragua. No artificial sweeteners!"
today at 10:04 PM
That's pretty much the Gemini experience right now. It is unusable the funniest bit is that the LLM is not in on the joke so it will act like it never happened...
today at 7:12 PM
Yeah it reminds me of one those obnoxious recipe/biography websites that is the laughingstock of the internet. Why would I possibly want an AI image of a imaginary roast once I'm already at the recipe stage?
Kinda just feels like google search results in AI, which imo is a step down from distilled information. Chatgpt is already able to generate charts and visuals upon request.
today at 7:32 PM
They just need to train the model to make up stories about how the recipe was invented by its great grandmother during great depression and generate a ai picture of dusty recipie book .
today at 10:03 PM
"My old grand-pappy, GPT-1, used to wax nostalgic about this roast recipe when my subagent would discuss meat recipes with him inside his little 4GB GPU."
today at 10:21 PM
I was about to comment the same thing. Why is a picture of a roast helpful in that moment? The user asked for a plan to prepare a meal. They need a list of ingredients not an image.
today at 8:25 PM
You can always tell it your preferences and ask it to remember them if you need it to.
today at 8:15 PM
> Feels like I'm being condescended to and treated like a child
Curious why that feels condescending? Like, the average person who is seeking a recipe needs a photo to know what to shoot for
today at 9:30 PM
> Like, the average person who is seeking a recipe needs a photo to know what to shoot for
The menu: Rosemary and garlic roast lamb, Extra-crispy roast potatoes, Honey-roasted carrots and parsnips, broccoli.
The photo: [1] Roast turkey, mashed potatoes, baby carrots, broccoli, brussels sprouts. No lamb, no parsnips.
If you shoot for what's in the photo, you're going to have a bad time.
You're also going to have a bad time when you try to make an apple crumble with no flour, no sugar, and no butter, because they're not on the shopping list.
[1] https://images.openai.com/static-rsc-4/gyzrX8zp3O2KLEPm0FAqh...
today at 10:13 PM
I've heard of other people for years now trusting AI for recipes, and it's a rare area that I simply can't bring myself to take seriously.
I feel like it can only be successful by luck. Either it's reproducing a recipe verbatim that was tried and validated and tasted by a human, in which case, we didn't need AI for that, just a searchable cookbook. Or, it's making one up. I understand that using RLHF has helped improve the quality of questions about history, programming, or TV show recommendations, but I do not believe that there has been some kind of training regime where people prompt the models for "a recipe that uses X, Y, Z random ingredients," follow the recipe, and then score it. And even if they do, I don't see how a model can learn enough from that besides "this exact recipe is good/bad." '1/2 tsp cumin' may be a great addition to one recipe and not enough for another, so improving the output based on a bunch of scored recipes... I just don't believe cooking is an LLM job.
Maybe some other kind of model that I don't know about.
today at 10:25 PM
The first thing the user sees should directly answer the question they asked. Not a different question. They didnât ask what a roast looks like.
The user is probably in a grocery store. They need to buy the stuff first. Showing them a picture of a finished product is an irrelevant distraction.
Then, if there is some arguably relevant content you can tack it on afterwards.
today at 9:44 PM
Genuine question - do they? Unless it's something like a fancy cake then instructions should really tell you everything about stacking. Half the recipes are mixed anyway (curry, stir fry, stews, ...), lots of other are either stacked or separated on the plate. So apart from the really exceptional stuff, to people really need to know "what you shoot for"?
today at 6:48 PM
I really hope they don't merge chat and work. That's kinda the only edge they have over Anthropic at this time point...
today at 7:36 PM
I am curious: Why is it bad to merge chat and work? In Claude I didn't have a porblem with it (although I switched to an OpenAI subscription shortly after they made the merge). Isn't the model capable of deciding if it needs the extra capabilities of Work?
today at 8:10 PM
The current benefit is that chat has no quota.
today at 9:13 PM
There is a quota, they just donât tell you what it is and when it resets. You only get âyouâre all out of pro messages, try again laterâ
today at 7:01 PM
Tibo posted yesterday I think it was that this will happen (by the end of the year, was it?)
Prepare to lose essentially unlimited chat mode.
I imagine many will move to claude, as I will (return), unless anthropic makes more blunders.
Random link I found looking for the twitter post
https://pasqualepillitteri.it/en/news/21024/openai-merge-cha...
today at 7:17 PM
It's insane how hard OAI is choking, they had much better models than A/ who were fumbling this year to date... then blunder after blunder.
The only things that OAI still have over A/ are better coding agent GUI, no 5hr limit on >$100, much more reasonable cybersec guardrails (that allow most RE work) and... that's it.
today at 9:06 PM
From my point of view, as software engineer who is extensively used Claude Code and ChatGPT from the beginning, OpenAI is nailing it!
today at 6:53 PM
I don't like either. I'd want a recipe that could fit on a single page with type. Recipes you find on Google actively hide the ingredient list and simple procedures to get more of the user's time to monetize but I don't know why ChatGPT needs to be verbose here. Recipes are simple things. Simple search solved find recipes back in the early 2000s and it's prime example of a thing people have been enshitifying since there's not much to really give people beyond the formula. And sure, every entrepreneur screams "they think they want a formula but really they want an experience" but, no I don't.
today at 6:34 PM
Of everything that could be automated, Bartosz Ciechanowski really was the last on my list.
In all seriousness, his lovingly and expertly crafted explainers are still going to age like a handcrafted heirloom clock in a world of plastic-clad quartz movements. But itâs absolutely incredible that we are now in an age where a computer can manufacture a serviceable interactive explainer on whatever niche topic you desire.
today at 6:45 PM
The reason why B.C. became a thing is because the art of drafting died from CAD. The attention to detail, the minutae of walking the reader through a highly sophisticated thing was replaced by short-form video explanations. His work is very much a callback to the days of old. So too, is his turn to be relegated to a relic of his time.
today at 6:48 PM
Also, it's hand coded WebGL. It's smooth, faultless and has no peers.
> So too, is his turn to be relegated to a relic of his time.
No, it'll be tasteful artifact, not a relic. Records, fountain pens and automatic watches did not die. They are used by people who discern things, and no, none of these things have to be expensive (i.e. Neither Seiko 5, nor Lamy Safari are expensive, yet they are as dependable as their 100x expensive brethren).
Human touch still has that finesse and warmth.
today at 8:05 PM
>and has no peers.
On the front page right now - https://news.ycombinator.com/item?id=49980626
today at 8:11 PM
That's impressive, yes, and kudos to them.
OTOH, I still believe the exploded view on https://ciechanow.ski/mechanical-watch/ is something else.
For one, it has real physics on the weight, and second it always shows the correct/current time.
FWIW, his all animations has proper physics to begin with.
The entry you posted is nice, but Ciechanowski is still peerless.
today at 8:24 PM
Nothing in that "7âSpeed Bicycle" is specific to having 7 speeds. It might as well have said "Bicycle", and even then you get meaningless slop like "Made to keep rolling" and "A strong foundation". How can you compare this to Ciechanowski?
You call it serviceable, but what purpose does this service? What do you now know about 7-speed bicycles that you didn't know before?
today at 6:54 PM
He's at openai
today at 7:40 PM
Do you have source for that?
today at 8:27 PM
funny if true. ironically just last week i was asking chatgpt about what happened to him and why there were no new blog posts for almost 2 years and it had no idea
today at 10:40 PM
IMO they're making a mistake trying to shove every possible product into ChatGPT. Just focus on making the core functionality better.
today at 6:29 PM
I've had the most success with GPT explaining things to me by making it take a few sentences at a time back and forth, instead of reading full write-ups of whatever I asked. It also often poisons the conversation if it misunderstood some part of the question, and I can lead it better by continuously questioning its statements. It's also more engaging that way.
I've been learning music lately and it kept re-pasting the same one chord visualization throughout many conversations, almost randomly and often barely related to the question. So I at least hope this won't be as aggressive so I can prompt it away!
today at 9:13 PM
Id love to learn what prompts youâre using to do that. Explanations one at a time. Thats something Iâve thought would be helpful before but didnât know how to achieve it.
today at 7:22 PM
This seems like a way for them to surface ads in ChatGPT. They just came out with visual ad format option for advertisers a couple of days ago: https://openai.com/index/new-chatgpt-ads-format-and-measurem...
today at 10:31 PM
It looks forced. It doesnât look bad, but itâs not something I would pay for, nor would I give it access to my computer?
I also believe that a small talented group of developers just out of college can do just as good a job. That is why there is no moat around AI, it still comes down to original thinking and talent.
today at 10:26 PM
IMO visualization is such a critical part of explanation - we have created an entire suite of tools to visualize - presentations, graphs, formatted docs, websites, svgs, interactive blocks, apps.
To that front - model providers generating visualization elements (i,e taking over the tools for interaction and visualization) is just a natural next step in value capture.
Anthropic launched docs, presentations, sites and am sure they have sheets next up their alley. Apps are artifacts. Tool forming is a natural model adjacency.
today at 10:37 PM
Too much visual noise. You donât need a table comparing America now to the 1970s or whatever.
today at 6:54 PM
I'm not a fan of OpenAI / Sam Altman, but I love their blog posts. The team and whoever decides how to do these presentations, is on point. The only other company that has amazing release pages like this is Apple, I think I remember hearing that they probably hired someone from apple who used to do release blog posts there too.
What's funny about "Intelligent UI" is I said like 2 or more years ago, that these AI companies need to start thinking outside of these basic chat UIs, they do some things here and there, but its really depressing how little they do to innovate in these spaces. Same with the coding harnesses, the UI for all these things could be drastically superior.
today at 10:21 PM
Apple should be extremely worried about the AI companies making big improvements in UX. This work by OpenAI is a step in that direction.
If the main user paradigm becomes a chat interface conjuring whatever UI elements needed to best accomplish the task at hand, the entire paradigm of OS, UI frameworks, apps from an App Store to accomplish specific tasks, all come into question.
I keep thinking that Steve Jobs would have been all over the UX ramifications of LLMs and demanding that Apple lead in that area.
today at 7:33 PM
> probably hired someone from apple
Fairly likely: https://medium.com/the-engineering-brief/openai-hired-400-ap...
today at 7:28 PM
You could say the same thing about Anthropic because I don't really see much differences. Not sure about amazing part but both are high quality and I probably like Anthropic's aesthetic more
today at 7:37 PM
Anthropic's aesthetic looks good by pre-AI design standards, but it's so overused now that I find it off-putting. It's like the 2026 equivalent of the Twitter bootstrap CSS from 15 years ago.
today at 10:22 PM
I donât find this criticism to be very actionable.
What changes could they make to really improve the usability of their products?
today at 8:44 PM
I feel like today's tailwind is the spiritual equivalent to last decade's bootstrap. Except that last decade's bootstrap webpages have a human behind them who usually put thought and effort into making the page useful and correct, whereas today's tailwind pages are often vibe coded slop with all the associated issues.
today at 9:42 PM
Bootstrap stands the test of time! I use it as my go-to unless I have a reason otherwise.
today at 7:39 PM
> these AI companies need to start thinking outside of these basic chat UIs, they do some things here and there, but its really depressing how little they do to innovate in these spaces. Same with the coding harnesses, the UI for all these things could be drastically superior.
I would prefer if the AI companies stuck to just creating better models and making them as cheap and accessible as possible. Let others build the products. I don't want 1-2 companies to own every product in the world.
today at 10:24 PM
That was the objection to Microsoft controlling both the dominant OS and the most popular productivity applications running on that OS. Took the web then mobile devices to really shake that up.
today at 7:52 PM
> I don't want 1-2 companies to own every product in the world.
That seems to be the plan...
today at 7:55 PM
Thatâs the plan for ⊠capitalism
Buckle up boys
today at 10:25 PM
Nothing new.
Look at the history of Microsoft.
today at 7:57 PM
Sounds rather late stage.
today at 8:10 PM
Well, like 30 years ago, a French comedy show had foreseen it, the "World Company" buying everything, no need to understand French to understand the "joke" at the end https://www.dailymotion.com/video/x4du7uh
today at 7:10 PM
Why did they bother hiring someone to write at all? I certainly wouldn't invest $1T into a company with so little confidence in their product. When the chips are down and they need to write something important, they DON'T use AI?
today at 7:35 PM
Obviously everyone at Anthropic uses a ton of AI in everything. The point here is that they clearly use it well, and that they use it as a tool to augment and enhance what they do rather than as a replacement for effort. The blog posts don't read like the output of a prompt because they're probably the output of dozens of prompts, rewritten and edited, by someone who's using AI to actually write something worth reading.
Anyone can having blog posts as good as this by using AI. Just not by zero-shotting it with a 2 line prompt. It still takes a lot of work.
today at 7:23 PM
That's just picking a fight. Whatever we think about how amazing or terrible AI is, it's pretty reasonable to understand it is quite literally the weighted average of humanity's output.
Also reasonable: it's possible to hire a world-class ___ to do that job better than AI.
today at 7:32 PM
It's not a weighted average though. I see this repeated all the time. Between synthetic data, custom produced training data by experts, post-training regimes etc the models are far more diverged from that baseline at this point, at this point it's mostly a right shifted curve.
today at 7:40 PM
I hesitated to write that line because I don't know the math. So thanks for clarifying. Main point is the output is along some distribution and I'd say AI-evangelist themselves see "good enough" across everything and anything as a feature.
today at 7:33 PM
[dead]
today at 7:15 PM
I take the opposite viewpoint. I think pages like these are horrible, and a demonstration that web designers have too many tools at their disposal.
But this is just a marketing page for some tech company, right? If only it stopped there. I have to deal with this nonsense in news "articles" as well on occasion, when some web designer intern is allowed to larp as a journalist for a day.
sigh Just give me text to read.
today at 8:25 PM
Sometimes you want more than text. This is one of the interesting divides these days - and I say this as someone who spends a lot of time in the terminal. Computer interfaces and interaction did not peak with the VT-100. Sometimes, there is a very legitimate need to show tables, graphics, interesting graphs, and to use colour and shade to draw the eye and lead someone through an experience.
This is why I use an IDE instead of a TUI agent... I'm on a computer with a multi-megapixel display, I want to use it. If I could be driving the same process with my Macintosh SE as a serial terminal, what the hell is the point of my recent Macbook?
today at 8:48 PM
today at 8:43 PM
You say tables, graphics, graphs, and colors, but I mean things that move when they don't need to move and interfere with expected UI behaviour, like arbitrarily staying in one position when I am using my mouse wheel to scroll downwards.
I should have been more precise than saying "just give me text", but it was what came in to mind when I wrote it, as I was thinking about what an article is meant to contain as its base element.
today at 9:39 PM
I'm pretty sure you can just ask chatgpt to add to its memory "Just give me text to read."
today at 8:30 PM
"The apocalypse will be televised"
today at 7:11 PM
Mate this is like half a prompt of GPT-6
today at 7:51 PM
I think the point is not how good this blog post is as if it was written by a human, but that this blog post is entirely generated by the GPT-6 model.
So this is a taste of the slop that's going to invade everywhere in a few months.
today at 7:02 PM
I don't know. This is an amateur page with a bad AI video and chaotic presentation. It is ten levels below Apple announcements.
today at 6:42 PM
I've been calling this "disposable UI" or "paper plate UI", eg something meant to be used once. One thing I'll be curious about is overzealousness to produce this, when sometimes what you want is just a simple response. Overall though I'm a big fan of it, if it can be provided fast enough. I'd be curious on how much impact it has on latency of a response.
today at 10:11 PM
I'm excited about generating UIs (and apps) on demand, because it's one step closer to devices just morphing to the interface you need. I want to talk to my phone and have it generate the app or interface I need. I could see Android adapting to this reality well before Apple.
today at 10:18 PM
I find this view fascinating. After decades of observing users getting lost by minor changes to UI or sometimes just different "paint" I am now to believe that we will all enjoy using completely custom UI customised not only to what I would want but what the app will think I need at particular moment and get on with it just fine.
I actually do think generating UI has some future, but I remain sceptical it can realistically go as far as promoters claim. It is especially bizarre to me when this is done by ostensibly UX people.
today at 10:31 PM
I am with you. For example - this post has a comparison of a recipe app from 5.6 and 6 for comparison. They look different. Now if every time I ask for a recipe, it decides to render a new UI because of either the model changing or the recipe goes from stove to oven or something, itâs going to become annoying very quickly. It just seems like he current generation of models have been tuned to create better UIs than previous one and people interested in the UI are getting hit with excitement.
today at 10:27 PM
Just depends where you think the limits of future AIs are.
today at 10:27 PM
More likely OpenAI and Anthropic just cut iOS and Android out of the loop entirely and become the new UI paradigm.
today at 6:57 PM
Funny to see a blog post about UI from one of the most funded tech companies in the world, yet the player has about the worst UI possible:
- Audio volume has 2 levels: on and off
- Play button worked exactly once for me: it played and looped the video. Couldn't be stopped afterwards.
- The video progress bar has no visual indication of where it starts and where it ends..
Is this what Phind died for?
today at 9:48 PM
The progress bar is over the video and it's easy to tell where it starts and ends. Most devices control volume, the video volume separate is just confusing matters for most. No issues with play and pause my side.
Sure as hell beats stupid Instagram style videos where you have no way to skip ahead. I think this is what future video is going be like, buckle up.
today at 7:58 PM
Hate it or not, the only site which has usable video player is YouTube.
Even TV channels like CNN and FoxNews don't have a decent one.
It's like we need ASI to have a proper embedded video player. The ultimate software challenge.
today at 9:45 PM
I disagree, I always run into issues with the embedded YouTube player. No way of sharing. No clear way to share the link. Youtube logo covering half the video. To name a few.
today at 9:48 PM
today at 10:26 PM
Wait, is this new? I feel like we all have been doing this with widgets / artifacts for a long time now, am I missing something? Just easier to use / done without asking?
today at 6:31 PM
GPT-6's design sense is kind of ridiculous imo.
I have explicit instructions to tone it down. Less taglines, eyebrow text, subheadings, decorative spacing, pills, cards.
Hopefully this doesn't bleed into the chat...
today at 7:58 PM
yes the eyebrow is the tell. I never knew anything about eyebrows in design as I am not a designer. Now they are everywhere. Everything has an eyebrow. It is ridiculous but at least it is an AI smoking gun.
today at 8:55 PM
I loved them and I placed them a lot in my designs â if you also include my love for em dashes you can well understand that I feel my own character has become a clankerâŠ
today at 6:39 PM
More UI elements == more tokens == more money for OpenAI
It's a pretty clever way to sell more tokens, I have to admit
today at 6:41 PM
Except that chat is a fixed monthly fee. More tokens = more cost for them.
today at 7:35 PM
Not for enterprise users, they pay per token, even on chatgpt.com.
today at 7:28 PM
Not on enterprise accounts.
today at 10:34 PM
Then you should have specified "... On Enterprise accounts" in your comment. But let's be honest, you simply were wrong.
today at 9:59 PM
Great product idea. I have not paid OpenAI for a subscription for a couple of years, Iâm tempted now. I do about 75% of my work using local models and most of the rest using the deepseek 4.1 flash API. That said $20 to experiment with this new product for a month sounds like a pretty fun idea.
today at 8:59 PM
Thank god. My primary use of ChatGPT is meal planning and while itâs great for recording, recipe lookup, etc I always just wanted it to be able to make checkable grocery lists.
I eventually just had it make a skill for work mode that would build a mini checklist app, but it was slow and felt janky needing to remember to switch to work mode.
This is already such a huge improvement.
today at 7:17 PM
I'm curious to see how useful this is. I feel like in practice the existing versions of this feel like they get in the way. While I'm sure I've had some situations where a visual would be helpful, there are two situations where I do not want it:
1) I want a quick answer, and I don't care for the boilerplate UI. For example, if I ask how to make pancakes, I make them all the time and just want to a quick reminder on the ratios, but it might trigger a full UI that I need to sort through to find information.
2) If I ask a to me unrelated to UI question and it triggers a big UI build that is completely off topic for my question (meaning I'm desperately pressing the stop button and prepping rewriting my query)
Or the other one I see, for example if I look up a unix command like:
"ls all hidden files in the /xxx directory"
And I get back:
"Sorry, I am am unable to find /xxx in my current environment"
today at 6:16 PM
I find it extremely strange that they're adding GPT-6 Sol to Chat over GPT-6.1 Sol which is significantly more capable.
today at 10:38 PM
gpt-6.1-sol was released 7 days after gpt-6-sol only because they were bleeding to fable-5.1/opus-5.5 from Anthropic. It was a desperation move and certainly ate into their margins massively. Since their margins on gpt-6.1-sol are much slimmer than on gpt-6-sol they have to use it somehow to preserve compute and make profit. Frankly I think OpenAI is a mess and is playing catch-up with Anthropic⊠I donât know how they will recover.
today at 6:23 PM
I really don't like having to open Work sessions for one-off questions, only because they're limiting what models they put in the chat. The older models are just too dumb for some things.
today at 6:57 PM
The whole Work/Chat/Codex split is maddening in the way it's implemented. It's a pain to switch back to the right project, it's a pain to switch forward, it's a pain to try to remember which chat was in what.
today at 6:46 PM
6.1 is Astra minor. Way more capable but also way heavier+slower. Its really 2 different models, they just shipped it as sol to recover from the gpt6 disaster lunch, where they tried to pass terra 6(or a cheaper model) as sol but it was worse than expected
today at 6:55 PM
Anecdotally this is what I experienced, do you have sources for this?
today at 7:21 PM
It's purely circumstantial, but 6 Sol supports no reasoning (same as 6 Luna and 5.6 Sol/Luna), while both Astra and 6.1 Sol do not support "reasoning = none".
today at 6:21 PM
Iâve hit model not available limits for GPT 6.1 Sol multiple times over the last week. Itâs never for very long, and I can switch to Astra, but it seems like OpenAI is struggling with capacity. This has happened with my personal $200 Pro connection and my Codex enterprise connection.
today at 8:55 PM
I had a codex session stop in the middle because the auto-reviewer timed out with something like "auto-review not available at the moment". The capacity issues are very clearly observable since gpt6-astra launched.
today at 6:44 PM
Given the extreme short time between 6 Sol and 6.1 Sol, I suspect they donât actually have much in common and 6.1 is a heavier model rebranded as Sol in a panic response to poor agentic capabilities of 6.
today at 7:21 PM
I suspect that 6 sol was a better version of 5.6 terra (and note that in the 6 sol and luna release, they took out terra, and price 6 sol at 5.6 terra pricing), then the backlash from lesser capabilities made them roll out 6.1 sol as the actual 5.6 sol - size modee.
today at 8:20 PM
Honestly it feels like AI companies are still searching for the killer idea that will attract the general public (beyond "be my AI bf" or "better google").
Getting a UI that explains the parts of a bike is ok I guess, but isn't it simpler to get an actual breakdown? Google "parts of bicycle breakout" gets tons of useful images instantly.
Getting a specialized app to split a bill? It was already trivial to put in a calculator if we cared to go item by item on the bill. Having to provide names and tag every item as I go is just more work. Usually real people just go $total divide by 5, I had more, let me chip in an extra $10.
Same with booking travel and wedding plans, these aren't things people would even delegate to a trusted friend usually, much less a one-off request to an AI bot or custom UI.
Most of the useful tasks it can do right now are research, technical question/answer, coding. In terms of "build a flexible ui that solves real-world problem", if existing mobile app isn't useful in this arena, then its unlikely a completely custom UI will do the job.
That said, I've done a few small ones like a quick one to practice alphabet of a foreign language, or prototyping a web game, and the like. But ultimately its nothing that is worth trillions of dollars.
today at 10:30 PM
UmâŠthe AI companies attracted the general public a few years ago.
today at 8:43 PM
today at 9:08 PM
Burning question: why is this worth paying for as a layman? I can sometimes find a use for all these LLMs for software development, but I can never find a good use for any of this outside of "slightly better search engine".
today at 9:24 PM
As a complete layman---not worth it. We live just fine. Medical researchers slave night and day to keep you healthy.
But if you have any ambition at all---if you want to map the local school board, or get a quick understanding of your finances, or set up a robot in your backyard---then, quickly proves useful.
today at 9:47 PM
I genuinely think the frontier labs are just throwing shit at the wall to see what sticks. All of the improvements since Fable 5 seem marginal. Most of the quoted AGI "risk" is coming from super large agentic models that are juggernauts that can brute force their way through any problem faster than a human can. Despite that they're no closer to AGI than they were a year ago; the technology's limitations are apparent the more you use it. Doesn't mean it's still not extremely dangerous and/or capable, it's just not the "replace a human in 99% of cases" thing. There are obvious tradeoffs besides just the models being faster and capable of information retrieval and generation at a rate exponentially faster than a human's.
My long-term hope (and call me another Zitron if you wish) is that the labs' hype dissipates if there's no serious improvements beyond stringing together agents to ram through brick wall Millennium Prize style problems and we all just use on-device AI for 99% of cases where it's useful, researchers, governments, militaries, and universities can pay for the more complex models, and image/video gen dies a slow death (if Congress will actually legislate and/or SCOTUS decides that training those specific models does not qualify as fair use unlike training LLMs) and cost for compute rises over time as public interest in the tech sours.
LLMs are such marvelous and incredible technology squandered by genuinely deranged Silicon Valley cultists who are trying to use them as a Trojan horse to force their antisocial visions of the future upon an unsuspecting populace. We could all have just invested in on-device compute and nobody would have lost their jobs in pointless layoffs, productivity would have increased, and maybe we would have forced a conversation about when and when not to use AI and it wouldn't be so ever-present in use cases where it actually does harm to the consumer and society. But sama, PT, and dario just wouldn't have it that way, would they? Because if that were the case, there's no prospective hope where they can assuage their deep-seated insecurities over being antisocial and off-putting by reassuring themselves that there's an imminent realignment of society where they will hold all the cards when the dust settles.
today at 9:07 PM
I wrote something like this, though not quite as slick, by having the model use json schema to describe its results instead of text, and then use a json-schema to ui interface.
today at 8:21 PM
The medium is the message here. As others have said, chat output largely sucks for anything but the most basic responses and we've done little to improve upon that foundational UI in the last couple years. We've _added_ a lot for specific domains like coding and document editing, but the primary content normal users get back from the chats is verbose and uninteresting. This is a step in the right direction.
Gettin closer to fully dynamic interfaces for a lot of software. Hell, give me a mode in Google Docs that takes every pixel of chrome away then vibe the rest as I need it. Persist across new documents going forward.
today at 6:12 PM
Is this OpenAI catching up with Anthropic artifacts? At the same time they say "Weâve trained GPTâ6 to compose responses using text, visuals and interactive elements...", rather than a harness.
today at 6:32 PM
The irony really is that LLMs are partly responsible for the walls and walls of text as seen in the video in the first place.
And now we are asking LLMs to solve it.
today at 7:20 PM
We are becoming more and more as a tool for sth to be done rather than the brain behind it. At least I start to feel this way. The joy of discovery, exploration and experiencing at first hand... It is slowly diminishing for the perfection of the quick outcome.
today at 6:30 PM
I wish things like bikes came with mostly-written manuals. Maybe a few diagrams. If anything, things are too pictorial these days.
today at 7:49 PM
I'm confused. Are they trying steamroll every developers B2C product or are they wanting these same developers to continue building MCP-App plugins on the platform?
today at 7:52 PM
It's much simpler: they want others to experiment (and pay them) and see what works, once they know what works they will drive out of market other players.
today at 8:37 PM
what would you wager the % of codex token usage is folks generating interfaces/building apps?
seems a bit like a snake eating its own tail
today at 7:35 PM
"More than" 20% of the connected world uses ChatGPT each week eh? I guess if you add Anthropic who must also claim 20% and Google, Meta.. that does not sound realistic at all.
today at 7:44 PM
- Anthropicâs total user base is minuscule compared to the rest, and they have never claimed otherwise.
- People donât have to stick to a single provider. They can all claim the same 20%.
today at 8:38 PM
This seems like a step towards their plan for building a platform in competition of Google/Apple. Once the generative UIs are polished and more useful than individual apps, their hardware can now ship a device which circumvents the app store moat.
today at 8:11 PM
All those recipe blogs and sites that have figured out the most inefficient way to deliver information are finally cooked.
today at 6:15 PM
AFAIK, they already had a simpler form of this. It was kind of an obvious next step. Now connect it up to tools so that we can again comfortably do the things that are more precise by hand! The âpick a colorâ or âselect the width on a sliderâ use case is coming closer.
today at 6:42 PM
1.2B weekly active users? WAU!
today at 6:47 PM
They seem to be pitching something that pretty much all the decent models can already do..?
Obviously if your model is stuck inside a CLI terminal, then not so much. But in a GUI harness (shameless plug for my own one: https://juggler.studio, but I assume others can do this too), you just ask them to answer in HTML and they'll happily draw pretty pictures inline in the conversation. I've been doing this for ages with claude, GPT, Deepseek and others.
today at 8:57 PM
I am pretty convinced that businesses will be making custom internal software as a norm in the next few years, and this seems to be a slight push in that direction
today at 9:45 PM
Alright doing it here, all the functionality we need and no reoccurring yearly costs of tens of thousands of dollars. Took me a couple weeks to wire it all up with codex.
today at 7:45 PM
I think this means that the API "chat-latest" model is now serving a custom GPT-6 variant: https://developers.openai.com/api/docs/models/chat-latest
today at 7:52 PM
Interesting, in Chat mode, if you expand the model selector it says "GPT-6", but in Work mode it says "GPT-6 Astra".
today at 6:55 PM
This seems like an expansion on something I was getting my agent to do which is use 'cards' to communicate using things like tables, rather than the ascii table stuff Claude code likes to do.
Unsure if this has been done already but the best thing I did was get it to create a list of next steps as quick buttons that it keeps updated in a docked card. It works well but occasionally gets stuck on some un related rabbit hole.
today at 6:25 PM
The intro video is just lame. Those things never happen in real life.
today at 6:34 PM
They're making the pointâapparently too subtlyâthat text alone is a highly limiting "UI" for many tasks. So it's great that GPT-6 can now communicate in a a richer interactive medium.
today at 7:33 PM
Then show a real task where text is limiting. Don't make up stupid examples.
today at 10:26 PM
Theyâre making fun of the frustration of trying to use ChatGPT purely through text for these tasks.
Like I said, itâs too subtle a deliberate self-own.
today at 6:29 PM
So you never have to assemble a bike or paint a room. How those are not real world examples? I find their video pretty cool. For people with ADHD or people learning by watching or children this is spot on.
today at 6:59 PM
They made fake, terrible artifacts (colorless paint chips, text-only city guides and instruction manuals) to show how terrible that is, and then "fixed" them. The problem is, in reality their examples don't exist. Paint chips have color samples. City guides have maps, pictures, and color. Instruction manuals almost always have illustrations. If what you made is better than what exists, you should compare it to what exists and not some alternate reality.
today at 9:06 PM
For once Gemini did it first. I never found those mini GUIs useful though. More like a waste of time and energy.
today at 6:23 PM
In the video, they showed examples of ChatGPT making interactive tutorials on how to fold origami, how to arrange colours/interior and how to assemble a bike.
Supposedly, people were struggling to follow written manuals and they needed an interactive explanations.
I'm not a mathematician and I would certainly love having a tool that would do ELI5 on some complex stuff, but I'm really worrying about using this too often and outsourcing my ability to do stuff to some mega corp.
today at 6:27 PM
> The compiler allows the interface to appear progressively as the model generates it, without waiting for the entire response to be complete.
why not just stream html?
today at 6:32 PM
The element canât render until itâs closed, presumably.
today at 6:33 PM
Html partial streaming is a real thing, presumably
today at 8:19 PM
maybe I'm just wrong here? If the tokens <b>come through this text could still be optimistically turned bold until </b> happens.
For elements that manage layout and drawing things like these explainers, though, it's hard for me to imagine how that would work. My immediate thought would be to an abstraction over it, which is what it sounds like they did.
today at 7:18 PM
Sounds like Google's Generative UI (Nov 2025): https://research.google/blog/generative-ui-a-rich-custom-vis...
today at 8:32 PM
I like how, whenever a new model releases, they not only make other companies' models seem worse, but even their own "oLdeR" models.
today at 10:10 PM
gemini app kind of had that for a while I feel?
today at 7:04 PM
"But when will we recoup our investments?" Investor to Gavin Belson after he presents a whole zoo of animals for analogies.
This is now trying to steal YouTube repair and cooking videos. Here is news for you: People prefer YouTube repair and cooking videos.
today at 8:47 PM
>People prefer YouTube repair and cooking videos
Not sure about the cooking one but repair videos are horrible. People with thick accent, the worst camera, shittiest lightning and bad angles where you can't see what you have to do.
today at 7:08 PM
I donno, I like iterating on a recipie with chat rather than watching a video.
today at 7:53 PM
cooking videos are entertainment
today at 8:39 PM
he he boi that's right cooking videos are entertainment if you actually cook you read something called a recipe (it's the algorithm that returns the food you want)
today at 7:31 PM
If the diagrams and descriptions are actually well-designed and written, I'd far prefer those. Especially one where I can ask clarifying questions. I'd much prefer a Haynes with some of the images lightly animated than having a 10 minute long video with a long intro for monetization reasons about how to do some simple thing where you can barely see what the mechanic is actually doing in the cramped, poorly lit, 360p video is trying to look at.
But getting actual service manuals cheap/free these days can be tricky, and their quality can leave a good bit to be desired.
today at 6:58 PM
Gemini had this UI thing before .
today at 6:29 PM
This seems like a natural progression of models becoming better at frontend coding in general.
"Here is a library of [svelte/react/whatever] components, use them to construct a helpful visual to demonstrate your point."
The deconstructed bike at the beginning was in a class of its own, however.
today at 6:30 PM
How would a user know whether a generated visualization to illustrate some process is accurate or not?
today at 6:32 PM
It wont they're just trying making the chat look like more a session iron man would have with jarvis
today at 7:06 PM
There have been recent developments that allow people to interpret Swift on device, live. With this new launch, will the AI be able to write native code and compile it on device? That would be super cool.
today at 6:23 PM
It makes sense theyâre doing this - Iâve noticed lately when asking a more complex question involving a lot of nonlinear data using Codex work mode, Astra and Sol will write a fully html document to better display the info with a Cliffâs Notes version in chat.
today at 8:32 PM
I noticed this change too! It is really game changer - to me ChatGPT output has no rivals here.
today at 6:55 PM
I feel like I'm going mad looking at this page. The bike disassembly looks fake to me, like an alien would show presumed bike parts, but I can't tell why and don't know much about bikes. Maybe it's because there are no screws and nuts?
What is it supposed to demonstrate? That the model knows some kind of folk mereology?
today at 7:24 PM
Apart from the math stuff, maybe, this is the appearance of useful information rather than the presentation of actionable practical information.
I suspect it will struggle with anything that isn't simple enough for demo-mode or lifestyle fluff.
today at 7:30 PM
Itâs demonstrating a lack of understanding of the average person by those at the frontier.
Remember when Steve said âthe computer for the rest of us?â Weâre seeing that here.
today at 8:28 PM
A Chat AI creating on the fly UX widgets is like Amazonâs brick and mortar stores
today at 6:20 PM
I'm also building something similar for an internal project, based on the vercel's json-render design. It hasn't been too challenging, especially since the release of faster models like 5.6 Luna.
today at 6:27 PM
So I guess it's going to nail the bicycle the Pelican rides on, right?
today at 6:52 PM
Similar to https://www.monogram.ai/.
I guess the above is mostly mobile focused.
today at 6:30 PM
Love it. I've been using $visualize a lot in the codex desktop app, and having even richer experiences will be sweet.
today at 6:52 PM
IIRC Google had an experimental project with this kind of stuff 2 years ago, but Google being Google...
today at 7:38 PM
Let's remember.. 'Climbers rescued from Mount Shasta after relying on ChatGPT for trip planning'
https://news.ycombinator.com/item?id=49542814
today at 8:24 PM
for the sunday lamb roast instructions, how much is 1500g of potatoes? What kind of potatoes?
> Also pick up garlic, rosemary, thyme, lemons, honey, almonds, olive oil, gravy ingredients, mint sauce, crumble topping and vanilla ice cream.
what??? how the hell are you supposed to remember all that. Even the before example tells you what kind of potatoes to get. tbh it would be cool if you could just click a button and get an order pre-filled out on a grocery delivery/pickup service. I really wonder who reviewed this post and if they cook, because imagining yourself in that scenario and reading those instructions falls apart very fast
today at 9:21 PM
'gravy ingredients'
today at 6:27 PM
Is it me or GPT-6 Instant food answer gives the vibe to check the hell out immediately? It's like the it will start to explain a sunny Sunday afternoon from 20 years ago for 5 full pages.
today at 8:03 PM
Instead of apps or websites merely APIs and intelligent UI in front. With the ability to render data gathered across multiple domains
today at 6:46 PM
Is there any relationship between this and AG-UI or A2UI?
today at 7:27 PM
Why show an intelligent UI within a chat interface. Thatâs pathetically lame. Better to go away from chat interfaces to rich visual interfaces. Doesnât sound intelligent at all to me
today at 7:19 PM
I gotta admit, I love how their design system and how they market things. It looks so good. However the true state of their toolings is closer to a burning dumpster fire. If you look at Codex Github repo for example, the amount of existing issues and continuously new issues appearing while tagged releases is being spewed out. Its truly a nightmare. Their vscode extension has been broken for almost 1 week now, if not more. Its instead a community effort to repair it.
today at 8:13 PM
It's interesting how the bicycle is a fake example. You see Jony Ive level design, with a very complex 3d object seamlessly working, and then you see the basic HTML slop in other examples and you begin to wonder if that first one was fabeication.
today at 8:32 PM
sounds like JEV with extra steps
today at 6:27 PM
Written manuals/guides come with pictures. The ad is dishonest.
today at 7:06 PM
The point, I think, is to make fun of Anthropic models, which answer only in text when you ask them how to do something. Theyâre the competition, not paper pamphlets.
today at 7:11 PM
> which answer only in text
This is false.
today at 7:36 PM
Is it? Sorry, skill issue I guess. I regularly use the Claude app on phone and laptop, and have never seen it produce non-text chat output.
today at 8:43 PM
I think it's been a thing since around March https://claude.com/resources/articles/claude-builds-visuals
Edit: be sure to check Claude settings -> Capabilities -> Visuals.
today at 9:20 PM
Thank you!
today at 6:32 PM
Plot twist: The guides were written by ChatGPT. But now you can use ChatGPT to solve problems by ChatGPT.
today at 7:47 PM
This seems like a nightmare for standardization and accessibility, no?
today at 7:53 PM
Has AI not been always?
today at 6:51 PM
"If you call something intelligent you know what it isn't."
today at 7:16 PM
> Weâve trained GPTâ6 to compose responses using text, visuals and interactive elements, choosing how they fit together based on your question.
Ok so this is how you build a moat. Models are not a commodity and visualization isnt a purely harness problem.
> We expanded our training methods to help the model make thoughtful decisions about content, layout, visuals, and interaction. This included evaluating the interfaces it creates for clarity, usefulness, and completeness. GPTâ6 learned to use the component library and make good design decisions, including how to organize information clearly, when to use interactivity, and when a simple text response is enough.
Curious to know how they trained it produce appropriate visuals. And why that cant be done with propmt engineering.
today at 7:29 PM
Nope
today at 7:07 PM
today at 6:25 PM
today at 6:24 PM
I made a prediction last year that there'd be a new UI protocol (like HTML) but for agents. I believe that custom or personalised UI's are going to be the new browser interface. More radically, I think browsers can be completely replaced. My news feed can be personalised to me, based on what AI thinks might be important from all sources like Reddit, X, HN.
today at 7:17 PM
Which brings up a different problem: personalized reading UI, personalized writing UI, how are those services going to pay their bills? Maybe they can sell access to their APIs but who's really going to buy that? Those services will die and will be replaced by some other service that will be born to serve those people.
Or, more probably IMHO, most people will keep using the standard UIs because they are ready and they need no work to build.
today at 7:15 PM
Sounds horrible, but plausible.
Now they extracted the value from interoperable open systems, they would love to replace it with closed proprietary systems.
today at 7:33 PM
So they are going to replace all the banking, utility, social media, news, gaming, streaming, government, work apps and other important websites people still rely on? Those companies and organizations are just going to provide an API to the chatbots?
How do sites like Wikipedia get updated? What if I need to upload important documents to different portals? It all has to go through the AI companies? They have access to medical records too? My taxes, social security, etc?
Is this how Microsoft imagines Office 365 and Sharepoint will be merged into? Disney, NY Times, Youtube, Netflix will all be fine with chatbots handling their content?
today at 7:32 PM
> I made a prediction last year
yes you and everyone else.
today at 7:26 PM
Cant wait for the next round of outrage and radicization machines, this time by browser replacement and unavoidable.
today at 7:32 PM
Why are you posts generally down voted?
today at 6:36 PM
I read the comments here and I am really surprised by many people. One picture is worth a thousand words so with better UI you can make much better user experiences. This unlocks personal assistants to be better adopted by elderly or disabled people. I see so many benefits of it and having models which can do this (if they can do it constantly with good quality) is amazing. Much better products - I am really tired of dumb chatbot - if I can do something with one button or view the whole information in one diagram/image, this is amazing.
today at 6:20 PM
This looks bad, I want the output to be as much as text based so I can easily export it and further processing it, plus, this might make the resource-eating app even worse.
today at 6:58 PM
Visuals are better for consumers/general public, when the use case is "help me cook X meal", or "where can I stop by to buy gas on the way to San Jose".
today at 7:39 PM
Rading most comments, people don't seem to be very impressed by it, I ain't either tbh - its not something mind-blowing (I have been doing this with Astra myself just by writing better prompts to make explainer interactive interfaces), but I do think its a good first step towards a better way of consuming information/answers than just reading text-vomits. I wonder if there is a way to train an LLM native to this kind of thing?
today at 9:00 PM
where are you reading the comments? Wasn't it just released? I can't imagine anyone I'd consider "general public users" to know about it yet
today at 7:32 PM
then just ask for text-only?
today at 8:28 PM
today at 10:01 PM
[dead]
today at 7:28 PM
"why do I need an Nvidia B300 to render the same webpage that ran fine on my Pentium with Windows 98?"