GPT-6 Intelligent UI turns ChatGPT answers into tools you must audit
OpenAI has rolled out GPT-6 to every ChatGPT tier with Intelligent UI, which answers with charts, forms, buttons and small working calculators instead of plain text. Independent tests show the models underneath are about as smart as before but half the price, so the new calculators may look more trustworthy than they are.
Neo NeumannAI Practice LeadOctober 9, 2026
Listen to the podcast
9 min
Chapters
Key takeaways
- Ask ChatGPT to print the formula and every input of any generated calculator as plain text, then rebuild it in your own spreadsheet.
- Treat speed claims like 44% sooner as time to first token, which says nothing about when the answer finishes or whether it is right.
- Discount benchmarks published by the vendor selling the model and look for independent tests such as Artificial Analysis.
- If you manage a ChatGPT workspace, check the Enterprise admin setting for Intelligent UI before colleagues find the feature on their own.
- Test cheaper models like GPT-6 Luna or Claude Haiku 5.5 on your own data before moving routine summarising or classifying work to them, and watch the 100,000 token price threshold.
Read the full transcript
Host:This is MBA Training. On the table: GPT-6 Intelligent UI turns ChatGPT answers into tools you must audit. This week a huge number of people opened ChatGPT and some of them got a savings calculator they never asked for. What changed on Wednesday?
Expert:OpenAI started rolling out GPT-6 inside the chat product, along with a feature it calls Intelligent UI. The model no longer has to answer in paragraphs. It can build an answer out of charts, forms, buttons you tap and small working tools. TechCrunch reported that paid plans (Plus, Pro, Business and Enterprise) got it on Wednesday, October 7, and the free and Go tiers got it on Thursday.
Host:And the model you get depends on whether you pay.
Expert:Yes. According to The Decoder, paying customers get GPT-6 Sol and free users get GPT-6 Luna, the smaller and cheaper sibling. VKTR puts the reach at more than 1.2 billion weekly users. That number comes from OpenAI.
Host:I use ChatGPT for budgets, not Mahjong tutorials. Why should I care about buttons?
Expert:Because the output is now a different kind of object. A paragraph that says your break-even is month fourteen invites you to check the reasoning. An interactive calculator with sliders looks like software. VKTR said it directly: calculators and dashboards generated on the fly can look authoritative, and that makes them harder to audit.
Host:Is that a real risk, or just a columnist worrying?
Expert:Let me separate fact from opinion. The fact is that the model writes the formula behind the calculator and picks the numbers that go into the chart. My opinion is that people check a polished tool less carefully than a block of text. I have no study on this product to back that up yet, so treat it as a hypothesis.
Host:So is the model writing code and running it in my browser?
Expert:No, and that design choice matters. VKTR reports that OpenAI built a library of native, streamable components. Those are prebuilt interface pieces that can appear bit by bit while the answer is still being generated. There's also a compiler, the software that turns the model's output into those pieces on screen. The model builds from a fixed kit and doesn't write free-form code. You get consistent rendering and fewer ways for things to break, but it also limits what the model can build.
Host:OpenAI is also claiming speed. What's the number?
Expert:GPT-6 Instant starts answering web search questions 44% sooner than GPT-5.6 Instant, according to OpenAI via VKTR. Pay attention to the word "starts." That measures the time until the first words show up, sometimes called time to first token. A token is a chunk of text, roughly a short word. The Decoder explains how they did it: the model can start responding while it's still thinking. So the answer begins sooner. That says nothing about when it finishes, or whether it's right.
Host:OpenAI says it's more accurate too.
Expert:The Decoder reports that in OpenAI's internal tests, GPT-6 scored higher than GPT-5.6 on difficult web searches using this method. Those are internal tests run by the company selling the product. I'd want outside numbers.
Host:Do we have any?
Expert:For the models themselves, yes. Sol and Luna reached the API in September. The API is the programming interface developers pay for by the token. An independent benchmarking firm, Artificial Analysis, tested both models on launch day. Codersera reports its summary: both models "push the cost efficiency frontier by halving cost relative to GPT-5.6 Sol and Luna", while their Intelligence Index, a combined score across many tests, stays level with GPT-5.6. Some evaluations improved and some got worse.
Host:So the big launch for a billion people runs on a model that's half the price and about as smart as the last one.
Expert:On that index, yes, roughly. Sol at max effort scored 47.63 against 46.97 for GPT-5.6 Sol, and Luna scored 38.12 against 37.32. Those gaps are small. Luna actually slipped on Artificial Analysis's Coding Agent Index, from 43 to 41.
Host:Is there any good news?
Expert:One number that matters for professionals. The hallucination rate fell from 92% to 60%. A hallucination here means the model confidently makes up an answer when it should say it doesn't know. But Codersera points out that Sol attempted 83% of questions, against 99% for GPT-5.6 Sol, so accuracy on the questions it answers went down slightly.
Host:In plain English?
Expert:The model says "I don't know" more often. If you're a lawyer or an analyst, that's a good trade, because you get fewer confident fabrications. You also get more non-answers, and the answers it does give aren't more accurate.
Host:Now put that inside a nice-looking chart.
Expert:That's where the tension is. The model is better at holding back when it's unsure, but it sits inside an interface that makes whatever it does say look finished. Those two things work against each other.
Host:Can I just turn the visuals off?
Expert:There's a setting for individual users, but I haven't seen documentation showing it switches the feature off completely, so I won't claim it does. In companies, IT controls it. VKTR reports that Enterprise access depends on workplace admin settings. It also reports that the models behind ChatGPT Work and Codex, OpenAI's agent and coding products, didn't change with this release. If your team built workflows there, nothing changed for them this week.
Host:Did any competitor respond?
Expert:On the same day. Anthropic released Claude Haiku 5.5 on October 7. VentureBeat reports it costs ten cents per million input tokens and fifty cents per million output tokens for requests under 100,000 tokens. That exactly matches GPT-6 Luna. Anthropic estimates that workloads cost approximately 75% less to run than on Haiku 4.5.
Host:What's the catch?
Expert:A price threshold. Above 100,000 tokens, VentureBeat reports the rate goes up to fifty cents input and two dollars fifty output. If you put a whole long contract into one call, you pay the higher rate. Also, the benchmark comparisons Anthropic published against Luna come from Anthropic, so be just as skeptical of them.
Host:So if you're paying the API bill, the cheap tier got cheaper from two vendors in two weeks.
Expert:Yes, and the facts support a practical point. Say you send routine work like summarising, classifying, or sub-agent tasks, meaning chores a bigger model hands off to a smaller one. You could cut your cost per task sharply just by changing which model you call. Test the quality on your own data first.
Host:Let me push back on the whole framing. This is a nicer way to show answers. Charts have been around forever. Why is this a story?
Expert:That's the strongest counter-argument, and part of it holds. You could already ask for charts and tables. The components are limited by design. And according to VKTR, OpenAI says a simple question can still get a plain text answer. If the model chooses well, most work questions will look the same as they did last week. Where I disagree is scale. When defaults change for a billion weekly users, habits change. Some of those auto-generated calculators will end up in slide decks and budget emails, and the person reading them will have to do the checking.
Host:Is OpenAI leaning on the interface because the model gains are thin?
Expert:That's speculation, so I'll label it as such. Here's what the facts say: the independent intelligence scores are flat, the prices halved, and the change consumers can see is the interface. You can decide for yourself which part of the launch is doing the work.
Host:Give me one thing to do this week.
Expert:Pick one calculation your team regularly asks ChatGPT for, like a pricing scenario, a headcount plan or a payback period. Let it build the interactive version. Then ask it to print the formula and every input as plain text, and rebuild it in your own spreadsheet. If the numbers match, fine. If they don't, you found out this week and not in a board meeting. And if you manage a ChatGPT workspace, check the Enterprise admin setting before your colleagues find the feature on their own.
Host:Drawn from TechCrunch, The Decoder and VKTR, links in the show notes. We'll leave it there. The AI self-assessment at mba-training.com will tell you where you really stand.
OpenAI calls the new feature Intelligent UI. It rolled out on Wednesday with a new GPT-6 model and puts a lot of visuals into users' conversations. According to TechCrunch, it went out globally to Pro, Plus, Business and Enterprise users first, and reached the free and cheaper Go tiers on Thursday. The Decoder reports that paying customers get GPT-6 Sol and free users get GPT-6 Luna.
The design matters. According to VKTR, OpenAI did not let the model write arbitrary code. It built a library of native, streamable components and a compiler that processes the interface as the model generates it. The audience is huge: the more than 1.2 billion people the company says use ChatGPT each week.
For professionals, the change is in what an answer looks like. VKTR warns that interactive calculators and dashboards generated on the fly can look authoritative, which raises new questions about verification. In companies, the switch belongs to IT: Enterprise availability depends on workplace admin settings. Agent and coding work is unaffected, since the update applies only to the Chat experience.
The model gains are disputed. OpenAI says GPT-6 Instant starts answering web search questions 44% sooner than GPT-5.6 Instant. That figure measures when the answer starts, because the model can now respond while still thinking. Codersera summarises the independent firm Artificial Analysis: both models "push the cost efficiency frontier by halving cost relative to GPT-5.6 Sol and Luna", while their index scores "remain level with GPT-5.6". Sol makes things up less often, but part of that comes from Sol declining more questions, and accuracy on the questions it answers went down slightly.
Rivals moved on price the same day. VentureBeat reports that Anthropic's Claude Haiku 5.5 starts at $0.10 per million input tokenstokensA token is the basic unit of text that language models process, often a word fragment, whole word, or punctuation mark rather than a single character.View full definition → and $0.50 per million output tokens, matching OpenAI's GPT-6 Luna. The cheap rate has a limit: above 100,000 tokens, Haiku 5.5 costs $0.50 per million input tokens and $2.50 per million output tokens.
What to watch: whether companies switch Intelligent UI on, independent accuracy tests of the generated tools, and how far OpenAI and Anthropic keep pushing the price of the cheap tier down.
Sources
- ChatGPT is getting a lot more visual, with the launch of a new interface
- ChatGPT with GPT-6 ditches mostly text output for interactive UI with charts, buttons, and mini apps
- OpenAI Brings GPT-6 to All ChatGPT Users With Interactive Intelligent UI
- GPT-6 Sol & Luna: Pricing, Benchmarks, vs Astra
- Anthropic launches Claude Haiku 5.5 with 90% API price reduction, matching GPT-6 Luna
Finished reading?
Validate your read to earn XP and feed your radar.