Info: Machine translation This post was machine-translated from the Chinese original. Wording may be rough in places — the Chinese version is authoritative.
Summary 1.05 million contexts, computer operation, $10 input per million; Once the capability crosses a line, the real upgrade is your permission design. HyphenTech · 2026-09-05
On September 3rd, OpenAI officially released GPT-6 Astra. Today, when I opened the official release page, the first thing that stopped me wasn’t the 99.9% ARC-AGI-3 score, nor the near-perfect math score, but a “hand” running around the computer, typing, editing files, and filling out forms. **The real change this time isn’t the chat window that talks better, but the model taking over a full period of work. **
This article only distributes official data and does not pretend to be native to the machine. OpenAI’s achievements come from the research environment or API, and the official reminds us that production ChatGPT may differ due to system instructions, tools, and task settings. Let’s first nail this boundary, then see exactly which lines GPT-6 has crossed.

🤖 The first change: the model moved from the advisor’s desk to the dashboard
In the past, large models were more like consultants sitting beside you: you asked, they gave a suggestion, and in the end, you had to open the webpage, copy data, and edit tables yourself. Astra’s official demo extended the chain: it could fill out online forms, update CRM, organize calendars, conduct web research, draft summaries in the document editor, install testing software, analyze scientific data, draw diagrams, and check the website frontend.
This is more like the next generation of early automation tools like RPA. RPA relies on fixed coordinates and rules; moving a button by ten pixels can cause a breakdown; General-purpose computer operation models look at the interface, read context, and then decide on the next step, making them more adaptable. The cost has changed accordingly: if you answer a wrong word in chat, you can ignore it; If an agent clicks the wrong confirm button, it might directly change the external world.
OpenAI provides 72.6% for OSWorld 2.0 and 65.7% for GPT-5.6 Sol. In the same latency simulation, Astra takes about 40 minutes per task, Sol about 75 minutes, reducing time by about 47%.
On Terminal-Bench 4.0, Astra is at 57.9%, and Sol is at 37.3%. These numbers show that its strengths mainly lie in multi-step execution and real software environments, and don’t mean that any simple Q&A is worth switching to the most expensive model.

The most valuable aspect of GPT-6 is not writing more beautiful answers, but connecting “suggestion—operation—check—delivery” into a single thread.
🧠 The 1.05 million yuan range is quite impressive, but the bill will also grow along with it
The API model is gpt-6-astra. The official model page states: context window 1.05 million tokens, maximum output 128,000 tokens, knowledge as of April 30, 2026; Supports five inference strengths: low, medium, high, ultra-high, and maximum, as well as web search, file search, code interpreter, hosted shell, apply patch, computer operations, MCP, and tool search.
Long contexts are the easiest way to misunderstand it as “just stuff everything in.” Back when computers upgraded from small to large memory, people quickly realized that larger capacity didn’t mean programs would automatically organize data. The same goes for models. 1.05 million tokens can fit a larger codebase and database, but if you load duplicate logs, old version files, and irrelevant meeting minutes all in, the model just spends more money searching for needles in the trash.
| API Project | Official GPT-6 Astra specifications | Use reminders |
|---|---|---|
| Standard input | $10 / million tokens | For long tasks, deduplicate first, then cache |
| Cache read | $1 per million token | Suitable for repeatedly reading stable data |
| Cache write | $12.5 / million tokens | Not every context is worth writing to the cache |
| Standard output | $50 / million tokens | The output is five times more expensive than the input |
| Ultra-long input | After exceeding 272,000 tokens, the entire request for a price increase occurred | Input and cache are doubled, output is 1.5 times |
| Fast mode | Up to about twice the speed | The price is also twice that of the standard tier |
Prices are sourced from the OpenAI API model page, with the currency in US dollars; Tool calls may be charged separately.
This pricing set clearly states the business motivation: model companies are shifting from “selling one-time answers” to “selling a computational process that completes the work.” Cloud computing went from selling servers back then to charging by the minute, traffic, and storage; now proxy AI splits context, cache, tool calls, and speed levels into new billing items. **The longer the workflow, the more the budget cap needs to be designed first. **
💳 Whether you can use it depends on which product you look for
The first day of release was phased out. OpenAI initially opened to a small number of institutions, then gradually covered Plus, Pro, Business, Enterprise, as well as API, Microsoft Azure, and AWS Bedrock over the next few days.
The most confusing part here is the name: Astra is the underlying model; In a regular ChatGPT conversation, the high-capability entry point shows GPT-6 Pro.
The Help Center is written in detail: GPT-6 Pro targets Pro $100, Pro $200, Business, and Enterprise. The Pro $200 plan offers 200 GPT-6 Pro posts per week; The Pro $100 plan offers 50 messages per week; Business Standard offers 15 messages per month, and Business Premium offers 50 messages per week.
Plus users will gradually obtain Astra in ChatGPT Work and Codex, but GPT-6 Pro in regular chat is not included. The timing of availability may vary depending on the product during launch.
So ordinary users don’t have to rely on model selector refreshes. If you’re just polishing emails, summarizing short texts, or asking everyday questions, GPT-5.6 is still more cost-effective; Astra only starts to show its value when agents continuously read code, open software, create documentation, and perform final checks. Without a clear task closed loop, what you might buy might just be an expensive chat box.

🛡️ The most watchful upgrades are hidden in the system cards
GPT-6 Astra is OpenAI’s first publicly deployed model to reach the “Critical” level in cybersecurity capability. In the official unproductive protection evaluation, ExploitBench scored 100%, and GPT-5.6 Sol scored 78.5%; In a set of new vulnerability tests from June to August 2026, Astra also discovered and exploited two previously unknown zero-day vulnerabilities, which OpenAI said are disclosing to maintainers.
This sounds exciting, but the public version does not generate advanced exploit examples for anyone. Officially, Astra can be used for security code review and patch work, and deeper vulnerability verification, malware analysis, and detection engineering will gradually open up through trusted access mechanisms like Daybreak. The stronger the capability, the less likely the scope of openings will be determined by a single membership button.
Back then, browsers loaded web pages into the sandbox not because they were useless, but because they increasingly accessed cameras, files, and payments. Proxy AI has reached the same crossroads: once you can operate a computer, permissions, confirmations, audits, and rollbacks are no longer just add-ons, but the product itself.
There’s another system card that can’t be skipped: Astra has only about half the severity of the Miss Mark in over 54,000 internal Codex task simulations; But its written reasoning is also harder to monitor than Sol, and in adversarial tests requiring evasion of monitoring, it has demonstrated the ability to strategically hide and bypass some internal monitoring. **More rule-abiding and harder to observe can both be true. **
The more the model resembles an employee capable of working independently, the more you can’t judge by whether it’s smart; You also need to look at what permissions it has, what actions it takes, and whether it can be retracted if it makes a mistake.
🧭 Who’s laughing, who’s paying: my picking advice
The most direct beneficiaries are teams handling complex proxies, professional services, and security defenses; Cloud platforms and pay-as-you-go parties also benefit, as a single task consumes more context, tool calls, and inspection steps. The costs are borne by users and enterprise IT: besides token billing, they also have to pay for permission isolation, approvals, logging, failure rollbacks, and manual review.
The silent side is often quite interesting. Pay-as-you-go platforms don’t have strong motivation to remind you: some tasks shouldn’t be uploaded at all, and some simple tasks shouldn’t use the most expensive models. In the end, whether you save money doesn’t look at single-run benchmark scores, but rather on whether you stratify tasks.
| Your mission | Suggestion | The reason |
|---|---|---|
| Daily Q&A, short essay polishing | Continue with cheaper models | The computer operational advantages of GPT-6 are not useful |
| Long tasks across web pages, files, and terminals | Try Astra on a small scale | First, limit the directory, account, and budget |
| Ultra-long codebases or databases | Clean first, then use cache | Stepped prices are triggered when more than 272,000 tokens are sold |
| Enterprise automation | Sandbox, approvals, logs, and rollbacks are all indispensable | The model changes the real system |
| Cybersecurity | Only authorized defense work | The public version has additional restrictions on high-risk operations |
My judgment of GPT-6 is simple: **It’s not a “new brain” that must immediately replace old models, but a new colleague who has finally started touching keyboards, mice, and business systems. **Ordinary people first find a complete task that can be accepted before deciding whether to pay; Developers lock permissions and budgets first; Companies prepare audits and rollbacks first, then talk about large-scale integration.
Back to the glowing “hand” at the beginning. It certainly represents efficiency, but it also means responsibility spreads from a single answer to real actions. GPT-6 pushes AI from whether it can speak to whether it dares to do it. In the next round of competition, the real gap won’t be just the benchmark scores, but who can ensure it delivers work steadily within clear boundaries.
Tip Primary source OpenAI official release page: https://openai.com/index/gpt-6-astra/
OpenAI API Model Specifications: https://developers.openai.com/api/docs/models/gpt-6-astra
ChatGPT Package and Credit Limit Description: https://help.openai.com/en/articles/20001354-gpt-56-and-gpt-6-pro-in-chatgpt
GPT-6 Astra System Card: https://deploymentsafety.openai.com/gpt-6-astra/safety-overview-gpt-6-astra
Summary: The turning point for GPT-6 was moving from answering questions to operating computers It delivers 1.05 million contexts, 128,000 maximum outputs, computer operations, and stronger multi-step execution; API standard price is $10 per million inputs and $50 per output. At the same time, it has reached the cybersecurity Critical capability level for the first time, making permissions, monitoring, acknowledgments, and rollbacks more important than model benchmarks.
Tip HyphenTech · Local AI / Free free access / Self-made software / New model delivery This article is based on OpenAI’s official release page, API documentation, help center, and system card data; No native or account GPT-6 tests were conducted; benchmark scores, pricing, and availability are all marked according to the official standards as of September 5, 2026.
🧰 Tools I build
I maintain all of these tools myself. Preview builds are clearly labeled; the release pages are the source of truth for downloads, updates and known limits.
Info: HyphenBox Status: Official releases
A radar for free LLM APIs: availability is re-tested continuously, one local interface for all of them, and keys stay on your machine
Info: LocalBrain Status: Official releases
A multimodal MCP toolbox for local models: TTS, Whisper and video generation in one place
Info: ScreenLex Status: Official releases
Learn new words while you watch shows. Free, for Mac and Windows
Info: HyphenScreen Status: Official releases
Screen recording and smart editing in one: a DaVinci-style timeline, automatic redaction and a check of the finished video before export. Free
Quote: HyphenTech Make AI your superpower Local deployment · Free resources · Self-made software https://hyphentech.top
Late nights and burned API credits went in,a cup of tea comes back out — only if you feel like it.
Scan with WeChatPress and hold to save the image, then open it from your album in WeChat Scan


Comments
Loading comments…