Everybody is testing GPT-6 Astra on code. I pointed it at the creative stack instead: Adobe Illustrator, After Effects, Figma and Photoshop.
I gave it my actual brand, my actual website, and my actual work, then let it drive the applications itself while I watched.
One task ran for 47 minutes straight. Another ran for an hour and 13 minutes and produced a full mobile app design I am shipping as my real product.
I have never opened After Effects in my life, and Astra made me a channel intro and an animated skit in it. Here is everything it built, what it cost, and the one thing about this launch that almost nobody is talking about.
Key Takeaways
What GPT-6 Astra Actually Is
Let me set the official facts before I show you my results.
OpenAI released GPT-6 Astra on September 3, 2026. In their own words:
"We're introducing GPT-6 Astra, the world's most intelligent and aligned model."
The developer docs are blunter: "Our most capable model, built for the hardest end-to-end work."
| Spec | Value |
|---|---|
| API model string | gpt-6-astra |
| Released | September 3, 2026 |
| Context window | 1,050,000 tokens |
| Max output | 128,000 tokens |
| Knowledge cutoff | April 30, 2026 |
| Reasoning effort | low, medium, high, xhigh, max |
| Input | Text and images |
| Output | Text only |
One thing to get straight early, because it confused me at first. In ChatGPT it appears as GPT-6 Pro, powered by GPT-6 Astra. The product label and the model name are different things.
Pricing And Who Actually Gets It
API pricing is $10 per million input tokens and $50 per million output. There is a detail in the fine print worth flagging: prompts over 272K input tokens get billed at 2x input and 1.5x output for the entire request, not just the overage.
On consumer plans, this is where people are going to get caught out:
If you are on Plus expecting to run the kind of hour-long creative sessions I am about to show you, check your access first.
Computer Use Is The Whole Story Here
This is what makes Astra different, and it is officially documented rather than a hack. From OpenAI's Computer Use documentation:
"With ChatGPT can see and operate graphical user interfaces on macOS or Windows… such as checking a desktop app… changing app settings."
It requires Screen Recording and Accessibility permissions on macOS. On Windows it runs foreground-only on the active desktop.
Now here is the honest caveat I want to put up front. OpenAI's official examples are KiCad, Blender, Unreal Engine 5, Excel and Power BI. Adobe Illustrator, After Effects, Photoshop and Figma appear nowhere in OpenAI's Astra materials.
So everything below is me testing undocumented territory. That cuts both ways: the results are more surprising, and they are also less guaranteed than an officially supported workflow.
Test 1: Illustrator, Vectorising My Own Brand
My first task was practical. I have AI-generated logos and mascots on promptslove.com for my three apps, OpenField, UIPitch and RanknestAI. They are raster images. I wanted proper vector files so I could reuse them in animation workflows.
Normally I would hire someone for this. Instead I asked Astra to open Illustrator and do it.
First browse promptslove.com see all 4 mascot for promptslove.com, uipitch, ranknestai and create detailed layered vector using [@Adobe Illustrator 2026]
It analysed my website in depth first, then rendered the Illustrator file itself.
What got me was the granularity. The eyes, the top emblem, the smile, each element has its own unique vector path. This is not an auto-trace of a PNG. Every element was designed inside Illustrator by the model.
You genuinely cannot tell which parts were AI-converted.
Test 2: After Effects, And I Have Never Opened It
This is the test that changed my mind about what this model is.
Some context so you understand the baseline. I hold an Adobe subscription only for Premiere Pro. I have never used After Effects in my life, because it is complicated and I have never had the time to learn it.
I asked Astra for two things: a 16:9 channel intro with an animated logo reveal and a blinking eye on my mascot, and a matching outro with all my social handles.
It worked for around 46 minutes.
The intro came back with the blinking eye animation and sound effects. The outro has a slot for my channel art and a next-video section, a loop animation on the logo, all my socials, and it holds my brand language throughout.
Here's Intro Video;
Here's outro Video;
I will be honest about my reaction. That blew my mind, because I have never touched this software.
Test 3: A Vox-Style Animated Skit
Then I pushed it harder. I wanted a Vox-style explainer skit, which meant chaining two different tools together: generate the images first, then animate them in After Effects with dialogue sync.
Create a VOX Style skit using [$imagegen] to create image first then use [@Adobe After Effects 2026] for added animation plus dialogue sync, Papercut style, animation and transition in 16:9 Here's the scenario Casual conversation between the two 2 guys one is worried about future who says "you know soon AGI will come and take over your job" and then other guy working computer "i doubt that bro, you think a machine can make design, create motions, vfx and 3d, it can only write text" and on date 4 september, 2026 AGI is born with openai logo and news says "GPT - 6 aASTRA is launched" and guy scratch his head (shot from behind - on pc screen reading about gpt-5 astra) ends there.
I gave it the scenario and the full context of what to build. It worked for 47 minutes straight and produced a 24-second finished piece.
What holds up when you actually inspect it:
That skit is the clip that opens my video, if you watched it.
For someone who has never touched After Effects, seeing this land means I could hand it a ten-minute video brief and offload the entire production.
Test 4: Figma, A Full Mobile App Design
This is the longest run and the most commercially useful output.
I asked Astra to design a mobile app for Promptslove in Figma, pulling all design material from my live website. I specified each screen, asked for curves, and told it to generate a working prototype with connectors between screens.
Design Mobile App using [@Figma] app, It's about promptslove get all the design material from the website itself but you need to design the following screen - Loader/opener - Login/signup screen - Dashboard with cards for prompts, generators, Automation, COMMUNITY, courses with user's greetings with status premium - Prompts page with library of prompts with search and categoru selector - Simple prompt page with title, description, prompt in code format and copy to clipboard open in chatgpt, claude icon - Generator page with list of generators like text prompt generator, image, video - generator single page with list of options and generated - Show animation while generating - Global menus for dashboard, prompt, generators and porofile at bottom - Community page with different boards - Profile settings - Courses page with list of searchable courses title featured image, oprogress bars - Show course single page with title orogress etc player as well - Show upgrade page - Show light dark switch on all - that ai bubble open like popover from bottom covering half, show like a chat ai helper what user want Add connectors and generate prototype using figma. loop animation, add mascot at top with loop animation with blinking continously. Research well with advance figma workarounds and intuitive mobile screen designings
It ran for 1 hour and 13 minutes.
What came back:
Doing this myself would have taken four or five hours, and honestly I could not have done it to this standard.
Here is the part that matters. I cannot imagine a better version of my own app than what it produced. I am using this as the final design and building the mobile app from it.
Test 5: Photoshop, Ad Creatives

Last test, and the fastest. I wanted ad creatives for the website, so I had Astra work in Photoshop alongside image generation, producing four different angles at set aspect ratios.
Now use Adobe photoshop and [$imagegen] as well for some asset and create Ad creatives for promptslove.com in 1:1, 9:16 3:4 4:5 for the following angle - Angle 1 - Before and after learning AI - Angle 2 - Robot (AI) as your servant - Angle 3 - RanknestAI - Rank on Chatgpt, claude, gemini not just Google - Angle 4 - Openfields - No Credits, Unlimited fun , Pay as you go GEN ai SUITE GUIDELINES - Minimalistic Ad creative, realistic representation, no over text, no over styling - Solid backgrounds for all - Gritty yet funny that generate curiosity
My guidelines were specific: minimalistic, realistic representation, no excess text, no over-styling, solid backgrounds, gritty yet funny, designed to create curiosity.
It took 26 minutes and delivered four creatives:
I am putting that last one into an actual advertising campaign.
One correction on my own video here. I referred to "Image N" as the image model. That is not an OpenAI product. The current model is GPT-Image-2, branded as ChatGPT Images 2.0 on consumer surfaces. Astra itself outputs text only and calls image generation as a separate tool.
The Benchmark That Explains All Of This
After running these tests I went back to OpenAI's published numbers, and the computer-use benchmarks line up with my experience closely.
| Benchmark | GPT-6 Astra | Claude Opus 5 | GPT-5.6 Sol |
|---|---|---|---|
| OSWorld 2.0 (computer use) | 72.6% | 70.2% | 65.7% |
| Agents' Last Exam | 59.3% | 55.5% | 53.6% |
| ScreenSpot-Pro (no tools) | 92.7% | Not published | Not published |
| Terminal-Bench 4.0 | 57.9% | 52.6% | Not published |
For context on that last row, Gemini 3.8 Flash scores 19.1% on Terminal-Bench 4.0. Claude Fable 5.1 sits at 55.8%, just behind Astra.
There is one benchmark where Astra loses, and I think it is worth being straight about: Humanity's Last Exam with tools, where Astra scores 57.2% against Fable 5.1's 65.0%. So this is not a clean sweep.
On speed, OpenAI's OSWorld latency simulation puts Astra at around 40 minutes per task, versus roughly 75 minutes for GPT-5.6 Sol. My runs landed between 26 and 73 minutes, which is right in that band.
OpenAI states its scores represent "the maximum at any effort" and that evals ran in their research environment rather than production ChatGPT. Worth knowing before you expect identical results.
The Thing Nobody Is Talking About
While researching this piece I found something in OpenAI's own documentation that deserves more attention than it is getting.
In the Path to Astra post, OpenAI writes:
"We now believe Astra meets the Critical cybersecurity capability threshold under our Preparedness Framework… It is the first model we are designating at this level."
The first model ever. It scored 100% on ExploitBench, and OpenAI states the model "discovered and used two zero-day vulnerabilities as part of an exploit chain."
Two things keep that in proportion. Those results reflect a research configuration with expanded access, not the default production model. And the shipping version refuses to create proof-of-concept exploits, with OpenAI reporting a 91.5% refusal rate on cyber jailbreak evaluations against 59% for Sol.
There is also a caveat that affects exactly the kind of work I did. OpenAI notes that safeguards may flag "tasks in which an agent is running for an extended period," pausing them in ChatGPT and Codex and stopping them outright in the API. My longest run was 73 minutes without interruption, but plan for it.
What People Are Building With GPT-6 Astra
My tests were all creative-tool driven, so I went through X to see where everybody else is pointing this thing. The dominant use case is 3D, overwhelmingly in Blender.
The post that set the tone, a one-shot 3D game in 45 minutes:
https://x.com/andytng28/status/2096191616407191819
His follow-up point is the useful one: the trick to getting good graphics out of it is image generation feeding the build.
Sci-fi megastructures in Blender, under five minutes each, one shot, on medium reasoning:
https://x.com/snapsnocaps/status/2096236638049304958
Architectural accuracy, which is the one that impressed me most. A model of the Rijksmuseum built from reference images, where Astra looked up the official floorplan to verify its dimensions:
https://x.com/jungle_jimjim/status/2096235789084479678
One image to an animated 3D warehouse with moving forklifts, workers and conveyors, built in Codex:
https://x.com/arjun_gupta95/status/2096236186306023601
A six-month-old bug that no frontier model could solve, one-shotted:
https://x.com/matthewmillerai/status/2096198340476076475
The Reactions That Push Back
I do not want to only show the highlight reel, because the response has not been universally positive.
A working Blender artist replied to one of the viral demos with genuine anger about AI-generated 3D work. I am not going to quote the language, but the sentiment is real and it is aimed directly at people like me posting results like these. I think that reaction deserves to be part of this conversation rather than filtered out of it.
There is also a fair critique that everyone is testing the flashy stuff:
https://x.com/jarvis0970/status/2096220842929586224
And a more measured take from someone using it in a real workflow, who found the computer use genuinely better than Sol while noting it still wanders off and spends tokens on things you did not ask for:
https://x.com/it_is_Randy/status/2096235760261255502
That last one matches my experience. These runs are not fire-and-forget. I watched them.
My Honest Verdict
I opened my video saying creative skills are going to die. Let me be more precise than that, because I do not fully believe the simple version.
What Astra genuinely does that I could not:
What it does not do:
The skill that survives is knowing what to ask for and recognising when the output is right. The skill that just got compressed is the software fluency, the thing I avoided for years by never learning After Effects.
I am not sad about that. I got a channel intro, an animated skit, a full app design and an ad campaign out of one afternoon.
Frequently Asked Questions (FAQs)
What is GPT-6 Astra?
GPT-6 Astra is OpenAI's model released September 3, 2026, which OpenAI calls "the world's most intelligent and aligned model." Its standout capability is Computer Use, letting it see and operate desktop applications on macOS and Windows. It has a 1,050,000 token context window and an April 2026 knowledge cutoff.
Can GPT-6 Astra really use Photoshop and After Effects?
In my testing, yes, and it produced usable output in both. But OpenAI does not officially document Adobe applications. Its published computer-use examples are KiCad, Blender, Unreal Engine 5, Excel and Power BI, so Adobe workflows are unsupported territory.
How much does GPT-6 Astra cost?
$10 per million input tokens and $50 per million output tokens on the API, with prompts over 272K tokens billed at 2x input and 1.5x output. On ChatGPT, Pro $200 gets 200 messages weekly, Pro $100 gets 50 shared with Sol Pro, and Plus does not get GPT-6 Pro in chat at all.
How long can GPT-6 Astra run autonomously?
OpenAI publishes no hard session limit, but its OSWorld latency simulation averages around 40 minutes per task. My runs went from 26 minutes to 1 hour 13 minutes. OpenAI does note that safeguards may pause agents running for extended periods.
Is GPT-6 Astra better than Claude Opus 5?
On computer use, yes by OpenAI's numbers: 72.6% versus 70.2% on OSWorld 2.0, and 59.3% versus 55.5% on Agents' Last Exam. But Claude Fable 5.1 beats Astra on Humanity's Last Exam with tools, 65.0% to 57.2%, so it depends entirely on the task.
Why is GPT-6 Astra classified as Critical for cybersecurity?
OpenAI designated it the first model to meet the Critical cybersecurity threshold under its Preparedness Framework after it scored 100% on ExploitBench and discovered two zero-day vulnerabilities during evaluation. Those results came from a research configuration, and the shipping model refuses to create proof-of-concept exploits.
Final Thoughts
The reason this launch feels different from the last few is not the benchmark chart. It is that the model stopped writing instructions for software and started using the software.
If you are testing Astra this week, do not just give it code. Point it at the application you have avoided learning, give it your real brand assets, and watch it work. Then decide for yourself which part of your job it actually touched.
I am going deeper on this next, and putting the full prompts from all five tests on promptslove.com along with everything else I use.





