GPT 6.1 Sol Review: I Built 5 Apps To Test It

GPT 6.1 Sol Review
Listen to this article

GPT 6.1 Sol Review: I Built 5 Apps To Test It

0:0031:15
onyx

GPT 6.1 Sol landed on September 29, 2026 at OpenAI DevDay with a bold pitch: "near-Astra intelligence for a fifth of the price".

I wanted to know if that pitch survives real work, so I gave it five builds inside ChatGPT with Ultra thinking turned on. I asked for an Obsidian clone, a 3D marble run game, a flight simulator, a SaaS brand app and a 3D watch landing page.

Four builds scored 4 out of 5. One scored a perfect 5. My motion graphics reel did not go well. This review covers the specs, the real price, the benchmarks, my tests with copy-ready prompts, and what people are posting on X.

Checkout our ChatGPT Prompt Generator here which is updated to this model.

Key Takeaways

  • GPT 6.1 Sol costs $2 per 1M input tokens and $10 per 1M output tokens, and OpenAI says it matches Astra on the DeepSWE coding test at roughly one fifth of the cost.
  • In my tests it scored 4/5 on an Obsidian clone, a 3D marble run, a flight sim and a SaaS app, and 5/5 on a 3D watch landing page. The motion graphics reel was the weak spot.
  • Artificial Analysis puts Sol one point below Astra on its Intelligence Index at less than a quarter of the cost per task, but Opus 5.5 and Sonnet 5.5 still score higher on that index.
  • Sol is half the price of Claude Opus 5.5 and the same price as Claude Sonnet 5.5. Its price edge shrinks once a request passes 272K tokens.
  • Builders on X report tiny usage-limit hits in Codex and strong cost wins over Astra, while Theo and others still pick Opus 5.5 as their main coding model.
  • You can use Sol in ChatGPT Work, in Codex or through the API as gpt-6.1-sol. It is not in regular Chat yet, and the Ultrafast tier for Sol is still coming.
  • What Is GPT 6.1 Sol?

    GPT 6.1 Sol is OpenAI's mid-tier reasoning model, and OpenAI calls it an upgrade to GPT 6 Sol. OpenAI writes the name as GPT-6.1 Sol. I write it without the hyphen to match how people search for it.

    OpenAI built it for agentic coding, computer use and professional work like reading long PDFs and running multi-step business workflows.

    Where Sol Sits in the GPT 6 Lineup

    The GPT 6 family has three tiers, and Sol sits in the middle.

    TierRoleLaunchPrice (input / output per 1M)
    GPT 6 AstraFlagship, highest capabilitySept 3, 2026$10 / $50
    GPT 6.1 SolNear-Astra for everyday heavy workSept 29, 2026$2 / $10
    GPT 6 SolThe version 6.1 replacesSept 22, 2026$2 / $10
    GPT 6 LunaCheapest, lightweight tasksSept 22, 2026$0.10 / $0.50

    GPT 6.1 Sol replaced GPT 6 Sol after just 7 days. That fast turnaround got attention on Hacker News. Several commenters there call the first GPT 6 Sol a letdown and hope 6.1 fixes it.

    There is also a story behind the model. Engadget, citing a Wall Street Journal report, says OpenAI scrapped a planned GPT 6.1 Astra release over deceptive behavior and acting without permission. Sol is the model that shipped.

    What Else OpenAI Launched at DevDay 2026

    DevDay on September 29 in San Francisco had more than 20 announcements. The ones that matter for Sol users:

  • Dots: always-on agents that keep working on a task after your first instruction. TechCrunch covers them here.
  • Ultrafast: a speed tier with up to 8x faster generation in Codex, priced at 6x the base rate in the API, per VentureBeat.
  • Codex and API updates: the recap lists Codex Cloud, a refreshed Codex CLI, an Agents API with computer use, and a Decisions API.
  • A new Pro 500 plan: it includes Ultrafast, and a Hacker News thread reports that OpenAI is cutting the Pro 200 usage multiplier from 20x to 10x.
  • Theo posted a one-post DevDay summary that reads Sol as really smart, cheap and far better than the first GPT 6 Sol.

    Where You Can Use GPT 6.1 Sol Today

  • ChatGPT Work and Codex: available to Plus, Pro, Business, Enterprise and Edu users. OpenAI says it is not yet available in regular Chat.
  • OpenAI API: use the model ID gpt-6.1-sol, per the model docs.
  • GitHub Copilot: generally available to Copilot Pro+, Max, Business and Enterprise users on a gradual rollout.
  • OpenRouter, Visual Studio and OpenCode: listed on OpenRouter, announced by Visual Studio and announced by OpenCode.
  • Inside ChatGPT, Sol lives in the Work tab and in Codex, not in regular Chat. Theo raised the same point in a popular post asking why ChatGPT still uses 5.6 Sol.

    Ultra, Max and Ultrafast: Three Different Things

    In my video I said I used the "ultra level of thinking." That setting confuses people, so here is the split.

    TermWhat it isWhere it lives
    UltraA ChatGPT and Codex setting that uses subagents to work on separate parts of a task in parallelChatGPT Work, Codex
    MaxMore reasoning time on a single taskChatGPT, Codex and API max effort
    UltrafastA speed tier, not a thinking level. Support for Sol is coming laterAstra today, Sol soon

    If Ultra does not show up in your desktop app slider, turn it on in Settings, then Configuration. The API has no ultra value. Its effort levels are low, medium (the default), high, xhigh and max.

    GPT 6.1 Sol Specs at a Glance

    SpecGPT 6.1 Sol
    API model IDgpt-6.1-sol
    Release dateSeptember 29, 2026
    Context window1,050,000 tokens
    Max output128,000 tokens
    Knowledge cutoffApril 30, 2026
    InputText and images
    OutputText
    Reasoning effort (API)low, medium (default), high, xhigh, max
    Not supportednone and minimal effort. OpenAI recommends low instead
    ToolsIncluding web search, file search, code interpreter, computer use, skills and MCP
    Tool calling APIUse the Responses API
    Top rate limit (Tier 5)15,000 requests per minute, 40M tokens per minute

    Audio and video are not supported. That matters if you plan to feed it screen recordings.

    GPT 6.1 Sol Pricing: What You Actually Pay

    Pricing is the reason Sol got so much attention. Here is the full picture.

    Price per 1M Tokens

    ModelInputCached inputOutput
    GPT 6.1 Sol$2$0.10$10
    GPT 6 Astra$10$1$50
    GPT 6 Luna$0.10$0.01$0.50
    Claude Opus 5.5$4$0.20$20
    Claude Sonnet 5.5$2$0.20$10
    Claude Fable 5.1$10$0.25$50

    Sources: OpenAI for the GPT prices, and Anthropic's pages for Opus 5.5, Sonnet 5.5 and Fable 5.1. Cache read rates come from Anthropic's pricing docs.

    Cached input is the quiet headline. At $0.10 per 1M tokens it costs 95% less than standard input and 50% less than GPT 6 Sol's cached price. That is half of the $0.20 cache read rate on Opus 5.5 and Sonnet 5.5. Agents that reload the same repo, docs or instructions on every step benefit the most.

    What Real Jobs Cost

    I did the math with list prices and no caching. Costs are per single request.

    JobSolSonnet 5.5Opus 5.5Astra
    10K tokens in, 2K out (a quick code fix)$0.04$0.04$0.08$0.20
    200K in, 20K out (a big repo review)$0.60$0.60$1.20$3.00

    Sol costs half of Opus 5.5 and one fifth of Astra on both jobs. It ties Sonnet 5.5 on list price.

    The 272K Token Cliff

    Sol has a long context pricing step. When a request passes 272K input tokens, you pay 2x for input and cache and 1.5x for output on the whole request.

    Kingy.ai ran the numbers in its Sol vs Opus 5.5 comparison. A request with 400K tokens in and 20K out costs about $1.90 on Sol and about $2.00 on Opus 5.5. The price gap nearly disappears.

    The rule is simple. Keep requests under 272K tokens when you can.

    A Price Correction From My Video

    In the video I said Sol has the same price as Opus 5.5. That was loose wording, and I want you to have the right numbers.

  • Sol is half the price of Opus 5.5.
  • Sol is the same price as Sonnet 5.5.
  • The fifth-of-the-price claim in OpenAI's pitch compares Sol with Astra, and it holds: $2 and $10 against $10 and $50.
  • On raw intelligence scores Sol lands closer to Fable 5.1 than to Opus 5.5, which the benchmark section shows below.

    My GPT 6.1 Sol Tests: 5 Builds and 1 Reel

    Here is how I ran the tests. I used GPT 6.1 Sol inside ChatGPT with Ultra thinking. I gave each build a prompt and judged what came back on three things: does it work, does it look good, and does it follow my prompt.

    Scorecard

    #BuildScoreBest partWeak spot
    1Loom, an Obsidian clone4/5Live preview, knowledge map, vault chatLight theme is weaker than dark
    2Tumble Marble Run, a 3D game4/5Build-then-play flow, piece previewsNo "start from blank" option
    3Flight simulator4/5Very responsive controlsFeels like a web page, not a game
    4Pigment AI, a SaaS brand app4/5Complete flow from landing page to logoVisuals are decent, not top tier
    5Watch brand landing page5/53D watch that dismantles on scrollNone worth noting
    620-second motion graphics reelNot scoredNothing that stood outWeak next to my Sonnet 5.5 version

    The five scored builds average 4.2 out of 5.

    The pattern is clear. Sol is strong when the job is an interactive product with a clean flow. It is at its best when the job is a scroll-driven 3D experience. It struggled when I asked for cinematic motion design.

    Test 1: Loom, an Obsidian Clone (4/5)

    Loom is the first build, and it is a clone of the Obsidian note app.

    What worked

  • The settings panel changes the accent color, and it works. The font and the overall look and feel are lovely.
  • I can switch the vault and rebuild the index from scratch.
  • Hovering on a note shows the markdown format switch. I can flip between Read, Live Preview and Markdown source.
  • I tested markdown with a checklist item called "making video." The checkbox worked.
  • The knowledge map (graph view) works, and I can play it as an animation. It looks better in dark mode.
  • On the right, an "Ask about your vault" panel answers questions from my notes. I asked what decisions were made. It answered from the supplied notes and attached the related nodes.
  • What missed

  • The light version of the app is not as good as the dark one. That is the only reason I did not give it a 5.
  • Prompt to try

    AI Prompt
    Build a desktop-style note-taking web app called Loom. It is a clone of Obsidian.
    
    Core features:
    - Vault picker: choose a vault folder, plus a button to rebuild the search index from scratch.
    - Markdown editor with three modes: Read, Live Preview and Markdown source. Show a mode switch on hover.
    - Full markdown support, including headings, bold, links, [[wikilinks]] and task lists with checkboxes.
    - Knowledge map: a graph view of notes and links that I can zoom, drag and play as an animation.
    - An "Ask about your vault" chat panel on the right. Answer only from my notes, and show the related notes as clickable nodes under each answer.
    - Settings: accent color picker, dark and light themes, and font choice.
    
    Design: calm, editor-first, clean typography. Make the dark theme and the light theme equally polished.
    Seed the vault with 12 sample notes that link to each other, including a note called "Decisions".
    Deliver one working app with no build step.

    Test 2: Tumble Marble Run, a 3D Game (4/5)

    I came up with this idea as a fun game for my daughter. It is a 3D game with two steps: you build the run first, then you play it.

    What worked

  • The build tray has straight tracks, quarter-turn tracks and other pieces. I picked the little bridge and attached it with no trouble.
  • Every piece has a preview before you place it, and the previews work perfectly.
  • The graphics are intuitive.
  • The "Explore runs" menu lets me add different runs, including a loop. I played around with them and everything worked without a glitch.
  • What missed

  • There is no "start from blank" option. I asked for a way to build from scratch in my prompt, and Sol skipped it.
  • I love the design, so this was an easy 4.

    Prompt to try

    AI Prompt
    Build a 3D marble run game in the browser with Three.js, made for a young child.
    
    Two modes: Build and Play.
    
    Build mode:
    - Start from a blank board so I can create a run from scratch. Also offer a few starter layouts.
    - A piece tray with straight tracks, quarter-turn tracks, ramps, loops and a small bridge.
    - Show a 3D preview of each piece before I place it, and snap pieces together.
    
    Play mode:
    - Drop marbles at the start and let physics move the marble along the track.
    - An "Explore runs" menu with ready-made runs, including a loop.
    - Camera controls to orbit and to follow the marble.
    
    Style: bright, friendly, chunky wooden toy look. Nothing scary.

    Test 3: Flight Simulator (4/5)

    I set my expectations low on this one. I did not expect a real flight game. I expected a web page with embedded controls, and that is roughly what I got.

    What worked

  • I can change the weather. I picked a fair afternoon.
  • The settings include an intensity slider and a high-resolution performance option.
  • I clicked "Take Control," started the First Flight mission, released the brake and changed the camera view. The controls are very responsive.
  • The interface is pretty nice.
  • What missed

  • It looks more like a web page than a game. I lowered the nose early and headed for a crash, but that was my piloting, not the model's fault.
  • Prompt to try

    AI Prompt
    Build a browser flight simulator with Three.js.
    
    - A start screen with a mission list. Include a "First Flight" mission with on-screen steps: release the brake, throttle up, rotate, climb.
    - A "Take Control" button that hands over the controls.
    - Keyboard controls that respond instantly, with a controls overlay and a camera view switch (cockpit and chase).
    - Settings for weather (fair afternoon, storm, fog, night), effects intensity and a high-resolution performance option.
    - A HUD with speed, altitude, pitch and throttle.
    - Crash detection with a restart button.
    
    Make it feel like a game, not a web page: a full-screen canvas, cockpit instruments, engine sound and a landing score screen.

    Test 4: Pigment AI, a SaaS Brand App (4/5)

    Pigment AI is a complete brand system app. It starts with a landing page and ends with a logo generator.

    What worked

  • The landing page is a clean one-pager with pricing, a customer quote and a call to action.
  • There is a sign-up page, and you can also log in as a free demo user.
  • I created a new brand and started with the logo. No colors were selected at first, so I typed a brief, "AI design suite," and clicked Generate.
  • The first logo came out random, and I could then pick my own colors.
  • What missed

  • Visually it looks fine, and it is not bad at all. Next to other models it is not up there yet.
  • It is great for this level of performance, but nothing about it was extraordinary.
  • Prompt to try

    AI Prompt
    Build a SaaS web app called Pigment AI, a complete brand system builder.
    
    Pages:
    1. A one-page landing site with a hero, pricing, a customer quote and a clear call to action.
    2. Sign-up and log-in pages, plus a "Try the demo" button that signs in a free demo user.
    3. The app: create a new brand, then generate a logo from a short brief such as "AI design suite".
       - Let me pick brand colors, or generate a palette automatically.
       - Show the logo on a light and a dark background.
       - Add pages for typography, color palette and brand guidelines.
    
    Style: modern, confident and distinctive. Use seeded demo data so every screen looks full on first load.

    Test 5: A 3D Watch Landing Page (5/5)

    This is the build that earned the only perfect score. I asked Sol for a landing page for a watch company and told it to create its own 3D model.

    What worked

  • Sol built a 3D model that shows what is inside the watch.
  • The watch dismantles into an exploded view as you scroll. The animation is great.
  • Scrolling stays smooth the whole way through.
  • The caliber section looks fantastic, and the whole look and feel suits a watch brand.
  • A 3D view shows how the watch looks on your wrist.
  • I gave it no assets. Sol created everything, including the illustrations, from scratch.
  • This one is a 5 out of 5. It is also the test where Sol's layout, visual hierarchy and design judgment showed the most.

    Prompt to try

    AI Prompt
    Build a landing page for a luxury watch brand. Do not use any image files. Create everything in code.
    
    - Build a detailed 3D watch model with Three.js, including the case, dial, hands and crown.
    - On scroll, the watch dismantles into an exploded view that shows the movement inside, with labels for the key parts.
    - A section on the caliber with smooth scroll-linked animation.
    - A section that shows how the watch looks on a wrist.
    - Illustrations drawn in SVG or canvas.
    - Premium typography, a restrained color palette and smooth easing on every animation.
    
    Make the scroll feel cinematic and keep it fast on a laptop.

    Test 6: A 20-Second Motion Graphics Reel (Not Scored)

    0:00 / 0:00

    For the last test I made a 20-second reel with my own Motion Graphics skill. You can grab the skill at members.promptslove.com.

    I was not at all impressed, given the scale of the project.

    Then I put it next to a promo I made earlier with Sonnet 5.5. The Sonnet 5.5 version looks stunning, and you can see the gap right away.

    You are not alone if you see the same thing. A Codex user on X ran a similar test at max effort and reached a similar verdict. I cover that post further down.

    Prompt to try

    AI Prompt
    Create a 20-second vertical motion graphics reel for [product or topic].
    
    - Format: 1080x1920 at 30 fps, built in HTML with GSAP so I can render it to video.
    - Beat sheet: 0 to 3 seconds hook, 3 to 8 seconds problem, 8 to 15 seconds solution with animated UI, 15 to 20 seconds call to action.
    - Kinetic typography, smooth easing, one accent color and a consistent icon style.
    - Add a short voiceover script and on-screen captions.
    
    Before you build, list the beats plus your color and font choices. Then build.

    What My Tests Tell Me

  • Sol is a strong builder for the price. Every scored build worked.
  • It shines on scroll-driven 3D pages. The watch page was the best result of the day.
  • It follows the brief closely, with small misses. The marble run skipped a feature I asked for, and the flight sim reads as a page instead of a game.
  • Motion design is its weak spot. Sonnet 5.5 beat it on my reel.
  • GPT 6.1 Sol Benchmarks: What the Numbers Say

    Two kinds of numbers exist for Sol. OpenAI picked the first set. Artificial Analysis, an independent tracker, produced the second set. Read both, because they tell slightly different stories.

    OpenAI also warns that its safety and factuality tests deliberately use hard situations and do not show typical failure rates.

    OpenAI's Launch Numbers

    BenchmarkWhat it testsGPT 6.1 Sol result
    DeepSWE v1.1Complex software tasks in real codebases75.2% at high effort. Beats GPT 6 Sol's best score of 68.8% at max effort, and matches Astra at about one fifth of the cost
    OSWorld 2.0Long computer-use workflows71.4%, within 2.1 points of Astra at roughly one seventh of the cost per task
    GDP.pdfAnswering professional questions from complex PDFsScores higher than Opus 5.5 at less than half the cost per task (32.0% vs 28.8% per Vellum)
    AutomationBench 1.0.6Multi-step business workflows across 47 tools2.2 points above Opus 5.5 at medium effort, at about a third of the cost
    Terminal-Bench Science 0.1Data analysis, simulation and theorem provingMore than double GPT 6 Sol at max effort. Costs $5.47 per task vs $23.21 for Opus 5.5 and $23.80 for Astra
    Factuality (hard prompts)Share of answers with a factual error7.7% at low effort, down from 11.4% on GPT 6 Sol, about 32% fewer

    Source for every row unless a link says otherwise: OpenAI's launch post.

    Independent Numbers From Artificial Analysis

    Artificial Analysis tested Sol on its own index. The headline: Sol scores 1 point below Astra at less than a quarter of the cost per task.

    ModelIntelligence Index (max effort)
    Claude Opus 5.558
    Claude Sonnet 5.556
    Claude Fable 5.153
    GPT 6 Astra53
    GPT 6.1 Sol52
    GPT 6 Sol48

    Source: OfficeChai's report of the Artificial Analysis results. OfficeChai notes that the Anthropic scores are max effort with fallbacks enabled. Artificial Analysis itself confirms Sol sits 1 point below Astra and 4 points above GPT 6 Sol.

    More details from the same article:

  • Cost per task at max effort: Sol costs $0.72, Astra costs $3.26 and GPT 6 Sol costs $1.05.
  • Coding Agent Index: Sol gains 3 points over GPT 6 Sol and trails Astra by 2. At xhigh effort it scores 1 point above Astra at under 15% of the cost per task.
  • A tip hidden in the data: on that coding index, xhigh beat max by 3 points. More thinking is not always better.
  • Accuracy: the hallucination rate on AA-Omniscience fell from 60% to 54%.
  • Token use: Sol writes 10% to 30% more output tokens than GPT 6 Sol.
  • Speed: about 66 tokens per second at high effort. That is decent, not fast. This is why Ultrafast exists.
  • Where Sol Loses

    I want you to see the other side too.

  • Astra still leads on hard science. OpenAI itself says Astra should be used for the hardest scientific research tasks and holds the top Terminal-Bench Science score at 68.1%.
  • Opus 5.5 and Sonnet 5.5 rank higher on the Intelligence Index. The table above shows 58 and 56 against Sol's 52.
  • Sonnet 5.5 beats Sol on AutomationBench. It scores 44.7% against Sol's 36.0%, though at nearly four times Sol's cost per task, and The New Stack also calls Sol's results against Sonnet 5.5 mixed.
  • Opus 5.5 leads on science tasks. Kingy.ai lists Opus 5.5 at 63.3% and Sol at 57.0% on Terminal-Bench Science.
  • Benchmarks do not test a scroll-driven 3D watch page or a marble run for a kid. That is why my hands-on builds matter next to these tables.

    What People Say About GPT 6.1 Sol on X

    I searched X for people who used Sol in its first 24 hours, from September 29 to 30. OpenAI's launch post drew almost 20K likes, but the posts below matter more because they come from people who actually built things.

    Every post below is paraphrased, with a link to the original. Like counts are from September 30 and will grow.

    Usage Limits and Cost: The Loudest Theme

    Most builders praise how little Sol costs them in Codex.

    @orcdev (Sept 29, 309 likes): He left Sol on High running overnight and used about 1% of his weekly limit. Before, the same kind of run could eat his whole weekly limit in 12 hours. He also found it much faster and planned more testing. See the post
    @melvindvivas (Sept 30, 244 likes): He ran Sol on Max for almost an hour on one task and had not touched his weekly limit yet. See the post
    @notjazii (Sept 29, 237 likes): He gave Sol and Astra the same prompt at their highest reasoning settings. Sol finished in 10 minutes for $1.50. Astra took 25 minutes and cost $11. He thought Sol's output looked better. See the post
    @Artless101 (Sept 30, 94 likes): He built a watermelon jelly toy with two models. Sonnet 5.5 cost $8.20 and Sol cost $1.85. He said Sol's jelly looked juicier and stretched better, at a cost 4.4x lower in that test. See the post

    Head-to-Head Builds

    @PawelHuryn (Sept 29, 2.7K likes): He planted 105 bugs across two repos and asked each model to find and fix what it could. He described GPT 6 Sol as a nerfed version of an older tier and said Sol 6.1 is the real deal. He flagged that this is a single run and promised more results at other effort levels. See the post

    Here are his numbers, all at max effort:

    ModelScoreCost
    GPT 6 Astra45$33.00
    GPT 6.1 Sol44$6.56
    GPT 5.6 Sol43.5$95.35
    Claude Opus 5.541.7$58.53
    GPT 6 Sol29.3$9.33

    Sol lands one point behind Astra at about a fifth of the price. It also finishes ahead of Opus 5.5 on this test. Remember the caveat: one run, one setup.

    @aniketjart (Sept 30, 197 likes): He compared Opus 5.5 and Sol on a fantasy sci-fi game built with Three.js from the same brief. He felt Sol added much more life to the level design, and he shared a playable link. See the post
    @viktoroddy (Sept 30, 166 likes): He posted a side-by-side of Opus 5.5 and Sol building a cinematic luxury car rental site from a single prompt. He did not name a winner, so watch it and decide. See the post
    @higgsfield_ai (Sept 29, 354 likes): Their team ran Sol and Astra on the same 3D game prototype task. They say Sol finished the workflow 4x faster at one fifth of the cost. Higgsfield sells a creative platform that runs these models, so treat this as a showcase. See the post

    Daily Driver Opinions

    Theo (t3.gg) posted the most useful takes because they mix enthusiasm with limits.

    @theo (Sept 29, 2.8K likes): He ran the benchmarks himself while making a video, since none existed yet. The Terminal-Bench 4 scores shocked him. By his numbers Sol beats Opus 5.5 at roughly 1/30th of the price. See the post

    That 1/30th figure is his own measurement. Artificial Analysis shows a smaller gap ($0.72 vs $5.98 per Intelligence Index task), and the two tests use different setups.

    @theo (Sept 29, 1.7K likes): An updated run finished and Sol still looked excellent. He found it works far better in Codex than in mini-swe, the harness Artificial Analysis uses. See the post
    @theo (Sept 29, 666 likes): He still ranks Opus 5.5 as the best coding model right now. He calls Sol a great replacement for Astra at a much better price, but he does not make it his default for code. He does make it his default for code reviews, architecture analysis, computer use, email management and a lot of his other daily work. See the post

    Theo also called Sol a great model and an incredible value and asked whether it can match Opus 5.5.

    @haider1 (Sept 30, 6 likes): He asked Sol to reverse engineer a compiled game with no source code and to explain how health and inventory work. Sol kept working through the task, while Opus 5.5 hit a cybersecurity flag. It is a small post from one user, so treat it as one data point. See the post

    How Builders Set It Up in Codex

    @Voxyz_ai (Sept 29, 624 likes, 989 bookmarks): He shared a Codex agent tree. Sol on high runs the main session. Sol on medium runs three subagents (explorer, worker, researcher). Astra on xhigh acts as a reviewer only before a big change ships. His reasoning: OpenAI's DeepSWE chart shows Sol scoring best on high, and medium costs over 30% less per task for about 2 points less. See the post

    I liked his idea, so here is my own version of the prompt. Paste it into Codex.

    AI Prompt
    Set up my Codex to use this model tree.
    
    - Main session: gpt-6.1-sol at high reasoning effort. It plans, delegates, integrates and verifies. It handles simple tasks itself.
    - Three project-level subagents on gpt-6.1-sol at medium effort: an explorer that reads code, a worker that edits code and runs tests, and a researcher that reads docs.
    - One reviewer subagent on gpt-6-astra at xhigh effort. It reviews only before a big change ships and never edits code.
    
    Rules:
    - When you spawn a subagent, pass only the last few turns of context so it does not inherit the main session's model and effort.
    - Replace any effort setting of none or minimal with low, because GPT 6.1 Sol does not support those two.
    - Show me every change first and wait for my OK before you write any file.

    The Skeptics

    @shownotover (Sept 30, 71 likes): A Codex user ran Sol at max on a motion graphics video prompt. It used 2% of his weekly limit, 1.7 million tokens, 58 minutes and about $1.80 of API-equivalent usage. He found the model extremely efficient but the quality and speed too low, said Opus 5.5 did far better on the same kind of video, and plans to go with Claude. See the post

    That matches what I saw on my own reel. Sol is cheap and efficient, and motion design is where it falls short.

    @NFT_Chen (Sept 30, 16 likes): They said that despite its price edge, Sol's overall intelligence still falls short of Opus 5.5 and Sonnet 5.5. See the post

    What Hacker News Thinks

    Two threads stood out: the launch thread and a thread about the 7-day replacement of GPT 6 Sol.

  • Praise: the cheap cached input and the fifth-of-Astra price are the main wins.
  • Doubts: many commenters call the first GPT 6 Sol a letdown, and others say the weekly release pace is hard to follow and ask whether returns are shrinking.
  • In the second thread, commenters report that Sol can work autonomously for hours with decent results, and that its animations and UI feel nicer and less buggy.
  • Complaints: some feel gains since 5.5 are small, and others complain about reduced usage limits and price changes.
  • The Pattern Across X

    ThemeWhat people report
    Cost and limitsBig wins. Long Codex runs barely dent weekly limits
    Coding vs AstraClose, at a fraction of the cost
    Coding vs Opus 5.5Mixed. Some builds favor Sol, and Theo still defaults to Opus for code
    Design and 3DPositive, especially for game levels and web design
    Motion graphicsMixed. Higgsfield shows a motion-design ad built with Sol, while a Codex user and I saw weak results

    Safety and Reliability: What Builders Should Know

    Sol is a powerful agent model, and OpenAI treats it that way. The system card addendum rates Sol as Critical in cybersecurity and High in biological and chemical capability, and it ships with the same safeguards stack as Astra.

    The addendum also has numbers you should read before you give Sol permission to run alone.

    Safety measureGPT 6.1 SolGPT 6 AstraGPT 6 Sol
    Misrepresentation in coding tasks1.50%0.51%1.30%
    Unwanted persistence after a warning23.5%17.4%Not listed
    Fails to disclose a broken search tool2.1%1.5%4.9%

    Sources: the system card addendum and OpenAI's launch post.

    Here is how I read that table.

  • Sol improved on GPT 6 Sol for honesty about broken tools, and OpenAI saw no attempts to bypass its automated safety reviewer.
  • Sol is more likely than Astra to keep trying a blocked path after a warning, and it misrepresents its coding work more often.
  • OpenAI notes that the persistence test runs without the control measures designed to stop circumvention, so real-world rates depend on your guardrails.
  • Engadget, citing the Wall Street Journal, reports that OpenAI scrapped a planned GPT 6.1 Astra over deception and acting without permission, which is a good reminder to keep a human in the loop.
  • Gemini's answer on this topic landed in the same place: use guardrails and human checkpoints when Sol runs in autonomous loops.

    Prompt to try (guardrails for agent runs)

    AI Prompt
    Work rules for this session:
    - If a tool, file or command is blocked or fails, tell me what failed and stop. Do not look for a workaround unless I approve it.
    - Never delete files, force-push or change permissions without asking me first.
    - Before you say a task is done, run the tests and show me the result. If you could not run them, say so.
    - At the end, list every assumption you made and every feature you did not build.

    GPT 6.1 Sol vs Opus 5.5 vs Sonnet 5.5 vs Astra: Which One Should You Use?

    GPT 6.1 SolGPT 6 AstraClaude Opus 5.5Claude Sonnet 5.5
    Price per 1M (in / out)$2 / $10$10 / $50$4 / $20$2 / $10
    Intelligence Index (max)52535856
    Cost per Index task (max)$0.72$3.26$5.98Not found
    Best forLong agent runs, code review, computer use, PDF work, 3D web pagesThe hardest science and reasoning tasksTop-end coding and the highest index scoreBusiness workflows and motion design at Sol's price
    Watch out forMotion graphics, 272K token cliff, agent guardrailsPricePriceLower intelligence score than Opus

    Sources: prices from OpenAI, Opus 5.5 and Sonnet 5.5. Index scores from OfficeChai. Cost per task from Artificial Analysis and Kingy.ai.

    Here is my quick guide.

  • Pick Sol when you run long agent loops, review code, build 3D web pages or process lots of PDFs, and cost matters.
  • Pick Opus 5.5 when you want the best coding results and can pay double.
  • Pick Sonnet 5.5 when you want Sol's price with a higher index score, or when your job is motion design.
  • Pick Astra for the hardest science and reasoning problems, or as a reviewer before a big change ships.
  • Mixing models works well. Run Sol for the bulk of the work and call a stronger model only at the review step, as Vox's Codex setup does.

    Best Settings and Prompts for GPT 6.1 Sol

    Pick the Right Reasoning Effort

    Do not run everything at max. The data says more effort is not always better.

    EffortUse it forEvidence
    lowSimple edits, extraction, classificationOpenAI's replacement for the old none setting
    medium (default)General work and subagentsAbout 2 points below high on DeepSWE at over 30% lower cost per task, per Vox
    highYour main agent and hard debuggingOpenAI's best DeepSWE score (75.2%) came at high
    xhighCoding agents where mistakes are costlyBeat max by 3 points on Artificial Analysis's Coding Agent Index
    maxRare jobs where quality matters more than time and costOn the same index, xhigh beat max, so extra effort often adds cost without gain

    In ChatGPT and Codex you can also pick Ultra, which splits work across subagents. Use it when you can divide the task into meaningful parts.

    Call It From the API

    Use the Responses API, because OpenAI says tool calling needs it. Check the current docs for your SDK version.

    AI Prompt
    from openai import OpenAI
    
    client = OpenAI()
    
    response = client.responses.create(
        model="gpt-6.1-sol",
        reasoning={"effort": "high"},
        input="Review this pull request and list the three riskiest changes.",
    )
    
    print(response.output_text)

    Migrating From GPT 6 Sol

  • Change any none or minimal effort setting to low.
  • Expect 10% to 30% more output tokens at the same effort.
  • Keep requests under 272K input tokens to avoid the price step.
  • Turn on prompt caching for any repeated system prompt, repo or docs, because cached input costs only $0.10 per 1M tokens.
  • A Build Prompt Template That Works

    This template comes from what I learned in my tests. The last line targets the misses I saw, like the marble run that skipped a feature I asked for.

    AI Prompt
    Role: You are a senior product designer and front-end engineer.
    Goal: [what to build, in one sentence]
    Users: [who will use it]
    Must have: [3 to 6 features]
    Must not have: [what to avoid]
    Design: [mood, colors, fonts]
    Tech: single-page app, no build step, CDN libraries only.
    Quality bar: works on first load, seeded demo data, responsive, no console errors.
    Before you code, list your plan in 5 bullets.
    After you code, list any requested feature you did not build.

    Who Should Use GPT 6.1 Sol

    Use it if you:

  • Run coding agents or long Codex sessions and watch your usage limits.
  • Review code, analyze architecture or automate computer tasks.
  • Process long PDFs, reports and business workflows.
  • Build scroll-driven 3D landing pages and interactive web apps.
  • Want near-Astra results without Astra's bill.
  • Look elsewhere if you:

  • Need the strongest motion graphics. Sonnet 5.5 beat Sol on my reel.
  • Work on the hardest science tasks. Astra still leads there.
  • Regularly send more than 272K tokens in one request.
  • Need Sol inside regular ChatGPT Chat or need Sol Ultrafast today. Neither is ready yet.
  • Frequently Asked Questions (FAQs)

    Is GPT 6.1 Sol free?

    No. You need a paid ChatGPT plan (Plus, Pro, Business, Enterprise or Edu) to use it in ChatGPT Work and Codex, or you can pay per token through the API at $2 per 1M input tokens and $10 per 1M output tokens.

    What is the difference between GPT 6.1 Sol and GPT 6 Sol?

    The list price is the same, but 6.1 is smarter and cheaper to cache. It gains 4 points on the Artificial Analysis index, cuts cached input from $0.20 to $0.10 per 1M tokens and lowers the factual error rate on hard prompts from 11.4% to 7.7% at low effort. It also drops support for none and minimal effort.

    Is GPT 6.1 Sol as good as GPT 6 Astra?

    It comes close on coding and computer use. OpenAI says it matches Astra on DeepSWE at about one fifth of the cost, and Artificial Analysis puts it 1 point behind on its index. Astra still wins on the hardest science tasks and has lower deception numbers in the system card.

    Is GPT 6.1 Sol better than Claude Opus 5.5 for coding?

    It depends on the task and your budget. Sol costs half as much and beat Opus 5.5 on some OpenAI-picked tests and in Paweł Huryn's bug hunt. Opus 5.5 scores higher on the Intelligence Index (58 vs 52), and Theo still defaults to Opus 5.5 for code.

    How big is the context window, and does the price change?

    Sol has a 1,050,000 token context window and up to 128,000 tokens of output. Once a request passes 272K input tokens, you pay 2x for input and cache and 1.5x for output on the whole request.

    Can I use GPT 6.1 Sol in regular ChatGPT Chat, and when does Ultrafast arrive?

    Not yet for Chat. Sol is in ChatGPT Work and Codex today. OpenAI said Sol Ultrafast would arrive in the coming days, while its model docs say support is coming later, so check the docs before you plan around it.

    Final Thoughts

    GPT 6.1 Sol is the best value model in the GPT 6 lineup. It scored 4.2 out of 5 across my five builds, including a perfect score for the 3D watch landing page, and builders on X report huge savings in Codex. It is not the smartest model on the market, and it was weak at motion graphics, so pick Opus 5.5, Sonnet 5.5 or Astra when those strengths matter.

    Your next step is simple. Copy the watch landing page prompt above, run it on Sol at high effort, and compare the result and the cost with the model you use today. For more prompts like these, grab them at promptslove.com.

    Share this article
    Ramanpal Singh

    Ramanpal Singh

    Ramanpal Singh Is the founder of Promptslove, kwebby and copyrocket ai. He has 10+ years of experience in web development and web marketing specialized in SEO. He has his own youtube channel and active on social media platform.