- Smarter with AI
- Posts
- SunBrief#96: ChatGPT Drops Its Closest Model to AGI
SunBrief#96: ChatGPT Drops Its Closest Model to AGI
Google boosts Gemini coding, Meta challenges GPT 5.6 Sol, and Nvidia buys Hugging Face for $12.9B

Welcome to the SunBrief
Today in SunBrief 🌞
Build a Company OS in Notion
OpenAI Releases GPT -6 Astra as Brockman Declares the “AGI Era”
Stock Updates
Google Launches Gemini 3.8 Flash With Stronger Coding at the Same Price
Meta’s Muse Spark 1.3 Matches GPT 5.6 Sol at a Lower Benchmark Cost
AI Highlights of the Week
Too Important to Miss
Stop Being the Human Glue: Build a Company OS in Notion
In the beginning, the founder often is the system. They answer every question, approve every decision, remember every customer detail, and connect every moving part. That level of involvement may feel necessary early on, but over time it turns the founder into the company’s biggest bottleneck. A lightweight company operating system helps replace constant interruptions with shared context. Instead of asking where the latest deck lives, what was decided in last week’s meeting, or who owns a launch, the team can self-serve from one central workspace. Weekly priorities, project owners, meeting notes, onboarding resources, customer context, and key decisions all have a clear home. The result is not just better organization. It is more autonomy. When the team knows where to find information and how work moves forward, founders spend less time repeating themselves and more time making high-leverage decisions. A clear operating system gives early-stage companies the structure they need without slowing them down. Spend less time answering repeat questions and more time building.
OpenAI Releases GPT 6 Astra as Brockman Declares the “AGI Era”
Stronger computer skills and standout benchmarks arrive with better memory, premium pricing, and tighter safety limits
OpenAI has unveiled GPT 6 Astra, putting computer use at the centre of the release with tasks such as building spreadsheets, filling tax forms, and laying out circuit boards. Trained on more than 100,000 GPUs, Greg Brockman marked its arrival by declaring: “Welcome to the AGI era.”
Key Points:
Computer Work Gets Faster: Astra handled desktop tasks with better results in 47% less time than Sol in OpenAI’s OSWorld simulations.
Big Scores, Important Context: Astra reached 97.6% on FrontierMath Tier 4 and 99.9% on ARC AGI 3 with OpenAI’s adapter, versus 62.7% using the standard setup.
Less Repeating Yourself: Experimental Codex memory saves project notes and makes earlier conversations searchable, helping Astra recover requirements and test results during longer assignments.
Power Comes With Restrictions: Astra scored 100% on ExploitBench without safeguards; its Critical cyber rating brings tighter controls, despite monitoring weaknesses found in adversarial tests.
A Higher API Bill: Pricing is $10 input and $50 output per million tokens, 2.5 times Sol’s promotional rates, with Fast mode costing double.
Access Opens Gradually: Selected organizations go first, followed by Plus, Pro, Business, and Enterprise users within existing allowances, alongside API, Azure, and Bedrock access.
Why It Matters:
Astra makes the case for judging AI by the work it finishes, not just the answers it gives. For businesses, the useful measure will be cost per completed task, including the time people spend reviewing results, rather than the price of tokens alone.
Does GPT 6 Astra feel like the beginning of the “AGI era” to you? |
Stock Updates

Google Launches Gemini 3.8 Flash With Stronger Coding at the Same Price
The update tackles tougher software and business tasks, but extra reasoning can still mean a bigger bill
Google has released Gemini 3.8 Flash, focusing on software projects and business tasks that take more than a quick answer. Its launch announcement keeps 3.7’s introductory pricing, while the model card flags an important tradeoff: the model works harder on difficult requests, sometimes using more tokens to finish them.
Key Points:
Tools to Get Work Done: Agents can search, run code and navigate computer interfaces rather than just suggest steps, though computer control remains in preview.
Stronger Benchmark Results: Google reports 61.4% on Finance Agent v2 and 54.9% on HLE Verified, up from 3.7’s 59% and 53.6%, respectively.
Room for Larger Projects: It retains a 1 million token context window for text, images, audio and video, with text responses up to 64K tokens.
Same Introductory Rates: API pricing stays at $0.75 input and $3.75 output per million tokens through December 2026, before doubling in January.
Safety Results Are Mixed: Automated tests found weaker safety outside English, although Google’s human testing found no major concerns and met child safety thresholds.
Available Across Google: Developers get API, AI Studio, Antigravity and enterprise access; Google AI Pro and Ultra subscribers can use it in Gemini and Search.
Why It Matters:
The appeal is getting more useful work from an affordable model rather than paying flagship rates for every task. But a low token price is only part of the equation: businesses still need to measure how much reasoning, waiting and human correction each finished job requires.
Would you choose Gemini 3.8 Flash over a flagship model for everyday work? |
Meta’s Muse Spark 1.3 Matches GPT 5.6 Sol at a Lower Benchmark Cost
Independent tests put Meta alongside OpenAI, while its strongest reasoning mode remains in limited preview
Meta has released Muse Spark 1.3, giving developers another reason to look beyond OpenAI and Anthropic. Independent testing puts its available version level with GPT 5.6 Sol on an overall intelligence benchmark, with a lower estimated cost to complete the tests.
Key Points:
How It Compares: Spark 1.3 at xhigh scores 61, matching Sol at max effort, while Claude Opus 5 at max leads with 63.
A Smaller Benchmark Bill: Artificial Analysis estimates $0.55 per task, compared with $0.95 for Sol, despite Meta keeping its token prices unchanged.
Less Coding Overhead: Meta’s internal comparisons found roughly 20% fewer tool calls and 25% fewer tokens than Spark 1.2 on coding work.
Better Judgment During Projects: Meta says it follows detailed instructions more reliably, asks for help when stuck, and confirms before consequential actions.
Check Which Version You Get: Muse Code and API access are live, but max reasoning remains in limited preview while additional safety testing continues.
Why It Matters:
Meta does not need to beat every rival to give developers a reason to switch; matching useful performance at a lower cost can be enough. That puts pressure on premium models to justify their prices with better results on actual jobs, not just leaderboard positions.
What does Muse Spark 1.3 signal about the AI race? |
AI Highlights of the Week
Nvidia Agrees to Buy Hugging Face for $12.9B
Nvidia has agreed to acquire Hugging Face for $12.93 billion, gaining one of the world’s biggest platforms for open AI models and datasets.
Hugging Face will remain an open platform, with developers free to use any models, clouds, or hardware without requiring Nvidia chips.
Google Launches WeatherNext 3 for More Accurate Forecasts
Google introduced WeatherNext 3, its most advanced global weather AI model, using real-time satellite data to generate new forecasts every hour.
It delivers forecasts up to 5× sharper than WeatherNext 2 and is now powering weather across Search, Gemini, Maps, and Google Cloud
AI Just Tackled a 350-Year-Old Math Problem
Anthropic says Claude produced a complete, computer-checked formalization of Fermat’s Last Theorem in just 11 days.
The result spans 13 million lines of Lean code, turning the famous proof into something a computer can verify step by step.
Runway Introduces Solaris for AI Generated Apps
Runway unveiled Solaris, a new Interface World Model that can generate interactive apps and websites in real time as users click, drag, and type.
Instead of relying on fixed screens and coded interactions, Solaris renders the interface frame by frame, letting software adapt continuously to what the user does.
Too Important to Miss
Last Week’s Poll Result
Would you trust an AI agent to operate physical equipment without constant human supervision?
Yes, if safeguards are proven → 40.00%
Maybe, for low-risk equipment → 26.67%
No, human oversight should always be required → 33.33%

Is AI video finally becoming a real editing tool instead of just a generation tool?
Yes, clearly → 100.00%
Would you feel comfortable with AI assisting surgeons during a brain operation?
Yes, if surgeons remain in control → 80.00%
No, the risks feel too high → 20.00%
Feedback
We’d love to hear from you!How did you feel about today's SunBrief? Your feedback helps us improve and deliver the best possible content. |
Know someone who may be interested?
And that's a wrap on today’s SunBrief!




Reply