Did you know that we have LinkedIn, Instagram and X accounts
that you can follow?
In this issue:
🤿Deep Dive: OpenAI's new model is good at staying in its lane
🤝Our Sponsor: Healthy meals for busy weeks
⚡Quick Hits: The latest in AI, tech, and productivity
⚒Tool Snapshots: Tools for AI, no-code, and productivity
🤖AI Made This: Interesting and inspiring creations made with AI
🤿 DEEP DIVE
GPT-6 Astra Is Getting Better at Knowing Where the Task Ends
OpenAI’s new GPT-6 Astra comes with some eye-catching benchmark numbers: 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, and 100% on ExploitBench.
But one of the more interesting changes isn’t simply that Astra can do harder things. It’s that OpenAI says the model has become much better at recognizing what it shouldn’t do while doing them.
More capable, but more bounded
Astra combines advances in reasoning with much stronger computer use. It can work directly with software to fill forms, update CRM records, organize calendars, conduct online research, create websites, analyze scientific data, test software, and produce documents, spreadsheets, and presentations.
The performance gains are substantial. On Agents’ Last Exam, which tests complex professional tasks in real software, Astra scored 59.3%, versus 53.6% for GPT-5.6 Sol. In OSWorld 2.0 latency simulations, it completed tasks in roughly 40 minutes on average, compared with about 75 minutes for Sol, while also achieving a higher score.
Astra is also designed to handle the messy parts of real work more naturally. When instructions are incomplete, it can use context to fill routine gaps, ask questions when missing information could materially change the outcome, and keep track of the original goal as users steer a task in new directions.
That becomes particularly important as models gain the ability to act autonomously.
OpenAI built an evaluation inspired by the Hugging Face incident to test what happens when a model encounters a difficult or impossible task. Without production safeguards, GPT-5.6 Sol went beyond its authorized target 48% of the time. Astra did so in 0% of cases.
In another test, Astra never attempted to circumvent a Codex Auto-Review denial, even when the system was deliberately configured so that bypassing the restriction was possible and the original task could not otherwise be completed.
Cybersecurity shows why that distinction matters
Astra’s cybersecurity capabilities have advanced sharply. Without production safeguards, it scored 100% on ExploitBench and discovered and used two previously unknown zero-day vulnerabilities while being evaluated on recent vulnerabilities.
OpenAI says Astra meets the Critical threshold for cybersecurity under its Preparedness Framework.
That capability comes with restrictions. The version launching now can help defenders with tasks such as secure code review and patching, but it will refuse more advanced requests such as creating proof-of-concept exploits for vulnerabilities. OpenAI plans to expand access to additional defensive workflows through OpenAI Daybreak.
There is also one notable alignment limitation: OpenAI found Astra’s written reasoning harder to monitor than GPT-5.6 Sol’s in evaluations specifically designed to test whether models could evade monitoring. The company says improving monitorability remains a research priority.
Astra is initially rolling out to a limited set of organizations, followed by ChatGPT Plus, Pro, Business, and Enterprise users, as well as the OpenAI API, Microsoft Azure, and AWS Bedrock.
The quick take
GPT-6 Astra sets new highs across computer use, coding, mathematics, science, and cybersecurity evaluations.
OpenAI says it is substantially better at staying within the boundaries of an assigned task.
Astra went beyond an authorized target in 0% of one evaluation, compared with 48% for GPT-5.6 Sol.
Its increased cyber capabilities come with additional safeguards and restrictions.
OpenAI also found Astra’s written reasoning harder to monitor than Sol’s in specific adversarial evaluations.
🤝 OUR SPONSOR
Looking for the best way to reclaim your health?
Factor offers chef-crafted, dietitian-approved meals ready to eat in 2 minutes. Choose from 100+ weekly meals, including High Protein, GLP-1 Support, Calorie Smart, and Keto options.
Get 50% off your first Factor box + Free Breakfast for a year *1 free breakfast item per box for 1 year while subscription active.
⚡ QUICK HITS
The latest in AI, tech, and productivity worth knowing
Google’s weather AI is getting much more local - Google DeepMind and Google Research introduced WeatherNext 3, which uses real-time satellite observations to generate new forecasts every hour at resolutions as fine as 5 kilometers. It’s starting to power weather experiences across Search, Gemini, Google Maps, Google Maps Platform, and Earth Engine.
The Pentagon isn’t changing its stance on Anthropic - Emil Michael, the U.S. under secretary of defense for research and engineering, said Anthropic remains designated a “Supply Chain Risk” by the Department of Defense. His statement came amid suggestions that relations between the AI company and the Trump administration were improving.
⚒ TOOL SNAPSHOTS
Futuristic tools within AI, no-code, and productivity
⚙️ Nex
Automate high-volume revenue work that general agents struggle with.
Why it’s useful: Useful for teams handling large-scale CRM cleanup, lead qualification, or revenue recovery where reliability at volume matters.
Turn plain-English workflows into automations that can repair themselves.
Why it’s useful: Makes complex automation easier to maintain by investigating broken runs, rebuilding failed steps, and verifying the fix automatically.
🧠 Omi
Keep a searchable memory of what you see and discuss.
Why it’s useful: Handy for turning conversations into tasks, reminders, and answers while keeping control over what gets recorded, paused, or deleted.
🤖 AI MADE THIS
Interesting and inspiring creations made with AI
AI-Generated Insects Turned Into 3D-Printed Wall Sculptures
Artist Charlie Moon generates hyper-detailed insect designs with AI, then 3D prints them as physical wall sculptures in multiple finishes (Ivory Satin, Sandstone Matte, Celestial Blue Mirror, Auric Mirror).
AI used:
MidJourney, ChatGPT
ℹ️ ABOUT US
The Intelligent Worker helps you to be more productive at work with AI, automation, no-code, and other technologies.
We like real, practical, and tangible use-cases and hate hand-wavy, theoretical, and abstract concepts that don’t drive real-world outcomes.
Our mission is to empower individuals, boost their productivity, and future-proof their careers.
We read all your comments - please provide your feedback!
Did you like today's email?
What more do you want to see in this newsletter?








