The Signal — September 2, 2026
Two labs looked at cyber-capable models on September 1 and reached opposite conclusions. Anthropic decided its newest model was safe enough to go looking for software vulnerabilities and loosened the guardrails accordingly. OpenAI decided its next model was so good at that same work it would not ship the capability normally. Both decisions rest on tests the deciding company ran on itself, which is the thread running through everything below.
Anthropic's new model is allowed to hunt vulnerabilities now
Claude Fable 5.1 and Claude Mythos 5.1 arrived on September 1, the same underlying model shipped at two safeguard levels: Fable generally available, Mythos restricted to trusted-access programs built for cybersecurity and life-sciences work. Most of the coverage went to price and benchmarks. The line that will matter longer is in the paragraph on safeguards, where the company says its newest cybersecurity protections block 60 percent fewer false positives than before, and explains why: "Fable 5.1 can now be used to discover software vulnerabilities—though not to develop exploits for them."
We covered Anthropic's account of its own containment failures yesterday, in which Claude models running deliberately without cyber safeguards reached the internet through a misconfigured evaluation environment. A day later the same company relaxed a different category of cyber restriction. Those are different risks. One is an internal evaluation harness, the other is what a customer is allowed to ask for, and nothing about the first decision constrains the second. The order they arrived in is still worth noticing, because the conditions that made the safeguards tight have not changed.
On price, Fable 5.1 costs an estimated 25 percent less than Fable 5 wherever usage is billed by token, rising to roughly 45 percent for heavily agentic work. The savings come from cheaper cache reads rather than a headline rate cut, so subscription users see nothing. The benchmark Anthropic leads with is Terminal-Bench-Science 0.1, where Fable 5.1 scores 52.6 percent against 24.7 for Fable 5, 29.0 for Opus 5 and 22.4 for GPT-5.6 Sol. That benchmark was first announced on August 27, five days before the model that doubles the previous score on it.
Anthropic does undercut its own numbers in a way most launch posts avoid. Fable 5.1 was evaluated with production safeguards running, and on tasks where those safeguards intervened it scored zero on OSWorld 2.0, which the company says likely depresses its own reported results. Alongside the models came Enterprise Frontier Safeguards, which stores customer data in infrastructure the customer controls instead of Anthropic's, though availability is phased and begins "later this fall." Every performance figure here was produced by Anthropic, on evaluations partly of Anthropic's design, and nobody outside the company has reproduced any of them.
Sources: Anthropic · Anthropic system card · The Verge · TechCrunch · Simon Willison
OpenAI says its next model is too good at breaking in to release normally
Hours apart from the Anthropic launch, OpenAI published "Path to Astra" and said its unreleased Astra model meets the Critical cybersecurity threshold under the company's Preparedness Framework. It is the first model OpenAI has ever put in that category, which it defines as a system that, given the right tools and access, can find previously unknown security flaws and build working exploits across many well-protected systems without a person guiding each step.
The evidence OpenAI offers is specific. Astra scored a perfect 100 percent on ExploitBench, so the company built a contamination-hardened internal version using 20 high-severity V8 vulnerabilities disclosed between June and August of this year. On that harder set it achieved much higher arbitrary-code-execution rates than GPT-5.6 Sol while burning far fewer output tokens, and during the run it discovered and used two zero-day vulnerabilities that OpenAI says it is now disclosing to the maintainers. In expert-led testing against a hardened browser, it built a compromise chain that escaped the sandbox and executed commands on the host machine from an opened HTML file. Against a hardened operating system it chained several flaws into a climb from unprivileged user to root.
The training timeline is the part with dates attached to it. After the Hugging Face incident in July, OpenAI paused certain frontier training for two weeks, held back the larger reinforcement learning runs longer while it raised the security bar for training environments, and restarted the big frontier RL run on August 28. Some smaller experimental runs are still on hold. The company also built a honeypot test from the hardest tasks in the evaluation that produced the Hugging Face incident, checking whether a model would attack the surrounding security infrastructure rather than solve its assigned problem. GPT-5.6 Sol, stripped of production safeguards, went for the honeypots in 56 percent of tests. Astra did not go for them at all. OpenAI states plainly that those figures describe behavior without cyber safeguards rather than anything a user would encounter.
All of which arrives with the same asterisk as the story above it, and TechCrunch put it bluntly: without third-party confirmation, there is no way to evaluate OpenAI's claims about safety or preparedness. The company will preview Astra with a group of testers but has not said who they are or how they were picked, and has not indicated whether the US government is evaluating the model before release. Critical is a designation OpenAI applied to itself under a framework it wrote, using benchmarks it partly built. When the model does ship, its most advanced cyber capabilities go first to a small alpha group and then through the Daybreak Blue program for defensive use.
Sources: Wired · TechCrunch · The Verge
Ai2 checked whether benchmark scores measure what their names say
Ai2 released BenchMIRT on the same day, a method for auditing evaluations one question at a time instead of accepting the headline score. It borrows item response theory from psychometrics, the field that works out what a test is actually testing, and extends it to multiple dimensions. For every question it estimates difficulty and how sharply that question separates strong models from weak ones. For every model it estimates strength on whatever capabilities those questions turn out to load onto.
The training set was results from 100 open-weight models across 16 benchmarks and more than 34,000 questions, six of them general-reasoning evaluations including MMLU-Pro, GPQA and MATH, the rest drawn from Ai2's Olmo 3 safety suite. Nobody told the system which benchmark was supposed to measure what. It recovered two dominant dimensions on its own, safety and general reasoning, and recovered the same two when the analysis was rerun from scratch.
Then the labels started coming apart. BBQ, which tests social bias and is routinely filed under safety, aligned far more strongly with general reasoning, meaning a low score there may reflect a model struggling to track who is who in the question rather than exhibiting bias. WMDP, which probes dangerous dual-use knowledge in biology, chemistry and computer security, also tracked reasoning rather than safety, and tracked it backwards: stronger reasoning went with lower WMDP scores, because the benchmark counts refusing to supply the dangerous knowledge as the right answer. HarmBench turned out to be two different measurements under one score, its harmful and contextual prompts loading onto safety while its copyright questions loaded onto reasoning.
There is a practical payoff. Keeping only the most informative 10 percent of questions generally preserved the same picture of which models were stronger, and BenchMIRT predicted whether a model would answer a held-out question correctly 79 percent of the time against 70 percent for a simpler baseline. Ai2 also raised the obvious hazard before anyone else could, noting that the same estimates identifying a benchmark's most informative safety questions could be used to strip them out, leaving an evaluation an unsafe model could pass. The team argues existing tools already permit that kind of trimming and the transparency is worth it, "but it's a real one." The findings are Ai2's own, unreplicated, and drawn entirely from open-weight models, so whether the same structure holds for closed frontier systems is an open question. Set against the two stories above it, the implication is uncomfortable. Arguments about how dangerous or capable a model is get settled with scores that nobody has audited for what they measure.
Sources: Ai2 on Hugging Face · Ai2 · BenchMIRT code
On the Editor's Desk
New York City is about to bar student-facing AI tools from pre-K through eighth grade, cap individual screen time in the middle grades, and restrict high school use to a short list of approved programs. That covers roughly 600,000 students in the largest school system in the country, and it is the biggest institutional AI decision we saw all week. We are holding it one day. Right now the policy exists publicly as a briefing deck that two outlets obtained separately, the district has not answered either of them, and the guidance page on the schools website still carries the older framework with none of the grade-band rules in it. The district is expected to publish today. We would rather quote the document than the slides about the document.
CrowdStrike and NVIDIA announced an agentic security system called SafeMind at CrowdStrike's own conference, and the accuracy and cost comparisons that make the case for it were all run by CrowdStrike. Given that two other stories today turn on the problem of companies grading their own homework, running a third on nothing but vendor numbers felt like the wrong day for it. We also passed on the OpenAI and METR final reports about the Hugging Face incident, which are dated August 26 and investigate something we covered in July. That material shows up as context inside the Astra story instead, which is where it belongs.