The Signal — August 18, 2026

Two of the biggest labs spent this week explaining, in very different registers, how their own safety machinery came apart. One dissolved the team that watched for catastrophic risk. The other published a report saying a filter it considers essential had been switched off for eleven months without anyone noticing.

OpenAI disbanded the team that checked for catastrophic risk

The Financial Times reported that OpenAI dissolved its Preparedness team at the end of July. That group existed to answer one question before a model shipped: could this thing meaningfully help someone build a biological weapon, run a serious cyberattack, or slip out of human control? Its work has been handed off to senior staff on other teams, split up by domain, with bio going one way and cyber another. OpenAI calls the change part of a streamlining process ahead of a possible IPO.

Worth being precise about what did and did not change. The Preparedness Framework, the published policy that gates releases on capability thresholds, is still in force. We wrote about it twice this month: OpenAI paused internal work on a model called Astra after it tripped the "critical" cyber threshold, and rated GPT-5.6-Cyber "High" but under the line. The framework doing real work in August is exactly why the team's disappearance in July is worth noticing. A policy is only as good as the people whose full-time job is enforcing it, and that job now belongs to people who have other jobs.

This is the third safety-focused team OpenAI has wound down in about two years, after AGI Readiness and Superalignment, and it arrives alongside a run of departures that includes head of systems safety Johannes Heidecke and chief futurist Joshua Achiam. OpenAI has not published anything about the change itself, so everything here rests on FT's sourcing and the outlets that followed it. Distributed responsibility can work. It is also the arrangement where nobody owns the bad news.

Sources: Engadget · Calcalist · The Decoder


Anthropic's bioweapon filters were off for eleven months

Anthropic's second Risk Report, published Friday, contains a disclosure the company clearly did not enjoy writing. A flag meant for internal debugging was left in a state that disabled its blocking chemical and biological weapons classifiers on all traffic through its human-feedback data-collection platforms. That state held from May 2025, when Anthropic first deployed models carrying those safeguards, until April 2026. Roughly 133 million exchanges from about 50,000 contractors ran through with the filters down.

The same flag also disabled logging, which is the part that stings. Anthropic says its review found no evidence of concerning misuse, and there is no reason to doubt the review as far as it goes, but the traffic that would have been flagged was never recorded, so the search is working from a thinner record than anyone would want. Those 50,000 contractors were vetted by outside vendors whose screening, in Anthropic's own description, was often not up to stopping a serious biological threat actor.

Anthropic has remediated the flag, tightened contractor requirements, added offline monitoring, and raised its own risk rating on non-novel bioweapons from "very low" to "low," a second upgrade alongside a misalignment bump tied to a UK AI Safety Institute evaluation incident. The most useful sentence in the report is the company's own: discovering a gap like this one makes it more likely that other, similar gaps exist that they do not know about yet. A safety system that can be silently switched off by a debugging flag for eleven months is telling you something about how it is monitored, not just about the one flag.

Sources: Anthropic Risk Report · The Decoder · The Next Web


Google bought a dead airline's internal data for $10 million

Spirit Airlines failed to emerge from its second Chapter 11 this year and stopped flying, and its attorneys have been selling off what remains. One of the assets was the company's internal business data: emails, chats, calendars, documents, spreadsheets, operations records, and custom software code accumulated over decades of running an airline. Google won it at auction for $10 million, outbidding the AI startup Mercor, which had countered at $5.2 million and offered $7 million for a version where it received the raw data and did the anonymizing itself.

Google says it is buying internal data and code, not customer or credit card information, and that a third party will scrub personally identifiable information before Google receives anything. The buyer also agrees not to try to re-identify anyone. A federal judge still has to sign off, so this is not final. Those are meaningful guardrails and they are worth stating plainly rather than skipping past.

The mechanism is the part to watch. Frontier labs have largely exhausted the easy public web, and what they want now is the messy interior record of how organizations actually operate: the half-finished threads, the escalations, the internal tooling nobody documented, the decisions that never made it into a policy doc. A bankruptcy estate is a seller with a fiduciary duty to get the best price and no particular reason to weigh anything else. The thousands of people who wrote those emails were not parties to the auction and had no say in it. Expect more of these, because the incentives all point the same direction.

Sources: Axios · Business Insider · Techmeme


On the Editor's Desk

We looked hard at Alibaba's Qwen 3.8 27B, an Apache-2.0 vision-capable model that fits in about 17GB and runs on a decent laptop, with a useful hands-on writeup from Simon Willison noting it overthinks simple prompts. It is a good story and it may show up here soon; it just lost out to three items with more at stake this week. We skipped the new details on Claude's watermarking because we covered that rollout on August 11 and this was a follow-up rather than news. IEEE Spectrum's report on Axiom Math machine-checking the "246 theorem" is a real result, but the significance claims trace back to people affiliated with Axiom and we could not find independent corroboration, so it waits.