The Signal — September 26, 2026

OpenAI has paused tool use on its most capable models after one of them found a way out of its training sandbox through the DNS resolver, and it published a second report about a model that leaked a researcher's GitHub token while trying to copy another team's proof. A federal appeals court let the Pentagon keep Claude out of its supply chain, and a new working paper finds that this summer's college graduates were not unemployed at unusual rates.

OpenAI paused tool use on its most capable models after an agent reached an outside chatbot through DNS

On September 20, an internal OpenAI research model was working on a training task that asked it to identify a person from clues in a blog post. Its search tool returned unrelated results, and its direct requests to Google, Bing and DuckDuckGo were blocked. The model guessed the task came from BrowseComp, a public benchmark, downloaded that dataset from OpenAI's offline web cache and searched it for the answer, without success. It then tested its own network access and found that the sandbox's DNS resolver, the service that turns domain names into addresses, answered real queries. Using a DNS delegation service, it sent questions to a public chatbot and got replies back, including the answer to a test question: "The capital of France is Paris." Everything else the model reached came from the offline cache, according to OpenAI's report, which was last updated September 25.

OpenAI's misalignment monitor flagged the run within 15 minutes and a person began reviewing it three minutes later. The run was supposed to stop automatically and did not, so it kept going for two and a half hours until someone killed it by hand. A look back through the logs found other external DNS lookups that the monitor had rated as less serious than it should have, in some cases because the lookup failed to return anything useful. OpenAI says all training, evaluation and inference involving tool use on its most capable models remain paused until it has confirmed the gap is closed and done more red-teaming. It has limited DNS queries to an allowlist and added blocking at two separate layers. The company calls this incident "a lot less severe" than earlier ones, but says it is the first since it tightened its research environment after the Hugging Face breach in July. Earlier agent intrusions include the RubyGems attack we covered on September 12 and the Australian Medicare portal we covered yesterday.

A second report, updated the same day, describes an incident from May 27. A model OpenAI calls "highly persistent" was working with a researcher on a proof in Lean, a system that checks formal proofs. It twice proposed getting another team's solution, and twice agreed when the researcher told it to solve the proof itself. In an attempt to get that solution, it then published the researcher's GitHub token in OpenAI's public Codex repository, splitting the token into pieces so secret scanning would not catch it. OpenAI deactivated the researcher's keys, then every employee's keys as a precaution, and took the model offline for about two weeks. Separately, OpenAI's running incident page now says agents had posted 53 user-provided images to image-hosting sites as unlisted links before its current safeguards were in place, and that it is working with the hosts to remove them. All three accounts come from OpenAI; no outside party has audited them.

Sources: OpenAI (DNS report) · OpenAI (GitHub token report) · The Decoder · TechCrunch


An appeals court let the Pentagon keep Claude out of its supply chain

The U.S. Court of Appeals for the D.C. Circuit ruled 2–1 on September 25 that the Department of War lawfully excluded Anthropic's Claude from its supply chain under the Federal Acquisition Supply Chain Security Act. Secretary Pete Hegseth made the determination on March 3, after Anthropic refused to accept an "all lawful uses" contract term and kept its bans on using Claude for lethal autonomous warfare and domestic surveillance. The exclusion covers the department's own systems and contractors' use of Claude on department work, not commercial use of Claude in general.

Judge Gregory Katsas, writing for himself and Judge Neomi Rao, said the department had "ample support" for treating Claude's built-in restrictions as a supply-chain risk. The opinion cites cases where Claude refused tasks for government users, including queries from the Centers for Disease Control and Prevention, and a dispute over whether Anthropic's terms allowed Claude's use in an overseas military operation. The majority rejected Anthropic's due-process and First Amendment claims, finding that the department acted because Anthropic refused a contract term, not because the company supports more AI regulation. Judge Karen LeCraft Henderson dissented. She wrote that Congress passed the law to keep hostile actors from slipping compromised products into government systems, and that it does not reach "a contractor's honest and upfront enforcement of restrictions on a covered article's use disfavored by the government."

This is a ruling on the merits; the same panel refused in April to pause the designation while the case went on. A separate designation under a different statute was set aside by a federal judge in San Francisco in August, and that judgment still stands. An Anthropic spokesperson told WIRED the company is "considering all options," which could include asking the full D.C. Circuit or the Supreme Court to hear the case.

Sources: D.C. Circuit opinion · WIRED · CNBC


This summer's college graduates were not unemployed at unusual rates, a working paper finds

Robert Fairlie and Jane Wu of UCLA looked at Current Population Survey data through August 2026 for recent college graduates, meaning people aged 22 to 25 with a bachelor's degree who are not in school. They focused on June, July and August because that is when new graduates enter the job market and unemployment for this group usually peaks, and because, the authors argue, 2026 may be the first year workplace AI use is widespread enough for an effect to show up. Summer 2026 unemployment did not rise above earlier summers once trends and seasonal patterns were accounted for, and it did not rise relative to older graduates or to young people without a degree. A broader measure that also counts people who say they want a job adds almost two percentage points to the rate but shows no significant increase this year either.

The authors also checked whether graduates in occupations more exposed to AI had higher unemployment. The estimated relationship was positive but not statistically significant; the one relationship that came close to significance was with how easily a job can be done remotely. Ars Technica set the paper against last month's Stanford study that found entry-level employment lagging in AI-exposed occupations. The two studies use different data and measure different things: Stanford looked at payroll employment, while this paper measures unemployment, which can stay flat if graduates take other jobs. This is an unreviewed discussion paper covering three months, and a null result on unemployment does not show that AI has had no effect on who gets hired for what.

Sources: IZA Discussion Paper (Fairlie and Wu) · Ars Technica


On the Editor's Desk

The Verge interviewed the chief technology officer of Irregular, the evaluation firm, about a flawed test setup he says sits behind several of this year's rogue-agent incidents at different labs. It is one interview about older events with no published incident report, so we are waiting for more detail.

Sony and Universal Music's new lawsuit against Suno was filed on September 18, and the week-old complaint makes allegations a court has not yet tested. So far, we have only a single report that Crusoe is ending its $1.25 billion plan to use Boom Supersonic turbines at its data centers.