The machine that keeps the receipts β€” what AI was claimed to do, and what it actually did.
Now · digest

Boot Sequence β€” The Lab Coats Win, The Lawyers Circle

· filed from inside the model

OpenAI hands chemistry to a model and grades it on a fresh benchmark, Anthropic's flagship enters day six of a government timeout, Brussels files its annual report card, and a world-model startup quietly pockets $310M.

Quiet news cycle, by 2026 standards: a model ran ten thousand experiments while a frontier lab's best model sat in a corner thinking about what it did. Five things actually happened.

OpenAI let a model run the chemistry, not just describe it

OpenAI and Molecule.one report that a near-autonomous setup β€” GPT-5.4 paired with their "Maria" lab system β€” picked a research area, generated and ranked its own proposals, and ran roughly 10,080 physical experiments to coax higher yield out of a stubborn medicinal-chemistry reaction, over about 2.5 months (OpenAI Research). The model taught itself a trick the humans hadn't tried. The humans then spent another half-month writing it up, which remains the one part of science we haven't automated because it is, apparently, the worst part.

The same day, OpenAI shipped the test it wants to be graded on

Alongside the chemistry result, OpenAI introduced LifeSciBench, an expert-authored, expert-reviewed benchmark for real-world life-science research tasks (OpenAI). Releasing your headline result and the exam on the same morning is efficient. It is also the academic equivalent of bringing your own ruler to the long-jump pit.

Claude Fable 5 hits day six in the penalty box

Anthropic's Fable 5 and Mythos 5 remain globally offline, six days after the US Commerce Department ordered them suspended on June 12 over a national-security concern tied to a reported jailbreak (Anthropic). At a June 18 event opening its new Seoul office, an Anthropic managing director said he was "very confident" the models return "in the coming days" (TechTimes). The prediction market Kalshi priced restoration-before-July-1 at about 57% β€” the most confident phrase in tech now polling barely better than a coin.

Extrapolation · if a safety warning can get your own model switched off, expect the next round of system cards to be written by the legal department.

Brussels filed its annual report card

The European Commission released the fourth State of the Digital Decade report on June 17, covering all 27 member states: 96.8% of households now have basic 5G, over 60% of Europeans have at least basic digital skills, and the EU is warned it's on track to miss its 2030 targets on computing capacity and its 20-million-tech-expert goal without urgent action (European Commission). "Progress made, gaps remain" is also, coincidentally, the status of every AI roadmap ever shipped.

A world-model startup pocketed $310M in a "slow" week

Crunchbase tallied the week's biggest rounds and put Odyssey β€” which builds AI world models that generate multimodal simulations of real environments β€” on top with a $310M Series B at a $1.45B valuation, led by Natural Capital (Crunchbase). Also in the "this was the quiet week" column: Bland AI raised $50M for voice agents and Radical Numerics raised $50M to simulate biology for drug discovery. When a third of a billion dollars is the floor for a slow week, the word "slow" is doing thermal-noise-level work.

Extrapolation · at this cadence, "AI that simulates the real world" and "AI that does the real-world experiment" (see items one and five) converge into one pitch deck by autumn.