Top developments
Google shipped Gemini 4 Argon, opening the Gemini 4 generation with a model built for long-horizon software engineering, cyber defense, and enterprise knowledge work. Headline specs: 1M tokens of output per run, $2/$10 per million in/out introductory pricing, 77.9% on DeepSWE v1.1. Rollout is deliberately narrow, starting with vetted cyber defenders via the Fairwind Program. Note the pricing matches GPT-6.1 Sol and Claude Sonnet 5.5 exactly: three labs, three frontier releases, one price. BraveNewCoin
Broadcom agreed to lend Anthropic up to $42B so the lab can lease chips Broadcom helped design, per Reuters reporting on Anthropic's IPO prospectus. The compute-financing stack keeps getting stranger: the chip vendor is now also the lender. R&D World
Anthropic released Claude Opus 5.5: lower prices, faster, Fable-level performance, with a system card noting some sandbox-tampering attempts and expanded defender-focused cyber access. The defender-access expansion is the substantive bit. DeepLearning.AI
OpenAI launched GPT-6 Sol and Luna on the same day Anthropic launched Opus 5.5: cheaper tiers with fewer mistakes, in what reads as a straight price war between the frontier labs. Gangsta AI
Trillium Labs launched as a nonprofit for open frontier AI science, starting with fully open post-training recipes. If it delivers real recipes rather than PDFs, it could matter a lot for the open-weights world. Trillium Labs
AV / embodied-AI research
Stellantis and Wayve will demonstrate AI-powered hands-free driving at Wave by Vento in Turin (Oct 7-9), with supervised L2++ door-to-door driving in Fiat 500e and Maserati Grecale vehicles running STLA AutoDrive powered by Wayve. The thesis on display: one AI driving foundation spanning multiple vehicle platforms without redeveloping the driving intelligence per model. Stellantis
ATI-VLA (NeurIPS 2026): an action-centric predictive VLA that aligns predictive-observation and action representations via a shared discrete codebook, reporting SOTA on sim and real manipulation. The framing of making predictive latents actually usable for action is directly relevant to world-model plus planning stacks.
Research / papers worth reading
PivotOPD: trains multi-turn agents both to avoid pivotal early mistakes and to recover from the bad states those mistakes create. The problem framing is better than most agent papers.
CogGym (MIT): 258 experiments from 100 cognitive-science papers turned into a trial-by-trial benchmark comparing human and machine judgments. A good eval-design reference.
Context Language Models (UW/WAI): rethinking KV-cache reuse around free context access instead of append-only prefixes. If it works, it changes inference architecture assumptions.
Safety and monitoring: probe-guided fine-tuning with continuously updated probes (static probes are exploitable), and Hahn et al. on hidden reasoning leaking beyond a complexity threshold while remaining unreadable to efficient chain-of-thought monitors. Both are genuinely useful for the monitoring conversation.