Top developments
GPT-6 lands in ChatGPT with "Intelligent UI." Interactive charts, buttons, and calculators built directly in chat; GPT-6 Sol for paid tiers, Luna for Free and Go. OpenAI is pushing ChatGPT past text toward a post-text product surface. Whether the interactivity proves genuinely useful or stays demo gloss is worth testing firsthand. gHacks
The small-model price war continues. Anthropic followed Haiku 5.5 by halving Sonnet 5.5 cache-read pricing to $0.10 per million tokens and adding monthly API credits, making the $0.10/M tier the contested ground. For anyone running agents at scale, this is the line item to re-benchmark. The Decoder
Google rolls out a Workspace-wide Gemini agent. One agent across Gmail, Drive, Docs, Sheets, and Calendar, with MCP support. The technically interesting part is MCP support, which makes the agent extensible rather than a closed demo; the agentic-workspace race against whatever OpenAI ships next is now on. Reuters
Frontier labs
Gemini 4 Argon takes #1 on APEX-Agents. It leads the overall leaderboard at roughly 82.4% Pass@1 and is the first model above 90% on any single domain (management consulting, 90.3%). The narrowing gap between frontier labs says more about benchmark saturation than a decisive lead. APEX-Agents
Anthropic launches a Cyber Mission and Critical Infrastructure Defense Program, plus a free open-source scanner. The program brings frontier models, on-site engineers, and threat research to OT defenders with 11 founding partners; the free OSS scanner could see real adoption. Positioning as much as product, but the scanner is concrete. Anthropic
Liquid AI ships its first open-weight on-device decision models. d1-3B runs at 8ms on an RTX 4090, outputting probability distributions in a single forward pass with zero generated tokens. Decision models, not chat models, as open weights for edge deployment is a genuinely new category, and one to watch for on-vehicle inference budgets. Liquid AI
Perplexity releases MIT-licensed multimodal retrievers (pplx-embed-v2-late). Two ColBERT-style late-interaction retrievers that handle text, images, and rendered PDF pages without OCR. A permissively licensed multimodal embedding model is genuinely useful for RAG pipelines. Perplexity
Nous Research raises $90M at a $1.5B valuation. The Series B was led by Robot Ventures with Nvidia and Samsung participating, and Hermes is reportedly at roughly 2.5% of global tokens. More funding signal than technical signal, but that token share is striking for an open-weights lab. TechCrunch
AV / embodied
A streaming reconstruction cluster. S2Tok uses persistent latent spatial tokens for streaming 3D Gaussian reconstruction; DynStream does online 4D Gaussians from unposed video; DeltaSplat does iterative Gaussian refinement for pose-free feed-forward splatting. The throughline: online neural reconstruction primitives are maturing fast, and they are directly relevant to vehicle perception stacks.
Video Prediction Policy 2 claims zero-shot generalization in video prediction and action generation for world-action models. "Predict better, act better": if the zero-shot claim holds, it is a step toward generalist world models. Verify against the evals before getting excited.
Sensor security: perspective-shift attacks. TUM researchers demonstrated an optical film that shifts camera and LiDAR fields of view by 20 degrees, leaving lanes visible but displaced by over a meter; the work won Best Paper at VehicleSec 26. A companion SoK on image-processing pipeline security and a reported ACM Asia CCS 26 paper could not be matched to citable sources, so treat those as unconfirmed. TUM
Research worth reading
Video Prediction Policy 2, S2Tok, DynStream, DeltaSplat, and DecepEval, a benchmark for deception in LLM agents.