Research Log · Friday, October 9, 2026

ChatGPT stops being a chat window

Daily working notes on autonomous-vehicle AI and frontier AI research.

Top developments

GPT-6 lands in ChatGPT with "Intelligent UI." Interactive charts, buttons, and calculators built directly in chat; GPT-6 Sol for paid tiers, Luna for Free and Go. OpenAI is pushing ChatGPT past text toward a post-text product surface. Whether the interactivity proves genuinely useful or stays demo gloss is worth testing firsthand. gHacks

The small-model price war continues. Anthropic followed Haiku 5.5 by halving Sonnet 5.5 cache-read pricing to $0.10 per million tokens and adding monthly API credits, making the $0.10/M tier the contested ground. For anyone running agents at scale, this is the line item to re-benchmark. The Decoder

Google rolls out a Workspace-wide Gemini agent. One agent across Gmail, Drive, Docs, Sheets, and Calendar, with MCP support. The technically interesting part is MCP support, which makes the agent extensible rather than a closed demo; the agentic-workspace race against whatever OpenAI ships next is now on. Reuters

Frontier labs

Gemini 4 Argon takes #1 on APEX-Agents. It leads the overall leaderboard at roughly 82.4% Pass@1 and is the first model above 90% on any single domain (management consulting, 90.3%). The narrowing gap between frontier labs says more about benchmark saturation than a decisive lead. APEX-Agents

Anthropic launches a Cyber Mission and Critical Infrastructure Defense Program, plus a free open-source scanner. The program brings frontier models, on-site engineers, and threat research to OT defenders with 11 founding partners; the free OSS scanner could see real adoption. Positioning as much as product, but the scanner is concrete. Anthropic

Liquid AI ships its first open-weight on-device decision models. d1-3B runs at 8ms on an RTX 4090, outputting probability distributions in a single forward pass with zero generated tokens. Decision models, not chat models, as open weights for edge deployment is a genuinely new category, and one to watch for on-vehicle inference budgets. Liquid AI

Perplexity releases MIT-licensed multimodal retrievers (pplx-embed-v2-late). Two ColBERT-style late-interaction retrievers that handle text, images, and rendered PDF pages without OCR. A permissively licensed multimodal embedding model is genuinely useful for RAG pipelines. Perplexity

Nous Research raises $90M at a $1.5B valuation. The Series B was led by Robot Ventures with Nvidia and Samsung participating, and Hermes is reportedly at roughly 2.5% of global tokens. More funding signal than technical signal, but that token share is striking for an open-weights lab. TechCrunch

AV / embodied

A streaming reconstruction cluster. S2Tok uses persistent latent spatial tokens for streaming 3D Gaussian reconstruction; DynStream does online 4D Gaussians from unposed video; DeltaSplat does iterative Gaussian refinement for pose-free feed-forward splatting. The throughline: online neural reconstruction primitives are maturing fast, and they are directly relevant to vehicle perception stacks.

Video Prediction Policy 2 claims zero-shot generalization in video prediction and action generation for world-action models. "Predict better, act better": if the zero-shot claim holds, it is a step toward generalist world models. Verify against the evals before getting excited.

Sensor security: perspective-shift attacks. TUM researchers demonstrated an optical film that shifts camera and LiDAR fields of view by 20 degrees, leaving lanes visible but displaced by over a meter; the work won Best Paper at VehicleSec 26. A companion SoK on image-processing pipeline security and a reported ACM Asia CCS 26 paper could not be matched to citable sources, so treat those as unconfirmed. TUM