Top developments
OpenAI opened a public beta of its Decisions API: typed structured answers (predicates, choices, scores) via POST /v1/decisions, at $0.10 per million input tokens with free output tokens, plus HIPAA and zero-data-retention support. The company claims 10x the speed of GPT-6 Luna. This productizes the 'LLM as judge' pattern that has been duct-taped together in eval pipelines for two years; if the latency claim holds, it obsoletes a category of eval infrastructure.
Anthropic unlocked Claude for offensive penetration testing through a three-tier Cyber Verification Program for verified security professionals. Claude Opus 5.5 completed 34 of 50 offensive tasks in Anthropic's own testing. Gating offensive capability behind verification is the responsible way to release this, and 34/50 is a real marker for how far automated red-teaming has come.
Anthropic also shipped Claude Dashboards and Motion: live charts generated from BigQuery, Snowflake, and Salesforce, plus code-based editable animations with MP4 export. Dashboards is the enterprise wedge; Motion is a genuinely new artifact type. Both point the same direction: Claude as a working surface, not a chat window.
Frontier AI labs
OpenAI released GPT-6.1 Sol and halved Sol and Luna API prices. The frontier price war is now the industry's main competitive axis; the Pro 200 plan was reworked with lower limits, which reads as margin defense while API prices race down.
Google's Gemini 4 Argon debuted at number one on APEX-Agents at 82.2% Pass@1. The agentic benchmark crown keeps rotating between labs, which says more about benchmark saturation than about a durable lead.
Samsung Labs released LittleBit-2, pushing LLM compression to 0.1 to 1.0 bits per weight. If near-1-bit models stay usable, the on-device story changes completely; read the quality numbers before believing the headline.
OpenAI's Lean-verified Navier-Stokes proof is being challenged: researchers show the formalization proves a weaker statement than the English manuscript. Formal does not mean faithful, and this is a useful stress test for what 'verified' claims actually cover.
TokenRouter (arXiv 2610.12242) proposes token-level LLM routing as a serving system. Fine-grained routing is where the remaining cost and latency wins live.
A quantized GLM now runs at 71.8 tokens per second on two desktop DGX Sparks (model card). Frontier-class inference on consumer hardware keeps getting closer.
AV / embodied-AI research
DreamTrue (arXiv 2610.12468) introduces action-faithful multi-view robot world models with counterfactual post-training, taking rank 1 in the AgiBot World Challenge 2026 world-model track. Action-faithfulness is the right objective for world models used in planning; this one deserves a full read.
VersaCamVLA (arXiv 2610.12451), accepted to NeurIPS 2026, maps arbitrary posed RGB views to fixed latent scene tokens for camera-configurable VLA policies. Camera-agnostic policies would remove a large chunk of per-vehicle integration work.
SDPAD (arXiv 2610.11583) is the first fully spike-driven end-to-end driving planner: 0.40m L2 error and 86.3 PDMS on NAVSIM at under 2% of the energy of ANN baselines. Spiking networks for driving are still exotic, but that energy figure matters for onboard compute budgets.
WOVEN (arXiv 2610.12417) proposes visual transition reasoning as a training primitive and finds frontier multimodal models systematically deficient at it. 'Your model cannot do this basic thing' papers are the useful kind.
BridgeGuard adds safety-constrained diffusion planning with a learned distance field (+2.9 driving score on Bench2Drive), and PlanWAM shapes future representations with planning objectives (93.8 PDMS on NAVSIM). Both point the same way: shaping world-model latents with planning and safety objectives beats pure prediction.
Research / papers worth reading
Also worth your time: VGGTWorld-VLA (intent-conditioned 3D world evolution for driving), CoCam4D (Bayesian cooperative camera-only 4D perception with 35-byte object primitives for vehicle-to-everything communication), GLIO2 (GPU-parallelized LiDAR-inertial-GNSS at 25 Hz on a Jetson Orin NX), LiteNWM (latent navigation world model with 128x end-to-end speedup), PLaW-VLA (2610.12285), Caught in the Act (2610.12445, deception probes at 98.8% AUC with a released FIBS dataset), LVSPM (2610.10960) and Slot3R (2610.12282) via Bluesky.