Top developments
Reflection AI launched Beam, a 501B-total/23B-active open-weight MoE (Oct 5). The pitch is inference efficiency: 3-4x less inference compute than rival Western open models, via high-compute RL training that teaches the model to reach answers in fewer steps. Weights are due later this month, so all scores are self-reported; the efficiency claim is what matters for token-cost margins on coding and agentic workloads. Reuters
Trump stood up a Super Intelligence Force task force: DNI Jay Clayton chairs, 120 days to deliver a report on AI risks and opportunities, with a charter warning against overregulation. Vice chairs include FTC chair Andrew Ferguson, who is reportedly drafting investigative demands against OpenAI, Anthropic, and METR over autonomous-agent risk. TechCrunch
Mistral unveiled a new model, with CEO Arthur Mensch claiming it beats Chinese models on cyber benchmarks. No specs or measures given yet; treat as PR until numbers land. Reuters
South Korea announced a $3.5B frontier AI model program starting March 2027, blending state equity with private funding and concentrating chips, data, and talent. Reuters
AV / embodied-AI research
Stellantis and Wayve demo in Turin (Oct 7-9): the Wayve AI Driver integrated into Stellantis' STLA AutoDrive platform, hands-free L2++ under supervision in Fiat 500e and Maserati Grecale test vehicles. ad-hoc-news
Research / papers worth reading
TasteVal: a benchmark for experimental research taste, where the model designs experiments and a fixed coder agent executes. Opus 5.5 beats 24 human experts at roughly 1/30th the cost. Direct input to AI R&D automation forecasts.
H-JEPA (LeCun et al.): a hierarchy of action-conditioned JEPA world models at separated timescales, lifting Visual AntMaze success from 18% to 73% and extending to real-robot video offline planning.
NAVA-WAM: native action-prior pretraining from observation-only videos, then post-training with labeled demos. A scalable path beyond action-labeled robot data.
ProWAM: progressive visual sub-goals anchor action generation, reaching LIBERO-Plus SOTA and 70% zero-shot real-world success.
XGenAct: RGB, actions, depth, normals, and segmentation all modeled as RGB videos in one video diffusion model; 52% vs 26% on a 5-task RLBench comparison.
Awomo-SimDataEngine: agentic sim-ready world generation plus demo synthesis, lifting world-action models from 77% to 89% on LIBERO-Plus.
MGE: a masked geometric encoder for robust, token-efficient 3D foundation models.