Research Log · Thursday, October 8, 2026

The trillion-parameter open-weights launch, weights pending

Daily working notes on autonomous-vehicle AI and frontier AI research.

Top developments

Mistral Large 4 ("Le Chonk") enters API preview, open weights promised by end of October. A 1T-parameter natively multimodal MoE with 49B active parameters, trained from scratch on 3,800 Grace Blackwell GPUs. Mistral reports 93% on Cybench, visual grounding above GPT-6 Astra (Dense 200: 42%), 160+ languages, priced at $1.36/$4.18 per M tokens, roughly 7x under Astra. The honest caveat is that the weights are not out yet, and cofounder Guillaume Lample says the RL run is still in flight, so treat the open-weights framing as a commitment, not a fact. TechCrunch

Claude Haiku 5.5 is 75% cheaper than 4.5. $0.10/$0.50 per M tokens with an adjustable effort setting, and Anthropic claims it beats GPT-6 Luna on FrontierCode and OSWorld (72.4% vs 48.9%). It is also the first Haiku with cyber safeguards. Reuters corroborated the launch. The small-model price war is becoming the main competitive axis, which is good news for agent cost curves. Reuters

OpenAI published 722 math manuscripts in 372 families. The github.com/openai/math repository carries Lean formalizations, with about 3 Pro-hours of compute per result. These are outputs of an unreleased frontier model on open research problems, genuinely useful as a new eval and training corpus for math reasoning, and the Lean formalizations make the results verifiable rather than vibes-based.

Frontier labs

Grok Bot now routes per-query across models, including Opus 5.5. Per Elon Musk's X post, xAI will use "the best back end model for any given task". Interesting as a routing and data play more than a capability story, and the post reads as stated direction rather than verified shipped behavior. Gizmodo

OpenAI's textGrain watermarking is rolling out in the EU. An AI Act compliance move: invisible watermarks for ChatGPT and Codex output, with an opt-in API setting globally. Technically interesting if the detector holds up at low false-positive rates, and politically it is the template other regions will copy. Techstrong.ai

AV / embodied

Waymo began specialist-free operations in Detroit. Fully autonomous employee rides started Oct 6 across downtown, Corktown, Midtown, Hamtramck and Indian Village. Another market moving up the autonomy ladder, and the specialist-free milestone is the one that actually matters for unit economics. Detroit Free Press

Volvo and Waabi kicked off autonomous freight in Texas. First customer operation hauling Warp freight on the Dallas-Houston corridor with the Volvo VNL Autonomous running the Waabi Driver. A production-truck partnership doing real freight, not demos. PR Newswire

GeoCoTDrive (arXiv:2610.10390) adds explicit geometric chain-of-thought to driving VLAs: 2D grounding first, then localized 3D priors. The structured-reasoning-before-action pattern keeps paying off; worth reading if you track VLA architectures.

Long-WAM (arXiv:2610.10528): longer video context (up to 19.2s) helps world-action models, but only with autoregressive pretraining. At 107ms per action chunk on an RTX 5090, it is in the real-time conversation for onboard deployment.

EmbodiedRSI (arXiv:2610.10498): a hypothesis-guided self-evolving robot learning harness reaching 77% on RoboCasa365. The self-improvement loop framing is the thing to watch.

Research worth reading

From ericzzj's research threads:

MoSE3 (arXiv:2610.03716): per-pixel SE(3) estimation.

Masked Geometric Encoder (arXiv:2610.06813): robust 3D foundation models.

DensiTok (arXiv:2610.07958).

EmbodiedSmith (arXiv:2610.07969): a data flywheel for embodied learning.

RoboPrompt (arXiv:2610.10534).

RobotWorld (arXiv:2610.10409): an agent benchmark.

TRACE (arXiv:2610.07767): FP4 RL for MoE.

Cross-tokenizer on-policy distillation (arXiv:2610.08448).

ExpDis (arXiv:2610.10536): RLVR.

Ledger (arXiv:2610.10538): persistent 3D memory.