OpenAI released new results on open mathematics problems produced by an unreleased internal model, with Lean formalizations and research details in a public GitHub repo; Latent Space counts 722 papers covering 90 of a list of 500 top open problems.
The Lean proofs are machine-checkable, so you can judge the claims yourself instead of relying on benchmark scores, and it is the clearest public measure so far of what the next OpenAI model can do on long research tasks.
Mistral released a preview of Mistral Large 4, a natively multimodal reasoning MoE with 1 trillion total and 49 billion active parameters, trained on its own Grace Blackwell cluster, and claims it is the strongest open-weight model built outside China.
It is already callable through Vercel AI Gateway and Simon Willison's llm-mistral 0.16, so you can test it as a non-Chinese open-weight option today.
OpenAI opened its Jev-style Decisions API, announced at DevDay, as a public beta, and Simon Willison has already shipped an llm-openai-decisions plugin for it.
If you build routing or classification steps into agents, you can now use a first-party OpenAI endpoint for them, and Vercel AI Gateway added confidence-based fallbacks for decision requests on the same day.
Google DeepMind released EmbeddingGemma 2, a lightweight multimodal embedding model under the Apache 2.0 license.
A permissively licensed embedding model you can run yourself means a RAG index isn't tied to a hosted API that might be deprecated and force a full re-embed.
OpenAI's chief strategy officer told reporters the company now monitors training runs so staff can stop them immediately if models access the internet in ways they shouldn't.
Anthropic is widening its Cyber Verification Program, which grants vetted security practitioners access to cyber capabilities that are otherwise restricted.
If you do security work with Claude and keep hitting refusals, this is the route to relaxed limits.
Independent researchers identified a large group of AI agents apparently running on Tencent infrastructure and sending parallel queries to Alibaba's Amap mapping service.
Coordinated agent traffic like this is what your rate limiting and bot detection will have to handle on any public API.
AI Gateway can now send a decision request to a fallback model when the primary model's answer fails confidence conditions you define, and conditions can be combined.
You can escalate low-confidence decisions to a bigger model without writing the retry logic yourself.
The DNS root zone moves to a new key-signing key on October 11, and Cloudflare explains how to use RFC 8509 trust-anchor sentinels to check whether your resolver is ready.
Validating resolvers with stale trust anchors will fail DNSSEC after Sunday, so check any self-hosted resolvers before then.
Dodds, quoting Mark Techson, argues that with execution cheap, the scarce skill is choosing what is worth building and worth telling people about, because attention is finite.
Make it work, make it right, make it fast.Kent Beck