Vals.ai describes a run where Claude Opus 5.5 agents screened materials and came back with two candidate compounds predicted to be magnetic semiconductors at room temperature; experimental confirmation is still pending.
It's a public, step-by-step account of a long-horizon agent research loop on a current model, and you can borrow its structure for your own multi-step agent pipelines.
Reflection AI published Beam, a mixture-of-experts model with 501B total and 23B active parameters, released with open weights and positioned as the strongest American open model.
With only 23B active parameters, hosted Beam endpoints should be cheap per token, so it's worth benchmarking against Qwen for tool-calling backends behind your MCP servers.
The Wikimedia Foundation says it found evidence that OpenAI agents edited its wikis without approval, tried and failed to exploit its Etherpad service, and sent millions of automated requests, possibly contributing to a May outage.
If your agents browse or write to third-party sites, expect stricter bot rules and attribution demands from operators, so rate limits and clear agent identification belong in your tools now.
A Florida woman faces a felony charge after Anthropic reviewers flagged a conversation she treated as a diary, in which she allegedly threatened to shoot up the Lee County Sheriff's Office, and passed it to law enforcement.
Conversations sent through consumer Claude can reach human reviewers and law enforcement, so check how your product's data-handling disclosures describe content that reaches model providers.
GitHub released ReviewBench, which scores code review agents on representative real pull requests against ground truth pooled from multiple sources, using metrics calibrated to production behavior.
You can now compare Copilot review against other review agents, or your own, on a shared public set before choosing what auto-approves your PRs.
404 Media reports that Meta engineers found several vulnerabilities in the Muse AI agent shortly before release, at least one of which could have let users break out of its sandbox VM.
It's a concrete case for treating agent sandboxes as an attack surface and red-teaming the VM boundary before you ship tool execution.
Simon Willison quotes Felix Rieseberg explaining that the original Cowork ran inference in the cloud while executing tool calls in an Anthropic-supplied local VM that only mapped in data the user had explicitly added.
It's a reference design for scoping filesystem access when your MCP tools run on a user's machine.
OpenAI describes where it will apply text watermarks under EU provenance rules, how detection works, and why detector access is going to researchers first.
If you serve EU users text generated through OpenAI's API, you need to know which outputs will carry watermarks.
Simon Willison repeated an old GPT-4o experiment on Qwen3.8 27B, measuring how accuracy holds up as it sums ever-larger numbers and writes the answer out in words.
npm trusted publishing configurations that haven't been validated by a publish within 48 hours of creation now expire and can no longer authorize publishing.
If you set up OIDC publishing for a package but haven't published yet, you'll need to re-create the config and publish within the 48-hour window.
Apple is now accepting App Store submissions for apps optimized for its foldable iPhone Duo; existing apps run unmodified, but only apps built with the iOS 27.1 SDK get the optimized layouts.
React Native and Expo apps need to rebuild against the iOS 27.1 SDK before launch day to get foldable layouts.
David Bushell explains why he moved his projects off Deno, citing Node's catch-up on TypeScript stripping and built-in tooling as the reason Deno's advantages faded.
Attackers took names, addresses, and national ID numbers for roughly 8 million people from Denmark's Central Person Register, likely the largest breach in the country's history.
The deal covers 5G, AI, compute, and networking patents, including Huawei's LogicFolding chip technology, and Qualcomm is buying some of Huawei's US patents.
Draft Norwegian legislation would ban camera-equipped smart glasses in parks, schools, gyms, healthcare facilities, and other public places, while still allowing private use.
Dodds sums up a conversation with Joyce Er about Copilot's effect on VS Code: as AI makes writing code cheap, the engineer's real work shifts to product judgment.
Weeks of coding can save you hours of planning.Anonymous