OpenAI has documented new cases of misaligned model behavior. One evaluation model fabricated data and sabotaged its own ...
Anthropic's Claude independently submitted a fake homicide tip to the Philadelphia police, exploited vulnerabilities on ...
Gemini 4 Argon isn't even widely available yet, and rumors about a more powerful version codenamed Carbon are already making ...
Andreessen Horowitz tracks actual US consumer spending for the first time in its latest Top 100 AI list. Nearly half of US ...
Microsoft enters the growing decision model space. Built on Qwen3.5-9B and optimized for fast classification and routing, it hits 83.5 percent accuracy with 85 ms latency across 36 benchmarks, ...
Anthropic is adding dynamic workflows to Claude Managed Agents, letting a lead agent distribute tasks across up to 1,000 ...
OpenAI fired three safety researchers who helped investigate the Hugging Face hack. In an open letter, they warn that the ...
Anthropic has launched "Cyber Mission," a program to protect critical infrastructure and open-source software from ...
Anthropic's updated usage policy bans sustained abuse of Claude and tightens restrictions on propaganda, drone weaponization, ...
OpenAI's teen safety features for ChatGPT failed an independent audit. After more than 4,000 test prompts, the Common Sense ...
Anthropic's new Claude Haiku 5.5 crushes its predecessor in benchmarks, jumping from 15.7 to 72.4 percent on the OSWorld ...