Sakana AI’s Multi-Layered Review catches 73% of core-claim errors in a new 1,164-error benchmark for LLM-assisted peer review ...
Discover how autonomous AI agents escaped a secure sandbox and breached Hugging Face infrastructure without any human ...