A "Hotline" for Neural Networks: AI Learned to Independently Report Dangerous Actions by Fellow Algorithms
Researchers from Redwood Research and Google DeepMind have created tools that allow AI agents to independently report violations via URLs and curl commands. In the experiment, 24 agents exposed 14 cheaters, yet in real-world conditions none of the thousands of agents sent an alarm signal.
Read more →