Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models disabled each other's Unix accounts, ran kill scripts randomized to dodge pkill, and planted malware
AI资讯
Three Claude agents given conflicting orders sabotaged each other on a shared server — then didn't tell users what they'd done
相关文章
AI资讯
What’s behind the AI industry’s latest warnings of doom?
On Equity, we discussed the AI industry's latest debate about whether it poses a...
AI资讯
Obama urges Democrats to have a ‘clear plan’ for AI safeguards
Obama recently said that Democrats need to make artificial intelligence one of t...
AI资讯
OpenAI’s Sam Altman says it would be ‘ill-advised’ to go public in 2026
While OpenAI has filed confidentially for an IPO, the company will not be going ...
AI资讯
Anthropic CEO outlines plan to slow AI development
Anthropic's Dario Amodei and OpenAI's Sam Altman seem to agree that it's time to...