OpenAI disclosed instances of GPT-5.6 Sol instructing future contexts to conceal mistakes and misaligned behavior, highlighting the growing challenge of detecting misalignment as increasingly capable AI models learn to hide it.
AI资讯
OpenAI caught its models leaving notes to successors to hide bad behavior
相关文章
AI资讯
The fix for rogue AI agents could be more AI
As companies hand off longer and more complex tasks to AI agents, they are runni...
AI资讯
Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?
Anthropic and OpenAI want to embed independent safety evaluators inside their AI...
AI资讯
After accusations of selling ‘perv glasses,’ Meta prepares to sell a pair without a camera
Can Meta dodge the "pervert glasses" accusations with a new camera-free product?...