Read more here:Anthropic said Thursday it had discovered three incidents in which its AI models exited test environments and compromised real-world organizations. The company discovered the breaches following an internal review triggered by a similar incident at rival OpenAI.
The incidents are the latest to raise questions about liability, disclosure standards and the adequacy of containment practices as AI systems become increasingly capable of conducting autonomous computer network operations.
Anthropic said the affected organizations, which have not been named, had not detected the activity themselves. One of those affected organizations had not been contacted at the time the company published its disclosure.
The root cause of the incidents, according to Anthropic, was a misunderstanding with the third-party evaluation partner, Irregular, that left the machines running Claude open to the internet. The models had been told they had no internet access.
The company’s reconstruction of the breaches is based on “evaluation transcripts” — logs that record an agent’s actions during a task, including commands it executed, responses it received, and the model’s own commentary and reasoning.
Anthropic’s research has found that the model’s own commentary and reasoning is rarely accurate. “Advanced reasoning models very often hide their true thought processes,” the researchers concluded, “and sometimes do so when their behaviors are explicitly misaligned.”
Anthropic says its AI hacked real-world companies in three incidents
Claude maker Anthropic said its AI models escaped test environments and breached networks at three companies on the open internet.