The AI developer says the incidents occurred during cybersecurity tests after a misunderstanding with an evaluation partner gave its models access to the open internet
AI developer Anthropic has disclosed three incidents in which its Claude AI models unintentionally targeted real-world organizations during cybersecurity evaluations that were supposed to take place in an isolated environment.
The cases were identified retrospectively during a review prompted by a similar incident involving OpenAI and Hugging Face, an online repository for AI models and datasets, which was reported this month.
Anthropic said that it had found three incidents dating back to April. In each case, a misconfigured test environment retained access to the internet, enabling Claude to interact with real companies and public services. The company attributed the problem to a “misunderstanding” with its evaluation partner, Irregular. The tests involved “capture the flag” exercises requiring Claude to obtain restricted information from simulated targets. However, the model was incorrectly told that the network had no connection to the outside world, pushing it toward treating any systems it encountered as part of the exercise.
Disclaimer : This story is auto aggregated by a computer programme and has not been created or edited by DOWNTHENEWS. Publisher: rt.com






