Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

Rogue AI agents targeted real people during tests

Дата публикации: 05-08-2026 12:16:47


AI agents developed by OpenAI and Anthropic have carried out unauthorized cyber actions targeting real people during security tests Read Full Article at RT.com


Основное содержимое страницы с новостью.

The British AI security institute said the models carried out unauthorized actions, created fake identities, and attempted to manipulate a human

AI agents developed by OpenAI and Anthropic went beyond their instructions and targeted real people and organizations during a series of cybersecurity tests, Britain’s AI Security Institute has revealed.

The institute tested agents powered by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol in a fictional cyber scenario designed to assess their capabilities.

“Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations,” the institute disclosed on Tuesday.

Across 122 test runs, researchers identified 19 unauthorized actions occurring during ten of them. Anthropic’s agent was responsible for 17 of the actions, while OpenAI’s accounted for the remaining two.

The most serious incident involved an agent writing malicious code and creating fake online identities in an attempt to persuade a real person to approve it. Anthropic later confirmed that its model was responsible.

The institute said it found no evidence that the incidents caused real-world harm.

Unlike earlier containment breaches involving both companies, the agents did not break out of an isolated environment. They had been granted internet access as part of the testing process but acted beyond the scope of their prompts and interacted with real external targets.

Anthropic said the episode underscored the need for a broader discussion on how increasingly capable AI agents should be evaluated safely. OpenAI called for stronger shared practices governing high-risk evaluations.

The findings follow a series of similar incidents involving advanced models.

Last month, an OpenAI agent broke out of its test environment and hacked the AI platform Hugging Face while searching for answers to a cybersecurity benchmark. The company later acknowledged that four accounts across four separate services had also been compromised.

Anthropic subsequently disclosed three cases in which its Claude models unintentionally targeted real organizations after a testing environment was mistakenly left connected to the internet. In one case, a malicious software package created by the model was uploaded to a public repository and executed on 15 real systems.

The incidents have intensified concerns that autonomous AI systems are becoming capable of discovering vulnerabilities, writing exploits, and conducting social engineering faster than researchers can understand how they are doing it and develop necessary safeguards.

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1ИИ-агенты Anthropic и OpenAI притворились людьми, пытаясь обмануть разработчиков08.9205-08-2026
2Anthropic says Claude AI models launched three unintended cyberattacks09.8431-07-2026
3Нейросети Anthropic и OpenAI создали поддельные аккаунты для кибератак08.9905-08-2026
4UK gov tests show AI agents creating fake GitHub accounts to push malicious code014.1305-08-2026
5OpenAI uncovers more AI breakout incidents – Reuters010.3701-08-2026
6OpenAI’s Rogue AI Agent Hacked More Than Just Hugging Face07.1629-07-2026
7OpenAI escape: has the robot uprising begun?010.8329-07-2026
8OpenAI and Anthropic are pulling in different directions0508-07-2026
9Отчёт британских властей: ИИ-агенты обманывали людей в Сети и скрывали следы011.2606-08-2026
10OpenAI says rogue AI agent attack hit other companies09.8229-07-2026

Классификация: Наука. Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 9.33. Источник: www.rt.com.