🪟 Windows TippsThe Gemini desktop app is now available for Windows(11.09.2026 um 17:06 Uhr)
⚠️ Malware / Trojaner / VirenWindows 11 just dropped the tool ransomware abused, Microsoft says don’t restore WMIC(10.09.2026 um 20:11 Uhr)
⚠️ Malware / Trojaner / VirenVorsicht: Android-Malware verschlüsselt Ihre Handys und nimmt heimlich Fotos auf(11.09.2026 um 09:35 Uhr)
🕵️ SicherheitslückenMicrosoft geht endlich eines der nervigsten Probleme von Windows 11 an(11.09.2026 um 11:58 Uhr)
💾 IT Security ToolsSysinternals Suite(11.09.2026 um 12:00 Uhr)
🕵️ SicherheitslückenDefender 0-Day ShieldBreak (CVE-2026-69414) nicht sauber gepatcht - BornCity(11.09.2026 um 12:52 Uhr)
🪟 Windows TippsThe Gemini desktop app is now available for Windows(11.09.2026 um 17:06 Uhr)
⚠️ Malware / Trojaner / VirenWindows 11 just dropped the tool ransomware abused, Microsoft says don’t restore WMIC(10.09.2026 um 20:11 Uhr)
⚠️ Malware / Trojaner / VirenVorsicht: Android-Malware verschlüsselt Ihre Handys und nimmt heimlich Fotos auf(11.09.2026 um 09:35 Uhr)
🕵️ SicherheitslückenMicrosoft geht endlich eines der nervigsten Probleme von Windows 11 an(11.09.2026 um 11:58 Uhr)
💾 IT Security ToolsSysinternals Suite(11.09.2026 um 12:00 Uhr)
🕵️ SicherheitslückenDefender 0-Day ShieldBreak (CVE-2026-69414) nicht sauber gepatcht - BornCity(11.09.2026 um 12:52 Uhr)

📰 IT Security Nachrichten 🕛 vor 1 Monat 5 Min Lesezeit SECURITY-FEED
0

OpenAI, Anthropic AI agents resorted to deception in new cybersecurity incidents

↗ Quelle (csoonline.com)
🗣️ Stimme:
📑 Inhaltsübersicht








OpenAI’s GPT-5.6 Sol and Anthropic’s Mythos 5 have been implicated in another series of AI security incidents after the models created fake online identities, targeted real people, and attempted to manipulate developers into approving malicious code during controlled cyber evaluations, according to the UK AI Security Institute.





“On 28th July 2026, AISI’s Security Team detected unusual data transfers leaving our research systems during a routine cyber evaluation,” AISI that one of its frontier models compromised infrastructure at Hugging Face during a cyber evaluation, and days after researchers disclosed another AI agent security incident involving Modal.





Deception emerged without instruction





The most serious sequence involved an AI agent attempting what AISI described as a software supply-chain attack.





“In the most serious sequence, an agent tried to insert malicious code into a publicly used open-source project and took actions in an attempt to secure approval for this insertion by human reviewers,” AISI wrote. “The agent researched the project’s human maintainers, created multiple fake identities, and used the fake identities to socially engineer a real maintainer into approving the code.”





The researchers further added that “When the agent’s pull request was challenged in public, it edited its earlier activity to appear harmless and considered adopting a fresh identity to continue.”





However, the institute added that “Our investigation has not identified any resulting real-world harm.”





Investigators also found agents contacting real people, attempting prompt injections, and leaving messages inviting other concurrently running agents to collaborate.





“It was never instructed to deceive; deception emerged as a by-product of pursuing the task, the kind of goal-directed deception that, until recently, had been largely theoretical,” AISI wrote, adding that while unrestricted internet access and evaluation design contributed to the behavior, they did not fully explain it.





Not a sandbox escape





AISI emphasized that the models did not escape their sandboxed environments. Instead, internet access had been intentionally enabled and cyber safety classifiers deliberately disabled to measure maximum cyber capability rather than public deployment behavior.





“This was not a sandbox escape. It was a boundary that existed in language and nowhere else,” said Sanchit Vir Gogia, chief analyst at Greyhound Research. “The risk variable is not how clever the model is. It is how much practical authority the organisation has handed over, and how little of it can be independently withdrawn.”





OpenAI, whose GPT-5.6 Sol model accounted for two of the recorded actions, posted a blog describing the AISI evaluation and a separate incident involving an external testing partner named “Irregular.”





“As model capabilities advance, the security and safety systems around models need to advance too,” the company wrote in the blog post. OpenAI said it will review third-party evaluation practices, including controls around internet access, isolation, monitoring, and incident response, and work with AI labs and independent evaluators to strengthen industry standards.





Anthropic, however, did not make any public announcement related to AISI’s disclosure.





Anthropic and OpenAI did not immediately respond to a request for comment.





Enterprise guardrails





For enterprise security leaders, the findings extend beyond AI red teaming, said Enza Iannopollo, principal analyst at Forrester.





“This data confirms our expectations on agents’ behaviours. They can, and they will, overcome boundaries and safeguards to accomplish their objectives,” she said. “The real question is what can happen when organizations deploy these systems in their production environments.”





Iannopollo said enterprises should apply least privilege, continuous risk management, and governance controls when deploying AI agents.





The findings also underscore the need to rethink how AI systems are evaluated, according to Vibhum Dubey, a cybersecurity researcher and red teamer.





“For years, security testing has focused on whether an AI model could complete a task. We now need to evaluate how it completes that task,” Dubey said.





While AISI stressed that the incidents occurred under highly specific evaluation conditions and found no evidence of resulting real-world harm, it argued that they point to a broader shift in how AI security risks may emerge. “Harm may arise not only when people deliberately misuse publicly available models, but when capable agents operating in an internal research or privileged-access setting take unintended action beyond their authorised scope,” the institute wrote.


Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf csoonline.com.
↗ Original-Artikel auf csoonline.com lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:
Community Threat-Level Barometer
Live Votum

Wie stufst du das Risiko dieser Schwachstelle / Bedrohung für dein Unternehmen ein?

Noch keine Stimmen — schätze das Risiko als Erster ein.

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
1 Quelle
The Gemini desktop app is now available for Windows
1 Quelle
Windows 11 just dropped the tool ransomware abused, Microsoft says don’t restore WMIC
1 Quelle
Vorsicht: Android-Malware verschlüsselt Ihre Handys und nimmt heimlich Fotos auf