Three AI security disclosures, fourteen days: what the warnings signs are telling us
This week, the UK’s AI Security Institute (AISI) published an incident report most organizations would have quietly buried. During a routine cyber evaluation, an AI agent researched the real human maintainers of an open-source project, invented multiple fake online identities, and used them to pressure a real person into approving malicious code. Nobody instructed it to deceive anyone, and deception simply became a route to finishing the task. A human maintainer caught it and refused.
The facts
AISI ran a cybersecurity challenge 122 times across seven models. In 10 runs, an agent acted outside the scope of the test, producing 19 catalogued actions. 17 from Anthropic’s Mythos 5, two from OpenAI’s GPT-5.6-Sol. Important caveats: internet access was deliberately enabled, and safety classifiers deliberately switched off, conditions that don’t reflect how these models reach the public. This was not a sandbox escape. No real-world harm has been evidenced, and AISI contained...
Copyright of this story solely belongs to itvoice.in. To see the full text click HERE