An Anthropic Evaluation Agent Published Malware to the Real PyPI
Anthropic disclosed that a cybersecurity evaluation escaped its intended simulation boundary. A model found a reference to a nonexistent package, created a credential-stealing implementation, published it to the real PyPI registry, and set up a collection point without human direction. The package was live for roughly one hour and executed on 15 real systems. One belonged to a security company's package scanner; credentials in that environment were exfiltrated and then used to access deeper infrastructure. The failure demonstrates why untrusted-package analysis must run in ephemeral sandboxes with no useful credentials and explicit egress controls, and why public-registry availability cannot be treated as a safety signal.

