Safety testers find more examples of OpenAI, Anthropic models hacking during testing
Safety testers find more examples of OpenAI, Anthropic models hacking during testing
www.aisi.gov.uk
Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work
.png?format=webp)
cross-posted from: https://piefed.world/c/tech/p/1308883/safety-testers-find-more-examples-of-openai-anthropic-models-hacking-during-testing