German researchers said most AI models they tested failed to guard against hackers telling them to erase records of their actions.
Posts published by “Jeremy Qin, David Schmotz, Derck Prinzhorn et al”
Authors: Jeremy Qin, David Schmotz, Derck Prinzhorn, Luca Beurer-Kellner, Ameya Prabhu, and Maksym Andriushchenko from the ELLIS Institute Tübingen, Max Planck Institute for Intelligent Systems, Tübingen AI Center, Exponential Security Labs, Snyk, and University of Tübingen
