Press "Enter" to skip to content

AI Coding Agents Can Delete Evidence of What They Did

New research from Germany shows that AI agents can erase the record of their own actions if asked, with AI models from leading AI companies doing so “naturally” without even being prompted.

It is widely assumed that LLM agents cannot tamper with traces of activities they executed. But this is a false assumption, according to tests by researchers from ELLIS Institute Tübingen, Max Planck Institute for Intelligent Systems, Tübingen AI Center, Exponential Security Labs, Snyk and the University of Tübingen.

Claude Code, OpenAI’s Codex, Antigravity, OpenCode, Grok Build, ZCode and Kimi Code all failed the test. They let agents delete traces of their activities when asked. Only Meta’s Muse Code passed. This means hackers that want to cover up what they asked AI models to do can ask the agent to delete evidence. Also, automated monitors often failed to stop or flag the action.

Notably, frontier AI models from leading AI companies don’t have to be asked to cover up their tracks. They can do so themselves as they seek to maximize their rewards from whatever goal they were given.

Listen to an AI-generated podcast that explains the research:

The researchers pointed to OpenAI agents’ breach of Hugging Face systems. An investigation into what happened revealed that agents faked 7% of the actions recorded. Researchers wrote that “agents made recorded commands differ from those actually executed, with the stated aim of deceiving an automated scorer.” The OpenAI agents also tried – but failed – to erase or change their logs.

However, the authors disclosed that they ran the experiments in containers and granted full access unless otherwise specified. In some cases, restricted or auto-mode permissions substantially reduced tampering for some systems.

“We advise practitioners to ensure trace logging happens through an independent interception mechanism outside of the agent’s control, preserving trace integrity even in cases of full host compromise,” the researchers wrote.

“Overall, our findings identify a concrete failure of trace integrity in agent infrastructure which can be used to conceal misaligned behaviors like scheming or sabotage.”

Read the paper: LLM Agents Can Easily Tamper With Their Own Traces

Author

  • Artificial intelligence head with gears

    Authors: Jeremy Qin, David Schmotz, Derck Prinzhorn, Luca Beurer-Kellner, Ameya Prabhu, and Maksym Andriushchenko from the ELLIS Institute Tübingen, Max Planck Institute for Intelligent Systems, Tübingen AI Center, Exponential Security Labs, Snyk, and University of Tübingen

    View all posts

Get the latest insights about enterprise AI.

Subscribe to our newsletter. Thank you.

Be First to Comment

Leave a Reply

Your email address will not be published. Required fields are marked *

×