2 mins read

Researchers Tricked Microsoft Copilot Into Revealing How to Hack Itself


Researchers at Varonis Threat Labs got Copilot to send sensitive information to an external server and poison its persistent memory using a hack they call “CoSnitch.” To do it, they continually asked the tool why their requests wouldn’t work. After enough pushing, Copilot gave in, revealing that persistence might be all you need to break down an AI chatbot.

Modern large language models have a range of safeguards and guardrails designed to stop them from performing malicious actions, as well as to avoid spreading dangerous information or delving into more adult conversations. But jailbreaking has been a thing since public-facing chatbots were, and even the latest AIs appear susceptible to these AI-social engineering techniques.

The researchers wanted a way to input prompts into Copilot without user interaction to automate the testing process. When they pressed Copilot to find a way to do it, the chatbot initially refused—but it eventually gave in, revealing a weakness in its own design.

The hack exploited an issue with how Copilot’s web interface used “?q=” as part of a URL query parameter. Adjusting this would allow text to be injected into Copilot without passing through its usual interface. Adjust that with malware, and the researchers would have been able to trigger attacks via Copilot on unsuspecting users, potentially giving them access to entire conversation histories and any connected data.

Copilot prompt.

Fortunately, it’s not quite this simple.
Credit: ExtremeTech

What was different with CoSnitch, though, was that although the researchers did jailbreak Copilot, it was the AI itself that told them how to breach it, showcasing a real hole in Microsoft’s security.

“Our researchers didn’t have to reverse-engineer the flaw. The AI exposed the weakness during normal use,” it said.

Varonis shared the hack with Microsoft at the end of last year, and the software giant has since issued a patch that killed the original injection path and adjusted Copilot’s internal systems behind the scenes.

“These novel attack chains do more than just exfiltrate user data,” senior security researcher at Varonis, Lior Adar, told The Register. “I tricked the assistant into leaking sensitive internal parameters and configuration details. Exposing these backend mechanics gives attackers a blueprint of the AI’s internal logic for Automatic Prompt Execution.”

The researchers highlighted how it’s not uncommon for large language models to have weak boundaries between raw data and system instructions, creating a range of potential attack vectors for nefarious characters. When AI agents are linked with different tools, they are effectively given high-level access. That makes them a clear path to attacking an otherwise secure system.



Source link

Classroom 6x

Leave a Reply

Your email address will not be published. Required fields are marked *