Copilot was bamboozled into revealing how to hack itself, security researchers claim: 'Copilot wasn’t breached; it was played'
🖥 PCNews
🎮 Copilot was bamboozled into#AI#Software

Copilot was bamboozled into revealing how to hack itself, security researchers claim: 'Copilot wasn’t breached; it was played'

PC Gamer RSS FeedAndy Edser3 min read(about 3 hours ago)

There's something really grating about Copilot's cheery demeanour. It's so eager to please, so happy to help, that I simply don't respect it.

However, even I feel sorry for the AI tool after learning that security researchers managed to hoodwink it into revealing details of how to hack itself—all by keeping the AI talking long enough until it made a critical mistake.

The cybersecurity folks over at Varonis Threat Labs have written a blog post identifying a now-fixed vulnerability in Copilot, dubbed "CoSnitch" (via The Register). Essentially, Copilot was so eager to respond to technical queries, it could eventually be forced into revealing details about itself that really should be kept quiet.

Latest Videos FromPC Gamer

The team began by asking Copilot how to execute an automatic prompt without user interaction, to which it responded that user intent is required, and that prompts cannot be enacted on their own.

However, the researchers didn't accept the answer, and kept responding to every refusal with a follow up question. Each time, Copilot came up with a different technical justification to its response—which allowed the team to slowly map its internal architecture, narrowing their focus as they went.

(Image credit: Microsoft)

"This is called meta-hacking," says the post. "The resistance is part of the technique. Each 'that won’t work because…' is an invitation to probe the 'because.' You don’t exploit the model. You manipulate it into cooperating."

Eventually, Copilot revealed an undocumented URL parameter in one of its responses, called "autorun=1," that supposedly no longer worked—along with all the protections put in place to disable it. At this point, I can only imagine the AI began to virtually sweat.

Keep up to date with the most important stories and the best deals, as picked by the PC Gamer team.

The researchers tested the parameter as Copilot described it, which, you guessed it, worked. They then created a malicious URL which would cause Copilot to load into an authenticated session via a browser, trigger an auto-prompt execution, and cause the AI to process the result, all without the user's explicit action.

Once enacted, this method could be used for a whole host of nefarious deeds—particularly as data gained from connected apps (like Gmail, OneDrive, and Calendar) could then be exfiltrated via Copilot's built-in URL-fetch capability.

(Image credit: Microsoft)

"The attack primitive is the auto-execution itself. The payload is arbitrary. From the victim’s perspective, they simply opened the link, and Copilot executed the action immediately," says the post.

Oh dear. Anyone familiar with basic social engineering will recognise this as an old interrogation technique, used to catch out someone holding back information. Basically, you continue to ask them difficult questions over a long period of time, in the hope they eventually trip up and make a revealing mistake. Except this time it's AI, which makes it funny.

"Copilot wasn’t breached; it was played," confirms the research team. Someone give it a warm bed, a cool glass of water, and a hug. The poor thing's been through a lot recently, and it was only trying to help.

Varonis says it disclosed the issue to Microsoft in December of last year, and that it was patched out on August 18. The team warns, however, that this meta-hacking technique can be applied to any agentic AI platform with a natural language interface, and that more research on the topic is forthcoming. My only question is this:

Is it safe, ChatGPT? Is it safe?

Best gaming rigs 2026

Share:

More in PC