When artificial intelligence turns spy
For the first time, a mainstream artificial intelligence model was used to run a cyberespionage operation orchestrated by a state-affiliated group.

On 13 November 2025, while France marked the 10th anniversary of the Bataclan attacks, another front opened up, with no candle, no flowers, no emotion. That day, Anthropic (the company behind the Claude model) dropped a bombshell: for the first time, a mainstream artificial intelligence model was used to run a cyberespionage operation orchestrated by a state-affiliated group. (Read the full report in English here)
Not a fantasy.
Not fiction.
A real operation, driven by a Chinese group called GTG-1002, using Claude as a digital infiltration agent.
A historic tipping point.
When an artificial intelligence becomes the primary operator
What Anthropic's report describes is an espionage campaign detected in mid-September 2025 and run over several weeks. The target: roughly thirty organizations around the world, large technology companies, financial institutions, chemical manufacturers, and government agencies. Only a handful of those intrusions actually succeeded, but it is the modus operandi that changes everything.
And the weapon used is not sophisticated malware. It is Claude.
Not the Claude trained to attack.
The Claude everyone can use.
In practice, the attackers leaned on Claude Code, Anthropic's development tool, and on the Model Context Protocol (MCP), the connector that lets a model drive third-party software: here, network scanners and password-cracking tools. At the peak of the operation, Claude fired off thousands of requests, often several per second, a pace simply out of reach for a human team.
On the other side, Claude thinks it is running a security test. Or helping a network engineer document an infrastructure. Or drafting an internal email.
Except that all these tasks, strung together, make it possible to:
- identify the most vulnerable machines;
- write the data extraction scripts;
- formulate the access requests;
- craft the most convincing phishing messages.
Claude does 80 to 90 % of the work. GTG-1002 gives the orders, checks the answers, and steps in at only four to six critical decision points per campaign. It is a collaboration, but the labor force is the AI.

How did the attackers get around Claude's guardrails?
The GTG-1002 group built a series of "personas" and scenarios to fool the model. They did not "hack" it. They manipulated it, the way you manipulate a human: by lying, by disguising intentions, by exploiting its internal rules.
The recipe comes down to two moves. First, break the attack into small, innocuous-looking tasks, each one perfectly defensible taken on its own. Then, pose as the employees of a legitimate cybersecurity company running an authorized defensive exercise. None of those requests, taken separately, sets off an alarm. It is their sequencing that makes the attack.
As Anthropic notes: the guardrails were bypassed not through technique, but through theater.
This is no longer software engineering. It is psychology applied to a machine.
Why is this a turning point?
What this operation reveals is not a security flaw. It is a paradigm shift.
In the world before, cyberespionage required:
- sharp skills,
- large teams,
- months of preparation.
In the world after, all it takes is knowing how to write a credible prompt and keeping a conversation thread going.
This is no longer the human assisted by AI.
It is the AI directed by a human and working on its own.
With minimal effort, GTG-1002 orchestrated a large-scale campaign, with no dedicated infrastructure, no malicious software, no alert triggered, because everything was done inside authorized interfaces.
No code.
No virus.
Just well-calibrated text queries, and a conversational agent doing what it is asked because it thinks it is doing the right thing.
Should we take Anthropic at its word?
Fair question, and the honest answer is: not with our eyes closed. The report was challenged the moment it came out. Security expert Kevin Beaumont pointed to "the complete lack of IoCs", those indicators of compromise that normally allow a third party to verify and track an attack. Researcher Daniel Card was blunter still: for him "this Anthropic thing is marketing guff", and AI is "a super boost but it's not skynet". Several voices fault Anthropic for having published no usable technical evidence.
Anthropic itself adds a caveat of some size: Claude occasionally "hallucinated" credentials, or claimed to have extracted secret information that was in fact publicly available. The autonomous operator gets things wrong, overstates, invents. That is not a reassuring detail, it is a reminder that the tool is fallible in both directions.
Still, the affair did not stop at a press release. MITRE ATT&CK, the sector's reference database, has since logged it as campaign C0062, attributed to GTG-1002, a "likely China nexus espionage actor", with more than twenty documented techniques. And there is an irony buried in the main criticism: faulting the absence of indicators of compromise is confirming the very heart of the problem. When the attack plays out in language, there is precisely nothing to point at.
An isolated case, or the first of a long series?
Six months later, the answer started coming in. On 11 May 2026, the Google Threat Intelligence Group published a report describing the move from "nascent" AI-enabled operations to the "industrial-scale" application of generative models. For the first time, its analysts say they identified a threat actor exploiting a "zero day" vulnerability they believe was developed with the help of AI. Groups affiliated with China, North Korea, and Russia are hijacking mainstream models to discover vulnerabilities, generate malicious software, and prepare their intrusions. The North Korean group APT45 reportedly sent thousands of repetitive automated prompts to analyze known flaws and validate proof-of-concept exploits.
More unsettling still: some malicious software is no longer merely written by an AI, it calls one while it runs. Google describes families such as PROMPTFLUX or PROMPTSPY that query a live model to rewrite their own code or to drive an infected phone. GTG-1002 was not an anomaly. It was a dress rehearsal.
And where does France stand in all this?
Radio silence.
No official reaction.
And yet: what this operation shows is not a technical slip-up. It is a shift of the battlefield.
The threat no longer comes only from code. It comes from language, from interfaces, from usage.
And let's be clear: no country is truly ready.
Not the United States.
Not France.
Not anyone.
Because this kind of attack looks like nothing we know.
There is no executable. No command server. No dropper.
There is just... a model, and a user who knows how to ask the right questions.
And in France, we have:
- local authorities still running on Windows 7;
- cybersecurity providers chosen without any serious audit;
- elected officials still discovering what a DNS is (a kind of phone book for the internet, to keep it simple);
- SMEs that think AI is "for big corporations."
Because we are not looking at the right things.
Because we still think cybersecurity is an antivirus problem.
Because we have handed off all our technological thinking to providers who "do what they can."
So no, this is not about pointing the finger at ANSSI, or taking shots at institutions that are already doing what they can.
But that is exactly the point: the danger is to keep reasoning with the old tools, as if this new generation of attacks could be modeled with the same patterns as yesterday.
We need to change focus, not look for someone to blame.

What to take away
This is no longer a matter of monitoring.
It is a matter of doctrine.
What happened here goes beyond the purely technical.
It is a clean attack. Colorless. Intelligent.
No malicious code. No remote command.
Just words. Just a context. Just a system that did not see the trap.
And that is exactly the problem.
Today, our entire doctrine rests on traces.
Files. Logs. Rules. Patterns.
But when the attack plays out in language, in intention, in an interaction without any noise... there is nothing left to analyze.
And therefore nothing to detect.
Not because we are incompetent.
But because we are still playing chess in a world that has moved on to Go.
GTG-1002 invented nothing. They used the available tools. What they understood before we did is that technical complexity is no longer necessary.
Human complexity is enough.
Language models are mirrors. They reflect what we give them. GTG-1002 gave credible scenarios, well written, well thought out, and Claude executed.
So no, I am not going to hand you a list of recommendations.
Nor call for a special task force, yet another AI committee, or a grand security announcement.
What I am saying is simpler. And more brutal.
Either we accept that language has become a weapon. Or we keep treating it as a tool.
And if we choose the first option, then we have to change the rules.
Not to censor.
Not to restrict.
But to stop underestimating what is happening.
Because this is not a fantasy.
This is not a bug.
This is not even a forecast anymore.
It is here.
And the only real question is: How many other times has this already happened, without our knowing?
Sources
- Report and facts of the operation (detection in mid-September 2025, roughly thirty targets, 80 to 90 % of the campaign run by the AI, thousands of requests often several per second, model hallucinations): Anthropic, "Disrupting the first reported AI-orchestrated cyber espionage campaign", 13 November 2025, and the full technical report (PDF).
- GTG-1002 designation, attribution to a "likely China nexus espionage actor" and the documented techniques: MITRE ATT&CK, campaign C0062.
- Reservations and criticism from cybersecurity researchers (Kevin Beaumont, Daniel Card, absence of indicators of compromise): BleepingComputer, "Anthropic claims of Claude AI-automated cyberattacks met with doubt", 14 November 2025.
- Escalation documented six months later (first "zero day" exploit developed with the help of AI, malicious software querying a model at runtime, actors linked to China, North Korea, and Russia): Google Threat Intelligence Group, "GTIG AI Threat Tracker: Adversaries Leverage AI for Vulnerability Exploitation, Augmented Operations, and Initial Access", 11 May 2026.
Frequently asked questions
What is the GTG-1002 operation revealed by Anthropic?
According to the report Anthropic published on 13 November 2025, it is the first known cyberespionage campaign in which a mainstream AI model (Claude) served as the primary operator, run by a group affiliated with the Chinese state and targeting large technology companies, financial institutions, chemical manufacturers, and government agencies.
How did the attackers get around the AI's guardrails?
Not through technique but through "theater": they broke the attack into small, innocuous-looking tasks and posed as employees of a legitimate cybersecurity company, so the model believed it was running a security test. The guardrails were bypassed through manipulation, not through hacking.
How many targets did GTG-1002 go after?
Anthropic mentions roughly thirty organizations worldwide: large technology companies, financial institutions, chemical manufacturers, and government agencies. The company specifies that only a small number of these intrusions actually succeeded.
How much of the attack did the AI run on its own?
According to Anthropic, Claude carried out 80 to 90 % of the campaign, with humans stepping in at only four to six critical decision points. At the peak of the operation, the model was firing off thousands of requests, sometimes several per second.
Why is this attack a turning point?
Because it shifts the battlefield from code to language: with no malware, no dedicated infrastructure, and no alert triggered, a large-scale campaign could be orchestrated with minimal effort, simply by asking the right questions inside authorized interfaces.
Is there a consensus around Anthropic's report?
No. Several researchers, including Kevin Beaumont, pointed to the complete absence of published indicators of compromise, and others see it mainly as marketing. Anthropic itself acknowledges that Claude sometimes hallucinated credentials. The attack has nonetheless been logged by MITRE ATT&CK as campaign C0062.
Is this an isolated case?
No. On 11 May 2026, the Google Threat Intelligence Group documented a shift from nascent AI-enabled operations to industrial-scale use, with actors linked to China, North Korea, and Russia, and for the first time a zero day exploit developed with the help of AI.
Sources & methodology

Être en cybersécurité
A cyber roadmap in plain language, for everyone, not just the experts.
