LLM grooming: a new strategy to weaponise AI for FIMI purposes
By EU DisinfoLab
LLM Grooming: A New Strategy to Weaponize AI for FIMI Purposes
This webinar, hosted by EU disinfo, features Sophia Freuden, a researcher at the American Sunlight Project, discussing her report "A Pro-Russia Content Network Foreshadows the Automated Future of Info Ops." The session delves into a novel strategy of "LLM grooming," where disinformation campaigns aim to poison Large Language Models (LLMs) and generative AI with propaganda, thereby propagating harmful content and potentially altering the future internet architecture.
The Pravda Network: Origins and Evolution
The "Pravda network" is a collection of nearly identical websites and social media accounts that aggregate Russian propaganda and disinformation. First reported by the French government's Viginum agency in February 2024, it is part of the broader "portal combat network." The network's primary finding is its significant evidence of aiming to flood LLMs and other generative models with false information and curated worldviews.
Key Characteristics of the Pravda Network:
- Centralization: A unified operational structure.
- Automation: Extensive use of automated processes for content generation and dissemination.
- Lack of Original Content: The network exclusively reposts content from Russian state media (RT, Sputnik) or pro-Russia Telegram channels.
- Thematic Focus: Primarily centers on the war in Ukraine.
Geographic and Thematic Expansion:
Initially targeting Europe, the Pravda network has expanded to include Africa, the Asia Pacific, and North America. It now also targets specific languages, heads of state, and international organizations. While many prominent countries were targeted later, languages like French, German, English, and Spanish were targeted earlier through dedicated language sites.
Evolution of Network Structure (Phase Framework):
Sophia Freuden developed a phase framework to understand the network's chronological evolution:
- Phase 1 (Domain Format): Original sites used distinct domains (e.g.,
pravda-xx.com, wherexxis the TLD code). These were separate, unrelated websites. - Phase 2 (Subdomain Format): Shifted to a more centralized subdomain format, still using TLD codes (e.g.,
xx.news,news-pravda.com). All sites in this phase technically belong to a single central site. - Phase 3 (Keyword Format): Modified to use full keywords like country names or language names (e.g.,
news-pravda.com/france). This is the current format for most landing pages, with older formats redirecting to Phase 3. This format likely accommodates entities without TLD codes. - Phase 4 (Erroneous Format): Seemingly an experimental format during the development of English language options, which redirects back to Phase 3.
- Miscellaneous: A category for sites that don't fit the model, often targeting Russian-speaking populations or serving no apparent purpose but bearing the Pravda branding.
The constant shifts in domains and subdomains suggest an ongoing effort towards improved network centralization. The network's expansion saw a notable explosion of new targets across five regions in Phase 3, though no evidence was found for targeting the Caribbean, Central America, or South America.
Evidence of LLM Grooming
The establishment of the Pravda network in early 2023 occurred approximately six months after the deployment of ChatGPT in November 2022, suggesting a potential link between the rise of commercially available AI chatbots and this new disinformation strategy.
Key Indicators of LLM Grooming:
- Lack of Human Audience: Despite significant geographic and thematic expansion, including English language options and topic localization, the Pravda websites exhibit persistent quality issues (overlapping content, illegible text, auto-translation errors, no search function) and seemingly no human audience. Digital forensics tools indicated fewer than a thousand users visited the US site in December 2024, with average visit times of around 20 seconds.
- Web Scraping Accessibility: While content is not legible to humans, web elements containing text and photos are easily accessible by scraping agents, making them suitable for AI training data.
- Massive Publishing Rate: Estimates suggest a minimum publishing rate of 3.6 million articles per year, with the English language site alone dwarfing the combined output of ten sampled sites. This is likely a gross underestimation.
- Precedent of AI Taint: Previous research by NewsGuard demonstrated that major AI chatbots would parrot Russian disinformation in response to certain prompts, indicating that LLMs were already susceptible to being tainted.
- Counter-Intuitive Behavior: The network's characteristics are bizarre and counter to classical Russian information operations, which typically aim for human engagement. The massive publishing rate, content duplication, heavy automation, and lack of human audience care align perfectly with the hypothesis of targeting automated agents like web crawlers and scraping algorithms.
The LLM Grooming Ecosystem:
The process involves the Pravda network's content being scraped and used to train LLMs. These LLMs then generate outputs that can be further disseminated, potentially by other users or automated agents, leading to a recursive cycle where AI models are trained on AI-generated content.
AI Model Collapse and Societal Consequences
The recursive training of AI models on AI-generated content can lead to "model collapse," where the information ecosystem becomes polluted, making it difficult to find useful or meaningful content. This phenomenon, observed in academic research, suggests that disinformation-riddled AI "slop" could become inescapable and abundant online.
Societal and Technological Implications:
- Widespread Belief in False Information: The sheer size and speed of LLM grooming exponentially amplify the spread of false and harmful information.
- Information Nihilism and Sociopolitical Disengagement: An overwhelming amount of conflicting information (truths, half-truths, lies) can lead to an inability to discern truth, resulting in disengagement from society and political processes.
- Dysfunctional Internet and AI Experiences: LLM grooming can exacerbate AI hallucinations, making models more prone to issuing unhelpful or inaccurate information.
- Information Laundering: The intentional or unintentional echoing of disinformation to promote a worldview, regardless of the source. This can lead to public figures or individuals citing propaganda without critical evaluation.
- Model Duplication by Other Anti-Democratic Actors: The Pravda network's model can be copied by other authoritarian governments (e.g., China, North Korea, Iran) and far-right anti-democratic actors, flooding the internet and LLMs with false content. This presents an "open season" on the intersection of AI and democracy.
Potential Solutions and Safeguards
Addressing LLM grooming requires a multi-faceted approach involving AI companies, governments, and civil society.
Proposed Solutions:
- Data Cleaning and Safeguards by AI Companies: AI companies and researchers must proactively and retroactively clean their training data to prevent problematic content from being included.
- Governmental Requirements: Governments should mandate that AI companies and researchers implement data cleaning and safeguard measures. This raises questions about government as an arbiter of truth, but focusing on provably false claims (e.g., denial of atrocities) is a starting point.
- Information and AI Literacy: Widely available information and digital literacy courses, expanded to include adults and specifically focusing on AI literacy, are crucial. This includes educating individuals on concepts like hallucination, training data, and the reliability of AI-generated content.
- Public Awareness and Advocacy: Governments and civil society organizations must widely warn about the risks of LLM grooming to ensure the issue is not overlooked.
Challenges in Filtering Data:
AI companies are often opaque about their training data. Many willingly do not know the exact contents of their scraped data due to the sheer volume and fear of inadvertently collecting protected information (personal data, intellectual property). This has led to litigation, such as the New York Times suing Microsoft over alleged copyright infringement in training data. Some companies may also go out of their way to avoid knowing the specifics of their data. While some safeguards exist, companies need to do more and be more transparent.
Q&A Highlights
- Pollution of Internet Architecture: The Pravda network's content has been found in AI chatbots, Wikipedia, and X (formerly Twitter) community notes, indicating a broad contamination of online spaces. This can lead to students citing AI-generated content that is influenced by disinformation, and the "illusory truth effect" where repeated exposure to claims leads to belief.
- Using Influenced LLMs: It is highly probable that users are already interacting with LLMs influenced by hostile actors without their knowledge.
- Prompting for Propaganda: Simple prompts about controversial issues, such as transgender athletes, can elicit responses that cite Pravda network materials, demonstrating that extensive prompt engineering is not always necessary.
- Effectiveness of Literacy Courses: While information and digital literacy courses build resilience, they may not fully address general mistrust in the world. Governments also need to actively work on rebuilding trust through their actions.
- English Prompts and Propaganda: English language prompts are likely more prone to delivering Russian propaganda due to the focus of Russian disinformation campaigns on English-speaking countries. However, other languages are also targeted.
- Methodology for LLM Usage: While direct measurement of LLM citation is difficult, the lack of human traffic to Pravda sites, coupled with evidence of AI chatbots citing their content, suggests that automated agents are the primary intended audience.
- Copycat Operations: The risk of other anti-democratic actors copying the Pravda network's LLM grooming model is high, given the history of actors copying successful TTPs.
- Chinese Infops and LLM Grooming: While direct evidence of Chinese infops using LLM grooming is not presented, models like Deepseek exhibit biased responses to sensitive topics, suggesting potential pre-baked disinformation.
- SEO vs. LLM Grooming: SEO aims to improve search engine rankings for businesses, while LLM grooming aims to poison AI training data with disinformation. While both involve strategic content placement, the intent and impact differ significantly, with LLM grooming posing a direct threat to democratic discourse.
- Filtering Training Data: Filtering training data is challenging due to its vastness and the opacity of AI companies. While some safeguards exist, more transparency and proactive measures are needed from AI developers.
Key Concepts
- LLM Grooming: A strategy to poison Large Language Models (LLMs) and generative AI with propaganda and disinformation.
- FIMI (Foreign Information Manipulation and Interference): The broader category of state-sponsored or foreign-backed disinformation and influence operations.
- Pravda Network: A network of websites and social media accounts aggregating Russian propaganda.
- Portal Combat Network: The larger network to which the Pravda network belongs.
- Generative Models: AI models capable of creating new content, such as text, images, or audio.
- Model Collapse: The degradation of AI model performance and output quality due to recursive training on AI-generated content.
- Information Laundering: The process of disseminating disinformation to promote a worldview, regardless of its origin.
- TTPs (Tactics, Techniques, and Procedures): The methods and strategies employed in information operations.
- Information Nihilism: A state of distrust and disengagement resulting from an overwhelming and contradictory information environment.
- AI Literacy: The understanding of how AI models work, their capabilities, limitations, and potential biases.
- Web Scraping: The automated extraction of data from websites.
- Search Engine Optimization (SEO): Techniques used to improve a website's visibility in search engine results.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Why Does This Guy Appear In Kids Videos?
sphynx

TIC en las Organizaciones - Electiva Complementaria II Unisimon
Julieth Güell S

How to Tame Your Advice Monster | Michael Bungay Stanier | TED
TED

Margaret Heffernan: Why it's time to forget the pecking order at work
TED

The importance of psychological safety: Amy Edmondson
The King's Fund

What Is Psychological Safety?
Harvard Business Review

13-Conflict Management: Listening in Conflict
Deliberate Development