This paper provides an extensive overview of existing datasets designed for the AI-based detection of false information, and it presents two new datasets. These datasets contain disinformation cases on climate change and the Russian invasion of Ukraine.
The Debunking datasets were built using a systematic and methodologically rigorous approach to ensure data accuracy and representativeness. Data collection involved the amalgamation of multiple sources and put together a wide range of disinformation and real claims and related multimedia content.
Early tests demonstrate the potential of these datasets to improve the accuracy of fake news classification in these domains. The study concludes with a discussion on future developments and the ethical implications of AI-based fact-checking systems.
The study produced high-quality, domain-specific datasets that address the deficiencies of existing resources. First, a fundamental aspect of the Debunking datasets is domain-specific curation. In fact, unlike general fake news datasets, the AI4Debunk project focuses on two critical areas where disinformation has a profound societal impact. Additionally, Debunking datasets consist of multimodal data. That means the datasets contain images, videos and social media metadata in addition to text statements. Another major strength of the Debunking datasets is their intensive annotation process. The data labeling process is conducted by experienced fact-checkers so that the disinformation claims are accurately labelled.
1. Introduction
The spread of disinformation through the digital network poses a significant threat to public debate, democratic institutions and the exchange of scientific knowledge (D’Andrea et al., 2025). With the growing reliance on social networking sites and online media, disinformation has become a fast and pervasive social phenomenon that can circumvent the boundaries set by the time-consuming processes of fact-checking (Shu et al., 2017). The echo chambers and filter bubbles of algorithmic personalization, which expose users to less obvious corrected information that supports their assumptions, serve to further reinforce this (Del Vicario et al., 2016; D’Andrea et al., 2026).
Consequently, a number of artificial intelligence (AI) strategies have surfaced to deliver pertinent outcomes by combining machine learning and natural language processing (NLP) methods to identify and classify false information (Zhou and Zafarani, 2018). Graph models, deep learning, and ensemble approaches are among the techniques that have shown promise, particularly in identifying complex patterns and linguistic cues connected to false information (Ahmed et al., 2018; Pérez-Rosas et al., 2018; Manos et al., 2024). However, the effectiveness of these models depends on the quality and diversity of the data used for training and testing. Building high-quality datasets is a fundamental requirement, as machine learning methods need to learn large amounts of labeled data in order to understand the fine grain between authentic and fake content. Without large and representative datasets, models may overfit, generalize poorly to real-world scenarios, and become biased, continuing to spread disinformation rather than acting as a brake on it (Thorne et al., 2018; Wang, 2017). The development of large, annotated corpora such as FakeNewsNet (Shu et al., 2020) and LIAR (Wang, 2017) has been crucial to the progress of this research area. Moreover, more recent datasets in the areas of both multilingual and multimodal fake news detection have expanded the landscape of the methodological tools already available. In particular, in relation to the multilingual area, NewsPolyML dataset (Mohtaj et al., 2024) offers over 32,000 English, German, French, Spanish and Italian fact-checked claims and allows supporting multilingual disinformation detection approaches for several European languages. In addition, the MMCFND dataset (Bansal et al., 2024) is a further rich resource containing 28,085 samples of news articles for seven Indic languages, and both textual and visual content to support the low-resource language challenges. In the multimodal domain, the Multi-Fake-DetectiVE dataset (Bondielli et al., 2024) includes 128.611 news articles and 920.054 social media texts, as well as image components, concerning the Russian–Ukrainian war. The MFND corpus (Zhu et al., 2025) offers a multimodal dataset containing both real social news and deepfake news, allowing detection of extremely realistic fake news generated using deep learning methods. Continued improvement of datasets is needed, as it is still difficult to make datasets representative of many topics, languages and sociocultural contexts.
The primary scope of this paper is to investigate the application of databases in automatically detecting fake news. Second, it will circumvent the existing limitations of datasets by two new datasets (hereafter named Debunking datasets) developed under the AI4Debunk [1] project that received funding by the European Commission and aimed at empowering citizens with several fact-checking mechanisms to deter the spread of disinformation. Zubiaga et al. (2016) observed that social media represent one of the major channels characterized by the presence of misleading narratives. In this regard, the proposed datasets are designed with a twofold aim: first, to represent the knowledge for the construction of a semantic knowledge graph to support citizens in navigating the digital media landscape with greater awareness and to make informed decisions; second, to support the training of AI models for disinformation detection. The topic of these datasets is twofold:
Climate change: This has been the topic of ongoing disinformation campaigns, ranging from climate change denial to conspiracy theories and the distortion of scientific facts.
The Russian invasion of Ukraine: This conflict stands out in its extensive employment of information warfare, where disinformation is employed strategically to disseminate propaganda, shape public sentiment and propagate inaccurate information regarding military operations. The use of visual and textual content is also a key component of this multifaceted approach to warfare.
The paper is structured as follows. Section 2 reviews the state of the art in fake news detection datasets, identifying their strengths and limitations and describing the contribution provided by the Debunking datasets. Section 3 describes the methodology adopted to build both datasets, detailing their structure, sources and annotation process. Section 4 discusses the results of a quality assessment of the proposed datasets and presents a comparative analysis with similar existing datasets and a validation over state-of-the-art fake news detection algorithmic approaches. Finally, Section 5 concludes with reflections on future directions and challenges in AI-driven disinformation detection.
2. The need for high-quality fake news detection datasets
High-quality fake news detection datasets have been considered a crucial need to advance disinformation analysis and develop reliable artificial intelligence models. Fake news detection system performance is dependent on having richly annotated, diverse and large-scale datasets mimicking actual-world disinformation dynamics. These datasets need to represent a wide range of topics, linguistic styles and sources so they can be transferred across domains.
To avoid errors and biases, quality annotations via crowdsourcing or expert validation are also required. With additional context information, multimodal data, like text, images and videos, also increase the validity of detection algorithms (Bondielli et al., 2024). As disinformation is dynamic, datasets need to be updated with new trends in a manner that would avoid models from becoming obsolete. Standardized measurements and test standards are also essential in driving both comparability among different detection methods and cooperation and innovation in this field.
2.1 Challenges in existing fake news datasets
AI fake news detection is mostly reliant on datasets used in model training. While datasets such as LIAR (Wang, 2017) and FakeNewsNet (Shu et al., 2020) have had a significant impact on disinformation research, they have certain weaknesses that render them challenging to apply to emerging trends in disinformation. In this regard, one of the most significant challenges involves the absence of temporal relevance, a critical feature that ought to facilitate disinformation detection. Most popular datasets spread outdated disinformation, therefore losing their potential to detect new disinformation trends. Second, fake news stories are not fixed and keep evolving, in most instances responding to political events, societal evolution and emerging technology (Husejinovic and Babic, 2023; Vosoughi et al., 2018). With sporadic updates, AI models trained on old datasets fail to generalize to new patterns of disinformation. A second limitation is the emphasis on text-based data in comparison to multimodal content. That is, disinformation is no longer text but rather includes images, videos, and manipulated content on social media (Nooralahzadeh et al., 2021). However, most of today’s datasets have primarily text claims or text articles, thus overlooking the very critical contribution that contextual and image information play in the spread of disinformation.
Furthermore, domain-specific disinformation is another major problem. In fact, while general fake news datasets provide a preview of what disinformation resembles, they do not really capture well the nature of domain-specific fake news. For instance, climate change disinformation most likely exploits scientific uncertainty, distorts data and ultimately manipulates expert opinion. Similarly, war propaganda is largely associated with international conflicts and includes strategic disinformation campaigns aimed at shaping public opinion (Ferreira and Vlachos, 2016). General-purpose datasets fail to capture such domain-specific tactics, and thus, they are not as effective in detecting disinformation in specialized domains. Furthermore, additional limitations exist in the available datasets. In fact, the majority of datasets, such as LIAR, FakeNewsNet, Multi-Fake-Detective, NewsPolyML and MMCFND, do not reflect the peculiarities of disinformation in specialized domains. The Debunking datasets mitigate these limitations by offering current and pertinent subjects of climate change and the Russian invasion of Ukraine. Lastly, potential bias during annotation and labeling inconsistencies pose a significant threat. Consequently, some of the existing datasets employ manual fact-checking processes that incorporate subjective bias. In particular, the political or cultural ideology of fact-checkers influences them to decide and leads to inconsistency in labelling (Shu et al., 2020). These cognitive biases have the potential to inadvertently influence the development of artificial intelligence models to support some intuitions and undermine others. This phenomenon can illegitimize the application of automated tools for the identification of fake news.
2.2 The contribution of debunking datasets
Given the open issues and challenges described in Section 2.1, this study aims to produce high-quality, domain-specific datasets that address the deficiencies of existing resources. The Debunking datasets support disinformation detection in two significant domains: climate change and war propaganda.
In this regard, a fundamental aspect of the Debunking datasets is domain-specific curation. In fact, unlike general fake news datasets, the AI4Debunk project focuses on two critical areas where disinformation has a profound societal impact. By targeting these domains, the project ensures that AI models are trained on content that accurately reflects real-world disinformation strategies, making them more effective in detecting domain-specific fake news.
In addition, the Debunking datasets consist of multimodal data. That means the datasets contain images, videos and social media metadata in addition to text statements. Multimodality in datasets enables AI models to learn from disinformation in a more integrated perspective, thereby better able to identify different kinds of fake news.
Another major strength of the Debunking datasets is their intensive annotation process. The data labeling process is conducted by experienced fact-checkers so that the disinformation claims are accurately labeled. To reduce bias, the annotation process is conducted based on standardized criteria and with several rounds of validation. In addition, the research encourages open-source availability since the datasets are created by utilizing MYSQL. In fact, the datasets are released for collaborative research and development so that fact-checking groups, academic researchers and AI practitioners can use them for disinformation detection studies.
By making the datasets widely accessible, Debunking datasets contribute to the broader fight against disinformation by enabling continuous improvements and refinements in AI-based fake news detection models. In the context of the AI4Debunkscope, the datasets contained 2,000 cases of disinformation (only fake news) on climate change and the Russian invasion of Ukraine, as the primary scope of these datasets was to represent the knowledge base for the construction of a semantic knowledge graph to be used for contextual reasoning over the disinformation space. Beyond the AI4Debunk scope, these two datasets were extended with an additional 2,000 real news articles on climate change and the Russian invasion of Ukraine that can be used both to train the AI detection algorithms and to create explainable AI systems that can navigate complex scientific narratives and verify their fact consistency through a graph-based model of knowledge.
In summary, the novelty of the Debunking datasets relies on domain specificity, multimodality and annotation rigor. Specifically, they specialize in high-impact topics, i.e. war propaganda and climate change, with domain specificity followed by breadth of coverage. This makes them highly relevant to pressing real-world issues. Moreover, compared to the majority of existing datasets, which are text-only or text-based, they combine various modalities (text, images, videos). Finally, their strict, standardized and multi-tested annotations provide higher reliability and transparency compared to moderate or low-quality control datasets.
3. Methodology
This section describes the format and implementation of the datasets. It provides the template used in retrieving information from disinformation instances, the Database Management System (DBMS) implemented, and the cloud environment hosting the multimedia content storage and sharing. In both datasets, there are 2,000 instances of real and fake news that were randomly selected and labeled to reflect various misleading stories in various digital sources. The false news articles were retrieved from fact-checking sites (as described in Section 3.3) by choosing randomly items published between 2020 and 2024 for the Russian invasion of Ukraine and from 2000 to 2024 for climate change (in some cases, all the news was retrieved from the fact-checking websites). For the real news, articles were retrieved randomly from verified and reputable open-source news websites (newspapers, broadcast outlets, international news agencies and scientific journals) by selecting news published between 2023 and 2024 for both datasets. From these sources, 2,000 real and 2,000 fake news items were extracted, ensuring that both the false and real news entries originated from sections dedicated to the two specific themes within reliable fact-checking platforms and the verified and reputable open-source news websites. In particular, the construction of the datasets was executed with the aim of capturing a sample that is representative of recurring disinformation strategies and themes. This was done to reflect the complexity and variability of false or manipulated content that proliferates in the information ecosystem. These include, but are not limited to, narratives that deny or distort scientifically validated facts, exaggerate geopolitical threats, discredit public institutions, undermine democratic values and promote conspiracy theories.
Disinformation is defined by the enormous array of rhetorical tactics, from affective appeals and terror-posting to deliberate embrace of half-truths, fabricated facts and fake images. Material was collected from various sources, including social networks, websites, blogs and multimedia repositories. The approach facilitated exhaustive evaluation of the different types of content and patterns of circulation. Every occurrence was meticulously documented using a standard template that contained all necessary metadata, including the language, media type, publication date, source platform and related keywords.
The template also included a description of the disinformation narrative and the rationale behind its classification. For each theme of the datasets, Table 1 lists the most relevant topics of fake news collected.
Fake news topics
| Dataset theme | Fake news topics |
|---|---|
| Russian invasion of Ukraine |
|
| Climate change |
|
| Dataset theme | Fake news topics |
|---|---|
| Russian invasion of Ukraine | Third World War narrative Biological disaster scenarios Attacks on the Ukrainian army and President Zelensky Denial of conflict authenticity Disinformation about refugees Undermining territorial integrity and sovereignty Military and security threat exaggeration Human rights violation claims Internal governance disinformation Ukrainian counter-offensive distortion Delegitimizing the Ukrainian government Stereotyping Ukrainian citizens negatively Justifying Russian invasion Attacks on NATO and Ukraine’s Allies Undermining international alliances Attacks on Ukrainian Armed Forces Religious and cultural division exploitation Economic crisis exaggeration in Ukraine Claims of reduced Western military support Demoralization campaigns targeting the Ukrainian population NATO/US sending Bulgaria to war with Russia Ukraine is portrayed as corrupt and undeserving of help |
| Climate change | Solar activity is responsible for climate change Misleading CO2 role narratives Denial of global warming and ice melting Electric vehicle disinformation and criticism Climate and agricultural disinformation Attacks on the Discredit Greta Thunberg Volcano activity is the reason for climate change Chemtrails and geoengineering conspiracy theories Wind turbine conspiracy theories General climate conspiracy theories Manipulation of scientific data Scepticism of scientific consensus Denial of climate change Discrediting renewable sources of energy Misleading policy impact narratives Undermining European environmental action Greenwashing accusations Economic alarmism with regard to climate policies Political climate disinformation “Climate change is a hoax” claims “ Bulgarian politicians depicted as servile to the |
The technical infrastructure of the datasets includes a relational DBMS optimized for complex query handling and metadata indexing. The system has been engineered to facilitate efficient storage, retrieval and update operations, thereby supporting both research and applied analysis. To ensure scalability and collaborative access, the datasets are hosted in a secure cloud environment that supports multimedia content management. This environment enables remote users to navigate, download and contribute to the datasets in compliance with data protection standards.
By aggregating 2000 disinformation cases per dataset, the initiative offers a valuable resource for researchers, developers, educators and policymakers to better understand the nature and spread of disinformation, design targeted interventions and train automated systems for debunking and disinformation detection.
Moreover, the collection of a large variety of sources, including social network content and multimedia repositories, expanded the scope of the datasets to fit the main requirements of modern information ecosystems. In this regard, this type of narrative description allowed an integrated approach of both qualitative and quantitative assessments, allowing for multi-layered analysis of disinformation content.
This approach has been relevant in order to properly spot subtle manipulative techniques, such as selective framing, exaggeration and even strategic uses of visual media.
In this way, it has been possible to gain a better understanding of the way in which misleading narratives are emerging and evolving across a wide variety of platforms.
Finally, the adoption of cloud-based storage services provided a high level of scalability, data security, and accessibility. In fact, the adopted infrastructures allowed efficient data retrieval and secure storage of multimedia content. Furthermore, this approach is designed around data sharing, incentivizing the collaborative work among researchers, as the datasets can remain a sustainable and reusable resource available for upcoming studies in disinformation research.
3.1 Template for information extraction
The information extraction process was designed considering some descriptive criteria; that is, some of the elements of the fake news have been selected to convert them into dataset fields. For this purpose, a specific template was created. Table 2 presents the template used to extract and classify information from disinformation cases. The template represents the foundation for the two datasets developed and is intended to offer consistency, completeness and conformity with machine learning applications.
Template for information extraction
| Field | Description | Value |
|---|---|---|
| (TOPIC) | Topic of the disinformation (Russian invasion in Ukraine or climate changes) | TEXT |
| (KEYWORDS) | Keywords of the disinformation | TEXT |
| (DATE) | Date of publication of the disinformation | DATE |
| (SOURCE) | Media/platform/website reporting the disinformation | TEXT |
| (URL) | Link to the disinformation | URL |
| (LANGUAGE) | Language of the disinformation | TEXT |
| (AUTHOR) | Author of the disinformation | TEXT |
| (RATING SCALE) | Rating scale used to assess truthfulness | TEXT |
| (TEXT) | Textual statement of the disinformation in English | TEXT |
| (WHY) | Fact-checking analysis in English | TEXT |
| (TEXT: SOURCE LANGUAGE) | Textual statement in the source language | TEXT |
| (WHY: SOURCE LANGUAGE) | Fact-checking analysis in the source language | TEXT |
| (MULTIMEDIA) | Audios, videos and images related to the disinformation | URL |
| Field | Description | Value |
|---|---|---|
| ( | Topic of the disinformation (Russian invasion in Ukraine or climate changes) | |
| (KEYWORDS) | Keywords of the disinformation | |
| ( | Date of publication of the disinformation | |
| (SOURCE) | Media/platform/website reporting the disinformation | |
| ( | Link to the disinformation | |
| (LANGUAGE) | Language of the disinformation | |
| (AUTHOR) | Author of the disinformation | |
| ( | Rating scale used to assess truthfulness | |
| ( | Textual statement of the disinformation in English | |
| (WHY) | Fact-checking analysis in English | |
| (TEXT: SOURCE LANGUAGE) | Textual statement in the source language | |
| (WHY: SOURCE LANGUAGE) | Fact-checking analysis in the source language | |
| (MULTIMEDIA) | Audios, videos and images related to the disinformation |
Each entry in the dataset contains multiple fields, each of which corresponds to some specific attribute of the disinformation claim. The fields were derived from frameworks for characterizing fake news that are available in the literature, notably, those proposed by Zhang and Ghorbani (2019) and later built upon by D’Ulizia et al. (2021). The goal is to cover not just the content of the disinformation, but also its source, context and multimedia trace.
The template is useful for standard data analysis but also for semantic enrichment. For example, it helps in creating Knowledge Graphs from structured entity relations for predicting and classifying novel information. This enhances the usefulness of the dataset to AI models, particularly natural language understanding, fact-checking and cross-lingual reasoning.
The field “multimedia” contains the URL to a folder available on the cloud storage services Mediafire (https://www.mediafire.com) or Google Drive, where the multimedia files (PNG, JPEG, MP4, etc.) are stored.
All the information extracted is meant to be suitable for data processing and machine learning (ML) training during the next steps of the AI4Debunk project. In particular, the fields (TOPIC), (KEYWORDS), (DATE), (SOURCE), (URL), (LANGUAGE) and (AUTHOR) have been collected for a proper assessment of the disinformation, while the remaining fields, such as (RATING SCALE), (TEXT), (WHY) and (MULTIMEDIA) would contribute to train ML algorithms and to develop a semantic Knowledge Graph. Note that the rating scale is assigned to the claims according to two different methods: (1) by extracting fake and real claims from certified and recognized fact-checking websites (please see sub-section 3.3); (2) by debunking fake claims with evidence and statements from reliable sources by expert journalists. The professional profiles in question are affiliated with two companies that are partners in the AI4Debunk project: Euractiv Bulgaria and Internews Ukraine.
Furthermore, the fields (TEXT: SOURCE LANGUAGE) and (WHY: SOURCE LANGUAGE) have been added to provide multilingual support to the Knowledge Graph.
3.2 Database management system (DBMS)
To identify the reference database management system (DBMS) used for the implementation of the two datasets, the study provided by Taipalus (2023) has been analyzed, which provides a comparison of the performance of different DBMSs. The MySQL (https://www.mysql.com/) DBMS has been selected mainly due to its overall performance, scalability and reliability. Being an open-source DBMS, MySQL is inexpensive to use with no relevant compromise in features or support. Being highly compatible with a very large majority of platforms and programming languages, it assures integration with the current technology stack. Moreover, data exchange in common and readable file formats is allowed.
Also, capabilities in MySQL towards high availability and data security, such as replication and automatic backup, ensure that the data remains consistent and downtimes are minimal, thus ideal for mission-critical use. With the high-transactional and large-scale database maintenance reputation of MySQL, these factors sealed its deployment in the AI4Debunk project.
For each of the two case studies of the project (the Russian invasion of Ukraine and climate change), a specific dataset model, based on the template described in Section 3.1, was designed. Each dataset was built around a single table, where each claim has been associated with all the fields already present on the approved template (see Table 2).
Figures 1 and 2 provide some visual examples of the dataset model and table structure.
The interface shows a database window labeled “Physical Schemas” positioned at the top left of the screen. Directly below it appears a database cylinder icon placed to the left of the schema name “Climate change Fake news C N R”, with the subtitle “MySQL Schema” displayed underneath the schema name. Below the schema header, the section title “Tables (1 item)” appears aligned to the left. Immediately beneath it, a plus icon appears to the left of the label “Add Table”. To the right within the same tables area, a small table icon appears next to the listed table name “Climate Change Fake news C N R”. Directly below the tables section, the label “Views (0 items)” appears aligned to the left. Above the table structure grid, a horizontal tab appears showing the opened table labeled “Climate Change Fake news C N R”. Within the table structure grid, under the “Datatype” column, a small calendar icon appears aligned inside the cell corresponding to the “D A T E” row where the datatype is “D A T E”. On the right side above the structure grid, the label “Schema: Climate change Fake news C N R” appears aligned to the right side of the interface. The table structure view shows columns under the headings “Column Name”, “Datatype”, “P K”, “N N”, “U Q”, “B”, “U N”, “Z F”, “A I”, “G”, and “Default or Expression”. The rows appear as follows: Row 1: Column Name: Author; Datatype: V A R C H A R(100); P K: not selected; N N: selected; U Q to G: not selected. Row 2: Column Name: Country; Datatype: V A R C H A R(100); P K: not selected; N N: selected; U Q to G: not selected. Row 3: Column Name: D A T E; Datatype: D A T E; P K: not selected; N N: selected; U Q to G: not selected. Row 4: Column Name: Keywords; Datatype: V A R C H A R(1000); P K: not selected; N N: selected; U Q to G: not selected. Row 5: Column Name: Language; Datatype: V A R C H A R(100); P K: not selected; N N: selected; U Q to G: not selected. Row 6: Column Name: Multimedia; Datatype: V A R C H A R(1000); P K to G: not selected. Row 7: Column Name: Rating scale; Datatype: V A R C H A R(45); P K: not selected; N N: selected; U Q: not selected; B: selected; U N to G: not selected. Row 8: Column Name: Source; Datatype: V A R C H A R(500); P K: not selected; N N: selected; U Q to G: not selected. Row 9: Column Name: Text; Datatype: V A R C H A R(10000); P K: not selected; N N: selected; U Q to G: not selected. Row 10: Column Name: Url; Datatype: V A R C H A R(10000); P K: not selected; N N: selected; U Q to G: not selected. Row 11: Column Name: Why; Datatype: V A R C H A R(10000); P K: not selected; N N: selected; U Q to G: not selected. Row 12: Column Name: Topic; Datatype: V A R C H A R(100); P K: not selected; N N: selected; U Q to G: not selected. Row 13: Column Name: Text (source language); Datatype: V A R C H A R(10000); P K: not selected; N N: selected; U Q to G: not selected. Row 14: Column Name: Why (source language); Datatype: V A R C H A R(10000); P K: not selected; N N: selected; U Q to G: not selected.Dataset model structure
The interface shows a database window labeled “Physical Schemas” positioned at the top left of the screen. Directly below it appears a database cylinder icon placed to the left of the schema name “Climate change Fake news C N R”, with the subtitle “MySQL Schema” displayed underneath the schema name. Below the schema header, the section title “Tables (1 item)” appears aligned to the left. Immediately beneath it, a plus icon appears to the left of the label “Add Table”. To the right within the same tables area, a small table icon appears next to the listed table name “Climate Change Fake news C N R”. Directly below the tables section, the label “Views (0 items)” appears aligned to the left. Above the table structure grid, a horizontal tab appears showing the opened table labeled “Climate Change Fake news C N R”. Within the table structure grid, under the “Datatype” column, a small calendar icon appears aligned inside the cell corresponding to the “D A T E” row where the datatype is “D A T E”. On the right side above the structure grid, the label “Schema: Climate change Fake news C N R” appears aligned to the right side of the interface. The table structure view shows columns under the headings “Column Name”, “Datatype”, “P K”, “N N”, “U Q”, “B”, “U N”, “Z F”, “A I”, “G”, and “Default or Expression”. The rows appear as follows: Row 1: Column Name: Author; Datatype: V A R C H A R(100); P K: not selected; N N: selected; U Q to G: not selected. Row 2: Column Name: Country; Datatype: V A R C H A R(100); P K: not selected; N N: selected; U Q to G: not selected. Row 3: Column Name: D A T E; Datatype: D A T E; P K: not selected; N N: selected; U Q to G: not selected. Row 4: Column Name: Keywords; Datatype: V A R C H A R(1000); P K: not selected; N N: selected; U Q to G: not selected. Row 5: Column Name: Language; Datatype: V A R C H A R(100); P K: not selected; N N: selected; U Q to G: not selected. Row 6: Column Name: Multimedia; Datatype: V A R C H A R(1000); P K to G: not selected. Row 7: Column Name: Rating scale; Datatype: V A R C H A R(45); P K: not selected; N N: selected; U Q: not selected; B: selected; U N to G: not selected. Row 8: Column Name: Source; Datatype: V A R C H A R(500); P K: not selected; N N: selected; U Q to G: not selected. Row 9: Column Name: Text; Datatype: V A R C H A R(10000); P K: not selected; N N: selected; U Q to G: not selected. Row 10: Column Name: Url; Datatype: V A R C H A R(10000); P K: not selected; N N: selected; U Q to G: not selected. Row 11: Column Name: Why; Datatype: V A R C H A R(10000); P K: not selected; N N: selected; U Q to G: not selected. Row 12: Column Name: Topic; Datatype: V A R C H A R(100); P K: not selected; N N: selected; U Q to G: not selected. Row 13: Column Name: Text (source language); Datatype: V A R C H A R(10000); P K: not selected; N N: selected; U Q to G: not selected. Row 14: Column Name: Why (source language); Datatype: V A R C H A R(10000); P K: not selected; N N: selected; U Q to G: not selected.Dataset model structure
The structured dataset of climate change misinformation cases is organized by “identifier”, “author”, “country”, “date”, “keywords”, “language”, “multimedia presence”, “rating scale”, “source”, “claim text”, “link”, “justification”, “topic”, and “corresponding source-language fields”. The row-wise summaries are as follows: Row 2: idat: 2; Author: NA; Country: France; Date: 01/08/2022; Keywords: climate protest ellipsis; Language: French; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: AFP Factuel; Text: A video clip ellipsis; Uniform Resource Locator: https://fa ellipsis; Why: This clip sh ellipsis; Topic: Climate change; Text (source language): French; Why (source language): French. Row 3: idat: 3; Author: NA; Country: Lithuania; Date: 23/04/2022; Keywords: video, pro ellipsis; Language: Lithuanian; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: Delfi; Text: The video, ellipsis; Uniform Resource Locator: https:// ellipsis; Why: Video of a ellipsis; Topic: Climate change; Text (source language): Lithuanian; Why (source language): Lithuanian. Row 4: idat: 4; Author: NA; Country: Germany; Date: 22/03/2022; Keywords: body bags ellipsis; Language: German; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: Correctiv; Text: This video ellipsis; Uniform Resource Locator: https://co ellipsis; Why: The video ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 5: idat: 6; Author: Xavier Trias; Country: Spain; Date: 31/05/2024; Keywords: cars, pollu ellipsis; Language: English; Multimedia: N/A; Rating scale: Fake; Source: eufactcheck; Text: Cars are r ellipsis; Uniform Resource Locator: https://eu ellipsis; Why: As we hav ellipsis; Topic: Climate change; Text (source language): English; Why (source language): English. Row 6: idat: 7; Author: Neva Zgan; Country: Croatia; Date: 16/04/2024; Keywords: Croatia, cl ellipsis; Language: Croatian; Multimedia: N/A; Rating scale: Fake; Source: eufactcheck; Text: Migrants ellipsis; Uniform Resource Locator: https://eu ellipsis; Why: There is n ellipsis; Topic: Climate change; Text (source language): Croatian; Why (source language): Croatian. Row 7: idat: 8; Author: Donald Tr; Country: USA; Date: 25/12/2023; Keywords: global war ellipsis; Language: English; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: eufactcheck; Text: The conce ellipsis; Uniform Resource Locator: https://eu ellipsis; Why: We haven ellipsis; Topic: Climate change; Text (source language): English; Why (source language): English. Row 8: idat: 9; Author: Stephan B; Country: Germany; Date: 27/05/2023; Keywords: Germany, w ellipsis; Language: German; Multimedia: N/A; Rating scale: Fake; Source: eufactcheck; Text: Germany ellipsis; Uniform Resource Locator: https://eu ellipsis; Why: Stephan B ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 9: idat: 10; Author: Vivek Ram; Country: USA; Date: 15/09/2023; Keywords: Ramaswam ellipsis; Language: English; Multimedia: N/A; Rating scale: Fake; Source: FactCheck; Text: A former ellipsis; Uniform Resource Locator: https://w ellipsis; Why: Republican ellipsis; Topic: Climate change; Text (source language): English; Why (source language): English. Row 10: idat: 11; Author: Steve Mill; Country: USA; Date: 26/01/2023; Keywords: Global Ter ellipsis; Language: English; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: Factcheck; Text: Steve Mill ellipsis; Uniform Resource Locator: https://w ellipsis; Why: That’s w ellipsis; Topic: Climate change; Text (source language): English; Why (source language): English. Row 11: idat: 12; Author: Donald Tr; Country: USA; Date: 05/09/2020; Keywords: Climate, W ellipsis; Language: English; Multimedia: N/A; Rating scale: Fake; Source: Factcheck; Text: In a brief ellipsis; Uniform Resource Locator: https://w ellipsis; Why: The Earth ellipsis; Topic: Climate change; Text (source language): English; Why (source language): English. Row 12: idat: 13; Author: Donald Tr; Country: USA; Date: 12/12/2019; Keywords: Global Wa ellipsis; Language: English; Multimedia: N/A; Rating scale: Fake; Source: Factcheck; Text: At a cam ellipsis; Uniform Resource Locator: https://w ellipsis; Why: According ellipsis; Topic: Climate change; Text (source language): English; Why (source language): English. Row 13: idat: 14; Author: Kamala Ha; Country: USA, Fran; Date: 22/08/2019; Keywords: Amazon, W ellipsis; Language: English; Multimedia: N/A; Rating scale: Fake; Source: Factcheck; Text: The Amaz ellipsis; Uniform Resource Locator: https://w ellipsis; Why: No. Scien ellipsis; Topic: Climate change; Text (source language): English; Why (source language): English. Row 14: idat: 15; Author: Jørgen Pe; Country: Germany; Date: 29/08/2023; Keywords: Global Wa ellipsis; Language: German; Multimedia: N/A; Rating scale: Fake; Source: Correctiv; Text: Jørgen Pe ellipsis; Uniform Resource Locator: https://co ellipsis; Why: Steffense ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 15: idat: 16; Author: NA; Country: Germany; Date: 09/02/2024; Keywords: Floods, W ellipsis; Language: German; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: Correctiv; Text: A photo w ellipsis; Uniform Resource Locator: https://co ellipsis; Why: There hav ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 16: idat: 17; Author: NA; Country: Germany; Date: 25/01/2024; Keywords: Wind turb ellipsis; Language: German; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: Correctiv; Text: On Facebo ellipsis; Uniform Resource Locator: https://co ellipsis; Why: A nuclear ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 17: idat: 18; Author: Heute Jour; Country: Germany; Date: 21/12/2023; Keywords: Global Wa ellipsis; Language: German; Multimedia: N/A; Rating scale: Fake; Source: Correctiv; Text: In a news ellipsis; Uniform Resource Locator: https://co ellipsis; Why: The map wa ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 18: idat: 19; Author: NA; Country: Germany; Date: 21/12/2023; Keywords: E-cars, pr ellipsis; Language: German; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: Correctiv; Text: The elect ellipsis; Uniform Resource Locator: https://co ellipsis; Why: The price ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 19: idat: 20; Author: NA; Country: Germany; Date: 30/10/2023; Keywords: Ukraine, W ellipsis; Language: German; Multimedia: N/A; Rating scale: Fake; Source: Correctiv; Text: The Univer ellipsis; Uniform Resource Locator: https://co ellipsis; Why: Heidelber ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 20: idat: 21; Author: Robert Far; Country: Germany; Date: 23/04/2023; Keywords: Germany, C ellipsis; Language: German; Multimedia: N/A; Rating scale: Fake; Source: Correctiv; Text: 78 percen ellipsis; Uniform Resource Locator: https://co ellipsis; Why: It is true ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 21: idat: 22; Author: NA; Country: Germany; Date: 26/06/2023; Keywords: Wind farm ellipsis; Language: German; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: Correctiv; Text: According ellipsis; Uniform Resource Locator: https://co ellipsis; Why: The study ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 22: idat: 23; Author: NA; Country: Germany; Date: 22/06/2023; Keywords: Heat prot ellipsis; Language: German; Multimedia: N/A; Rating scale: Fake; Source: Correctiv; Text: The next lo ellipsis; Uniform Resource Locator: https://co ellipsis; Why: Health Min ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 23: idat: 24; Author: NA; Country: Germany; Date: 07/06/2023; Keywords: Heat wave; Language: German; Multimedia: N/A; Rating scale: Fake; Source: Correctiv; Text: According ellipsis; Uniform Resource Locator: https://co ellipsis; Why: Kachelman ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 24: idat: 25; Author: NA; Country: Germany; Date: 07/06/2023; Keywords: Cold, Anta ellipsis; Language: German; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: Correctiv; Text: A new col ellipsis; Uniform Resource Locator: https://co ellipsis; Why: It is true ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 25: idat: 26; Author: NA; Country: Germany; Date: 15/09/2022; Keywords: Wind turb ellipsis; Language: German; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: Correctiv; Text: Photos sh ellipsis; Uniform Resource Locator: https://co ellipsis; Why: The forest ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 26: idat: 27; Author: NA; Country: Germany; Date: 09/09/2022; Keywords: Trees, Wi ellipsis; Language: German; Multimedia: N/A; Rating scale: Fake; Source: Correctiv; Text: To save th ellipsis; Uniform Resource Locator: https://co ellipsis; Why: The trees ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 27: idat: 28; Author: Harald La; Country: Germany; Date: 08/09/2022; Keywords: Alpine gla ellipsis; Language: German; Multimedia: N/A; Rating scale: Fake; Source: Correctiv; Text: A simulati ellipsis; Uniform Resource Locator: https://co ellipsis; Why: The retreat ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 28: idat: 29; Author: Dieter Nu; Country: Germany; Date: 09/08/2022; Keywords: CO2 footp ellipsis; Language: German; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: Correctiv; Text: Dieter Nu ellipsis; Uniform Resource Locator: https://co ellipsis; Why: The quote ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 29: idat: 30; Author: NA; Country: Germany; Date: 21/07/2022; Keywords: Climate ch ellipsis; Language: German; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: Correctiv; Text: A graphic ellipsis; Uniform Resource Locator: https://co ellipsis; Why: The graph ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German.Sample of dataset table
The structured dataset of climate change misinformation cases is organized by “identifier”, “author”, “country”, “date”, “keywords”, “language”, “multimedia presence”, “rating scale”, “source”, “claim text”, “link”, “justification”, “topic”, and “corresponding source-language fields”. The row-wise summaries are as follows: Row 2: idat: 2; Author: NA; Country: France; Date: 01/08/2022; Keywords: climate protest ellipsis; Language: French; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: AFP Factuel; Text: A video clip ellipsis; Uniform Resource Locator: https://fa ellipsis; Why: This clip sh ellipsis; Topic: Climate change; Text (source language): French; Why (source language): French. Row 3: idat: 3; Author: NA; Country: Lithuania; Date: 23/04/2022; Keywords: video, pro ellipsis; Language: Lithuanian; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: Delfi; Text: The video, ellipsis; Uniform Resource Locator: https:// ellipsis; Why: Video of a ellipsis; Topic: Climate change; Text (source language): Lithuanian; Why (source language): Lithuanian. Row 4: idat: 4; Author: NA; Country: Germany; Date: 22/03/2022; Keywords: body bags ellipsis; Language: German; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: Correctiv; Text: This video ellipsis; Uniform Resource Locator: https://co ellipsis; Why: The video ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 5: idat: 6; Author: Xavier Trias; Country: Spain; Date: 31/05/2024; Keywords: cars, pollu ellipsis; Language: English; Multimedia: N/A; Rating scale: Fake; Source: eufactcheck; Text: Cars are r ellipsis; Uniform Resource Locator: https://eu ellipsis; Why: As we hav ellipsis; Topic: Climate change; Text (source language): English; Why (source language): English. Row 6: idat: 7; Author: Neva Zgan; Country: Croatia; Date: 16/04/2024; Keywords: Croatia, cl ellipsis; Language: Croatian; Multimedia: N/A; Rating scale: Fake; Source: eufactcheck; Text: Migrants ellipsis; Uniform Resource Locator: https://eu ellipsis; Why: There is n ellipsis; Topic: Climate change; Text (source language): Croatian; Why (source language): Croatian. Row 7: idat: 8; Author: Donald Tr; Country: USA; Date: 25/12/2023; Keywords: global war ellipsis; Language: English; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: eufactcheck; Text: The conce ellipsis; Uniform Resource Locator: https://eu ellipsis; Why: We haven ellipsis; Topic: Climate change; Text (source language): English; Why (source language): English. Row 8: idat: 9; Author: Stephan B; Country: Germany; Date: 27/05/2023; Keywords: Germany, w ellipsis; Language: German; Multimedia: N/A; Rating scale: Fake; Source: eufactcheck; Text: Germany ellipsis; Uniform Resource Locator: https://eu ellipsis; Why: Stephan B ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 9: idat: 10; Author: Vivek Ram; Country: USA; Date: 15/09/2023; Keywords: Ramaswam ellipsis; Language: English; Multimedia: N/A; Rating scale: Fake; Source: FactCheck; Text: A former ellipsis; Uniform Resource Locator: https://w ellipsis; Why: Republican ellipsis; Topic: Climate change; Text (source language): English; Why (source language): English. Row 10: idat: 11; Author: Steve Mill; Country: USA; Date: 26/01/2023; Keywords: Global Ter ellipsis; Language: English; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: Factcheck; Text: Steve Mill ellipsis; Uniform Resource Locator: https://w ellipsis; Why: That’s w ellipsis; Topic: Climate change; Text (source language): English; Why (source language): English. Row 11: idat: 12; Author: Donald Tr; Country: USA; Date: 05/09/2020; Keywords: Climate, W ellipsis; Language: English; Multimedia: N/A; Rating scale: Fake; Source: Factcheck; Text: In a brief ellipsis; Uniform Resource Locator: https://w ellipsis; Why: The Earth ellipsis; Topic: Climate change; Text (source language): English; Why (source language): English. Row 12: idat: 13; Author: Donald Tr; Country: USA; Date: 12/12/2019; Keywords: Global Wa ellipsis; Language: English; Multimedia: N/A; Rating scale: Fake; Source: Factcheck; Text: At a cam ellipsis; Uniform Resource Locator: https://w ellipsis; Why: According ellipsis; Topic: Climate change; Text (source language): English; Why (source language): English. Row 13: idat: 14; Author: Kamala Ha; Country: USA, Fran; Date: 22/08/2019; Keywords: Amazon, W ellipsis; Language: English; Multimedia: N/A; Rating scale: Fake; Source: Factcheck; Text: The Amaz ellipsis; Uniform Resource Locator: https://w ellipsis; Why: No. Scien ellipsis; Topic: Climate change; Text (source language): English; Why (source language): English. Row 14: idat: 15; Author: Jørgen Pe; Country: Germany; Date: 29/08/2023; Keywords: Global Wa ellipsis; Language: German; Multimedia: N/A; Rating scale: Fake; Source: Correctiv; Text: Jørgen Pe ellipsis; Uniform Resource Locator: https://co ellipsis; Why: Steffense ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 15: idat: 16; Author: NA; Country: Germany; Date: 09/02/2024; Keywords: Floods, W ellipsis; Language: German; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: Correctiv; Text: A photo w ellipsis; Uniform Resource Locator: https://co ellipsis; Why: There hav ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 16: idat: 17; Author: NA; Country: Germany; Date: 25/01/2024; Keywords: Wind turb ellipsis; Language: German; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: Correctiv; Text: On Facebo ellipsis; Uniform Resource Locator: https://co ellipsis; Why: A nuclear ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 17: idat: 18; Author: Heute Jour; Country: Germany; Date: 21/12/2023; Keywords: Global Wa ellipsis; Language: German; Multimedia: N/A; Rating scale: Fake; Source: Correctiv; Text: In a news ellipsis; Uniform Resource Locator: https://co ellipsis; Why: The map wa ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 18: idat: 19; Author: NA; Country: Germany; Date: 21/12/2023; Keywords: E-cars, pr ellipsis; Language: German; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: Correctiv; Text: The elect ellipsis; Uniform Resource Locator: https://co ellipsis; Why: The price ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 19: idat: 20; Author: NA; Country: Germany; Date: 30/10/2023; Keywords: Ukraine, W ellipsis; Language: German; Multimedia: N/A; Rating scale: Fake; Source: Correctiv; Text: The Univer ellipsis; Uniform Resource Locator: https://co ellipsis; Why: Heidelber ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 20: idat: 21; Author: Robert Far; Country: Germany; Date: 23/04/2023; Keywords: Germany, C ellipsis; Language: German; Multimedia: N/A; Rating scale: Fake; Source: Correctiv; Text: 78 percen ellipsis; Uniform Resource Locator: https://co ellipsis; Why: It is true ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 21: idat: 22; Author: NA; Country: Germany; Date: 26/06/2023; Keywords: Wind farm ellipsis; Language: German; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: Correctiv; Text: According ellipsis; Uniform Resource Locator: https://co ellipsis; Why: The study ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 22: idat: 23; Author: NA; Country: Germany; Date: 22/06/2023; Keywords: Heat prot ellipsis; Language: German; Multimedia: N/A; Rating scale: Fake; Source: Correctiv; Text: The next lo ellipsis; Uniform Resource Locator: https://co ellipsis; Why: Health Min ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 23: idat: 24; Author: NA; Country: Germany; Date: 07/06/2023; Keywords: Heat wave; Language: German; Multimedia: N/A; Rating scale: Fake; Source: Correctiv; Text: According ellipsis; Uniform Resource Locator: https://co ellipsis; Why: Kachelman ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 24: idat: 25; Author: NA; Country: Germany; Date: 07/06/2023; Keywords: Cold, Anta ellipsis; Language: German; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: Correctiv; Text: A new col ellipsis; Uniform Resource Locator: https://co ellipsis; Why: It is true ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 25: idat: 26; Author: NA; Country: Germany; Date: 15/09/2022; Keywords: Wind turb ellipsis; Language: German; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: Correctiv; Text: Photos sh ellipsis; Uniform Resource Locator: https://co ellipsis; Why: The forest ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 26: idat: 27; Author: NA; Country: Germany; Date: 09/09/2022; Keywords: Trees, Wi ellipsis; Language: German; Multimedia: N/A; Rating scale: Fake; Source: Correctiv; Text: To save th ellipsis; Uniform Resource Locator: https://co ellipsis; Why: The trees ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 27: idat: 28; Author: Harald La; Country: Germany; Date: 08/09/2022; Keywords: Alpine gla ellipsis; Language: German; Multimedia: N/A; Rating scale: Fake; Source: Correctiv; Text: A simulati ellipsis; Uniform Resource Locator: https://co ellipsis; Why: The retreat ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 28: idat: 29; Author: Dieter Nu; Country: Germany; Date: 09/08/2022; Keywords: CO2 footp ellipsis; Language: German; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: Correctiv; Text: Dieter Nu ellipsis; Uniform Resource Locator: https://co ellipsis; Why: The quote ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German. Row 29: idat: 30; Author: NA; Country: Germany; Date: 21/07/2022; Keywords: Climate ch ellipsis; Language: German; Multimedia: https://drive.g ellipsis; Rating scale: Fake; Source: Correctiv; Text: A graphic ellipsis; Uniform Resource Locator: https://co ellipsis; Why: The graph ellipsis; Topic: Climate change; Text (source language): German; Why (source language): German.Sample of dataset table
The dataset model, shown in Figure 1, is built on one universal table where one row represents a single case of disinformation. The structure of the dataset follows from a specific template, described earlier in this document, calling for a range of fields meant to encompass all the relevant facts of a claim of disinformation. These types include common metadata such as topic, keywords, publication date, source, URL, language and author, and more advanced information, such as the English translation of the entire text of the disinformation, the corresponding fact-checking analysis in both English and the original languages and references to associated multimedia content, such as images or videos.
Figure 2 displays an example of how the instances of disinformation are labeled in the Debunking datasets. It shows a snapshot of the original data table used in the project, taking on the structural pattern outlined in Figure 1. The rows and columns in this image show how each instance of disinformation is arranged and indexed. The columns are the various parameters specified in the schema of the dataset and each row represents a single instance of disinformation.
3.3 Data collection
The Debunking datasets were constructed using a systematic and methodologically rigorous approach to ensure data accuracy and representativeness. Data collection involved the amalgamation of multiple sources and put together a wide range of disinformation claims, which include:
Fact-checking sites, such as PolitiFact [2], Climate Feedback [3] and EUvsDisinfo [4], that provide verified analysis of disinformation. Many of the claims made on these websites are published on social media platforms, such as Twitter, Facebook and Telegram.
Open-source data, such as Ukrainian news, providing actual news on the invasion of Ukraine by Russia, and News sources and climate change, providing real news on climate change.
This collection process ensures that Debunking datasets properly reflect real-world disinformation patterns accurately, hence being best suited for training and testing AI.
3.4 Annotation and validation
The annotation process followed a structured methodology to maintain high-quality labeling standards. Each data instance underwent multiple stages of verification:
Initial Filtering: Irrelevant, low-quality or redundant content was removed to maintain the integrity of the dataset.
Labeling Framework: The data instances were classified into the following two categories:
True (factually accurate claims verified by credible sources)
False (clearly false or made-up claims)
To reduce subjectivity, annotators were trained to use only evidence from official and highly reputable primary sources, databases of verified claims and well-known fact-checking organizations, as described in Section 3.3. Before completing each label, annotators cross-checked it against several different independent sources. To improve inter-annotator agreement and reliability, a subset of the data was also double-annotated by different experts, and disagreements were settled through consensus discussions.
Nevertheless, some limitations of this strict annotation procedure still exist. First, for claims that are only partially true or context-dependent, the True/False binary classification may be reductive and may oversimplify complex information. Moreover, some claims that were true at the annotation time might become false later due to advancements in science, new evidence, or shifting conditions. Finally, human judgment can never be completely free from subjective interpretation or implicit biases, even if consensus mechanisms are applied.
Despite that, the following annotation process guarantees that Debunking datasets have a homogeneous and impartial basis on which the models for detecting disinformation can be trained and tested. Through the use of professional fact-checking, minimizing bias and ensuring high inter-annotator agreement, the datasets improve the quality and effectiveness of AI-powered fake news detection tasks.
Fake news multimedia contents were downloaded from reference fact-checking websites and available databases that are owned by two International Media Organizations, i.e. EURACTIV Bulgaria (https://euractiv.bg/) and Internews Ukraine (https://internews.ua/en). These files were stored online instead of being directly inserted into the database. This decision was made to have more resources (e.g. computation and storage capabilities) for executing queries efficiently and for making data processing and analytics more efficient.
Moreover, despite the adopted quality control measures, several potential ethical issues have been taken into account. In this regard, the study acknowledges a wide range of phenomena, such as annotator bias, cultural and political subjectivity, that could lead to challenges in labeling sensitive topics or privacy concerns when collecting a large variety of data (including those coming from social media).
For these reasons, the low score reported in the “Ethical standards” parameter in DQAF assessment would require plans for formal annotation guidelines (see Section 4.1).
Two online file storage services were utilized, both of which offer free versions. These services have been designed to maintain high-performance levels of the datasets over time. The cloud storage services used were Mediafire (https://www.mediafire.com) and Google Drive (https://www.google.com/intl/it_it/drive/).
Each dataset is delivered in the following file formats:
A CSV file with semicolon separator;
The model file (.mwb), containing the physical schema of the dataset and the populated table.
These file formats allow easy editing, processing and managing data, with the possibility to connect the database models to internal or external servers. Figure 3 shows a sample portion of the CSV file of the climate change dataset.
The display presents a structured dataset of climate change misinformation cases organized by “idat”, “Author”, “Country”, “Date”, “Keywords”, “Language”, “Multimedia”, “Rating scale”, “Source”, “Text”, “Url”, “Why”, “Topic”, “Text (source language)”, and “Why (source language)”. The row-wise are as follows: Row 2: France, 01/08/2022 — Climate protest video claim in French; rated Fake by AFP Factuel with supporting explanation. Row 3: Lithuania, 23/04/2022 — Protest-related video in Lithuanian; rated Fake by Delfi. Row 4: Germany, 22/03/2022 — Body bags video misattributed to climate topic; rated Fake by Correctiv. Row 5: Spain, 31/05/2024 — Cars and pollution claim in English; rated Fake by eufactcheck. Row 6: Croatia, 16/04/2024 — Migrant-related climate narrative; rated Fake by eufactcheck. Row 7: USA, 25/12/2023 — Global warming concept claim; rated Fake by eufactcheck. Row 8: Germany, 27/05/2023 — Germany-related climate statement; rated Fake by eufactcheck. Row 9: USA, 15/09/2023 — Political climate-related statement; rated Fake by FactCheck. Row 10: USA, 26/01/2023 — Global temperature claim; rated Fake by Factcheck. Row 11: USA, 05/09/2020 — Climate and wildfire briefing claim; rated Fake by Factcheck. Row 12: USA, 12/12/2019 — Campaign rally climate statement; rated Fake by Factcheck. Row 13: USA/France, 22/08/2019 — Amazon production claim; rated Fake by Factcheck. Row 14: Germany, 29/08/2023 — Global warming and ice claim; rated Fake by Correctiv. Row 15: Germany, 09/02/2024 — Flood-related image claim; rated Fake by Correctiv. Row 16: Germany, 25/01/2024 — Wind turbine and nuclear narrative; rated Fake by Correctiv. Row 17: Germany, 21/12/2023 — Global warming map claim; rated Fake by Correctiv. Row 18: Germany, 21/12/2023 — E-car pricing claim; rated Fake by Correctiv. Row 19: Germany, 30/10/2023 — Ukraine and CO2 claim; rated Fake by Correctiv. Row 20: Germany, 23/04/2023 — 78 percent CO2-related statement; rated Fake by Correctiv. Row 21: Germany, 26/06/2023 — Wind farm study claim; rated Fake by Correctiv. Row 22: Germany, 22/06/2023 — Heat protection and lockdown claim; rated Fake by Correctiv. Row 23: Germany, 07/06/2023 — Heat wave statement; rated Fake by Correctiv. Row 24: Germany, 07/06/2023 — Cold record claim in Antarctica; rated Fake by Correctiv. Row 25: Germany, 15/09/2022 — Wind turbine and forest claim; rated Fake by Correctiv. Row 26: Germany, 09/09/2022 — Trees and global climate narrative; rated Fake by Correctiv. Row 27: Germany, 08/09/2022 — Alpine glacier simulation claim; rated Fake by Correctiv. Row 28: Germany, 09/08/2022 — CO2 footprint quote claim; rated Fake by Correctiv. Row 29: Germany, 21/07/2022 — Climate graphic claim; rated Fake by Correctiv.Sample of the CSV file of the climate change dataset
The display presents a structured dataset of climate change misinformation cases organized by “idat”, “Author”, “Country”, “Date”, “Keywords”, “Language”, “Multimedia”, “Rating scale”, “Source”, “Text”, “Url”, “Why”, “Topic”, “Text (source language)”, and “Why (source language)”. The row-wise are as follows: Row 2: France, 01/08/2022 — Climate protest video claim in French; rated Fake by AFP Factuel with supporting explanation. Row 3: Lithuania, 23/04/2022 — Protest-related video in Lithuanian; rated Fake by Delfi. Row 4: Germany, 22/03/2022 — Body bags video misattributed to climate topic; rated Fake by Correctiv. Row 5: Spain, 31/05/2024 — Cars and pollution claim in English; rated Fake by eufactcheck. Row 6: Croatia, 16/04/2024 — Migrant-related climate narrative; rated Fake by eufactcheck. Row 7: USA, 25/12/2023 — Global warming concept claim; rated Fake by eufactcheck. Row 8: Germany, 27/05/2023 — Germany-related climate statement; rated Fake by eufactcheck. Row 9: USA, 15/09/2023 — Political climate-related statement; rated Fake by FactCheck. Row 10: USA, 26/01/2023 — Global temperature claim; rated Fake by Factcheck. Row 11: USA, 05/09/2020 — Climate and wildfire briefing claim; rated Fake by Factcheck. Row 12: USA, 12/12/2019 — Campaign rally climate statement; rated Fake by Factcheck. Row 13: USA/France, 22/08/2019 — Amazon production claim; rated Fake by Factcheck. Row 14: Germany, 29/08/2023 — Global warming and ice claim; rated Fake by Correctiv. Row 15: Germany, 09/02/2024 — Flood-related image claim; rated Fake by Correctiv. Row 16: Germany, 25/01/2024 — Wind turbine and nuclear narrative; rated Fake by Correctiv. Row 17: Germany, 21/12/2023 — Global warming map claim; rated Fake by Correctiv. Row 18: Germany, 21/12/2023 — E-car pricing claim; rated Fake by Correctiv. Row 19: Germany, 30/10/2023 — Ukraine and CO2 claim; rated Fake by Correctiv. Row 20: Germany, 23/04/2023 — 78 percent CO2-related statement; rated Fake by Correctiv. Row 21: Germany, 26/06/2023 — Wind farm study claim; rated Fake by Correctiv. Row 22: Germany, 22/06/2023 — Heat protection and lockdown claim; rated Fake by Correctiv. Row 23: Germany, 07/06/2023 — Heat wave statement; rated Fake by Correctiv. Row 24: Germany, 07/06/2023 — Cold record claim in Antarctica; rated Fake by Correctiv. Row 25: Germany, 15/09/2022 — Wind turbine and forest claim; rated Fake by Correctiv. Row 26: Germany, 09/09/2022 — Trees and global climate narrative; rated Fake by Correctiv. Row 27: Germany, 08/09/2022 — Alpine glacier simulation claim; rated Fake by Correctiv. Row 28: Germany, 09/08/2022 — CO2 footprint quote claim; rated Fake by Correctiv. Row 29: Germany, 21/07/2022 — Climate graphic claim; rated Fake by Correctiv.Sample of the CSV file of the climate change dataset
4. Database quality assessment, comparative analysis and validation
A database quality assessment, as well as a comparative analysis and a validation across several existing datasets and state-of-the-art fake news detection algorithms, is performed. In particular, the database quality assessment is relevant to ensure a proper grade of reliability, accuracy and efficiency of data systems. This process considers the evaluations of several specific aspects, such as data integrity, completeness, consistency and relevance. Once these elements are correctly assessed, it is possible to identify potential issues and improve the overall functionality of a dataset.
Then, the comparative analysis aims to further improve this process by presenting a comparison of the Debunking datasets to five existing datasets by highlighting their strengths and weaknesses. Finally, the validation aims to evaluate the effectiveness of the Debunking datasets in debunking disinformation by considering four state-of-the-art algorithmic approaches to fake news detection.
In the following sub-section, the quality assessment and the comparative analysis to optimize data management practices are described in detail.
4.1 Database quality assessment
To evaluate the Debunking datasets, the study employed the Data Quality Assessment Framework (DQAF) (Sebastian-Coleman, 2012), which evaluates data quality along five relevant dimensions:
Assurance of integrity, which assesses the completeness and trustworthiness of the data.
Methodological soundness, which examines whether the database is coherent with rigorous and consistent analytical methodologies.
Correctness and dependability, which evaluate the consistency and accuracy of data retrieval and performance over time.
Serviceability, which measures the database’s responsiveness to user needs, including its adaptability and support infrastructure.
Accessibility, which ensures that the database is user-friendly and readily available to authorized users.
Each dimension also comprises a set of elements of good practice and several related indicators that allow measuring the level of accomplishment of the elements. In this direction, a thorough evaluation based on the selected criteria can provide a comprehensive understanding of the database’s strengths and limitations related to the quality of the collected data.
For each above-mentioned dimensions, we reported the elements of good practice and the indicators specific to evaluating our two datasets, as shown in Table A.1 in Appendix A. A qualitative scale that goes from Practice Observed (≥80%) to Practice Not Observed (<20%) is used to score each indicator.
The datasets fared well with regard to professionalism and transparency, to start based on integrity. For the sake of promoting objectivity and reducing bias, data gathering was carried out independently by specialists. The data generation process was made open to ensure that there was no lack of transparency. The ethical standards score was low since there were no formalized behavior norms for the workers who were undertaking data annotation. In fact, even professional fact-checkers could be inclined to implicit cognitive, cultural or political biases that could potentially affect decisions. These risks are even higher in particular sensitive domains, such as those related to armed conflict or climate change, as these topics could expose annotators to emotional stress or well-being concerns. In addition, multilingual contexts could amplify these potential issues due to the subtle semantic differences or cultural connotations during the interpretation. For these reasons, future dataset releases would better deal with current limitations through the implementation of explicit ethical protocols in the attempt to provide proper solutions, such as annotator training on bias mitigation and clear decision reports to improve both transparency and reproducibility.
Methodologically sound, the datasets were of high international standards of conformity. Classification frameworks and conceptual structure employed conform to best practice in the field. Both thematic focus and structural organization of the datasets signify the presence of an open and standard approach. Accuracy and reliability dimensions yielded unbalanced results. To their credit, information sources were ample and trustworthy due to their being based on good fact-checking websites.
On the downside, the absence of statistical tools for data processing and the absence of revision studies compromised performance in this area. While intermediate outcomes were internally consistent and validated, there are no current tools for updating or monitoring claim status as new evidence appears.
From a serviceability point of view, the datasets follow a pre-scheduled release pattern and are clearly labeled in such a way that initial and final versions of data can be readily distinguished. This facilitates the timeliness and transparency dimensions. There are no mandated procedures to obtain longitudinal consistency or trend observation over dataset versions.
Finally, the availability was rated highly. Both datasets are presented in a welcoming way and with good documentation, such that they can be reused and properly interpreted with ease. Nevertheless, despite extensive documentation, there are no detailed descriptions of statistical processing techniques and deviations from international practices.
After defining the DQAF for assessing the data quality of our datasets, we evaluated the scores of the indicators listed in the third column of Table A.1 in Appendix A. The scores and motivation for each indicator, resulting from the assessment, as well as the total score of the dimensions, are shown in Table A.2 in Appendix A.
The results of the data quality assessment we performed on the two datasets are summarized in Table 3. The results show a similar pattern: both datasets are strong in integrity, soundness of methodology and availability, but have reciprocal areas of improvement in statistical processing, temporal consistency and tracing revisions.
Data quality assessment framework – summary results
| Dimensions/elements | Datasets | |
|---|---|---|
| The Russian invasion of Ukraine dataset | Climate change dataset | |
| 1. Integrity | ||
| 1.1 Professionalism | O | O |
| 1.2 Transparency | O | O |
| 1.3 Ethical standards | LO | LO |
| 2. Methodological soundness | ||
| 2.1 Concepts and Definitions | O | O |
| 2.2 Scope, Classification and Recording | O | O |
| 3. Reliability | ||
| 3.1 Source data | O | O |
| 3.2 Statistical Techniques | NO | NO |
| 3.3 Assessment and validation of intermediate data and statistical outputs | O | O |
| 3.4 Revision studies | NO | NO |
| 4. Serviceability | ||
| 4.1 Periodicity and Timeliness | O | O |
| 4.1 Consistency | NO | NO |
| 4.2 Revision policy | O | O |
| 5. Accessibility | ||
| 5.1 Data accessibility | O | O |
| 5.2 Metadata accessibility | LO | LO |
| 5.3 User support | O | O |
| Dimensions/elements | Datasets | |
|---|---|---|
| The Russian invasion of Ukraine dataset | Climate change dataset | |
| 1. Integrity | ||
| 1.1 Professionalism | O | O |
| 1.2 Transparency | O | O |
| 1.3 Ethical standards | ||
| 2. Methodological soundness | ||
| 2.1 Concepts and Definitions | O | O |
| 2.2 Scope, Classification and Recording | O | O |
| 3. Reliability | ||
| 3.1 Source data | O | O |
| 3.2 Statistical Techniques | ||
| 3.3 Assessment and validation of intermediate data and statistical outputs | O | O |
| 3.4 Revision studies | ||
| 4. Serviceability | ||
| 4.1 Periodicity and Timeliness | O | O |
| 4.1 Consistency | ||
| 4.2 Revision policy | O | O |
| 5. Accessibility | ||
| 5.1 Data accessibility | O | O |
| 5.2 Metadata accessibility | ||
| 5.3 User support | O | O |
Note(s): Key to symbols: O = Practice Observed; LO = Practice Largely Observed; LNO =Practice Largely Not Observed; NO = Practice Not Observed; NA = Not Applicable. Practice observed: Current practices generally meet or achieve the objectives of DQAF internationally accepted statistical practices without any significant deficiencies. Practice largely observed: Some departures, but these are not seen as sufficient to raise doubts about the authorities’ ability to observe the DQAF practices. Practice largely not observed: Significant departures, and the authorities will need to take significant action to achieve observance. Practice not observed: Most DQAF practices are not met. Not applicable: Used only exceptionally when statistical practices do not apply to the country’s circumstances
4.2 Comparative analysis with existing databases
In this section, an analysis is conducted to compare the Debunking datasets with the following five existing fake news detection datasets: LIAR (Wang, 2017), FakeNewsNet (Shu et al., 2020), Emergent (Ferreira and Vlachos, 2016), Climate-Fact (Nooralahzadeh et al., 2021) and Multi-Fake-DetectiVE (Bondielli et al., 2024). In order to facilitate a comparative analysis of the datasets, a selection was made based on the ten qualitative criteria outlined in D’Ulizia et al. (2021). These criteria assume a relevant role for evaluating fake news datasets and their suitability for training and benchmarking AI models in disinformation detection.
Besides, the selection of these requirements was made to identify the reliability and comparability of data sources in the context of disinformation detection. The requirements were later re-developed and re-adapted to meet the specific features and objectives of the Debunking datasets. This allowed for the development of a more context-aware assessment framework. The following criteria are the subject of comparison:
Homogeneity of news length: Most of the traditional datasets (e.g. LIAR, Emergent) contain special items, e.g. political quotes or article titles, which render the text lengths heterogeneous. In contrast to this, Debunking datasets ensure consistency by collecting full news articles or detailed summaries presented in a uniform template, to ensure a high level of comparability and consistency of input across samples.
Homogeneity in news domain: Debunking datasets are developed based on domain homogeneity by focusing solely on two high-impact domains: the Russian invasion of Ukraine and climate change. This targeted approach contrasts with other datasets characterized by a general-purpose approach, such as FakeNewsNet, LIAR or NewsPolyML, which include a wide variety of themes, which could make domain-specific model training less effective.
Homogeneity in type of disinformation: In a different way from mixed-type datasets (i.e. Emergent, characterized by stance classification), Debunking datasets are exclusively focused on fabricated or misleading news articles or social network posts. They exclude other types, such as satirical content, resulting in a more coherent corpus centered on disinformation.
Homogeneity in application purpose: Debunking datasets have been specifically developed for the detection and debunking of disinformation through machine learning and knowledge graph reasoning. This aligns with fact-checking applications, whereas some datasets (i.e. Emergent) are designed with a particular focus on stance detection or rumor tracking.
Availability of both fake and real news: In contrast to balanced datasets, such as LIAR or FakeNewsNet, Debunking datasets focus exclusively on disinformation cases. In this direction, its primary objective would be to enable debunking and disinformation profiling rather than offer binary classification.
Textual format availability: Debunking datasets provide textual transcriptions of disinformation content in both the original and translated (English) versions. All claims are documented in structured text format, making them susceptible to semantic enrichment procedures. Multimodal content (image, video) is linked but not necessary for text processing.
Public availability: Debunking datasets are publicly available in usable form, along with extensive documentation. This output is in accordance with current literature, e.g. Climate-Fact, and conflicts with current literature, with limited or partial data availability.
Multilingual availability: Debunking datasets support multilingual input by offering content in both the original language and English. This can assure a better grade of versatility with respect to most English-only datasets (i.e. LIAR, Emergent and MFND), making Debunking datasets better suited for cross-lingual training or multilingual evaluation settings.
Verifiability of ground truth: All claims collected in Debunking datasets are annotated and validated by expert fact-checkers from AI4Debunk partner organizations (Euractiv Bulgaria and Internews Ukraine), with references to primary debunking sources. This ensures high verifiability of the ground truth, in contrast with crowdsourced or automatically labeled datasets, where annotation methodologies may be different.
Belonging to a predefined time frame: The Debunking datasets were not collected within a defined project timeline, and this results in a lack of temporal coherence. Nevertheless, this approach is not consistent with the methods employed in datasets that adhere to best practices for benchmarking and trend detection. Rather, it mirrors the approach of legacy datasets that aggregate content across multiple years without implementing temporal filters.
A summary of the performed comparison is briefly reported in Table 4. In this comparative overview, the Debunking datasets’ strengths and limitations are highlighted.
Comparative analysis of datasets
| Database | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| Criterion | Debunking | LIAR | FakeNewsNet | Emergent | Climate-fact | Multi-fake-DetectiVE | NewsPolyML | MMCFND | MFND |
| News length homogeneity | ✔ | ✕ | ✕ | ✕ | ✔ | ✔ | ✕ | ✕ | ✕ |
| Domain homogeneity | ✔ | ✕ | ✕ | ✕ | ✔ | ✕ | ✕ | ✕ | ✔ |
| Disinformation type homogeneity | ✔ | ✕ | ✕ | ✕ | ✔ | ✔ | ✕ | ✕ | ✕ |
| Application purpose alignment | ✔ | ✔ | ✔ | ✕ | ✔ | ✔ | ✔ | ✔ | ✔ |
| Fake and real news | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ |
| Textual format availability | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ |
| Public availability | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ |
| Multilingual content | ✔ | ✕ | ✕ | ✕ | ✔ | ✕ | ✔ | ✔ | ✕ |
| Verifiability of ground truth | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ |
| Predefined time frame | ✕ | ✕ | ✕ | ✕ | ✔ | ✔ | ✔ | ✕ | ✕ |
| Database | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| Criterion | Debunking | FakeNewsNet | Emergent | Climate-fact | Multi-fake-DetectiVE | NewsPolyML | |||
| News length homogeneity | ✔ | ✕ | ✕ | ✕ | ✔ | ✔ | ✕ | ✕ | ✕ |
| Domain homogeneity | ✔ | ✕ | ✕ | ✕ | ✔ | ✕ | ✕ | ✕ | ✔ |
| Disinformation type homogeneity | ✔ | ✕ | ✕ | ✕ | ✔ | ✔ | ✕ | ✕ | ✕ |
| Application purpose alignment | ✔ | ✔ | ✔ | ✕ | ✔ | ✔ | ✔ | ✔ | ✔ |
| Fake and real news | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ |
| Textual format availability | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ |
| Public availability | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ |
| Multilingual content | ✔ | ✕ | ✕ | ✕ | ✔ | ✕ | ✔ | ✔ | ✕ |
| Verifiability of ground truth | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ |
| Predefined time frame | ✕ | ✕ | ✕ | ✕ | ✔ | ✔ | ✔ | ✕ | ✕ |
Table 4 provides a comprehensive overview of a multitude of information that might facilitate the selection and comprehension of datasets for disinformation research. The comparison confirms that the Debunking datasets offer a high degree of specialization and methodological rigor in specific aspects, such as domain specificity, multilingualism, structured annotation and multimodality.
Unlike large-topic general-purpose datasets like LIAR, FakeNewsNet or NewsPolyML, which span a wide range of subjects and sources, the Debunking datasets focus on two subjects of high social impact (climate change and the Russian invasion of Ukraine), where disinformation has severe social consequences. That specificity renders the AI models trained on them more useful and valuable by enabling improved adaptation to the domain and less noise from non-relevant material.
From the perspective of diversity of content, Debunking datasets are unique in that they include both true and false news stories, making them balanced training environments. In addition, full-text transcriptions in original and English translation modes support enhanced multilingual abilities as well as cross-lingual transfer learning, which are crucial in combating disinformation that crosses national and linguistic borders. This feature is lacking or available in an intermittent fashion in common datasets that are mostly English-biased and monolingual.
Public access and transparency are other strengths of the Debunking datasets. They appear in open-access formats (CSV and MWB) with comprehensive documentation so that it is a simple matter to reuse and integrate into AI pipelines. Additionally, each claim passes rigorous verification by expert fact-checkers, which ensures high reliability and reduces the likelihood of annotation bias, compared to crowdsourced or weakly labeled datasets. Integration of organized metadata and multimedia links renders the datasets more valuable for multimodal AI tasks, like image-based and video-based fake news detection.
Even if not yet time-sensitive, the focus of the Debunking datasets on near-history and impactful events contributes to temporal urgency. Overall, the comparison strongly confirms that the Debunking datasets not only pick up where others left off but set a new high-water mark of quality, organization and usability for AI-based disinformation detection.
4.3 Validation for fake news detection
To evaluate the effectiveness of the Debunking datasets in debunking disinformation, we evaluated four state-of-the-art algorithmic approaches to fake news detection: Logistic Regression, Random Forest, Bidirectional LSTM (BiLSTM) and a fine-tuned BERT transformer model. We trained and evaluated these models on both Debunking datasets (climate change and Russian invasion of Ukraine with real and fake news examples) and the five established benchmarking datasets introduced in Section 4.2.
Each dataset was pre-processed using standard NLP techniques, including tokenization, stop-word removal and TF-IDF vectorization for traditional machine learning models. For deep learning models (BERT and BiLSTM), token embedding techniques and pre-trained language representations were used. An 80/20 stratified train/test split was used to preserve class balance, and evaluation metrics were averaged across 5-fold cross-validation for stability.
The accuracy, precision, recall and F1 scores are presented in Figures 4–7, respectively. The Debunking datasets performed well, with a high degree of efficacy, especially when tested using deep learning models. For instance, on the Debunking Climate dataset, the BERT model had a 94% accuracy, 95% precision, 93% recall and 94% F1 score. The same performance was observed on the Debunking Ukraine dataset, though with slightly poorer but still great accuracy of 93% and an F1-score of 93%.
The horizontal axis lists “Debunking Climate”, “Debunking Ukraine”, “L I A R”, “FakeNewsNet”, “Emergent”, “Climate-Fact”, “Multi-Fake-DetectIVE”, “NewsPolYML”, “M M C F N D”, and “M F N D”. The vertical axis ranges from 0.00 to 1.00 in increments of 0.10. The legend is positioned at the bottom center below the graph. It lists four models: “L R”, “R F”, “BiLSTM”, and “B E R T”, each represented by a distinct colored bar. For “Debunking Climate”: L R: 0.84, R F: 0.87, BiLSTM: 0.90, B E R T: 0.94. For “Debunking Ukraine”: L R: 0.83, R F: 0.86, BiLSTM: 0.89, B E R T: 0.93. For “L I A R”: L R: 0.66, R F: 0.70, BiLSTM: 0.74, B E R T: 0.79. For “FakeNewsNet”: L R: 0.72, R F: 0.75, BiLSTM: 0.78, B E R T: 0.83. For “Emergent”: L R: 0.68, R F: 0.71, BiLSTM: 0.75, B E R T: 0.80. For “Climate-Fact”: L R: 0.81, R F: 0.84, BiLSTM: 0.88, B E R T: 0.92. For “Multi-Fake-DetectIVE”: L R: 0.82, R F: 0.85, BiLSTM: 0.89, B E R T: 0.93. For “NewsPolYML”: L R: 0.74, R F: 0.77, BiLSTM: 0.81, B E R T: 0.86. For “M M C F N D”: L R: 0.71, R F: 0.75, BiLSTM: 0.79, B E R T: 0.84. For “M F N D”: L R: 0.76, R F: 0.79, BiLSTM: 0.83, B E R T: 0.88.Accuracy scores resulting from the validation
The horizontal axis lists “Debunking Climate”, “Debunking Ukraine”, “L I A R”, “FakeNewsNet”, “Emergent”, “Climate-Fact”, “Multi-Fake-DetectIVE”, “NewsPolYML”, “M M C F N D”, and “M F N D”. The vertical axis ranges from 0.00 to 1.00 in increments of 0.10. The legend is positioned at the bottom center below the graph. It lists four models: “L R”, “R F”, “BiLSTM”, and “B E R T”, each represented by a distinct colored bar. For “Debunking Climate”: L R: 0.84, R F: 0.87, BiLSTM: 0.90, B E R T: 0.94. For “Debunking Ukraine”: L R: 0.83, R F: 0.86, BiLSTM: 0.89, B E R T: 0.93. For “L I A R”: L R: 0.66, R F: 0.70, BiLSTM: 0.74, B E R T: 0.79. For “FakeNewsNet”: L R: 0.72, R F: 0.75, BiLSTM: 0.78, B E R T: 0.83. For “Emergent”: L R: 0.68, R F: 0.71, BiLSTM: 0.75, B E R T: 0.80. For “Climate-Fact”: L R: 0.81, R F: 0.84, BiLSTM: 0.88, B E R T: 0.92. For “Multi-Fake-DetectIVE”: L R: 0.82, R F: 0.85, BiLSTM: 0.89, B E R T: 0.93. For “NewsPolYML”: L R: 0.74, R F: 0.77, BiLSTM: 0.81, B E R T: 0.86. For “M M C F N D”: L R: 0.71, R F: 0.75, BiLSTM: 0.79, B E R T: 0.84. For “M F N D”: L R: 0.76, R F: 0.79, BiLSTM: 0.83, B E R T: 0.88.Accuracy scores resulting from the validation
The horizontal axis lists “Debunking Climate”, “Debunking Ukraine”, “L I A R”, “FakeNewsNet”, “Emergent”, “Climate-Fact”, “Multi-Fake-DetectIVE”, “NewsPolYML”, “M M C F N D”, and “M F N D”. The vertical axis ranges from 0 to 1.0 in increments of 0.1. The legend is positioned at the bottom center below the graph and lists four models: “L R”, “R F”, “BiLSTM”, and “B E R T”, each represented by a distinct colored bar. For “Debunking Climate”: L R: 0.83, R F: 0.86, BiLSTM: 0.91, B E R T: 0.95. For “Debunking Ukraine”: L R: 0.82, R F: 0.85, BiLSTM: 0.90, B E R T: 0.94. For “L I A R”: L R: 0.64, R F: 0.68, BiLSTM: 0.72, B E R T: 0.77. For “FakeNewsNet”: L R: 0.70, R F: 0.73, BiLSTM: 0.76, B E R T: 0.81. For “Emergent”: L R: 0.66, R F: 0.69, BiLSTM: 0.73, B E R T: 0.78. For “Climate-Fact”: L R: 0.80, R F: 0.83, BiLSTM: 0.87, B E R T: 0.91. For “Multi-Fake-DetectIVE”: L R: 0.81, R F: 0.84, BiLSTM: 0.88, B E R T: 0.92. For “NewsPolYML”: L R: 0.73, R F: 0.76, BiLSTM: 0.80, B E R T: 0.85. For “M M C F N D”: L R: 0.70, R F: 0.74, BiLSTM: 0.78, B E R T: 0.83. For “M F N D”: L R: 0.75, R F: 0.78, BiLSTM: 0.82, B E R T: 0.87.Precision scores resulting from the validation
The horizontal axis lists “Debunking Climate”, “Debunking Ukraine”, “L I A R”, “FakeNewsNet”, “Emergent”, “Climate-Fact”, “Multi-Fake-DetectIVE”, “NewsPolYML”, “M M C F N D”, and “M F N D”. The vertical axis ranges from 0 to 1.0 in increments of 0.1. The legend is positioned at the bottom center below the graph and lists four models: “L R”, “R F”, “BiLSTM”, and “B E R T”, each represented by a distinct colored bar. For “Debunking Climate”: L R: 0.83, R F: 0.86, BiLSTM: 0.91, B E R T: 0.95. For “Debunking Ukraine”: L R: 0.82, R F: 0.85, BiLSTM: 0.90, B E R T: 0.94. For “L I A R”: L R: 0.64, R F: 0.68, BiLSTM: 0.72, B E R T: 0.77. For “FakeNewsNet”: L R: 0.70, R F: 0.73, BiLSTM: 0.76, B E R T: 0.81. For “Emergent”: L R: 0.66, R F: 0.69, BiLSTM: 0.73, B E R T: 0.78. For “Climate-Fact”: L R: 0.80, R F: 0.83, BiLSTM: 0.87, B E R T: 0.91. For “Multi-Fake-DetectIVE”: L R: 0.81, R F: 0.84, BiLSTM: 0.88, B E R T: 0.92. For “NewsPolYML”: L R: 0.73, R F: 0.76, BiLSTM: 0.80, B E R T: 0.85. For “M M C F N D”: L R: 0.70, R F: 0.74, BiLSTM: 0.78, B E R T: 0.83. For “M F N D”: L R: 0.75, R F: 0.78, BiLSTM: 0.82, B E R T: 0.87.Precision scores resulting from the validation
The horizontal axis lists “Debunking Climate”, “Debunking Ukraine”, “L I A R”, “FakeNewsNet”, “Emergent”, “Climate-Fact”, “Multi-Fake-DetectIVE”, “NewsPolYML”, “M M C F N D”, and “M F N D”. The vertical axis ranges from 0 to 1 in increments of 0.1. The legend is positioned at the bottom center below the graph and lists four models: “L R”, “R F”, “BiLSTM”, and “B E R T”, each represented by a distinct colored bar. For “Debunking Climate”: L R: 0.85, R F: 0.88, BiLSTM: 0.89, B E R T: 0.93. For “Debunking Ukraine”: L R: 0.84, R F: 0.87, BiLSTM: 0.88, B E R T: 0.92. For “L I A R”: L R: 0.67, R F: 0.72, BiLSTM: 0.76, B E R T: 0.81. For “FakeNewsNet”: L R: 0.73, R F: 0.76, BiLSTM: 0.80, B E R T: 0.85. For “Emergent”: L R: 0.69, R F: 0.72, BiLSTM: 0.77, B E R T: 0.82. For “Climate-Fact”: L R: 0.82, R F: 0.85, BiLSTM: 0.89, B E R T: 0.93. For “Multi-Fake-DetectIVE”: L R: 0.83, R F: 0.86, BiLSTM: 0.90, B E R T: 0.94. For “NewsPolYML”: L R: 0.75, R F: 0.78, BiLSTM: 0.82, B E R T: 0.87. For “M M C F N D”: L R: 0.72, R F: 0.76, BiLSTM: 0.80, B E R T: 0.85. For “M F N D”: L R: 0.77, R F: 0.80, BiLSTM: 0.84, B E R T: 0.89.Recall scores resulting from the validation
The horizontal axis lists “Debunking Climate”, “Debunking Ukraine”, “L I A R”, “FakeNewsNet”, “Emergent”, “Climate-Fact”, “Multi-Fake-DetectIVE”, “NewsPolYML”, “M M C F N D”, and “M F N D”. The vertical axis ranges from 0 to 1 in increments of 0.1. The legend is positioned at the bottom center below the graph and lists four models: “L R”, “R F”, “BiLSTM”, and “B E R T”, each represented by a distinct colored bar. For “Debunking Climate”: L R: 0.85, R F: 0.88, BiLSTM: 0.89, B E R T: 0.93. For “Debunking Ukraine”: L R: 0.84, R F: 0.87, BiLSTM: 0.88, B E R T: 0.92. For “L I A R”: L R: 0.67, R F: 0.72, BiLSTM: 0.76, B E R T: 0.81. For “FakeNewsNet”: L R: 0.73, R F: 0.76, BiLSTM: 0.80, B E R T: 0.85. For “Emergent”: L R: 0.69, R F: 0.72, BiLSTM: 0.77, B E R T: 0.82. For “Climate-Fact”: L R: 0.82, R F: 0.85, BiLSTM: 0.89, B E R T: 0.93. For “Multi-Fake-DetectIVE”: L R: 0.83, R F: 0.86, BiLSTM: 0.90, B E R T: 0.94. For “NewsPolYML”: L R: 0.75, R F: 0.78, BiLSTM: 0.82, B E R T: 0.87. For “M M C F N D”: L R: 0.72, R F: 0.76, BiLSTM: 0.80, B E R T: 0.85. For “M F N D”: L R: 0.77, R F: 0.80, BiLSTM: 0.84, B E R T: 0.89.Recall scores resulting from the validation
The horizontal axis lists “Debunking Climate”, “Debunking Ukraine”, “L I A R”, “FakeNewsNet”, “Emergent”, “Climate-Fact”, “Multi-Fake-DetectIVE”, “NewsPolYML”, “M M C F N D”, and “M F N D”. The vertical axis ranges from 0 to 1 in increments of 0.1. The legend is positioned at the bottom center below the graph and lists four models: “L R”, “R F”, “BiLSTM”, and “B E R T”, each represented by a distinct bar. For “Debunking Climate”: L R: 0.84, R F: 0.87, BiLSTM: 0.90, B E R T: 0.94. For “Debunking Ukraine”: L R: 0.83, R F: 0.86, BiLSTM: 0.89, B E R T: 0.93. For “L I A R”: L R: 0.65, R F: 0.70, BiLSTM: 0.74, B E R T: 0.79. For “FakeNewsNet”: L R: 0.71, R F: 0.74, BiLSTM: 0.78, B E R T: 0.83. For “Emergent”: L R: 0.67, R F: 0.70, BiLSTM: 0.75, B E R T: 0.80. For “Climate-Fact”: L R: 0.81, R F: 0.84, BiLSTM: 0.88, B E R T: 0.92. For “Multi-Fake-DetectIVE”: L R: 0.82, R F: 0.85, BiLSTM: 0.89, B E R T: 0.93. For “NewsPolYML”: L R: 0.74, R F: 0.77, BiLSTM: 0.81, B E R T: 0.86. For “M M C F N D”: L R: 0.71, R F: 0.75, BiLSTM: 0.79, B E R T: 0.84. For “M F N D”: L R: 0.76, R F: 0.79, BiLSTM: 0.83, B E R T: 0.88.F1 scores resulting from the validation
The horizontal axis lists “Debunking Climate”, “Debunking Ukraine”, “L I A R”, “FakeNewsNet”, “Emergent”, “Climate-Fact”, “Multi-Fake-DetectIVE”, “NewsPolYML”, “M M C F N D”, and “M F N D”. The vertical axis ranges from 0 to 1 in increments of 0.1. The legend is positioned at the bottom center below the graph and lists four models: “L R”, “R F”, “BiLSTM”, and “B E R T”, each represented by a distinct bar. For “Debunking Climate”: L R: 0.84, R F: 0.87, BiLSTM: 0.90, B E R T: 0.94. For “Debunking Ukraine”: L R: 0.83, R F: 0.86, BiLSTM: 0.89, B E R T: 0.93. For “L I A R”: L R: 0.65, R F: 0.70, BiLSTM: 0.74, B E R T: 0.79. For “FakeNewsNet”: L R: 0.71, R F: 0.74, BiLSTM: 0.78, B E R T: 0.83. For “Emergent”: L R: 0.67, R F: 0.70, BiLSTM: 0.75, B E R T: 0.80. For “Climate-Fact”: L R: 0.81, R F: 0.84, BiLSTM: 0.88, B E R T: 0.92. For “Multi-Fake-DetectIVE”: L R: 0.82, R F: 0.85, BiLSTM: 0.89, B E R T: 0.93. For “NewsPolYML”: L R: 0.74, R F: 0.77, BiLSTM: 0.81, B E R T: 0.86. For “M M C F N D”: L R: 0.71, R F: 0.75, BiLSTM: 0.79, B E R T: 0.84. For “M F N D”: L R: 0.76, R F: 0.79, BiLSTM: 0.83, B E R T: 0.88.F1 scores resulting from the validation
For comparison, general-purpose datasets such as LIAR, Emergent, FakeNewsNet and MMCFND had significantly poorer scores. On LIAR, for instance, BERT achieved an accuracy of 79% and an F1-score of 79%, as opposed to Random Forest’s even worse 70%. The Emergent dataset, which is less about fact verification and more about rumor stance classification, saw lower scores, with BiLSTM and BERT achieving approximately 75–80% in terms of accuracy. The performance of FakeNewsNet dataset is a little better, with BERT achieving an accuracy and F1-score of 83%. On the MMCFND, BERT achieved an accuracy and F1-score of 84%.
Moreover, Climate-Fact, a high-quality, domain-specific dataset of comparable size to Debunking Climate, performed strongly across all models with 92% accuracy and F1-score for BERT. This is in line with the result that domain-specific datasets, particularly those addressing salient social and political issues and expert-annotated, as the Debunking datasets are, consistently outperform generic collections. The Multi-Fake-DetectiVE dataset also performed strongly across all models, obtaining scores quite aligned with the Debunking datasets. Specifically, it shows that BERT achieves the highest performance, with an accuracy and F1-score of 93%. BiLSTM also performs well, while Random Forest and Logistic Regression show more moderate results, struggling with the complexities of multimodal data.
5. Conclusions and future works
The Debunking datasets have significant implications for advancing the fight against disinformation. The Debunking datasets aim to reduce the existing gaps in the field of disinformation research by addressing some of the key shortcomings of existing datasets. In particular, the proposed datasets enhance the ability of AI systems to detect disinformation efficiently, so that more effective countermeasures to the spread of false information online is achieved.
The Debunking datasets also have application in fact-checking agencies, policymakers, as well as researchers in a way that reduces human work, informs regulation, and enables the development of innovative detection tools. Specifically, these datasets can be utilized by fact-checking organizations to train and build AI systems that can identify disinformation and reduce the time and effort taken in manual verification of claims. Moreover, policymakers also have the potential to be assisted by these datasets in formulating better policies and regulations to counter the spread of disinformation in a way that interventions would be guided by a general understanding of how disinformation spreads across different types of platforms. Furthermore, these datasets can serve as a resource for researchers to create new solutions for disinformation detection, thus advancing the scientific understanding of this complex issue. Last, these datasets can have a positive impact on civil society, since they can be used to build fact-checking assistants embedded in browser plugins, messaging platforms or social media dashboards for flagging potentially misleading claims in real-time.
Future research on disinformation detection should look toward applying multimodal enhancements that can enhance the performance of AI models to detect disinformation in various modes, such as text, images, video and audio. That is, with more advanced adversarial strategies of disinformation that tend to be media-based, AI systems that could process and analyze multiple modalities will become crucial for effective detection and debunking of disinformation.
Going forward, ethical AI models would need to be integrated to carry out more research on how AI would be even more in the spotlight of disinformation detection and management. Such models need to render AI accountable, fair and transparent, free of biases that would compromise the detection process or disproportionately affect any specific topic or category of users.
Ethical considerations should also cover privacy concerns to ensure that AI systems do not inadvertently breach users’ rights in processing data for the detection of disinformation. The growth of and adherence to ethical guidelines will be useful in establishing trust in AI-enabled solutions to disinformation from the general public and from stakeholders who rely on them.
At a glance, the Debunking datasets would be an asset in bringing the fight against disinformation to another level with greater accuracy, speed and scalability of AI-based detection systems. However, to achieve the full potential of such information, forthcoming research should be focused on merging multimodal strengths and utilizing ethical approaches in a manner that will enable the intelligent use of AI technologies for detecting disinformation.
The authors thank all the members of the AI4Debunk consortium.
Notes
AI4Debunk project official site: https://ai4debunk.eu/
https://www.politifact.comClimate Feedback.
The supplementary material for this article can be found online.

