Panning for gold: Lessons learned from the platform-agnostic automated detection of political content in textual data

Makhortykh, Mykola; de León, Ernesto; Urman, Aleksandra; Christner, Clara; Sydorova, Maryna; Adam, Silke; Maier, Michaela; Gil-Lopez, Teresa

Computer Science > Computation and Language

arXiv:2207.00489 (cs)

[Submitted on 1 Jul 2022]

Title:Panning for gold: Lessons learned from the platform-agnostic automated detection of political content in textual data

Authors:Mykola Makhortykh, Ernesto de León, Aleksandra Urman, Clara Christner, Maryna Sydorova, Silke Adam, Michaela Maier, Teresa Gil-Lopez

View PDF

Abstract:The growing availability of data about online information behaviour enables new possibilities for political communication research. However, the volume and variety of these data makes them difficult to analyse and prompts the need for developing automated content approaches relying on a broad range of natural language processing techniques (e.g. machine learning- or neural network-based ones). In this paper, we discuss how these techniques can be used to detect political content across different platforms. Using three validation datasets, which include a variety of political and non-political textual documents from online platforms, we systematically compare the performance of three groups of detection techniques relying on dictionaries, supervised machine learning, or neural networks. We also examine the impact of different modes of data preprocessing (e.g. stemming and stopword removal) on the low-cost implementations of these techniques using a large set (n = 66) of detection models. Our results show the limited impact of preprocessing on model performance, with the best results for less noisy data being achieved by neural network- and machine-learning-based models, in contrast to the more robust performance of dictionary-based models on noisy data.

Subjects:	Computation and Language (cs.CL); Computers and Society (cs.CY)
Cite as:	arXiv:2207.00489 [cs.CL]
	(or arXiv:2207.00489v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2207.00489

Submission history

From: Mykola Makhortykh [view email]
[v1] Fri, 1 Jul 2022 15:23:23 UTC (1,127 KB)

Computer Science > Computation and Language

Title:Panning for gold: Lessons learned from the platform-agnostic automated detection of political content in textual data

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Pfad - The Proxy pFad of © 2024 Garber Painting. All rights reserved.

Computer Science > Computation and Language

Title:Panning for gold: Lessons learned from the platform-agnostic automated detection of political content in textual data

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Pfad - The Proxy pFad of © 2024 Garber Painting. All rights reserved.