Sociolinguistically Informed Interpretability: A Case Study on Hinglish Emotion Classification

Tatariya, Kushal; Lent, Heather; Bjerva, Johannes; de Lhoneux, Miryam

Computer Science > Computation and Language

arXiv:2402.03137 (cs)

[Submitted on 5 Feb 2024]

Title:Sociolinguistically Informed Interpretability: A Case Study on Hinglish Emotion Classification

Authors:Kushal Tatariya, Heather Lent, Johannes Bjerva, Miryam de Lhoneux

View PDF

Abstract:Emotion classification is a challenging task in NLP due to the inherent idiosyncratic and subjective nature of linguistic expression, especially with code-mixed data. Pre-trained language models (PLMs) have achieved high performance for many tasks and languages, but it remains to be seen whether these models learn and are robust to the differences in emotional expression across languages. Sociolinguistic studies have shown that Hinglish speakers switch to Hindi when expressing negative emotions and to English when expressing positive emotions. To understand if language models can learn these associations, we study the effect of language on emotion prediction across 3 PLMs on a Hinglish emotion classification dataset. Using LIME and token level language ID, we find that models do learn these associations between language choice and emotional expression. Moreover, having code-mixed data present in the pre-training can augment that learning when task-specific data is scarce. We also conclude from the misclassifications that the models may overgeneralise this heuristic to other infrequent examples where this sociolinguistic phenomenon does not apply.

Comments:	5 pages, Accepted to SIGTYP 2024 @ EACL
Subjects:	Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:2402.03137 [cs.CL]
	(or arXiv:2402.03137v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2402.03137

Submission history

From: Kushal Jayesh Tatariya [view email]
[v1] Mon, 5 Feb 2024 16:05:32 UTC (336 KB)

Computer Science > Computation and Language

Title:Sociolinguistically Informed Interpretability: A Case Study on Hinglish Emotion Classification

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Pfad - The Proxy pFad of © 2024 Garber Painting. All rights reserved.

Computer Science > Computation and Language

Title:Sociolinguistically Informed Interpretability: A Case Study on Hinglish Emotion Classification

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Pfad - The Proxy pFad of © 2024 Garber Painting. All rights reserved.