Solving Label Variation in Scientific Information Extraction via Multi-Task Learning

Pham, Dong; Ho, Xanh; Ha, Quang-Thuy; Aizawa, Akiko

Abstract:Scientific Information Extraction (ScientificIE) is a critical task that involves the identification of scientific entities and their relationships. The complexity of this task is compounded by the necessity for domain-specific knowledge and the limited availability of annotated data. Two of the most popular datasets for ScientificIE are SemEval-2018 Task-7 and SciERC. They have overlapping samples and differ in their annotation schemes, which leads to conflicts. In this study, we first introduced a novel approach based on multi-task learning to address label variations. We then proposed a soft labeling technique that converts inconsistent labels into probabilistic distributions. The experimental results demonstrated that the proposed method can enhance the model robustness to label noise and improve the end-to-end performance in both ScientificIE tasks. The analysis revealed that label variations can be particularly effective in handling ambiguous instances. Furthermore, the richness of the information captured by label variations can potentially reduce data size requirements. The findings highlight the importance of releasing variation labels and promote future research on other tasks in other domains. Overall, this study demonstrates the effectiveness of multi-task learning and the potential of label variations to enhance the performance of ScientificIE.

Comments:	14 pages, 7 figures, PACLIC 37
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2312.15751 [cs.CL]
	(or arXiv:2312.15751v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2312.15751

Computer Science > Computation and Language

Title:Solving Label Variation in Scientific Information Extraction via Multi-Task Learning

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Pfad - The Proxy pFad of © 2024 Garber Painting. All rights reserved.