Web Survey: Data Annotation Bottleneck and Active Learning for Natural Language Processing in the Era of Large Language Models
收藏资源简介:
We publish a dataset of 144 community perspectives on the need for annotated training data in the field of Natural Language Processing (NLP). The dataset documents how large language models (LLMs) have impacted the persistent issue of the lack of annotated data. Additionally, we inquire about the current state of the "active learning" annotation method. Our target group consists of experts from academia, industry, and government institutions. The data was collected with a web survey that was open online to voluntary participants for 6 weeks, from December 15th, 2024, to January 26th, 2025. keywords: large language models, supervised learning, data annotation, active learning, practical application We publish a dataset of 144 community perspectives on the need for annotated training data in the field of Natural Language Processing (NLP). The dataset documents how large language models (LLMs) have impacted the persistent issue of the lack of annotated data. Additionally, we inquire about the current state of the "active learning" annotation method. Our target group consists of experts from academia, industry, and government institutions. The data was collected with a web survey that was open online to voluntary participants for 6 weeks, from December 15th, 2024, to January 26th, 2025. keywords: large language models, supervised learning, data annotation, active learning, practical application



