遇见数据集

BCD: A Large-Scale Linguistically Annotated Benchmark Dataset for Bangla Compound Word Decomposition

收藏
Zenodo2026-08-09 更新2026-08-13 收录
官方服务:

资源简介:

BCD (Bangla Compound Decomposition) is a large-scale linguistically annotated dataset designed for research on Bangla compound word decomposition, morphological analysis, and semantic relation modeling. The dataset contains 23,394 Bangla compound-word instances and provides structured linguistic information for each compound. The primary task represented in the dataset is the decomposition of a Bangla compound word into its constituent roots. For example: Compound word: বিদ্যালয়Decomposition: বিদ্যা + আলয় In addition to the original compound and its decomposition, BCD provides morphological, structural, semantic, etymological, and source-related annotations. This makes the dataset suitable not only for compound-word segmentation but also for developing and evaluating NLP models for morphological parsing, sequence generation, semantic relation classification, and linguistically informed language modeling. Dataset Features Each instance contains the following fields: id – Unique identifier for each dataset instance. compound_word – The original Bangla compound word. decomposition – Decomposition of the compound into its constituent parts. root_1 – First constituent/root of the compound. root_2 – Second constituent/root of the compound. compound_type – Structural classification of the compound, such as combinations of grammatical categories. semantic_relation – Semantic relationship between the constituent roots. origin – Linguistic origin of the compound, including categories such as Sanskrit. source – Source from which the compound was collected, including resources such as dictionaries, newspapers, and Wikipedia. Semantic Annotations A distinctive feature of BCD is its semantic-relation annotation. The dataset represents a range of relationships between compound constituents, including: LOCATION_OF PERSON_ROLE PART_OF MADE_OF PURPOSE_OF CONTAINER_OF SOURCE_OR_ORIGIN AGENT_OF PATIENT_OR_TARGET TIME_OF QUALITY_OR_ATTRIBUTE POSSESSION INSTRUMENT_FOR COLLECTION_OR_GROUP ABSTRACT_RELATION These annotations enable researchers to investigate the relationship between morphological structure and compound semantics, rather than treating compound decomposition as a purely character-level segmentation problem. Potential Applications BCD can be used for a variety of Bangla NLP and computational linguistics tasks, including: Bangla compound word decomposition Morphological analysis and parsing Sequence-to-sequence compound decomposition Character-level and byte-level language modeling Semantic relation classification Compound-type classification Multitask learning for Bangla morphology Evaluation of multilingual and language-specific Transformer models Development of rule-based and neural morphological analyzers Benchmarking Large Language Models (LLMs) for Bangla morphology The dataset is particularly suitable for investigating whether character-level and byte-level Transformer architectures, such as ByT5, can effectively learn the morphological structure of Bangla compounds. Intended Research Use BCD is intended to provide a standardized resource for researchers working on low-resource and morphologically rich language processing, particularly Bangla. Researchers can use the dataset to establish reproducible benchmarks, compare traditional linguistic approaches with neural architectures, and investigate how different modeling strategies handle Bangla compound formation. The dataset can also support future work on multi-task morphological modeling, where a model jointly predicts compound decomposition, constituent roots, compound type, and semantic relation. Dataset Summary Property Description Dataset name BCD – Bangla Compound Decomposition Language Bangla (Bengali) Number of instances 23,394 Main task Compound word decomposition Morphological information Root 1, Root 2, decomposition Structural information Compound type Semantic information Semantic relation Etymological information Origin Source information Source Primary application Bangla NLP and computational morphology BCD is intended to serve as a reusable benchmark and linguistic resource for advancing research in Bangla compound word decomposition and morphological analysis.

提供机构:
Zenodo
创建时间:
2026-08-09
二维码
社区交流群
二维码
科研交流群
商业服务