CarakaBinary
收藏资源简介:
This dataset contains a collection of 31,179 words annotated with categorical labels and corresponding numerical labels. The dataset has been preprocessed. As the name suggests, CarakaBinary is a dataset that is prepared from scratch to facilitate Sanskrit binary text classification, more precisely for the Sanskrit Ayurvedic compound and non-compound word classification. CarakaBinary includes 31,179 terms extracted from Carakasamhitā, an ancient Indian Ayurvedic text where the terms are annotated with the categorical labels, including "C" that stands for compound word and "NC" that indicates non-compound words. In addition to these terms, we provided numerical labels: 0 for non-compound words and 1 for compound words to make the function of computational analysis smoother. In addition to this, we provided the English translation of each term in the dataset.



