遇见数据集

Contrastive Phono-Lexical Dataset of Urdu and Bahasa Indonesia used in "Phono-lexical similarity between Bahasa Indonesia and Urdu: A corpus-based contrastive analysis study."

收藏
Zenodo2026-02-23 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains a corpus of 326 phono-lexically similar word pairs from Urdu and Bahasa Indonesia, compiled for the study titled “Phono-lexical similarity between Bahasa Indonesia and Urdu: A corpus-based contrastive analysis study.” The dataset systematically documents phonological form, semantic equivalence, syntactic category, and etymological origin of each lexical pair. Each entry includes Urdu script, Romanized transliteration, Bahasa Indonesia equivalent, IPA transcription, semantic comparison, grammatical classification, etymological background, and contrastive analysis coding. The primary aim of this dataset is to identify and classify cross-linguistic similarities and differences between Urdu and Bahasa Indonesia, particularly in relation to cognates, partial cognates, and false friends. Using a corpus-based contrastive analysis framework, the dataset contributes to research in applied linguistics, comparative linguistics, bilingual lexicography, and cross-linguistic transfer studies. This open-access dataset supports reproducibility, transparency, and further research on phonological resemblance and lexical convergence between typologically distinct languages. Associated manuscript under review. Keywords: phono-lexical similarity; contrastive analysis; cognates; partial cognates; false friends.

提供机构:
Zenodo
创建时间:
2026-02-23
二维码
社区交流群
二维码
科研交流群
商业服务