遇见数据集

IndoEng-MT Sentiment Dataset

收藏
Mendeley Data2026-09-08 收录
官方服务:

资源简介:

The IndoEng-MT Sentiment Dataset is a parallel Indonesian–English sentiment analysis dataset designed for cross-lingual NLP research, machine translation evaluation, and multilingual large language model (LLM) benchmarking. Each sample contains an original Indonesian text and its English machine-translated counterpart, enabling researchers to study sentiment preservation and consistency across languages. The dataset includes the following columns: - text: Original Indonesian text. - translated_text: English machine-translated version of the original text. - sentiment: Sentiment label associated with the text. - has_slang: Indicates whether the Indonesian text contains slang, informal expressions, or non-standard vocabulary. - has_codeswitch: Indicates whether the text contains code-switching between Indonesian and another language (e.g., English). - has_ambivalent: Indicates whether the text exhibits mixed, conflicting, or ambiguous sentiment cues that may complicate sentiment classification.

创建时间:
2026-08-07
二维码
社区交流群
二维码
科研交流群
商业服务