A Multi-Domain Turkish Aspect-Based Sentiment Analysis Dataset (10K Annotated Texts)
收藏资源简介:
This dataset introduces a large-scale Turkish Aspect-Based Sentiment Analysis (ABSA) resource consisting of 10,000 manually annotated texts collected from diverse domains, including e-commerce, accommodation, entertainment, education, and social media platforms. Each text is annotated with:• Aspect (target) terms preserved in their original morphological surface forms,• Sentiment polarity labels (positive, negative, neutral),• Domain category information,• Source platform metadata. The dataset was created through fully manual data collection and annotation procedures without automated scraping tools, ensuring compliance with platform terms of service, ethical standards, and privacy protection regulations. Annotation quality was ensured through a multi-stage review process involving cross-checking among annotators and consensus-based adjudication of disagreements. The dataset supports two main ABSA subtasks:• Aspect Term Extraction (ATE)• Aspect-Level Sentiment Classification (ALSC) Data are provided in CSV format for statistical analysis and machine learning workflows, and in XML format to preserve hierarchical annotation structure. Additional documentation includes annotation guidelines and usage notes to facilitate reproducibility. This resource is intended to support research in Turkish natural language processing, particularly fine-grained sentiment analysis in morphologically rich languages, and to serve as a benchmark dataset for future ABSA studies.



