Elgold intermediate: raw texts
收藏资源简介:
The dataset contains raw texts scrapped from various internet sources which were used for creating the Elgold dataset. The texts were collected from 7 main categories: "News", "Job offers", "Movie reviews", "Automotive blogs", "Amazon product reviews", "Scientific papers abstracts", and "Historic blogs". The Scientific Papers category was additionally divided into five subcategories: "Biomedicine", "Life Sciences", "Mathematics", "Medicine & Public Health", and "Science, Humanities and Social Sciences, multidisciplinary". The raw texts were collected from publicly available Internet sources by the group of 14 participants. Every category has 2-3 participants assigned. The dataset consists of approximately 100 texts for each category (and subcategory in the case of "Scientific papers abstracts").



