遇见数据集

Multilingual MigrationsKB: A Mulitlingual Knowledge Base of Migration related annotated Tweets

收藏
Zenodo2022-01-29 更新2026-04-07 收录
数据链接:
官方服务:

资源简介:

<strong>Multilingual MigrationskB (MGKB) </strong>is a mulitlingual extended version of English MGKB. The tweets geotagged with Geo location from 32 European Countries (<em><strong>Austria, Belgium, Bulgaria, Croatia, Cyprus, Czech, Denmark, Estonia, Finland, France, Germany, Greece, Hungary, Ireland, Italy, Latvia, Lithuania, Luxembourg, Malta, Netherlands, Poland, Portugal, Romania, Slovakia, Slovenia, Spain, Sweden, Iceland, Liechtenstein, Norway, Switzerland, the United Kingdom</strong></em>) are extracted and filtered by 11 languages (<em><strong>English, French, Finnish, German, Greek, Dutch, Hungarian, Italian, Polish, Spain, Swedish</strong></em>). Metadata information about the tweets, such as <strong>Geo information (place name, coordinates, country code)</strong> are included. <strong>MGKB</strong> contains <strong>sentiments, offensive and hate speeches, topics, hashtags, user mentions</strong> in RDF format. The schema of <strong>MGKB</strong> is an extension of TweetsKB for migration related information. Moreover, to associate and represent the potential economic and social factors driving the migration flows, the data from Eurostat and FIBO ontology was used. To represent multilinguality, the CIDOC Conceptual Reference Model (CIDOC-CRM) is used. The extracted economic indicators, i.e., GDP Growth Rate, Total Unemployment Rate, Youth Unemployment Rate, Long-term Unemployment Rate and Income per househould, are connected with each tweet in RDF using geographical and temporal dimensions. For this version, the Multilingual MGKB is delivered separated by year. The extracted topic words are also published. Code: https://github.com/migrationsKB/MRL Please contact Yiyi Chen (yiyi.chen@partner.kit.edu) for pretrained models (Sentiment analysis/hate speech detection/ETM) if necessary.

提供机构:
Yiyi Chen
创建时间:
2022-01-29
二维码
社区交流群
二维码
科研交流群
商业服务