遇见数据集

Slosky: A Pseudonymized Corpus of Slovene Bluesky Posts

收藏
Zenodo2026-04-20 更新2026-05-26 收录
官方服务:

资源简介:

Slosky is a pseudonymized corpus of 141,013 public Slovene-language posts from the Bluesky social network, produced by 432 authors and spanning August 2023 to April 2026. The corpus was constructed through a multi-stage ATProto pipeline combining:(1) network-wide discovery of Slovene-tagged posts,(2) full-history backfill of discovered authors via DID resolution and retrieval from their home Personal Data Servers, and(3) post-level language filtering using langid.py and langdetect. Manual validation was used to assess inclusion quality for the retained decision groups. The corpus is released in both JSONL and CSV formats. Author identities are pseudonymized. DIDs and handles are replaced with sequential identifiers, @-mentions in post text are replaced with consistent pseudonyms, and personal domains are redacted. This resource is intended as a research corpus of Slovene Bluesky content, not as a statistically representative sample of all Slovene-language activity on Bluesky. Inclusion depends on discovery through Slovene-tagged posts and subsequent author-history backfill. Files:- slosky_corpus_anon.jsonl — corpus in JSON Lines format- slosky_corpus_anon.csv — corpus in CSV format- README.txt — dataset documentation Notice and take-down:If you believe this dataset contains material that should not be reproduced here, please contact the author. Legitimate requests will be addressed in the next released version of the corpus.

提供机构:
Zenodo
创建时间:
2026-04-17
二维码
社区交流群
二维码
科研交流群
商业服务