openeurollm/propella-annotations
收藏资源简介:
propella annotations数据集包含由propella-1-4b模型生成的多语言文档注释,覆盖六个主要类别:核心内容、分类、质量与价值、受众与目的、安全与合规、地理相关性。这些注释可用于大规模筛选、选择和整理LLM训练数据。数据集包含多个子集,如fineweb-2、finepdfs、hplt-3等,每个子集都有详细的注释数量和语言分布。数据集还提供了使用示例、许可证信息(CC-BY-4.0)、引用和致谢部分。
The propella annotations dataset contains document annotations produced with propella-1-4b, a small multilingual LLM that annotates text documents across six categories: core content, classification, quality & value, audience & purpose, safety & compliance, and geographic relevance. The annotations can be used to filter, select, and curate LLM training data at scale. The dataset includes multiple subsets such as fineweb-2, finepdfs, hplt-3, etc., each with detailed annotation counts and language distributions. The README also provides usage examples, license information (CC-BY-4.0), citation, and acknowledgments.




