遇见数据集

Document Quality Scoring for Web Crawling - Scored OWS data

收藏
Zenodo2025-03-31 更新2026-05-29 收录
官方服务:

资源简介:

This repository contains quality scores for the OWS datasets listed in Table 1 in [1]. The scores are computed with the QT5-small model trained by Chang et al [2] as outlined in [1] (containerised approach). For storage efficiency, we provide only the quality scores, not the full metadata files. However, the folder structure is the same as in the original dataset (as identified with the unique ID provided by the OWLER dashboard) for compatibility. The scores are arranged in the same order as the documents in the metadata parquet-files, where a file 'scores_0.txt' contains the scores for the documents in 'metadata_0.parquet' in the same folder in the original dataset. It is to be noted that the quality scores denote the log-probability of the document being relevant to any query. [1] Pezzuti, F., Mueller, A., MacAvaney, S. & Tonellotto, N. (2025, April). Document Quality Scoring for Web Crawling. In The Second International Workshop on Open Web Search (WOWS). [2] Chang, X., Mishra, D., Macdonald, C., & MacAvaney, S. (2024, July). Neural Passage Quality Estimation for Static Pruning. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 174-185).

提供机构:
Zenodo
创建时间:
2025-03-30
二维码
社区交流群
二维码
科研交流群
商业服务