遇见数据集

Document texts extraction of zbMATH Open volumes (1-529)

收藏
Zenodo2025-07-23 更新2026-05-26 收录
官方服务:

资源简介:

The dataset represents the result of text extraction for the documents from zbMATH Open volumes. Document texts were extracted for 810,977 documents. Start/end indexes were obtained for 721,288 documents. Columns description: "id" - id of the document in the zbMATH Open volumes. "Abstract/Review/Summary" - extracted AI-enhanced document text. "first_char_info" - Position of the first character of the document text. 'Not reviewed' if it was not reviewed. 'No overlapping.' if was not found. "last_char_info" - Position of the last character of the document text. 'Not reviewed' if it was not reviewed. 'No overlapping.' if was not found. The position information is given in the following form {'file': '001/00000005.json', 'index': 3681}","{'file': '001/00000006.json', 'index': 1235} For the initial data, see also 10.5281/zenodo.14253145

提供机构:
Zenodo
创建时间:
2025-07-23
二维码
社区交流群
二维码
科研交流群
商业服务