Document texts extraction of zbMATH Open volumes (1-529)
收藏资源简介:
The dataset represents the result of text extraction for the documents from zbMATH Open volumes. Document texts were extracted for 810,977 documents. Start/end indexes were obtained for 721,288 documents. Columns description: "id" - id of the document in the zbMATH Open volumes. "Abstract/Review/Summary" - extracted AI-enhanced document text. "first_char_info" - Position of the first character of the document text. 'Not reviewed' if it was not reviewed. 'No overlapping.' if was not found. "last_char_info" - Position of the last character of the document text. 'Not reviewed' if it was not reviewed. 'No overlapping.' if was not found. The position information is given in the following form {'file': '001/00000005.json', 'index': 3681}","{'file': '001/00000006.json', 'index': 1235} For the initial data, see also 10.5281/zenodo.14253145



