遇见数据集

Raw Early Modern Bibliographic and Actor Data Retrieved from the Bibliothèque nationale de France SPARQL Endpoint (1454–1799)

收藏
Zenodo2026-08-10 更新2026-08-13 收录
官方服务:

资源简介:

Dataset overview The dataset consists of two complementary collections of raw data retrieved from the Bibliothèque nationale de France (BnF) SPARQL endpoint, covering the period 1454–1799. The first contains bibliographic edition data for the complete chronological range, while the second contains consolidated data on the distinct actors associated with those editions. Actor records include the entities identified during the acquisition process and were retrieved through individual actor-level queries before being merged into a single dataset. Together, the two collections provide the complete raw input for the subsequent harmonisation pipeline, supporting data cleaning, normalisation, entity reconciliation, bibliographic–actor linking, and RDF materialisation. The code for data retrieval was developed within the scope of research activities led at the Helsinki Computational History Group (COMHIS) and is available at https://github.com/ariannamorettj/Early_Modern_BnF_Harmonisation/tree/main/01_data_retrieval under an ISC license. Dataset structure Editions (bibliographic data) edition_data.zip contains: bnf_edition_data_raw.csv: complete compiled dataset covering all years from 1454 to 1799. edition_raw_data_by_year/: one CSV file per year containing the raw data; this provides an alternative organisation of the same data contained in the single compiled CSV file. sessionInfo_data_acquisition_20260714_054125.txt: session information and execution log for the bibliographic data acquisition process. It records the computational environment used for the acquisition, including the R version, operating system, locale, time zone, and the versions of the main R packages and dependencies. It also contains a batch-level execution log that indicates the processed year ranges, start and end times, duration, and completion status for each acquisition batch. Actors actor_data.zip contains: actor_data.csv : final merged dataset, recently updated (220 MB, including the 670 recovered actors) actor_queries_results/ : one CSV file for each actor (124,695 files) distinct_actors_cache.csv : complete list of distinct actors with their corresponding index last_processed_index.txt : processing progress checkpoint sessionInfo_data_acquisition_20260714_140220.txt, sessionInfo_data_acquisition_20260714_155802.txt, sessionInfo_data_acquisition_20260717_001545.txt : these files document the BnF actors data acquisition process, recording the R execution environment and the batch-by-batch retrieval of actor records. They cover the acquisition from the initial test batch through the subsequent large-scale runs, including completed and skipped batches, processing times, and the final acquisition of all 124,695 actors. Process monitor and report data monitor_data.zip contains: recover_missing_acquisitions_20260805_115752_R.txt : complete report of the latest recovery run, integrating data that might have been missed in the original retrieval process. query_editions_*_R.txt and query_agents_*_R.txt : reports from the original data acquisition runs resume_info.log : status and outcome of the most recent execution For data analysis, we suggest using the two complete, final datasets: actor_data.csv (actors) and bnf_edition_data_raw.csv (editions).

提供机构:
Zenodo
创建时间:
2026-08-10
二维码
社区交流群
二维码
科研交流群
商业服务