Automatically Generated FAIR Datasets: A Survey of Present and Potential Outdoor Self-Cleaning Coating Materials by an Advanced Literature Screening Strategy
收藏资源简介:
In this work, we use Natural Language Processing (NLP) strategies to generate open- and traceable-sourced FAIR (findable, accessible, interoperable and reusable) datasets about self-cleaning coating materials from 39,011 multi-disciplinary papers. The data are from function-based and property-based corpora for self-cleaning purposes. These datasets are presented in four different formats for diverse uses or combined uses: material frequency statistics, material dictionary, measurement value datasets for self-cleaning-related properties and optical properties, and sentiment statistics of material stability and durability. This provides a literature-based, comprehensive and highly-refined data resource for the development of self-cleaning coatings and also offers potential pathways for material discovery and prediction by machine learning.



