遇见数据集

Flash Flood Custom Named Entity Recognition Model: Dataset

收藏
DataCite Commons2025-06-02 更新2025-04-16 收录
官方服务:

资源简介:

This published dataset is a source for researchers and practitioners to improve upon the performances of the (Flash Flood) FF-NER model. FF-NER, a custom Named Entity Recognition (NER) model designed to mine relevant information from diverse and voluminous FF-related web-based texts. FF-NER extracts specific information for eight entities: Region, Street, Location, County, Water Body, Monetary Damage, Power Outages, and Infrastructure Services. To develop FF-NER, we curated a dataset of 2,670 FF-related web paragraphs and experimented with conventional and advanced NER techniques. The FF-related paragraphs are the ones labeled 1 under the Damage/Impact category in the Flash Flood BERT Text Classification Model: Dataset (https://doi.org/10.17603/ds2-p9jf-xe52 v1). Our final model outperforms the baseline, improving accuracy, precision, recall, and F1 score by 4.6%, 7.1%, 4.6%, 9.35%, respectively. The study is currently under publication. We have designed FF-NER to be jointly used with FF-IR (https://doi.org/10.1016/j.envsoft.2023.105734) and FF-BERT (https://doi.org/10.1016/j.aei.2023.102293). Engineers, researchers, and practitioners can use FF-NER to enhance existing databases, recognize patterns, and conduct detailed analysis of past FF events. This dataset contains six csv files, BIO_Dataset_Train.csv, BIO_Dataset_Test.csv, BIOES_Dataset_Train.csv, BIOES_Dataset_Test.csv, Prompt_FF_NER_Train.csv, and Prompt_FF_NER_Test.csv. These represent three different datasets (BIO-tagged, BIOES_tagged, and Prompt-based) and for each type of dataset, we have a Training file (2170 paragraphs) and a Testing file (500 paragraphs). We used the BIOES and BIO tagged datasets to train and test the encoder-based deep learning architectures, whereas we used the Prompt based dataset to fine-tune a decoder-based Large Language Model (i.e., LlaMA2-13B). We used the train file to train the models and the test file to compare the performances of the trained models using unseen test data. Details regarding the dataset are available in the readme file.

提供机构:
Designsafe-CI
创建时间:
2024-04-10
二维码
社区交流群
二维码
科研交流群
商业服务