遇见数据集

ClimaText

收藏
arXiv2021-01-03 更新2024-07-25 收录
官方服务:

资源简介:

ClimaText数据集由苏黎世联邦理工学院创建,专注于气候变化主题检测。数据集包含715条与气候变化相关的标注句子,来源于Wikipedia、美国证券交易委员会(SEC)10K文件及网络收集的气候变化声明。数据集通过图基于的启发式方法和人工标注过程构建,旨在通过机器学习模型识别气候变化相关文本,应用于内容过滤、情感分析、自动摘要、问答及事实核查等领域,以解决气候变化信息自动化提取的挑战。

The ClimaText dataset was created by ETH Zurich and focuses on climate change topic detection. It contains 715 annotated climate change-related sentences sourced from Wikipedia, U.S. Securities and Exchange Commission (SEC) 10-K filings, and climate change statements collected from the web. Constructed via graph-based heuristic methods and manual annotation procedures, this dataset is designed to enable machine learning models to recognize climate change-related text, and has applications in fields including content filtering, sentiment analysis, automatic summarization, question answering, and fact-checking, aiming to address the challenges of automated extraction of climate change information.

创建时间:
2020-12-01
二维码
社区交流群
二维码
科研交流群
商业服务