遇见数据集

Replication Data for: Detecting Formatted Text: Data Collection Using Computer Vision

收藏
DataONE2026-03-13 更新2026-05-19 收录
官方服务:

资源简介:

Research in political science has begun to explore how to use large language and object detection models to analyze text and visual data. However, few studies have explored how to use these tools for data extraction. Instead, researchers interested in extracting text from poorly formatted sources typically rely on optical character recognition and regular expressions or extract each item by hand. This letter describes a workflow process for structured text extraction using free models and software. I discuss the type of data best suited to this method, its usefulness within political science, and the steps required to convert the text into a usable dataset. Finally, I demonstrate the method by extracting agenda items from city council meeting minutes. I find the method can accurately extract sub-sections of text from a document and requires only a few hand labeled documents to adequately train.

政治学领域的研究已开始探索如何借助大语言模型(Large Language Model)与目标检测模型分析文本及视觉数据。然而,鲜有研究探讨如何利用这类工具开展数据抽取工作。与之相对,有志于从格式杂乱的数据源中抽取文本的研究者,通常会依赖光学字符识别(Optical Character Recognition, OCR)与正则表达式,或是手动逐条提取内容。本通讯介绍了一套利用免费模型与软件实现结构化文本抽取的工作流程。本文将探讨适配该方法的最优数据类型、其在政治学研究中的应用价值,以及将文本转换为可用数据集所需的具体步骤。最后,本文以从市议会会议纪要中抽取议程条目为例,对该方法进行了演示。研究结果表明,该方法可精准从文档中抽取文本子章节,且仅需少量人工标注的文档即可完成充分训练。

创建时间:
2026-04-07
二维码
社区交流群
二维码
科研交流群
商业服务