遇见数据集

Web-Scraped Nigerian Pidgin English Text Dataset from Digital News Platforms

收藏
Zenodo2026-02-24 更新2026-05-29 收录
官方服务:

资源简介:

This dataset consists of Nigerian Pidgin English text collected through web scraping of multiple Nigerian Pidgin news and media websites, capturing a wide range of contemporary topics, linguistic styles, and sociocultural expressions. The corpus was cleaned, normalised, and curated to ensure linguistic consistency and usability for downstream natural language processing tasks. It was used to fine-tune and evaluate quantised Large Language Models, enabling analysis of performance–efficiency trade-offs in low-resource deployment scenarios. The dataset is designed to support research and development of robust Nigerian Pidgin English language models for multilingual NLP, low-resource language modelling, and culturally grounded AI applications.

本数据集包含通过网络爬虫抓取多座尼日利亚皮钦英语(Nigerian Pidgin English)新闻与媒体网站所获取的文本,涵盖丰富的当代议题、语言风格与社会文化表达。该语料库经清洗、标准化与精选处理,以保障语言一致性,并满足下游自然语言处理(Natural Language Processing,NLP)任务的使用需求。本数据集被用于微调与评估量化大语言模型(Large Language Model,LLM),助力分析低资源部署场景下的性能与效率权衡问题。本数据集旨在支撑鲁棒的尼日利亚皮钦英语语言模型的研发,以服务于多语种NLP、低资源语言建模以及基于文化语境的人工智能(Artificial Intelligence,AI)应用研究与开发。

提供机构:
Zenodo
创建时间:
2026-02-24
二维码
社区交流群
二维码
科研交流群
商业服务