遇见数据集

Particle-Less Japanese Corpus

收藏
Zenodo2025-09-24 更新2026-05-26 收录
官方服务:

资源简介:

Description This dataset contains a large-scale, pre-processed Japanese text corpus derived from the Japanese Wikipedia (jawiki). The source is a full Wikipedia dump dated June 1, 2025. The final corpus consists of approximately 25.4 million sentences and has been specifically processed to be "particle-less". Its primary feature is the removal of grammatical particles and other non-semantic tokens to focus on content words. Corpus Creation Process The corpus was generated using a custom Python script with the following methodology: Source Data: The raw text was extracted from a Japanese Wikipedia XML dump. Tokenization: Morphological analysis and tokenization were performed using the MeCab tagger. Token Reconstruction Logic: A custom look-ahead algorithm was applied to preserve the integrity of conjugated words. Base forms of verbs (動詞) and adjectives (形容詞) were automatically concatenated with their subsequent auxiliary verbs (助動詞). For example, the token sequence 食べ (tabe) and た (ta) is reconstructed into the single token 食べた (tabeta). Filtering: A comprehensive filtering process was applied to remove tokens belonging to Part-of-Speech (POS) categories that are often considered grammatical or non-semantic noise. The following categories were excluded: 助詞 (Particles) 補助記号 (Supplementary Symbols: commas, periods, etc.) 接尾辞 (Suffixes) 接頭辞 (Prefixes) 感動詞 (Interjections) フィラー (Fillers) 記号 (General Symbols) 空白 (Whitespace) Data Format File: particle_less_japanese_corpus.txt Structure: The corpus is provided as a single text file. Each line in the file represents one processed sentence. Tokens within a sentence are separated by a single space. Example Line:釣り 対する 思い入れ 大きい

提供机构:
Zenodo
创建时间:
2025-09-24
二维码
社区交流群
二维码
科研交流群
商业服务