遇见数据集

FontVibe Kaomoji Dataset: 82,109 kaomoji merged from seven independent collections, with six-language semantic labels

收藏
Zenodo2026-09-27 更新2026-10-01 收录
官方服务:

资源简介:

82,109 kaomoji (Japanese text emoticons) merged and deduplicated from seven independent source families. Every entry carries the sources it came from, so the overlap statistics can be recomputed from the shipped data. The seven collections barely overlap: across the 21 pairs the median Jaccard overlap is 0.3%, four pairs share not a single entry, and only one clears 5%. 73,302 entries exist only in the Japanese IME dictionaries and appear in none of the English-side collections. 69,679 entries carry emotion, intent and subject labels drawn from a controlled vocabulary in English, Japanese, Chinese, Spanish, Portuguese and German, so equivalent terms across those languages return the same entries. A four-level tier split marks how safely each entry renders outside Japanese contexts. 106 entries are original work, drawn for concepts that had no kaomoji, and are released CC0. Coverage against each upstream source is verified at run time rather than asserted: scripts/verify_coverage.py downloads each upstream file when it runs and prints every entry it does not carry. Also distributed as kaomoji-dataset on npm and PyPI. A browsable version of the corpus is at fontvibe.ai/ja/tools/kaomoji. Published by Funovate Technology LLC (Japan Corporate Number: 9020003029274).

提供机构:
Zenodo
创建时间:
2026-09-24
二维码
社区交流群
二维码
科研交流群
商业服务