FontVibe Kaomoji Dataset: 82,109 kaomoji merged from seven independent collections, with six-language semantic labels
收藏资源简介:
82,109 kaomoji (Japanese text emoticons) merged and deduplicated from seven independent source families. Every entry carries the sources it came from, so the overlap statistics can be recomputed from the shipped data. The seven collections barely overlap: across the 21 pairs the median Jaccard overlap is 0.3%, four pairs share not a single entry, and only one clears 5%. 73,302 entries exist only in the Japanese IME dictionaries and appear in none of the English-side collections. 69,679 entries carry emotion, intent and subject labels drawn from a controlled vocabulary in English, Japanese, Chinese, Spanish, Portuguese and German, so equivalent terms across those languages return the same entries. A four-level tier split marks how safely each entry renders outside Japanese contexts. 106 entries are original work, drawn for concepts that had no kaomoji, and are released CC0. Coverage against each upstream source is verified at run time rather than asserted: scripts/verify_coverage.py downloads each upstream file when it runs and prints every entry it does not carry. Also distributed as kaomoji-dataset on npm and PyPI. A browsable version of the corpus is at fontvibe.ai/ja/tools/kaomoji. Published by Funovate Technology LLC (Japan Corporate Number: 9020003029274).



