遇见数据集

SIH/lindera-dicts

收藏
Hugging Face2026-05-14 更新2026-06-14 收录
官方服务:

资源简介:

该数据集是预构建的Lindera二进制词典,用于日语(IPADIC和UniDic)和韩语(ko-dic)的形态素分割。这些词典被打包供LDaCA文本分析工具网络应用按需下载。网络应用通过其polars-text Rust扩展使用这些词典——当用户首次对日语或韩语语料库运行标记化时,匹配的压缩包会从此数据集获取并缓存到每个操作系统的缓存目录中(macOS为~/Library/Caches/ldaca/lindera/,Linux为~/.cache/ldaca/lindera/,Windows为%LOCALAPPDATA%ldacalindera),并在后续调用中重复使用。将词典捆绑在wheel中会超过安装大小限制,因此我们以精简方式分发wheel,并在首次使用时从此处流式传输词典。

Prebuilt Lindera binary dictionaries for Japanese (IPADIC, UniDic) and Korean (ko-dic) morpheme segmentation, packaged for on-demand download by the LDaCA Text Analytics Tools web app. The web app uses these dictionaries through its polars-text Rust extension — the first time a user runs tokenisation on a Japanese or Korean corpus, the matching tarball is fetched from this dataset into the per-OS cache directory (~/Library/Caches/ldaca/lindera/ on macOS, ~/.cache/ldaca/lindera/ on Linux, %LOCALAPPDATA%ldacalindera on Windows) and reused for every subsequent call. Bundling the dicts in the wheel would push the install size past our hard limit, so we ship the wheel slim and stream the dict here on first use.

提供机构:
SIH
二维码
社区交流群
二维码
科研交流群
商业服务