Package profile
jieba
- Summary: Chinese Words Segmentation Utilities
- Author: Sun, Junyi
- License: MIT License
- Homepage: https://github.com/fxsjy/jieba
- Source: https://github.com/fxsjy/jieba
- Number of releases: 32
- First release: 0.20 on 2012-11-06
- Latest release: 0.42.1 on 2020-01-20
- Latest release size: 18.3 MB (sdist)
Dependencies
Jieba has no dependency.Dependent packages
| Package | Optional | Group |
|---|---|---|
| funasr | false | |
| crfm-helm | true | cleva |
| goose3 | true | chinese |
| inspect-evals | true | sevenllm |
| lighteval | true | multilingual |
Similar packages
- segtoksentence segmentation and word tokenization tools
- sudachipyPython version of Sudachi, the Japanese Morphological Analyzer
- mecab-python3Python wrapper for the MeCab morphological analyzer for Japanese
- pythainlpThai Natural Language Processing library
- wordninjaProbabilistically split concatenated words using NLP based on English Wikipedia uni-gram frequencies.
- langchain-text-splittersLangChain text splitting utilities
- semchunkA Python library for splitting text into smaller chunks while preserving as much local semantic context as possible.