Package profile
sentencepiece
- Summary: Unsupervised text tokenizer and detokenizer.
- Author: Taku Kudo <taku@google.com>
- License: Apache-2.0
- Homepage: https://github.com/google/sentencepiece
- Source: https://github.com/google/sentencepiece
- Number of releases: 34
- First release: 0.0.0 on 2017-08-28
- Latest release: 0.2.2 on 2026-07-12
- Latest release size: 2.1 MB (wheel)
1 known vulnerabilityLatest: GHSA-38vq-g6vr-w8wf — Sentencepiece has a a heap overflow issue View all →
Dependencies
Sentencepiece has 3 dependencies, 3 of which optional.View all 3 dependencies →Dependent packages
| Package | Optional | Group |
|---|---|---|
| aiperf | false | |
| allennlp | false | |
| amd-quark | false | |
| auto-gptq | false | |
| bpemb | false |
Similar packages
- segtoksentence segmentation and word tokenization tools
- untokenizeTransforms tokens into original source code (while preserving whitespace).
- textparserText parser.
- pytokensA Fast, spec compliant Python 3.14+ tokenizer that runs on older Pythons.
- tokenize-rtA wrapper around the stdlib `tokenize` which roundtrips.
- langchain-text-splittersLangChain text splitting utilities
- chardetUniversal character encoding detector