YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Tokenizer Inspector
A small CLI tool for SLM builders: inspect and compare Hugging Face tokenizers before you commit one to a training run.
What it tells you
- Vocab size and token type (byte-level BPE, SentencePiece Unigram, WordPiece, etc.)
- Special tokens โ what's added on top of the learned vocab
- Chars/token on your test corpus โ higher means more compact (fewer tokens to train on)
- Vocab utilization โ what fraction of the vocab your data actually touches
- Roundtrip check โ does
decode(encode(text))give you back the same text? - Top tokens โ what the tokenizer is actually producing for your data
- Side-by-side comparison of two or more tokenizers on the same corpus
Install
pip install tokenizers huggingface_hub
Usage
# Inspect a single tokenizer by repo ID
python3 tokenizer_inspector.py gpt2
# Compare two tokenizers side by side
python3 tokenizer_inspector.py gpt2 HuggingFaceTB/SmolLM2-135M --compare
# Tokenize your own text
python3 tokenizer_inspector.py gpt2 --text "Hello, world! This is a test."
# Show the first 30 tokens produced
python3 tokenizer_inspector.py gpt2 --show-tokens
# Use a local tokenizer.json file
python3 tokenizer_inspector.py ./my_tokenizer.json
# JSON output (for scripting)
python3 tokenizer_inspector.py gpt2 --json
Example output
============================================================
COMPARISON (2 tokenizers)
============================================================
Metric gpt2 HuggingFaceTB/Smol
-------------------- -------------------- --------------------
Vocab size 50,257 49,152
Token type byte-level BPE byte-level BPE
Special tokens 1 17
Chars/token 2.6 2.72
Corpus tokens 116 111
Vocab util 0.13% 0.12%
Longest tok 128 81
Roundtrip ok ok
When to use this
- Before choosing a tokenizer for training: compare chars/token on a sample of your actual data, not just English prose
- Debugging a bad tokenizer: if roundtrip fails, your tokenizer is lossy
- Vocab sizing: if vocab utilization is very low (<1%), your vocab is oversized for your data
- Special token audit: see exactly what control tokens are in the vocab
Notes
- The default corpus is mixed English + Python code. Use
--corpusor--textto test on your own data. - Chars/token is corpus-dependent. A tokenizer that's great for code may be worse for Turkish, and vice versa.
- This tool reads
tokenizer.json(the fast-rs format). For SentencePiece.modelfiles, it falls back to that format.
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support