YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Tokenizer Inspector

A small CLI tool for SLM builders: inspect and compare Hugging Face tokenizers before you commit one to a training run.

What it tells you

  • Vocab size and token type (byte-level BPE, SentencePiece Unigram, WordPiece, etc.)
  • Special tokens โ€” what's added on top of the learned vocab
  • Chars/token on your test corpus โ€” higher means more compact (fewer tokens to train on)
  • Vocab utilization โ€” what fraction of the vocab your data actually touches
  • Roundtrip check โ€” does decode(encode(text)) give you back the same text?
  • Top tokens โ€” what the tokenizer is actually producing for your data
  • Side-by-side comparison of two or more tokenizers on the same corpus

Install

pip install tokenizers huggingface_hub

Usage

# Inspect a single tokenizer by repo ID
python3 tokenizer_inspector.py gpt2

# Compare two tokenizers side by side
python3 tokenizer_inspector.py gpt2 HuggingFaceTB/SmolLM2-135M --compare

# Tokenize your own text
python3 tokenizer_inspector.py gpt2 --text "Hello, world! This is a test."

# Show the first 30 tokens produced
python3 tokenizer_inspector.py gpt2 --show-tokens

# Use a local tokenizer.json file
python3 tokenizer_inspector.py ./my_tokenizer.json

# JSON output (for scripting)
python3 tokenizer_inspector.py gpt2 --json

Example output

============================================================
  COMPARISON (2 tokenizers)
============================================================
  Metric                               gpt2   HuggingFaceTB/Smol
  -------------------- -------------------- --------------------
  Vocab size                         50,257               49,152
  Token type                 byte-level BPE       byte-level BPE
  Special tokens                          1                   17
  Chars/token                           2.6                 2.72
  Corpus tokens                         116                  111
  Vocab util                          0.13%                0.12%
  Longest tok                           128                   81
  Roundtrip                              ok                 ok

When to use this

  • Before choosing a tokenizer for training: compare chars/token on a sample of your actual data, not just English prose
  • Debugging a bad tokenizer: if roundtrip fails, your tokenizer is lossy
  • Vocab sizing: if vocab utilization is very low (<1%), your vocab is oversized for your data
  • Special token audit: see exactly what control tokens are in the vocab

Notes

  • The default corpus is mixed English + Python code. Use --corpus or --text to test on your own data.
  • Chars/token is corpus-dependent. A tokenizer that's great for code may be worse for Turkish, and vice versa.
  • This tool reads tokenizer.json (the fast-rs format). For SentencePiece .model files, it falls back to that format.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support