YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Repetition Diagnostic
A builder tool for tiny language model developers. Sweeps repetition_penalty values and measures how they affect generation quality.
Why
Tiny models (<50M params) are extremely prone to degenerate repetition loops. The default repetition_penalty=1.0 means no penalty โ the model can output the same token forever. Most training scripts don't test generation quality, so you discover this at inference time with no idea what to do.
This tool gives you a quick diagnostic: sweep the penalty, see where quality peaks, and know what to set in your inference config.
What it measures
| Metric | Description |
|---|---|
| Repetition rate | Fraction of output tokens that already appeared earlier in the sequence |
| Unique token ratio | Unique tokens / total tokens (diversity) |
| Max loop length | Longest repeated n-gram (how deep the loop goes) |
| Repeated trigrams | Count of 3-token sequences appearing 2+ times |
Usage
# Basic: sweep penalties on a single prompt
python3 repetition_diagnostic.py --model gpt2 --prompt "Once upon a time"
# Multiple prompts
python3 repetition_diagnostic.py --model gpt2 --prompts "Once upon a time|The cat sat|I want to"
# Custom penalty range
python3 repetition_diagnostic.py --model gpt2 --penalties "1.0,1.05,1.1,1.2,1.5"
# Save results as JSON
python3 repetition_diagnostic.py --model gpt2 --output results.json
Example output
Model: gpt2
Prompts: ['Once upon a time']
Penalties: [1.0, 1.1, 1.2, 1.5, 2.0]
Prompt RP Tokens Repet Unique MaxLoop RepTri
------------------------------------------------------------------------
Once upon a time 1.0 48 25.5% 72.9% 3 1
Once upon a time 1.1 48 0.0% 100.0% 0 0
Once upon a time 1.2 48 0.0% 100.0% 0 0
Once upon a time 1.5 48 0.0% 100.0% 0 0
Once upon a time 2.0 48 0.0% 100.0% 0 0
Interpretation
- rp=1.0 with high repetition rate: Your model has a repetition problem. Try rp=1.1 or 1.2.
- rp=1.1 already fixes it: Good โ your model is mostly fine, just needs a small nudge.
- Even rp=2.0 still shows loops: The model is severely undertrained or has a degenerate mode. No penalty will save it.
- rp=1.0 shows no repetition: Your model is robust. You can use the default.
Requirements
- Python 3.8+
torchtransformers
pip install torch transformers
Notes
- Uses the same seed for all penalty values so the comparison is fair (same random draws, different penalty).
- Generation is done manually (not
model.generate()) so the penalty is applied exactly as specified, without HuggingFace's internal logic. - Works with any model that
AutoModelForCausalLMcan load.
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support