Sol Lite Base
Lite Base reuses four transformer blocks on a second pass. That gives it fourteen block applications while storing ten blocks. We trained the 14,995,843-parameter model from scratch on 8,153,333,760 tokens.
The released weights are for text completion. There was no instruction tuning or assistant safety tuning. The repository includes a PyTorch loader and generation helper.
Architecture
| Setting | Value |
|---|---|
| Parameters | 14,995,843 |
| Training tokens | 8,153,333,760 |
| Context | 2,048 tokens |
| Tokenizer | 4,096-entry digit-aware byte-level BPE |
| Hidden width | 256 |
| Stored blocks / effective applications | 10 / 14 |
| Recurrent layout | 1 prelude, 4 middle blocks used twice, 5 coda blocks |
| Attention | 8 query heads, 2 KV heads, head dimension 32 |
| Q/K normalization | Per-head RMSNorm with RoPE |
| Recurrence | Learned pass embeddings and channel-wise refresh gates |
| FFN | Gated SiLU, width 1,465 |
| Local memory | 2,048-entry bigram/trigram EngramLite |
| Token embeddings | Tied to the output head |
| Weights | FP32 safetensors |
Grouped-query attention and loop conditioning support the repeated pass. XSA subtracts attention values, and EngramLite supplies local memory. Learned pass embeddings and channel-wise refresh gates distinguish the two uses of the middle blocks.
Generate text
Install the dependencies, then download the repository and import its standalone implementation:
pip install "torch>=2.5" "transformers>=5" safetensors huggingface_hub
from huggingface_hub import snapshot_download
import sys
model_dir = snapshot_download("solintellegence/Sol-Lite-Base")
sys.path.insert(0, model_dir)
from modeling_sol_lite import load_model, generate
model, tokenizer = load_model(model_dir, device="cpu")
text = generate(
model,
tokenizer,
"The future of efficient language models is",
max_new_tokens=64,
)
print(text)
Benchmark results
| Benchmark | Examples | Normalized accuracy |
|---|---|---|
| HellaSwag | 10,042 | 27.72% |
| ARC-Easy | 2,376 | 34.72% |
| ARC-Challenge | 1,172 | 22.87% |
| PIQA | 1,838 | 57.73% |
| ArithMark-3 | 1,000 | 34.20% |
The local Intelligence Index score is 8.799. We ran the complete zero-shot language-model splits with lm-eval 0.4.12. ArithMark-3 used independent tokenization and normalized accuracy. Raw evaluation outputs are saved in evals/.
Training sources
The English curriculum combined FineWeb-Edu, DCLM, UltraFineWeb levels 1-3, FineWeb-HQ, FineMath, and FinePhrase. We prepared the stream beforehand and kept the tokenizer fixed. Public benchmark examples weren't used to choose the checkpoint.
Files
The checkpoint is model.safetensors. Use modeling_sol_lite.py for SolForCausalLM, loading, and generation; config.json supplies the architecture settings. tokenizer.json and tokenizer_config.json are the matching tokenizer files. Run provenance is in training_state.json.
Intended use and license
Lite Base is for experiments with small language models, recurrence, memory, and distillation. Generated text can repeat itself, lose coherence, or give wrong answers. Its training doesn't cover instruction following.
The model is licensed under CC BY 4.0. Upstream datasets retain their own terms.
- Downloads last month
- 942
