Sol Lite Base

Sol Lite Base

Lite Base reuses four transformer blocks on a second pass. That gives it fourteen block applications while storing ten blocks. We trained the 14,995,843-parameter model from scratch on 8,153,333,760 tokens.

The released weights are for text completion. There was no instruction tuning or assistant safety tuning. The repository includes a PyTorch loader and generation helper.

Architecture

Setting Value
Parameters 14,995,843
Training tokens 8,153,333,760
Context 2,048 tokens
Tokenizer 4,096-entry digit-aware byte-level BPE
Hidden width 256
Stored blocks / effective applications 10 / 14
Recurrent layout 1 prelude, 4 middle blocks used twice, 5 coda blocks
Attention 8 query heads, 2 KV heads, head dimension 32
Q/K normalization Per-head RMSNorm with RoPE
Recurrence Learned pass embeddings and channel-wise refresh gates
FFN Gated SiLU, width 1,465
Local memory 2,048-entry bigram/trigram EngramLite
Token embeddings Tied to the output head
Weights FP32 safetensors

Grouped-query attention and loop conditioning support the repeated pass. XSA subtracts attention values, and EngramLite supplies local memory. Learned pass embeddings and channel-wise refresh gates distinguish the two uses of the middle blocks.

Generate text

Install the dependencies, then download the repository and import its standalone implementation:

pip install "torch>=2.5" "transformers>=5" safetensors huggingface_hub
from huggingface_hub import snapshot_download
import sys

model_dir = snapshot_download("solintellegence/Sol-Lite-Base")
sys.path.insert(0, model_dir)

from modeling_sol_lite import load_model, generate

model, tokenizer = load_model(model_dir, device="cpu")

text = generate(
    model,
    tokenizer,
    "The future of efficient language models is",
    max_new_tokens=64,
)

print(text)

Benchmark results

Benchmark Examples Normalized accuracy
HellaSwag 10,042 27.72%
ARC-Easy 2,376 34.72%
ARC-Challenge 1,172 22.87%
PIQA 1,838 57.73%
ArithMark-3 1,000 34.20%

The local Intelligence Index score is 8.799. We ran the complete zero-shot language-model splits with lm-eval 0.4.12. ArithMark-3 used independent tokenization and normalized accuracy. Raw evaluation outputs are saved in evals/.

Training sources

The English curriculum combined FineWeb-Edu, DCLM, UltraFineWeb levels 1-3, FineWeb-HQ, FineMath, and FinePhrase. We prepared the stream beforehand and kept the tokenizer fixed. Public benchmark examples weren't used to choose the checkpoint.

Files

The checkpoint is model.safetensors. Use modeling_sol_lite.py for SolForCausalLM, loading, and generation; config.json supplies the architecture settings. tokenizer.json and tokenizer_config.json are the matching tokenizer files. Run provenance is in training_state.json.

Intended use and license

Lite Base is for experiments with small language models, recurrence, memory, and distillation. Generated text can repeat itself, lose coherence, or give wrong answers. Its training doesn't cover instruction following.

The model is licensed under CC BY 4.0. Upstream datasets retain their own terms.

Downloads last month
942
Safetensors
Model size
15M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Space using solintellegence/Sol-Lite-Base 1

Collection including solintellegence/Sol-Lite-Base