OdiaGPT-30M (Pretrained from Scratch)

OdiaGPT-30M is an autoregressive decoder-only Transformer built and trained completely from scratch for the Odia language.

It does not use pretrained GPT-2, Llama, or Indic model weights, nor off-the-shelf tokenizers. All 31.5 million parameters began from random Gaussian initialization, and its 8,192-token BPE tokenizer was trained exclusively on a cleaned subset of the AI4Bharat Sangraha Odia corpus.

Model Summary

  • Organization / Creator: CodeHima
  • Architecture: Decoder-only Transformer (Llama-compatible)
  • Parameters: 31,465,984 (~31.5M)
  • Layers: 8
  • Hidden dimension: 512 (8 attention heads, head dimension: 64)
  • Feed-Forward: SwiGLU ($d_{ff} = 1536$)
  • Positional Encoding: Rotary Position Embeddings (RoPE, base 10000.0)
  • Normalization: RMSNorm ($\epsilon = 10^{-6}$)
  • Context Length: 512
  • Vocabulary: 8,192 (SentencePiece BPE with byte fallback)
  • Precision: FP16 (Safetensors format)

Training Dynamics

  • Dataset: AI4Bharat Sangraha (verified/ori) — 28.5M tokens / 110M cleaned characters
  • Tokens Seen: ~29.5 Million tokens (1,800 steps)
  • Initial Loss: 9.07 -> Final Validation Loss: 4.664 (Perplexity: 106.09)
  • Hardware: Single NVIDIA Tesla T4 on Google Colab (~35,200 tokens/sec)

Usage with Hugging Face Transformers

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "CodeHima/OdiaGPT-30M"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.float16).cuda()

prompt = "ଓଡ଼ିଶାର ଲୋକମାନେ"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")

outputs = model.generate(
    **inputs,
    max_new_tokens=64,
    temperature=0.8,
    top_k=50,
    top_p=0.9,
    repetition_penalty=1.15,
    do_sample=True
)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Sample Output

Prompt: ଆଜି ସକାଳେ
Completion: ଆଜି ସକାଳେ ସକାଳେ ଦୁଇ ଝିଅ ସହ ଆସି ଶୋଇପଡ଼ିଲା । ସେହି ସମୟରେ ବାପା ଘର ଲୋକ କୁ ଖବର ଦେଇ କହିଥିଲେ ଯେ, "ତେବେ ମାଆ, ତମେ ଆଉ କିଛି କହିଦେବୁ । ମୁଁ ଏ କଥା କିଛି ଶୁଣିନାହିଁ ।"

Limitations

  • Base foundation model, not instruction-tuned.
  • Developed as part of the OdiaGPT from-scratch educational curriculum.
Downloads last month
-
Safetensors
Model size
31.5M params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using CodeHima/OdiaGPT-30M 1