Caelis Neural Base 3

Caelis Neural Base 3 (CNB-3): A Foundational Protein Language Model

Hugging Face

Caelis Neural Base 3 (CNB-3) is a ~985-million-parameter autoregressive causal language model pre-trained for protein sequence modeling. It is designed to serve as a foundational Neural Base for downstream biotechnology, including protein generation, sequence completion, mutation scoring, and specialized fine-tuning.

Important: CNB-3 uses a custom Transformers architecture registered as cnb3. It is GPT-2-compatible at the implementation level, but it is not a generic GPT-2 checkpoint. The repository therefore includes custom Python architecture files and must be loaded with trust_remote_code=True.


⚠️ Out-of-Scope Use & Limitations

  • NOT FOR CONVERSATIONAL CHAT / TEXT NLP: CNB-3 and its CaelisBioTokenizer-2 tokenizer are designed for biological sequence modeling, not conversational text, chatbots, or general human-language NLP.
  • Biological Domain: The vocabulary and training setup are intended for amino-acid and biological sequence modeling.
  • Intended Uses: Protein sequence generation, completion, mutation scoring, protein scaffolding, enzyme/protein design research, and downstream adaptation.
  • Research Use: Generated sequences are model outputs and require appropriate biological validation. Model likelihood is not, by itself, proof of folding, activity, safety, or experimental performance.

Why "Neural Base"?

CNB-3 is designed as a starting checkpoint for downstream protein-language research:

  • General Protein Prior: A causal model trained to represent evolutionary sequence distributions.
  • Causal Autoregressive Backbone: Predicts subsequent tokens conditioned on the preceding sequence.
  • Generation & Completion: Can generate continuations from protein sequence seeds.
  • Downstream Adaptation: Can be fine-tuned or adapted for specialized protein families and tasks.

Technical Architecture

CNB-3 uses a custom architecture registered with the Transformers model type:

model_type = "cnb3"

The implementation is intentionally compatible with GPT-2's causal decoder architecture, while exposing its own configuration and model classes:

CaelisNeuralBaseConfig  -> GPT2Config
CaelisNeuralBase        -> GPT2LMHeadModel

Architecture specifications

Specification Value
Parameters ~985M
Architecture Custom CaelisNeuralBase causal LM
Transformers base implementation GPT2LMHeadModel
Configuration base GPT2Config
Decoder blocks 32
Hidden dimension 1536
Attention heads 24
Context length 1024 tokens
Vocabulary CaelisBioTokenizer-2
Model type cnb3
Weight format safetensors

The custom architecture exists so the checkpoint can identify itself as CNB-3 rather than as a generic GPT-2 model.


Custom Transformers Files

Two Python files in this repository are required for remote-code loading.

configuration_cnb3.py

from transformers import GPT2Config

class CaelisNeuralBaseConfig(GPT2Config):
    model_type = "cnb3"

This class defines the CNB-3 configuration. It inherits from GPT2Config, allowing CNB-3 to reuse GPT-2's configuration validation and serialization behavior while registering its own model type.

modeling_cnb3.py

from transformers import GPT2LMHeadModel
from configuration_cnb3 import CaelisNeuralBaseConfig

class CaelisNeuralBase(GPT2LMHeadModel):
    config_class = CaelisNeuralBaseConfig

This class defines the CNB-3 model implementation. It inherits the causal language-model implementation from GPT2LMHeadModel and associates it with CaelisNeuralBaseConfig.

Important import detail

The import in modeling_cnb3.py intentionally uses:

from configuration_cnb3 import CaelisNeuralBaseConfig

and not:

from .configuration_cnb3 import CaelisNeuralBaseConfig

When Transformers loads remote model code, these files can be loaded as isolated modules rather than as a conventional Python package. The absolute import avoids the resulting No module named configuration_cnb3 error.

Do not delete these files

configuration_cnb3.py and modeling_cnb3.py are required repository components. The model weights can remain intact if they are removed, but Transformers will no longer have the Python classes needed to instantiate the checkpoint through the configured auto_map.


config.json and auto_map

The model configuration must identify the custom architecture and map Transformers' Auto classes to the Python implementations.

The relevant configuration is:

{
  "model_type": "cnb3",
  "architectures": [
    "CaelisNeuralBase"
  ],
  "auto_map": {
    "AutoConfig": "configuration_cnb3.CaelisNeuralBaseConfig",
    "AutoModel": "modeling_cnb3.CaelisNeuralBase",
    "AutoModelForCausalLM": "modeling_cnb3.CaelisNeuralBase"
  }
}

The auto_map values follow this format:

filename_without_py.ClassName

The three entries serve different Auto classes:

  • AutoConfig β†’ CaelisNeuralBaseConfig
  • AutoModel β†’ CaelisNeuralBase
  • AutoModelForCausalLM β†’ CaelisNeuralBase

This mapping is what allows Transformers to locate the custom Python implementation when trust_remote_code=True is enabled.


Correct Loading with Transformers

Recommended loading method

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

repo_id = "InserloftResearch/CaelisNeuralBase-3"

tokenizer = AutoTokenizer.from_pretrained(
    repo_id,
    trust_remote_code=True,
)

model = AutoModelForCausalLM.from_pretrained(
    repo_id,
    trust_remote_code=True,
    torch_dtype=torch.float16,
    device_map="auto",
)

model.eval()

Why these settings matter

trust_remote_code=True

Required because CNB-3 provides its own Python architecture implementation. Without it, Transformers will not execute the repository's custom configuration/model classes.

AutoModelForCausalLM

Use the Auto class rather than directly instantiating GPT2LMHeadModel. Although CNB-3 is implemented on top of the GPT-2 causal-LM implementation, the intended public class is CaelisNeuralBase.

torch_dtype=torch.float16

The checkpoint can be stored/trained in FP32, while FP16 is useful for inference on compatible GPUs because it substantially reduces memory requirements.

device_map="auto"

Allows Accelerate/Transformers to place the model automatically on available hardware.


Verify the Custom Architecture

After loading, verify that the intended classes were instantiated:

print(type(model).__name__)
print(type(model.config).__name__)
print(model.config.model_type)
print(model.config.auto_map)

Expected values:

CaelisNeuralBase
CaelisNeuralBaseConfig
cnb3
{
    "AutoConfig": "configuration_cnb3.CaelisNeuralBaseConfig",
    "AutoModel": "modeling_cnb3.CaelisNeuralBase",
    "AutoModelForCausalLM": "modeling_cnb3.CaelisNeuralBase"
}

This is a useful diagnostic when troubleshooting an incorrectly configured repository.


Reusing CaelisBioTokenizer-2

CaelisBioTokenizer-2 is the tokenizer associated with CNB-3 and can also serve as a starting vocabulary for compatible downstream biological sequence architectures.

from transformers import AutoTokenizer

repo_id = "InserloftResearch/CaelisNeuralBase-3"

tokenizer = AutoTokenizer.from_pretrained(
    repo_id,
    trust_remote_code=True,
)

print("Vocabulary size:", len(tokenizer))

When training a new architecture from scratch, verify tokenizer compatibility with the new model before using the vocabulary as-is.


Quick Start: Protein Generation

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

repo_id = "InserloftResearch/CaelisNeuralBase-3"

tokenizer = AutoTokenizer.from_pretrained(
    repo_id,
    trust_remote_code=True,
)

model = AutoModelForCausalLM.from_pretrained(
    repo_id,
    trust_remote_code=True,
    torch_dtype=torch.float16,
    device_map="auto",
).eval()

seed = "MNSFSTSAFGPVAFSLGLLLVLPAA"

inputs = tokenizer(
    seed,
    return_tensors="pt",
    add_special_tokens=False,
).to(model.device)

with torch.no_grad():
    outputs = model.generate(
        **inputs,
        max_new_tokens=200,
        do_sample=True,
        temperature=0.7,
        top_k=50,
        top_p=0.95,
        repetition_penalty=1.2,
        pad_token_id=tokenizer.pad_token_id,
        eos_token_id=tokenizer.eos_token_id,
    )

generated = tokenizer.decode(
    outputs[0],
    skip_special_tokens=True,
)

generated_clean = "".join(generated.split())

print(generated_clean)

Generation recommendations

  • temperature=0.7 provides a practical balance between determinism and sampling diversity.
  • top_k=50 and top_p=0.95 constrain sampling to plausible token distributions.
  • repetition_penalty=1.2 can reduce repetitive motifs during longer generations.
  • Always remove tokenizer-introduced whitespace before treating the decoded output as a raw amino-acid sequence.
  • CNB-3 has a maximum context length of 1024 tokens. Longer sequences require a chunking/sliding-window strategy rather than passing the entire sequence to the model at once.

Sampling parameters are generation controls, not guarantees of biological activity, folding, or experimental performance.


Checking the Checkpoint Before Use

A simple sanity check can help detect an accidentally uploaded random initialization.

import torch

wte = model.transformer.wte.weight.detach().float()

print(f"Embedding std: {wte.std().item():.5f}")

The exact value depends on the training process and checkpoint. GPT-2-style random initialization can produce values around its initialization scale, while a trained checkpoint may exhibit substantially different statistics.

Do not use embedding standard deviation alone as proof that a model is trained. Combine it with checkpoint metadata, loss/evaluation results, generation sanity checks, and known training artifacts.


Mutation Scoring

For single amino-acid mutation analysis, a wild-type marginal approach can be preferable to comparing full-sequence likelihoods.

The idea is to evaluate the model's next-token distribution at the mutated position and compare the probability assigned to the mutant versus the wild-type residue.

import torch
import torch.nn.functional as F

@torch.no_grad()
def score_mutation(wt_seq, position, wt_aa, mut_aa):
    prefix = wt_seq[:position - 1]

    if len(prefix) < 5:
        return 0.0

    ids = tokenizer(
        prefix,
        return_tensors="pt",
        add_special_tokens=False,
    )["input_ids"]

    if ids.shape[1] > 1023:
        ids = ids[:, -1023:]

    ids = ids.to(model.device)

    logits = model(ids).logits[0, -1]
    logp = F.log_softmax(logits.float(), dim=-1)

    wt_ids = tokenizer(
        wt_aa,
        add_special_tokens=False,
    )["input_ids"]

    mut_ids = tokenizer(
        mut_aa,
        add_special_tokens=False,
    )["input_ids"]

    wt_lp = logp[wt_ids].logsumexp(0).item()
    mut_lp = logp[mut_ids].logsumexp(0).item()

    return mut_lp - wt_lp

Example:

wt = "MNGTEGPNFYVPFSNKTGVVRSPFEAPQYYLAEPWQFSMLAAY"

score = score_mutation(
    wt,
    15,
    "K",
    "R",
)

print(f"Score K15R: {score:.4f}")

Interpretation:

  • Positive score: the mutant is assigned greater model likelihood than the supplied wild-type residue under this scoring procedure.
  • Negative score: the mutant is assigned lower model likelihood.
  • Near zero: the model assigns similar likelihoods.

These scores are model-derived likelihood differences, not direct measurements of biological function or fitness.


Context Length and Long Proteins

CNB-3 has a context length of 1024 tokens.

For sequences that exceed the available context:

  • Do not pass the full sequence directly into the model.
  • Use a sliding-window/chunking strategy.
  • Preserve sufficient left context around the position being scored.
  • For mutation scoring, the relevant prefix must fit within the model's context window.

A CUDA device-side assert can also persist after a previous failed kernel execution. If a CUDA error occurs after an invalid sequence length or tensor operation, restart the Python/Jupyter process before retrying.


Fine-Tuning & Downstream Adaptation

CNB-3 can serve as a starting checkpoint for downstream protein-language research, including:

  • LoRA / parameter-efficient fine-tuning
  • Full-parameter fine-tuning
  • Specialized protein-family adaptation
  • Sequence-generation research
  • Mutation-effect modeling
  • Reward-guided optimization research

Example baseline training configuration:

from transformers import TrainingArguments

training_args = TrainingArguments(
    output_dir="./CNB-3-Specialized",
    learning_rate=3e-5,
    weight_decay=0.01,
    per_device_train_batch_size=4,
    gradient_accumulation_steps=8,
    fp16=True,
    lr_scheduler_type="cosine",
)

Training hyperparameters should be adjusted to the dataset, hardware, sequence lengths, objective, and adaptation method.


Repository Structure

The Hugging Face repository contains the model weights, configuration, tokenizer, and custom architecture implementation:

CaelisNeuralBase-3/
β”œβ”€β”€ .gitattributes
β”œβ”€β”€ README.md
β”œβ”€β”€ config.json
β”œβ”€β”€ configuration_cnb3.py
β”œβ”€β”€ modeling_cnb3.py
β”œβ”€β”€ generation_config.json
β”œβ”€β”€ model.safetensors
β”œβ”€β”€ tokenizer.json
β”œβ”€β”€ tokenizer_config.json
β”œβ”€β”€ LICENSE
└── Caelis_Neural_Base_3_HF_Banner.png

The two custom Python files are part of the executable model definition:

configuration_cnb3.py  β†’ configuration class
modeling_cnb3.py       β†’ model class
config.json            β†’ model_type + auto_map
model.safetensors      β†’ trained parameters

Deleting the custom Python files without also changing the model architecture/configuration will make the checkpoint incompatible with its current auto_map.


Common Errors

The checkpoint you are trying to load has model type cnb3 but Transformers does not recognize this architecture

Likely cause: trust_remote_code=True was not supplied, or the custom architecture files/configuration are missing.

Fix:

AutoModelForCausalLM.from_pretrained(
    repo_id,
    trust_remote_code=True,
)

No module named configuration_cnb3

Likely cause: modeling_cnb3.py uses a relative import.

Correct:

from configuration_cnb3 import CaelisNeuralBaseConfig

Avoid:

from .configuration_cnb3 import CaelisNeuralBaseConfig

Repo id must use alphanumeric chars

Likely cause: A local path was not resolved and Transformers interpreted it as a Hub repository identifier.

Fix: Verify the local path exists or use:

InserloftResearch/CaelisNeuralBase-3

Model loads but generation is nonsensical

Check, in order:

  1. type(model).__name__
  2. type(model.config).__name__
  3. model.config.model_type
  4. model.config.auto_map
  5. tokenizer/model compatibility
  6. checkpoint integrity
  7. whether the expected trained weights were uploaded
  8. generation parameters and context length

CUDA error: device-side assert triggered

Potential causes include an invalid sequence length or an earlier CUDA kernel failure. Keep inputs within the 1024-token context limit and restart the Python/Jupyter process after a device-side assert before retrying.


Pro-Tips

  1. Respect the 1024-token context limit. Use sliding windows for longer proteins.
  2. Clean decoded sequences. Use "".join(sequence.split()) before treating generated text as an amino-acid sequence.
  3. Use repetition controls for long generations. repetition_penalty=1.2 is a reasonable starting point.
  4. Treat temperature as a sampling parameter. 0.7 is a useful starting value, but there is no universal biological optimum.
  5. Validate generated sequences downstream. Sequence likelihood is not equivalent to folding, activity, stability, or experimental success.
  6. Keep the custom architecture files in the Hub repository. They are required by the current remote-code loading configuration.

Resources


Citation

If you use CNB-3 in research, software, or downstream model development, please cite the model repository and identify the checkpoint version used.

Caelis Neural Base 3 (CNB-3)
Inserloft Research / Inserloft

Downloads last month
292
Safetensors
Model size
1.0B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support