- Caelis Neural Base 3 (CNB-3): A Foundational Protein Language Model
- Technical Architecture
- Custom Transformers Files
config.jsonandauto_map- Correct Loading with Transformers
- Recommended loading method
- Why these settings matter
- Generation recommendations
The checkpoint you are trying to load has model type cnb3 but Transformers does not recognize this architectureNo module named configuration_cnb3Repo id must use alphanumeric chars- Model loads but generation is nonsensical
CUDA error: device-side assert triggered
- Recommended loading method
- Verify the Custom Architecture
- Reusing
CaelisBioTokenizer-2 - Quick Start: Protein Generation
- Checking the Checkpoint Before Use
- Mutation Scoring
- Context Length and Long Proteins
- Fine-Tuning & Downstream Adaptation
- Repository Structure
- Common Errors
- Pro-Tips
- Resources
- Citation
Caelis Neural Base 3 (CNB-3): A Foundational Protein Language Model
Caelis Neural Base 3 (CNB-3) is a ~985-million-parameter autoregressive causal language model pre-trained for protein sequence modeling. It is designed to serve as a foundational Neural Base for downstream biotechnology, including protein generation, sequence completion, mutation scoring, and specialized fine-tuning.
Important: CNB-3 uses a custom Transformers architecture registered as
cnb3. It is GPT-2-compatible at the implementation level, but it is not a generic GPT-2 checkpoint. The repository therefore includes custom Python architecture files and must be loaded withtrust_remote_code=True.
β οΈ Out-of-Scope Use & Limitations
- NOT FOR CONVERSATIONAL CHAT / TEXT NLP: CNB-3 and its
CaelisBioTokenizer-2tokenizer are designed for biological sequence modeling, not conversational text, chatbots, or general human-language NLP. - Biological Domain: The vocabulary and training setup are intended for amino-acid and biological sequence modeling.
- Intended Uses: Protein sequence generation, completion, mutation scoring, protein scaffolding, enzyme/protein design research, and downstream adaptation.
- Research Use: Generated sequences are model outputs and require appropriate biological validation. Model likelihood is not, by itself, proof of folding, activity, safety, or experimental performance.
Why "Neural Base"?
CNB-3 is designed as a starting checkpoint for downstream protein-language research:
- General Protein Prior: A causal model trained to represent evolutionary sequence distributions.
- Causal Autoregressive Backbone: Predicts subsequent tokens conditioned on the preceding sequence.
- Generation & Completion: Can generate continuations from protein sequence seeds.
- Downstream Adaptation: Can be fine-tuned or adapted for specialized protein families and tasks.
Technical Architecture
CNB-3 uses a custom architecture registered with the Transformers model type:
model_type = "cnb3"
The implementation is intentionally compatible with GPT-2's causal decoder architecture, while exposing its own configuration and model classes:
CaelisNeuralBaseConfig -> GPT2Config
CaelisNeuralBase -> GPT2LMHeadModel
Architecture specifications
| Specification | Value |
|---|---|
| Parameters | ~985M |
| Architecture | Custom CaelisNeuralBase causal LM |
| Transformers base implementation | GPT2LMHeadModel |
| Configuration base | GPT2Config |
| Decoder blocks | 32 |
| Hidden dimension | 1536 |
| Attention heads | 24 |
| Context length | 1024 tokens |
| Vocabulary | CaelisBioTokenizer-2 |
| Model type | cnb3 |
| Weight format | safetensors |
The custom architecture exists so the checkpoint can identify itself as CNB-3 rather than as a generic GPT-2 model.
Custom Transformers Files
Two Python files in this repository are required for remote-code loading.
configuration_cnb3.py
from transformers import GPT2Config
class CaelisNeuralBaseConfig(GPT2Config):
model_type = "cnb3"
This class defines the CNB-3 configuration. It inherits from GPT2Config, allowing CNB-3 to reuse GPT-2's configuration validation and serialization behavior while registering its own model type.
modeling_cnb3.py
from transformers import GPT2LMHeadModel
from configuration_cnb3 import CaelisNeuralBaseConfig
class CaelisNeuralBase(GPT2LMHeadModel):
config_class = CaelisNeuralBaseConfig
This class defines the CNB-3 model implementation. It inherits the causal language-model implementation from GPT2LMHeadModel and associates it with CaelisNeuralBaseConfig.
Important import detail
The import in modeling_cnb3.py intentionally uses:
from configuration_cnb3 import CaelisNeuralBaseConfig
and not:
from .configuration_cnb3 import CaelisNeuralBaseConfig
When Transformers loads remote model code, these files can be loaded as isolated modules rather than as a conventional Python package. The absolute import avoids the resulting No module named configuration_cnb3 error.
Do not delete these files
configuration_cnb3.py and modeling_cnb3.py are required repository components. The model weights can remain intact if they are removed, but Transformers will no longer have the Python classes needed to instantiate the checkpoint through the configured auto_map.
config.json and auto_map
The model configuration must identify the custom architecture and map Transformers' Auto classes to the Python implementations.
The relevant configuration is:
{
"model_type": "cnb3",
"architectures": [
"CaelisNeuralBase"
],
"auto_map": {
"AutoConfig": "configuration_cnb3.CaelisNeuralBaseConfig",
"AutoModel": "modeling_cnb3.CaelisNeuralBase",
"AutoModelForCausalLM": "modeling_cnb3.CaelisNeuralBase"
}
}
The auto_map values follow this format:
filename_without_py.ClassName
The three entries serve different Auto classes:
AutoConfigβCaelisNeuralBaseConfigAutoModelβCaelisNeuralBaseAutoModelForCausalLMβCaelisNeuralBase
This mapping is what allows Transformers to locate the custom Python implementation when trust_remote_code=True is enabled.
Correct Loading with Transformers
Recommended loading method
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
repo_id = "InserloftResearch/CaelisNeuralBase-3"
tokenizer = AutoTokenizer.from_pretrained(
repo_id,
trust_remote_code=True,
)
model = AutoModelForCausalLM.from_pretrained(
repo_id,
trust_remote_code=True,
torch_dtype=torch.float16,
device_map="auto",
)
model.eval()
Why these settings matter
trust_remote_code=True
Required because CNB-3 provides its own Python architecture implementation. Without it, Transformers will not execute the repository's custom configuration/model classes.
AutoModelForCausalLM
Use the Auto class rather than directly instantiating GPT2LMHeadModel. Although CNB-3 is implemented on top of the GPT-2 causal-LM implementation, the intended public class is CaelisNeuralBase.
torch_dtype=torch.float16
The checkpoint can be stored/trained in FP32, while FP16 is useful for inference on compatible GPUs because it substantially reduces memory requirements.
device_map="auto"
Allows Accelerate/Transformers to place the model automatically on available hardware.
Verify the Custom Architecture
After loading, verify that the intended classes were instantiated:
print(type(model).__name__)
print(type(model.config).__name__)
print(model.config.model_type)
print(model.config.auto_map)
Expected values:
CaelisNeuralBase
CaelisNeuralBaseConfig
cnb3
{
"AutoConfig": "configuration_cnb3.CaelisNeuralBaseConfig",
"AutoModel": "modeling_cnb3.CaelisNeuralBase",
"AutoModelForCausalLM": "modeling_cnb3.CaelisNeuralBase"
}
This is a useful diagnostic when troubleshooting an incorrectly configured repository.
Reusing CaelisBioTokenizer-2
CaelisBioTokenizer-2 is the tokenizer associated with CNB-3 and can also serve as a starting vocabulary for compatible downstream biological sequence architectures.
from transformers import AutoTokenizer
repo_id = "InserloftResearch/CaelisNeuralBase-3"
tokenizer = AutoTokenizer.from_pretrained(
repo_id,
trust_remote_code=True,
)
print("Vocabulary size:", len(tokenizer))
When training a new architecture from scratch, verify tokenizer compatibility with the new model before using the vocabulary as-is.
Quick Start: Protein Generation
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
repo_id = "InserloftResearch/CaelisNeuralBase-3"
tokenizer = AutoTokenizer.from_pretrained(
repo_id,
trust_remote_code=True,
)
model = AutoModelForCausalLM.from_pretrained(
repo_id,
trust_remote_code=True,
torch_dtype=torch.float16,
device_map="auto",
).eval()
seed = "MNSFSTSAFGPVAFSLGLLLVLPAA"
inputs = tokenizer(
seed,
return_tensors="pt",
add_special_tokens=False,
).to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=200,
do_sample=True,
temperature=0.7,
top_k=50,
top_p=0.95,
repetition_penalty=1.2,
pad_token_id=tokenizer.pad_token_id,
eos_token_id=tokenizer.eos_token_id,
)
generated = tokenizer.decode(
outputs[0],
skip_special_tokens=True,
)
generated_clean = "".join(generated.split())
print(generated_clean)
Generation recommendations
temperature=0.7provides a practical balance between determinism and sampling diversity.top_k=50andtop_p=0.95constrain sampling to plausible token distributions.repetition_penalty=1.2can reduce repetitive motifs during longer generations.- Always remove tokenizer-introduced whitespace before treating the decoded output as a raw amino-acid sequence.
- CNB-3 has a maximum context length of 1024 tokens. Longer sequences require a chunking/sliding-window strategy rather than passing the entire sequence to the model at once.
Sampling parameters are generation controls, not guarantees of biological activity, folding, or experimental performance.
Checking the Checkpoint Before Use
A simple sanity check can help detect an accidentally uploaded random initialization.
import torch
wte = model.transformer.wte.weight.detach().float()
print(f"Embedding std: {wte.std().item():.5f}")
The exact value depends on the training process and checkpoint. GPT-2-style random initialization can produce values around its initialization scale, while a trained checkpoint may exhibit substantially different statistics.
Do not use embedding standard deviation alone as proof that a model is trained. Combine it with checkpoint metadata, loss/evaluation results, generation sanity checks, and known training artifacts.
Mutation Scoring
For single amino-acid mutation analysis, a wild-type marginal approach can be preferable to comparing full-sequence likelihoods.
The idea is to evaluate the model's next-token distribution at the mutated position and compare the probability assigned to the mutant versus the wild-type residue.
import torch
import torch.nn.functional as F
@torch.no_grad()
def score_mutation(wt_seq, position, wt_aa, mut_aa):
prefix = wt_seq[:position - 1]
if len(prefix) < 5:
return 0.0
ids = tokenizer(
prefix,
return_tensors="pt",
add_special_tokens=False,
)["input_ids"]
if ids.shape[1] > 1023:
ids = ids[:, -1023:]
ids = ids.to(model.device)
logits = model(ids).logits[0, -1]
logp = F.log_softmax(logits.float(), dim=-1)
wt_ids = tokenizer(
wt_aa,
add_special_tokens=False,
)["input_ids"]
mut_ids = tokenizer(
mut_aa,
add_special_tokens=False,
)["input_ids"]
wt_lp = logp[wt_ids].logsumexp(0).item()
mut_lp = logp[mut_ids].logsumexp(0).item()
return mut_lp - wt_lp
Example:
wt = "MNGTEGPNFYVPFSNKTGVVRSPFEAPQYYLAEPWQFSMLAAY"
score = score_mutation(
wt,
15,
"K",
"R",
)
print(f"Score K15R: {score:.4f}")
Interpretation:
- Positive score: the mutant is assigned greater model likelihood than the supplied wild-type residue under this scoring procedure.
- Negative score: the mutant is assigned lower model likelihood.
- Near zero: the model assigns similar likelihoods.
These scores are model-derived likelihood differences, not direct measurements of biological function or fitness.
Context Length and Long Proteins
CNB-3 has a context length of 1024 tokens.
For sequences that exceed the available context:
- Do not pass the full sequence directly into the model.
- Use a sliding-window/chunking strategy.
- Preserve sufficient left context around the position being scored.
- For mutation scoring, the relevant prefix must fit within the model's context window.
A CUDA device-side assert can also persist after a previous failed kernel execution. If a CUDA error occurs after an invalid sequence length or tensor operation, restart the Python/Jupyter process before retrying.
Fine-Tuning & Downstream Adaptation
CNB-3 can serve as a starting checkpoint for downstream protein-language research, including:
- LoRA / parameter-efficient fine-tuning
- Full-parameter fine-tuning
- Specialized protein-family adaptation
- Sequence-generation research
- Mutation-effect modeling
- Reward-guided optimization research
Example baseline training configuration:
from transformers import TrainingArguments
training_args = TrainingArguments(
output_dir="./CNB-3-Specialized",
learning_rate=3e-5,
weight_decay=0.01,
per_device_train_batch_size=4,
gradient_accumulation_steps=8,
fp16=True,
lr_scheduler_type="cosine",
)
Training hyperparameters should be adjusted to the dataset, hardware, sequence lengths, objective, and adaptation method.
Repository Structure
The Hugging Face repository contains the model weights, configuration, tokenizer, and custom architecture implementation:
CaelisNeuralBase-3/
βββ .gitattributes
βββ README.md
βββ config.json
βββ configuration_cnb3.py
βββ modeling_cnb3.py
βββ generation_config.json
βββ model.safetensors
βββ tokenizer.json
βββ tokenizer_config.json
βββ LICENSE
βββ Caelis_Neural_Base_3_HF_Banner.png
The two custom Python files are part of the executable model definition:
configuration_cnb3.py β configuration class
modeling_cnb3.py β model class
config.json β model_type + auto_map
model.safetensors β trained parameters
Deleting the custom Python files without also changing the model architecture/configuration will make the checkpoint incompatible with its current auto_map.
Common Errors
The checkpoint you are trying to load has model type cnb3 but Transformers does not recognize this architecture
Likely cause: trust_remote_code=True was not supplied, or the custom architecture files/configuration are missing.
Fix:
AutoModelForCausalLM.from_pretrained(
repo_id,
trust_remote_code=True,
)
No module named configuration_cnb3
Likely cause: modeling_cnb3.py uses a relative import.
Correct:
from configuration_cnb3 import CaelisNeuralBaseConfig
Avoid:
from .configuration_cnb3 import CaelisNeuralBaseConfig
Repo id must use alphanumeric chars
Likely cause: A local path was not resolved and Transformers interpreted it as a Hub repository identifier.
Fix: Verify the local path exists or use:
InserloftResearch/CaelisNeuralBase-3
Model loads but generation is nonsensical
Check, in order:
type(model).__name__type(model.config).__name__model.config.model_typemodel.config.auto_map- tokenizer/model compatibility
- checkpoint integrity
- whether the expected trained weights were uploaded
- generation parameters and context length
CUDA error: device-side assert triggered
Potential causes include an invalid sequence length or an earlier CUDA kernel failure. Keep inputs within the 1024-token context limit and restart the Python/Jupyter process after a device-side assert before retrying.
Pro-Tips
- Respect the 1024-token context limit. Use sliding windows for longer proteins.
- Clean decoded sequences. Use
"".join(sequence.split())before treating generated text as an amino-acid sequence. - Use repetition controls for long generations.
repetition_penalty=1.2is a reasonable starting point. - Treat temperature as a sampling parameter.
0.7is a useful starting value, but there is no universal biological optimum. - Validate generated sequences downstream. Sequence likelihood is not equivalent to folding, activity, stability, or experimental success.
- Keep the custom architecture files in the Hub repository. They are required by the current remote-code loading configuration.
Resources
- Official Model Hub β model weights, configuration, tokenizer, architecture files, and discussions.
- Inserloft Research β Inserloft's research and developer resources.
Citation
If you use CNB-3 in research, software, or downstream model development, please cite the model repository and identify the checkpoint version used.
Caelis Neural Base 3 (CNB-3)
Inserloft Research / Inserloft
- Downloads last month
- 292