vulnerability-attack-technique-classification-roberta-base

Suggests MITRE ATT&CK (Enterprise) techniques from a free-text vulnerability description. This is a multi-label classifier (roberta-base with a sigmoid head, one output per technique) fine-tuned on CIRCL/vulnerability-attack-techniques, a gold dataset of 1,207 CVEs whose labels come from expert MITRE CTID mappings. Given a CVE description it scores 53 parent techniques (e.g. T1190 Exploit Public-Facing Application, T1505 Server Software Component), for use as a ranked list of candidate techniques for analyst review.

The model is trained with VulnTrain and runs in production on the public Vulnerability-Lookup instance operated by CIRCL, served locally by ML-Gateway: every vulnerability page has an ATT&CK tab with the model's suggestions (example: CVE-2021-44077).

The methodology, evaluation protocol, and the negative result on LLM-assisted label expansion are described in the paper Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion.

DOI: 10.57967/hf/9623

How to use

With VulnTrain, against a live CVE:

vulntrain-infer-attack-classification --cve CVE-2021-44077 \
    --model CIRCL/vulnerability-attack-technique-classification-roberta-base

With plain Transformers:

import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer

model_id = "CIRCL/vulnerability-attack-technique-classification-roberta-base"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
model.eval()

description = (
    "Zoho ManageEngine ServiceDesk Plus before 11306, ServiceDesk Plus MSP "
    "before 10530, and SupportCenter Plus before 11014 are vulnerable to "
    "unauthenticated remote code execution."
)
inputs = tokenizer(description, truncation=True, max_length=512, return_tensors="pt")
with torch.no_grad():
    probs = torch.sigmoid(model(**inputs).logits)[0]

for i in probs.argsort(descending=True)[:5]:
    print(f"{model.config.id2label[int(i)]}  {probs[i]:.4f}")

Technique IDs map to names via the ATT&CK Enterprise matrix; sigmoid >= 0.5 is the prediction threshold used in evaluation, but the model is most useful as a ranker (see the recall@k numbers below).

Intended uses & limitations

Intended: triage assistance — given a vulnerability description, surface a short ranked list of candidate ATT&CK techniques for a human analyst to confirm or reject. This is how it is deployed in Vulnerability-Lookup, where the UI explicitly flags the output as unverified AI-generated suggestions.

Limitations:

  • The label space is the 53 parent techniques with at least 5 training examples; sub-techniques are collapsed to their parent, and techniques outside this vocabulary can never be suggested.
  • The gold set skews toward exploited-in-the-wild CVEs (CTID's corpus and the KEV catalog), so coverage is best for the techniques common there.
  • Trained on 972 CVEs — deliberately label-quality-bound rather than data-bound (see the paper's gold-size scaling curve: every metric still improves monotonically with more curated rows).
  • English descriptions only; input is truncated at 512 tokens.
  • Scores are not calibrated probabilities.
  • Suggestions are not verified mappings; treat them as guidance, never as authoritative CTI.

Training and evaluation data

Labels come from the two public expert sources of CVE→ATT&CK mappings, both following the CTID Mapping ATT&CK to CVE for Impact methodology: the CTID attack_to_cve mappings and the CTID Mappings Explorer KEV mappings (ATT&CK 16.1). Technique IDs revoked since publication are remapped to their successors via the ATT&CK STIX revoked-by relationships. Descriptions are joined from CIRCL/vulnerability-scores (Vulnerability-Lookup). Full details in the dataset card and the VulnTrain documentation.

Splits: 972 train / 106 validation / 118 test examples. Checkpoint selection uses the validation split; the test split is touched once, for the numbers reported here. The paper documents why this matters: selecting the best checkpoint on the test split inflates and destabilizes small-test-split metrics enough to produce spurious conclusions.

Notably, this model was not trained on LLM-generated labels: the paper's controlled experiments show that folding in LLM-labeled CVEs at ≈0.39 agreement with the experts yields no reliable ranking improvement and measurably degrades rare-technique coverage at scale (the comparison checkpoint is published as ...-llm-expanded).

Evaluation

Held-out test split, this checkpoint (seed 42):

Metric Value
Recall@5 0.6440
Recall@3 0.5181
F1 micro 0.3899
F1 macro 0.1910
Precision micro 0.2740
Recall micro 0.6756

Recall@5 = 0.64 means that on average 64% of an unseen CVE's expert-assigned techniques appear in the model's top five suggestions — roughly double the zero-shot embedding-similarity baseline reported in the paper. Across five seeds under the identical protocol, the numbers of record are recall@5 0.673 ± 0.019, recall@3 0.536 ± 0.032, micro-F1 0.410 ± 0.006, macro-F1 0.177 ± 0.014. The complete trainer logs are published in the paper repository.

Training procedure

Binary cross-entropy over 53 sigmoid outputs, with per-label pos_weight balancing (capped at 20) to keep rare techniques trainable. Trained with vulntrain-train-attack-classification (VulnTrain).

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 1e-05
  • train_batch_size: 32
  • eval_batch_size: 32
  • seed: 42
  • optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: linear
  • num_epochs: 40
  • max_length: 512
  • loss: BCEWithLogitsLoss, balanced pos_weight (min 2.447, max 20.0)
  • checkpoint selection: best macro-F1 on the validation split

Training results

Training Loss Epoch Step Validation Loss F1 Micro F1 Macro Precision Micro Recall Micro Recall At 3 Recall At 5
0.9311 1.0 31 0.8413 0.1833 0.0401 0.1258 0.3376 0.1798 0.2611
0.8366 2.0 62 0.7876 0.1452 0.0234 0.1065 0.2278 0.1047 0.2086
0.8116 3.0 93 0.7717 0.2054 0.0518 0.1641 0.2743 0.1938 0.3234
0.7941 4.0 124 0.7576 0.3204 0.0804 0.2382 0.4895 0.3387 0.5009
0.7743 5.0 155 0.7435 0.3044 0.0841 0.2119 0.5401 0.3313 0.4696
0.7645 6.0 186 0.7290 0.3208 0.0906 0.2174 0.6118 0.3969 0.5391
0.7472 7.0 217 0.7163 0.3551 0.1130 0.2571 0.5738 0.4068 0.5741
0.7219 8.0 248 0.7056 0.3224 0.1079 0.2173 0.6245 0.4079 0.5521
0.7195 9.0 279 0.6933 0.3576 0.1495 0.2449 0.6624 0.4252 0.5663
0.6835 10.0 310 0.6845 0.3705 0.1665 0.2579 0.6582 0.4708 0.6090
0.6539 11.0 341 0.6768 0.4063 0.1810 0.2947 0.6540 0.5227 0.6318
0.6484 12.0 372 0.6725 0.3632 0.1734 0.2520 0.6498 0.4449 0.6200
0.6249 13.0 403 0.6664 0.3974 0.1782 0.2862 0.6498 0.5034 0.6396
0.6109 14.0 434 0.6585 0.3801 0.1721 0.2724 0.6287 0.4834 0.6491
0.6004 15.0 465 0.6539 0.3872 0.1678 0.2781 0.6371 0.4752 0.6347
0.5896 16.0 496 0.6502 0.4049 0.1777 0.2996 0.6245 0.4768 0.6397
0.5667 17.0 527 0.6478 0.3866 0.1682 0.2737 0.6582 0.4941 0.6472
0.5661 18.0 558 0.6425 0.4108 0.1910 0.3022 0.6414 0.5128 0.6667
0.5501 19.0 589 0.6394 0.3880 0.1861 0.2758 0.6540 0.4822 0.6561
0.5461 20.0 620 0.6377 0.4097 0.1804 0.3010 0.6414 0.5069 0.6687
0.5351 21.0 651 0.6338 0.4028 0.1715 0.3002 0.6118 0.4987 0.6624
0.5215 22.0 682 0.6351 0.4146 0.1964 0.3054 0.6456 0.4943 0.6875
0.5155 23.0 713 0.6315 0.4056 0.1737 0.3023 0.6160 0.4994 0.6553
0.5063 24.0 744 0.6269 0.4286 0.1879 0.3208 0.6456 0.5195 0.6923
0.5061 25.0 775 0.6264 0.4178 0.1869 0.3069 0.6540 0.5246 0.6656
0.4996 26.0 806 0.6301 0.4073 0.1792 0.3053 0.6118 0.5274 0.6958
0.4950 27.0 837 0.6225 0.4133 0.1771 0.3079 0.6287 0.5376 0.6593
0.4928 28.0 868 0.6228 0.4173 0.1768 0.3166 0.6118 0.5187 0.7048
0.4819 29.0 899 0.6242 0.4263 0.1871 0.3225 0.6287 0.5494 0.6970
0.4766 30.0 930 0.6194 0.4166 0.1862 0.3094 0.6371 0.5226 0.6871
0.4694 31.0 961 0.6213 0.42 0.1914 0.3175 0.6203 0.5399 0.6918
0.4781 32.0 992 0.6209 0.4292 0.1893 0.3281 0.6203 0.5439 0.7060
0.4636 33.0 1023 0.6218 0.4347 0.1962 0.3276 0.6456 0.525 0.6797
0.4641 34.0 1054 0.6216 0.4314 0.1887 0.3261 0.6371 0.5415 0.6797
0.4592 35.0 1085 0.6206 0.4313 0.1916 0.3282 0.6287 0.5466 0.6858
0.4526 36.0 1116 0.6208 0.4357 0.1926 0.3333 0.6287 0.5447 0.6863
0.4643 37.0 1147 0.6191 0.4218 0.1899 0.3196 0.6203 0.5392 0.6863
0.4501 38.0 1178 0.6191 0.4242 0.1885 0.3224 0.6203 0.5368 0.6929
0.4570 39.0 1209 0.6206 0.4350 0.1932 0.3326 0.6287 0.5281 0.6910
0.4436 40.0 1240 0.6199 0.4325 0.1916 0.3296 0.6287 0.5329 0.6882

Framework versions

  • Transformers 5.13.0
  • Pytorch 2.12.1+cu130
  • Datasets 4.8.5
  • Tokenizers 0.22.2

Related artifacts

Artifact Location DOI
Gold dataset (1,207 CVEs, CTID-curated labels) CIRCL/vulnerability-attack-techniques 10.57967/hf/9621
LLM expansion dataset (negative result) CIRCL/vulnerability-attack-techniques-llm-scaling 10.57967/hf/9622
LLM-expanded comparison model CIRCL/vulnerability-attack-technique-classification-roberta-base-llm-expanded 10.57967/hf/9624
Code vulnerability-lookup/VulnTrain —
Paper + trainer logs vulnerability-lookup/cve-attack-mapping-paper —

Citation

@misc{bonhomme2026cveattack,
  title  = {Mapping CVEs to MITRE ATT\&CK Techniques: A Curated Gold-Set
            Classifier and the Limits of LLM-Assisted Label Expansion},
  author = {Bonhomme, C{\'e}dric},
  year   = {2026},
  note   = {Preprint},
}

Acknowledgements

Developed at CIRCL in the context of the AIPITCH project, co-funded by the European Union.

Downloads last month
86
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for CIRCL/vulnerability-attack-technique-classification-roberta-base

Finetuned
(2366)
this model

Dataset used to train CIRCL/vulnerability-attack-technique-classification-roberta-base

Evaluation results