Instructions to use thealper2/graphcodebert-code-clone-detection with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use thealper2/graphcodebert-code-clone-detection with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="thealper2/graphcodebert-code-clone-detection")# Load model directly from transformers import AutoTokenizer, GraphCodeBERTForCloneDetection tokenizer = AutoTokenizer.from_pretrained("thealper2/graphcodebert-code-clone-detection") model = GraphCodeBERTForCloneDetection.from_pretrained("thealper2/graphcodebert-code-clone-detection", device_map="auto") - Notebooks
- Google Colab
- Kaggle
# Load model directly
from transformers import AutoTokenizer, GraphCodeBERTForCloneDetection
tokenizer = AutoTokenizer.from_pretrained("thealper2/graphcodebert-code-clone-detection")
model = GraphCodeBERTForCloneDetection.from_pretrained("thealper2/graphcodebert-code-clone-detection", device_map="auto")graphcodebert-code-clone-detection
Binary code-clone detection. Full fine-tune of
microsoft/graphcodebert-base on
PoolC/1-fold-clone-detection-600k-5fold,
using GraphCodeBERT's data-flow-aware pairwise architecture.
Output labels: 0 = not clone, 1 = clone.
Architecture
Not a generic sequence-pair classifier. The two snippets are encoded
separately by one shared GraphCodeBERT encoder, each with its own
graph-guided masked attention, and the two <s> vectors are concatenated for
classification:
Linear(2 x 768 -> 768) -> tanh -> Linear(768 -> 2)
Per-snippet input layout (length 640):
| segment | length | content | position_idx |
|---|---|---|---|
| code tokens | 512 | <s> + BPE code tokens + </s> |
2 .. n+1 |
| data-flow nodes | 128 | one slot per DFG variable node (<unk> id) |
0 |
| padding | remainder | <pad> |
1 |
A data-flow node's input embedding is the average of the embeddings of the
code tokens it was identified from. Graph-guided attention allows: code to
code; <s>/</s> to everything; node to the code tokens it comes from (and
back); node to adjacent nodes.
Preprocessing
The dataset contains Python snippets, so data flow is extracted with the
tree-sitter-python grammar via a port of GraphCodeBERT's DFG_python
extractor (comment/docstring stripping -> AST -> variable states ->
comesFrom / computedFrom edges).
code_length |
512 |
data_flow_length |
128 |
| total sequence length | 640 |
| distinct snippets featurised | 44,950 |
| mean data-flow nodes / snippet | 44.22 |
| snippets with empty data flow | 263 |
| total data-flow edges | 2,481,388 |
| extraction status counts | {"ok": 44930, "comment_strip_failed": 13, "dfg_failed": 7} |
No example was dropped: a snippet whose data flow could not be extracted is kept with an empty graph and counted above.
Data splits
The repository provides one of 5 predefined folds as train + val; those groups are disjoint and are kept as-is. val is partitioned further into validation/test along problem-group boundaries.
similar equals (code1_group == code2_group) for every row, so the group
columns are a perfect label proxy and are never used as features.
| split | source | pairs | positives | negatives | groups |
|---|---|---|---|---|---|
| train | train fold |
50,000 | 25,000 | 25,000 | 240 |
| validation | half of val by group |
20,000 | 10,000 | 10,000 | 29 |
| test | other half of val by group |
20,000 | 10,000 | 10,000 | 30 |
Train/validation/test share no problem group and no code snippet; this is asserted at runtime before training starts. 337,398 pairs of the held-out fold were dropped because their two snippets fell on opposite sides of the validation/test group boundary.
Class weighting: Measured majority-class share 0.5000 is within the 0.6 threshold, so weighted cross entropy is NOT used.
Training
| optimizer | adamw_torch |
| learning rate | 2e-05 |
| scheduler | linear with 0.1 warmup ratio (938 steps) |
| epochs | 3.0 |
| per-device batch size | 16 |
| gradient accumulation | 1 |
| effective batch size | 16 |
| weight decay | 0.01 |
| gradient clipping | 1.0 |
| mixed precision | fp16 |
| gradient checkpointing | False |
| seed | 42 |
| trainable parameters | 125,236,994 |
| training time | 1.923 h |
| GPU | NVIDIA GeForce RTX 5060 Ti (15.9 GB) |
| torch / transformers | 2.11.0+cu128 / 5.17.0 |
Checkpoint selection: best validation F1 (load_best_model_at_end=True,
metric_for_best_model="f1"). Best validation F1 = 0.8672.
The test split was scored once, after selection.
Results
| split | accuracy | precision | recall | F1 | TP | TN | FP | FN |
|---|---|---|---|---|---|---|---|---|
| validation | 0.8557 | 0.8032 | 0.9422 | 0.8672 | 9,422 | 7,692 | 2,308 | 578 |
| test | 0.8747 | 0.8410 | 0.9240 | 0.8805 | 9,240 | 8,253 | 1,747 | 760 |
Test confusion matrix ([[TN, FP], [FN, TP]]): [[8253, 1747], [760, 9240]]
Usage
This checkpoint uses a custom pairwise head and a graph-guided attention
mask, so AutoModelForSequenceClassification will not reproduce these
results. Use the repository's own model class and preprocessing:
import torch
from transformers import AutoTokenizer
from modeling import load_model # from this project
from preprocess import build_snippet_features, CloneCollator
from config import Config
cfg = Config()
tokenizer = AutoTokenizer.from_pretrained("thealper2/graphcodebert-code-clone-detection")
model = load_model("thealper2/graphcodebert-code-clone-detection").eval()
features = build_snippet_features(cfg, [code_a, code_b], tokenizer, num_proc=1)
collator = CloneCollator(features)
batch = collator([(0, 1, 0)]) # (snippet_a, snippet_b, dummy label)
with torch.no_grad():
logits = model(**{k: v for k, v in batch.items() if k != "labels"}).logits
label = int(logits.argmax(-1)) # 0 = not clone, 1 = clone
Limitations
- Trained on competitive-programming Python solutions grouped by problem; "clone" therefore means solves the same problem, which is closer to semantic (Type-4) similarity than to syntactic copy-paste detection.
- Data flow is extracted with the Python grammar only. Other languages need the
matching
tree_sitter_<lang>grammar andDFG_<lang>function. - Snippets longer than 512 BPE tokens are truncated; 885 of 44,950 distinct snippets hit that limit.
- Both directions of a pair are not explicitly symmetrised; the head sees
concat(<s>_1, <s>_2)in the given order.
- Downloads last month
- -
Model tree for thealper2/graphcodebert-code-clone-detection
Base model
microsoft/graphcodebert-baseDataset used to train thealper2/graphcodebert-code-clone-detection
Evaluation results
- f1 on PoolC/1-fold-clone-detection-600k-5foldself-reported0.880
- accuracy on PoolC/1-fold-clone-detection-600k-5foldself-reported0.875
- precision on PoolC/1-fold-clone-detection-600k-5foldself-reported0.841
- recall on PoolC/1-fold-clone-detection-600k-5foldself-reported0.924
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="thealper2/graphcodebert-code-clone-detection")