recICL

recICL is a In Context Learning recommender. We trained this checkpoint ourselves, from scratch, synthetic-data prior, and packaged it for inference. It takes the text embeddings of a user's recent items plus a few other users' sequences as context. From these it ranks a whole catalog in one forward pass, with no training on that catalog. It has 170M parameters and uses 1536-d item embeddings from gte-Qwen2-1.5B-instruct.

Status: v0 research checkpoint. On our 7 Amazon test sets it is clearly better than a mean-of-history embedding baseline. See Evaluation and Limitations.

How it works

  • Items are text embeddings only. No item IDs are learned, so any catalog whose items you can describe in text can be scored. The model expects L2-normalised 1536-d embeddings from Alibaba-NLP/gte-Qwen2-1.5B-instruct (settings below).

  • Input.

    • The user's last ≤ 14 items, oldest first.
    • Up to 8 context sequences (≤ 15 items each) from other users of the same catalog. pool sequences that contain one of the user's last 2 items, ranked by recency-weighted Jaccard overlap.
  • Model. RecPFN has 4 blocks. In each block:

    • causal self-attention runs over the user's sequence, with hard-ALiBi masks: 4 heads see only the 1–4 most recent positions, and 4 heads see the whole prefix;
    • the same layer also encodes the context sequences;
    • the user's positions then cross-attend to all context positions.
  • Output. Each item's score is the dot product of that vector with the item's embedding, ranked over the full catalog.

  • Training is prior-fitted. The model only ever sees synthetic data:

    • Every training batch is a fresh synthetic "environment" of 1,000 items, whose embeddings are drawn from a fixed bank of 800,000 text embeddings.
    • Each environment holds 16 query sequences and 128 context sequences, generated from a random transition structure: random graphs in stage 1, plus latent-factor (concept) structure in stage 2. Popularity bias, recency windows and repeats are also sampled.
    • To predict well, the model has to infer each environment's patterns from the context and the item embeddings. No real user data was used in training.

Intended use and vision

Intended use today: research and prototyping of next-item ranking on text-described catalogs, especially cold-start settings where no per-catalog model can be trained. Evaluate it on your own data before relying on it; see the limitations below.

Vision (roadmap, not current capability). This checkpoint is the first step towards a personalization layer for AI agents:

  • an API that ranks the options an agent is about to show a person, personalized from that person's first few actions, with no per-customer training;
  • it improves in two loops: per user instantly, through the context, and across users through post-training on real outcomes, gated offline before release;
  • each ranking comes with a calibrated confidence.

Roadmap gates, in order:

  1. beat the last-item baseline and the RecPFN paper's reported results;
  2. cold start on unseen domains;
  3. a free agent API for agentic shopping;
  4. the post-training loop.

v0 has not passed gate 1. It does not beat the last-item baseline on HR@10, and its HR@10 is below the paper's reported RecPFN numbers on 5 of the 7 datasets. That comparison is only indicative, because our data build differs from the paper's. v0 also has no confidence calibration (scores are raw dot products) and no post-training loop.

Quick start

# if the repo is private, set HF_TOKEN to a token that can read it
pip install huggingface_hub
python -c "from huggingface_hub import snapshot_download; snapshot_download('devtaji/recICL', local_dir='recICL')"
cd recICL
pip install -r requirements.txt      # torch==2.8.0, numpy==1.26.4, safetensors==0.5.3 (tested with Python 3.11)
python example.py

example.py loads config.json and model.safetensors and builds a toy catalog of 500 random vectors. It then retrieves context for a 4-item history and prints the top 10. We ran it in a clean Python 3.11 CPU container that had only requirements.txt installed (under 10 s):

loaded recICL: 169,918,464 parameters, D=1536, device=cpu
history: [3, 12, 7, 311]
context: 8 sequences, e.g. [[7, 311, 45, 90], [12, 7, 311, 45, 88], [311, 45, 88, 402]]
top-10 items (index, score):
   1. item 311  score   11.281
   2. item   7  score    8.659
   3. item  12  score    5.628
   4. item  45  score    4.301
   5. item 253  score    2.690
   ...

Item 45 comes 4th, right after the user's own recent items. In this random catalog its embedding is unrelated to the user's items; what links it to the user is the context, where 45 follows 311.

With your own catalog, the calls are the same as in example.py. Run this from the downloaded folder:

import numpy as np
import recicl as ir

model = ir.load_model(".")                                                     # config.json + model.safetensors
catalog = ir.ItemCatalog(np.load("item_embeddings.npy"), dim=model.input_dim)  # your [N, 1536] embeddings, row i = item i
other_users = [[5, 9, 3, 12], [12, 7, 311, 45, 88]]                            # other users' item indices, oldest first
retriever = ir.ContextRetriever(other_users, model.config)
history = [3, 12, 7, 311]                                                      # this user's items, oldest first
items, scores = ir.recommend(model, catalog, history, retriever.retrieve(history), k=10)

recommend(..., exclude_history=True) drops items the user already has. The reported evaluation does not do this, because repeat purchases count as hits there. ir.baseline_scores(catalog, history, mode="last" | "balanced") gives the two EmbKNN baselines. The model runs on the first CUDA device if there is one; set CUDA_VISIBLE_DEVICES="" to force the CPU.

Item embeddings

Embed item texts with the same encoder and settings as the training prior:

setting value
model Alibaba-NLP/gte-Qwen2-1.5B-instruct at revision a9af15a6372d7d6b25e9fb07c2ccb9e1fe645644
remote code trust_remote_code=True (required; the model's own bidirectional code and tokenizer)
prompt none (document mode)
max tokens 512
compute dtype bf16, output float32
pooling last token (`<
normalisation L2
item text in our Amazon evaluation Title: {title}; Brand: {brand}; Categories: {c1, c2, ...}

Vectors from any other encoder will not work without retraining.

Evaluation

We evaluated on the TEST split of 7 Amazon Reviews 2018 categories, which we rebuilt from the raw ratings and metadata. The protocol:

  • Split: by user, 70/10/20 (seed 42). At most 10,000 users are sampled into the test split, and users with a single interaction are then dropped. This is why Arts and Pantry evaluate about 3,000 of their 10,000 test users.
  • Task: predict each test user's last item from up to 14 preceding items.
  • Context: 8 sequences retrieved from train-split users.
  • Scoring: batch size 1, ranking over the full item catalog (no sampled candidates, repeats allowed), HR@10 and MRR@10.
  • Hardware: NVIDIA L4, fp32.

The two baselines use the same embeddings:

  • EmbKNN-bal scores items against the mean of the history embeddings. Its numbers are close to the "EmbKNN" column of the RecPFN paper's Table 2: within about 10% on the five datasets where our data matches the paper's (see †).
  • EmbKNN-last scores items against the last history item only. This is the EmbKNN that the paper's text describes.
Dataset Test users HR@10 this model HR@10 EmbKNN-bal HR@10 EmbKNN-last MRR@10 this model MRR@10 EmbKNN-bal MRR@10 EmbKNN-last
Appliances † 145 0.1034 0.0621 0.1310 0.0543 0.0191 0.0648
Arts 2,997 0.2206 0.2015 0.2179 0.1576â–² 0.1394 0.1535
Games 10,000 0.0722 0.0479 0.0736 0.0461â–² 0.0164 0.0432
Movies 10,000 0.1237 0.0695 0.1221 0.0683â–² 0.0248 0.0612
Pantry 3,106 0.2347 0.2315 0.2389 0.1954 0.1866 0.1955
Scientific † 5,619 0.0616▼ 0.0411 0.0749 0.0339▼ 0.0162 0.0376
Software 718 0.1574 0.1309 0.1713 0.0992 0.0390 0.1004
Mean (7) 0.1391 0.1121 0.1471 0.0936 0.0631 0.0937

â–² / â–¼ = significantly better / worse than EmbKNN-last (paired bootstrap over test users, 2,000 resamples, 95% CI excludes 0). Against EmbKNN-bal:

  • HR@10 is significantly better on 5 of 7 datasets (not on Appliances or Pantry);
  • MRR@10 is significantly better on all 7.

† Appliances and Scientific: our data differs from the paper's. Our EmbKNN-bal HR@10 is 42% (Appliances) and 50% (Scientific) below the paper's value for the same baseline, so our data build differs from the authors'. Their filtering is not fully specified. Appliances has only 145 test users.

Per-dataset numbers (HR/MRR @3/5/10, both baselines, paired CIs, McNemar p, and the context ablations under Limitations) are in eval/results.json. All the numbers come from our own data build and are not directly comparable to the paper's tables.

On a CPU the numbers can differ slightly: Software HR@10 is 0.1560 instead of 0.1574, and Appliances MRR@10 is 0.0681 instead of 0.0543. The cause is duplicated items. Some items have identical embeddings (Software: 21,632 items, 21,342 distinct embeddings; Appliances: 30,238 and 30,063), and CPU and GPU order exactly tied scores differently. The reported values lie within the range that tie order allows.

Limitations

  • A trivial baseline beats it on HR@10. EmbKNN-last, which just recommends items similar to the last item, has higher HR@10 on 5 of 7 datasets, and the 7-dataset mean is 0.1471 vs 0.1391. The gap is significant only on Scientific. On MRR@10 the two are level on average (0.0936 vs 0.0937): the model is significantly better on Arts, Games and Movies and worse on Scientific.

  • Context helps, but less than in the paper, and bad context hurts a lot. We ran two context ablations on this checkpoint (TEST, same protocol):

    • Removing the 8 context sequences lowers HR@10 by 1% to 47% per dataset (mean −19% over the 7; significant on 5 of 7, not on Games or Software). It lowers MRR@10 by 8% to 63% (mean −29%; significant on 6 of 7, not on Software). The RecPFN paper reports larger drops for its own model: −38% to −82% HR@10 on the four of these datasets it ablates.
    • Eight randomly chosen context sequences are far worse than none: HR@10 falls by 34% to 86%.

    So context should come from the same catalog and share items with the user's recent history, as the bundled retriever ensures. Significance here means the paired bootstrap 95% CI over test users excludes 0.

  • Evaluation data. For Scientific and Appliances, our data does not match the paper's (see †). The Appliances (145) and Software (718) test sets are small, so their numbers are noisy.

  • One seed, chosen from nine runs. We picked this checkpoint as the best of our 9 training runs on validation HR@10 (mean over the 7 Amazon validation splits: 0.1315, rank 1 of 9; MRR@10 0.0884, rank 2), so its test numbers may be slightly optimistic. Across 3 seeds of this configuration, test HR@10 ranges from 0.1034 to 0.1379 on Appliances and from 0.1490 to 0.1630 on Software. This checkpoint is the lowest of the three on Appliances.

  • Numerical-stability risk. The architecture has no bound on its attention logits: there is no q/k normalisation and the blocks are post-LN. In 1 of our 9 training runs (not this one), the logits reached about 1e9 and the fp32 fused attention backward pass on an H100 returned garbage gradients, ending in NaN weights. This checkpoint trained without incident. If you fine-tune it, skip steps with non-finite gradients or bound the logits.

  • Scope.

    • Only the last 14 history items and 15 items per context sequence are used.
    • It works only with the embedding model above.
    • Scores are uncalibrated dot products.
    • It was evaluated offline, on Amazon product categories only: no online tests, no fairness or safety evaluation.
    • The retriever needs a pool of other users' sequences from the same catalog. Without one the model still runs, but expect roughly the no-context numbers above rather than the table.

Training details

model n_layers=4, n_heads=8, D=1536, icl_module_type='alternating', positional_embedding_scheme='hard-alibi', max_sequence_len=15, dropout 0.2; 169,918,464 parameters, fp32
optimisation AdamW, lr 1e-4 (6 warm-up epochs, then cosine decay to 1e-6), batch 16 query sequences, gradient accumulation 2, full-softmax cross-entropy over each environment's items, training seed 0
context in training 128 synthetic sequences per batch (8 at inference)
epoch 500 training batches + 200 synthetic validation batches
early stopping patience 20 on the ratio HR@10(model) ÷ HR@10(EmbKNN-bal), measured on synthetic validation batches
stage 1 random-graph prior; 120 epochs (no early stop), best epoch 101; 2.57 h on 1× NVIDIA H200
stage 2 random-graph + latent-factor prior, started from the stage-1 checkpoint; stopped early after 47 epochs, best epoch 27; 1.10 h on 1× NVIDIA H100 80GB
total compute about 3.7 GPU-hours

The exact prior parameters and all other settings are in config.json (training).

Synthetic prior provenance. The environments' item embeddings are drawn from a bank of 800,000 texts that we sampled (seed 0) from four BEIR corpora and embedded with the encoder above:

corpus texts
FiQA-2018 (all non-empty documents and queries) 64,248
Quora 200,000
FEVER 200,000
MS MARCO 335,752

The model saw only vectors derived from these texts; neither the texts nor the vectors are in this repository. Each corpus has its own terms. In particular, MS MARCO is released for non-commercial research purposes only. If you plan commercial use, check whether weights trained on embeddings of MS MARCO passages are acceptable for your case. No user data of any kind was used in training.

Files

file content
model.safetensors weights, fp32, 64 tensors (sha256 02c4edc3a4944a9e7a0b6025217b3ea7af16e53c9483e85e9a573074a228a60c)
config.json architecture, inference and embedding settings, training configuration, evaluation protocol, checksums
example.py the tested quick start
eval/results.json the TEST results above, with all metrics, baselines and paired statistics
requirements.txt torch, numpy, safetensors (pinned, tested)
Downloads last month
8
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train devtaji/recICL