·
AI & ML interests
DEJAN - The AI Influence agency.
[DEJAN](https://dejan.ai/)
Recent Activity
repliedto their post 1 day ago Ox Alpha is GLM
https://dejan.ai/blog/ox-alpha/
A parameter-free k-nearest-neighbour classifier over Normalized Compression Distance (Lee et al., MobiSys ’24, Eq. 1, built on Jiang et al.'s gzip-based text classifier). NCD compares two texts by how well they compress together. C(s) is the gzip-compressed length of s. Text sharing an author's patterns compresses better together than text from a different author, so the method needs no model weights and no embeddings.
The reference corpus covers 60 prompts (essays, code, emails, dialogue, poetry) answered by five known models: GPT-5.5, Claude Opus 5, Gemini 3.7 Flash, Gemini 3.1 Pro Preview, and GLM-5.3, for 293 reference texts. ox-alpha answered the first 13 of those prompts, plus one additional novel prompt never given to the reference models beforehand, for 14 queries in total. Each query was classified against the reference corpus independently, with a k-nearest-neighbour vote (k=5):
Model ox-alpha samples matched
GLM-5.3 7 / 14
Claude Opus 5 3 / 14
Gemini 3.7 Flash 2 / 14
GPT-5.5 1 / 14
Gemini 3.1 Pro Preview 1 / 14
GLM-5.3 wins at every k tested: 7/14 at k=3, 7/14 at k=5, 6/14 at k=7, 7/14 at k=9. Claude Opus 5 is the consistent second place. repliedto their post 3 days ago Ox Alpha is GLM
https://dejan.ai/blog/ox-alpha/
A parameter-free k-nearest-neighbour classifier over Normalized Compression Distance (Lee et al., MobiSys ’24, Eq. 1, built on Jiang et al.'s gzip-based text classifier). NCD compares two texts by how well they compress together. C(s) is the gzip-compressed length of s. Text sharing an author's patterns compresses better together than text from a different author, so the method needs no model weights and no embeddings.
The reference corpus covers 60 prompts (essays, code, emails, dialogue, poetry) answered by five known models: GPT-5.5, Claude Opus 5, Gemini 3.7 Flash, Gemini 3.1 Pro Preview, and GLM-5.3, for 293 reference texts. ox-alpha answered the first 13 of those prompts, plus one additional novel prompt never given to the reference models beforehand, for 14 queries in total. Each query was classified against the reference corpus independently, with a k-nearest-neighbour vote (k=5):
Model ox-alpha samples matched
GLM-5.3 7 / 14
Claude Opus 5 3 / 14
Gemini 3.7 Flash 2 / 14
GPT-5.5 1 / 14
Gemini 3.1 Pro Preview 1 / 14
GLM-5.3 wins at every k tested: 7/14 at k=3, 7/14 at k=5, 6/14 at k=7, 7/14 at k=9. Claude Opus 5 is the consistent second place. repliedto their post 3 days ago Ox Alpha is GLM
https://dejan.ai/blog/ox-alpha/
A parameter-free k-nearest-neighbour classifier over Normalized Compression Distance (Lee et al., MobiSys ’24, Eq. 1, built on Jiang et al.'s gzip-based text classifier). NCD compares two texts by how well they compress together. C(s) is the gzip-compressed length of s. Text sharing an author's patterns compresses better together than text from a different author, so the method needs no model weights and no embeddings.
The reference corpus covers 60 prompts (essays, code, emails, dialogue, poetry) answered by five known models: GPT-5.5, Claude Opus 5, Gemini 3.7 Flash, Gemini 3.1 Pro Preview, and GLM-5.3, for 293 reference texts. ox-alpha answered the first 13 of those prompts, plus one additional novel prompt never given to the reference models beforehand, for 14 queries in total. Each query was classified against the reference corpus independently, with a k-nearest-neighbour vote (k=5):
Model ox-alpha samples matched
GLM-5.3 7 / 14
Claude Opus 5 3 / 14
Gemini 3.7 Flash 2 / 14
GPT-5.5 1 / 14
Gemini 3.1 Pro Preview 1 / 14
GLM-5.3 wins at every k tested: 7/14 at k=3, 7/14 at k=5, 6/14 at k=7, 7/14 at k=9. Claude Opus 5 is the consistent second place. View all activity Organizations
Updated • 9.25k
• 12
Token Classification
• 0.3B • Updated • 13
• 1
0.3B • Updated • 5
• 1
dejanseo/reverse-prompter
Text Generation
• 0.3B • Updated • 92
• 3
dejanseo/ecommerce-query-volume-classifier
Text Classification
• 0.2B • Updated • 18
• 1
Token Classification
• 0.3B • Updated • 46
• 9
dejanseo/ecommerce-taxonomy-classifier
Text Classification
• Updated • 4
Text Classification
• 11.7M • Updated • 45
• 3
Token Classification
• Updated • 325
• 9
dejanseo/confidence-distribution-threshold-detector
0.3B • Updated • 6
Token Classification
• Updated • 39
Text Generation
• 1B • Updated • 23
• 1
dejanseo/query-reformulation
60.5M • Updated • 5
dejanseo/query-reformulator-large
0.2B • Updated • 3
dejanseo/gemma-embed-large
Updated
Feature Extraction
• Updated • 2
dejanseo/gemma-embed-stage-3
dejanseo/universal-query-classifier-base
0.2B • Updated • 7
dejanseo/gemma-embed-stage-2
dejanseo/gemma-embed-stage-1
dejanseo/universal-query-classifier-large
0.4B • Updated • 5
dejanseo/universal-query-classifier-xsmall
70.6M • Updated • 5
dejanseo/universal-query-classifier-small
0.1B • Updated • 6
0.4B • Updated • 8
dejanseo/vec2vec-gemini-mxbai