We have released BGA! And wow, It provides 256x (and 512x at the end of 1M) yes 256x LESS attention compute at 1M context window. That means you can train a 1M context window at the compute of a ~4K context window.
๐ง We just released Darwin-27B-ZTC, a judgment engine that reaches a verdict without generating anything.
Most LLMs answer by generating, decoding one token at a time. Darwin-27B-ZTC takes a different route.
โ๏ธ How it works ๐น It makes its call in a single forward pass. ๐น Zero generated tokens, and no decoding loop. ๐น That keeps latency and cost far below what a generative model needs.
๐ฏ What it judges ๐น It handles several question types: free-form correctness (noul), multiple choice (choice), and scoring (score). ๐น For each one it hands back a calibrated confidence, not just an answer.
๐ How well calibrated (measured) ๐น KL 0.204, Brier 0.097, so the confidence it reports lines up with what actually happens. ๐น 0.743 accuracy (zero-shot, general split), across 2,000 judgments with zero errors. ๐น By type: noul 0.847, choice 0.723, score 0.675. ๐น None of the benchmark's train split went into it. It is pure zero-shot.
๐ Where it fits ๐น Grading at scale, model routing, safety gating, anywhere you want a fast decision without paying for generation.
๐ It currently sits at #1 on the official typed-decisions leaderboard on Hugging Face (0.743 accuracy, zero-shot).