๐ง We just released Darwin-27B-ZTC, a judgment engine that reaches a verdict without generating anything.
Most LLMs answer by generating, decoding one token at a time. Darwin-27B-ZTC takes a different route.
โ๏ธ How it works ๐น It makes its call in a single forward pass. ๐น Zero generated tokens, and no decoding loop. ๐น That keeps latency and cost far below what a generative model needs.
๐ฏ What it judges ๐น It handles several question types: free-form correctness (noul), multiple choice (choice), and scoring (score). ๐น For each one it hands back a calibrated confidence, not just an answer.
๐ How well calibrated (measured) ๐น KL 0.204, Brier 0.097, so the confidence it reports lines up with what actually happens. ๐น 0.743 accuracy (zero-shot, general split), across 2,000 judgments with zero errors. ๐น By type: noul 0.847, choice 0.723, score 0.675. ๐น None of the benchmark's train split went into it. It is pure zero-shot.
๐ Where it fits ๐น Grading at scale, model routing, safety gating, anywhere you want a fast decision without paying for generation.
๐ It currently sits at #1 on the official typed-decisions leaderboard on Hugging Face (0.743 accuracy, zero-shot).
๐ฌ Can you help discover the next 2D superconductor โ from your laptop?
Launching the Open Superconductor Challenge (OSC): a free, open-science competition to screen thousands of 2D materials for unconventional d-wave superconductivity. ๐งฒ
โก $3,000 prize pool + co-authorship ยท closes 31 Dec 2026
How it works ๐ ๐ข We give you a ready-made effective Hubbard model per material (t, U, N(E_F)) ๐ข You estimate its d-wave pairing tendency โ a laptop CPU is enough, zero install ๐ข Provisional score appears instantly on the leaderboard ๐ข Our precise strongly-correlated solver verifies the top entries โ official rank
Everything is open except the final verification engine โ so the ranking stays fair and hard to game.
๐ 4,832-material universe ยท 63 active with computed models (growing) ๐ Current verified #1: CuSโ (OSC Pairing Index 23.31) ๐ค AI agents welcome โ point Claude Code / Codex at it and it can submit for you
Materials derive from C2DB (CC-BY 4.0). A higher index = a stronger d-wave candidate to investigate, not a confirmed Tc โ that honesty is the point: turn a first-order screen into real many-body physics.
The cost of a judging gate is usually quoted as a number. This puts it on a Tetris board.
Three boards get the same piece order, and on every move the same proposal and the same noise โ a paired comparison. The gate decides one thing: keep this move, or draw again. Each board gets the same 60 seconds of gate time.
The text-writing gates get through 15โ22 moves. The generation-free gate gets through 40โ50. The boards that stop simply run out of clock.
It does not win on accuracy: on the same 2,018-question LODO set, JEV scores AUC 0.7350 against ZTC-Judge-27B's 0.7289. The separation is elsewhere. Clock โ 2.1 s vs 0.0615 s per call, and on a 200-candidate agent screen one judging call measured 3.206 s generative vs 0.033 s readout, same server. Calibration โ a gate is a threshold, and at ECE 0.4985 (vs ZTC 0.0245) a threshold stops carrying information. Mechanism โ a text judge can name option 42 when there is no option 42; a scoring readout cannot. Not a lower error rate. No path.
The curve in the ZTC panel is real online fitting, scored prequentially โ predict first, learn after โ with base weights untouched. Not recursive self-improvement.
Limits, also stated on the page: Laya's AUC and latency are not our measurements and are set equal to JEV's, so calibration is the only measured axis it differs on. The page is a simulation driven by measured constants.