capturezone's picture

capturezone

capturezone

AI & ML interests

None yet

Recent Activity

liked a model about 10 hours ago
FINAL-Bench/Darwin-27B-ZTC
liked a dataset about 10 hours ago
LocalLLaMA/typed-decisions
reacted to SeaWolf-AI's post with ๐Ÿ‘ about 10 hours ago
๐Ÿง  We just released Darwin-27B-ZTC, a judgment engine that reaches a verdict without generating anything. Most LLMs answer by generating, decoding one token at a time. Darwin-27B-ZTC takes a different route. โš™๏ธ How it works ๐Ÿ”น It makes its call in a single forward pass. ๐Ÿ”น Zero generated tokens, and no decoding loop. ๐Ÿ”น That keeps latency and cost far below what a generative model needs. ๐ŸŽฏ What it judges ๐Ÿ”น It handles several question types: free-form correctness (noul), multiple choice (choice), and scoring (score). ๐Ÿ”น For each one it hands back a calibrated confidence, not just an answer. ๐Ÿ“Š How well calibrated (measured) ๐Ÿ”น KL 0.204, Brier 0.097, so the confidence it reports lines up with what actually happens. ๐Ÿ”น 0.743 accuracy (zero-shot, general split), across 2,000 judgments with zero errors. ๐Ÿ”น By type: noul 0.847, choice 0.723, score 0.675. ๐Ÿ”น None of the benchmark's train split went into it. It is pure zero-shot. ๐Ÿš€ Where it fits ๐Ÿ”น Grading at scale, model routing, safety gating, anywhere you want a fast decision without paying for generation. ๐Ÿ† It currently sits at #1 on the official typed-decisions leaderboard on Hugging Face (0.743 accuracy, zero-shot). ๐Ÿ”— Links Model: https://huggingface.co/FINAL-Bench/Darwin-27B-ZTC Leaderboard: https://huggingface.co/datasets/LocalLLaMA/typed-decisions Curious to hear what you make of the single-pass, no-generation approach. ๐Ÿ™Œ
View all activity

Organizations

None yet