Thomas Kim PRO
AI & ML interests
Recent Activity
Organizations
dflash2
Pls nvfp4 ver.
Hi dipankarsarkar,
Thank you for taking the time to give my work a proper look.
I agree, the token cut should be the primary focus. I led with Terminal-Bench scores because it's been the "headline" result since the first 3.6 pi-tune release. My main goal was to target the effort to success imbalance observed in initial evals of the base model.
Both arms did not get the same number of attempts. Evaluations were performed on cloud infrastructure which led to task errors related to sandbox/compute infrastructure as well as runtime/task enviroment failures. For these tasks that errored and I considered to be beyond the agents control I would queue a recovery attempt. 2 was the largest number of recovery attempts so it became the upper bound for number of attempts. Valid failures were not retried and tasks were scored on first completion.
Turn/Token means are over only kept attempts. Yes I agree, I realized TB2.1 could no longer be a clean evaluation set after its use in the SFT evaluations. To compensate I introduced GPQA and SciCode for post-rl evaluation. These sets had the benefit of generally being reliable and faster and allowed for additional points of comparison.
Ideally I would like to rerun evaluations consistently under a controlled environment, but as a full-time student I am constrained in compute budget as well as time and felt compelled to release the model before the semester gets busy. This concern also extends to the GGUF evaluations but results were reported under a strict pass@1 score.
These are my main priorities after release and I'll update the results when they're done.
Thanks again for the feedback!
Qwen3.8-27B-pi: Effort-Ordered Reasoning for Agentic Coding
Qwen3.8-27B-pi: Effort-Ordered Reasoning for Agentic Coding
Trained on curated Pi coding sessions, then refined through reinforcement learning to balance task success with more economical reasoning at low and medium effort, while keeping xhigh focused on correctness.
79.78% on Terminal-Bench 2.1, versus Base model's 75.28%, with approximately 12% fewer output tokens.
Available in BF16, FP8, and GGUF, with serving guides and MTP/DFlash2 companions!
https://huggingface.co/collections/bytkim/qwen38-pi
bytkim/Qwen3.8-27B-pi-FP8
bytkim/Qwen3.8-27B-pi
bytkim/Qwen3.8-27B-pi-GGUF