Model Requests & Suggestions

#1
by IsValorum - opened

Suggest a model for a future handcrafted MiniPlus or NanoPlus release.

Please include:

  • Upstream model link
  • Intended use or strengths
  • Why it would be valuable at a MiniPlus or NanoPlus footprint

I will review requests based on demand, architecture, local-inference practicality, and whether the model would benefit from a custom tensor-by-tensor quantization.

IsValorum pinned discussion

CyberTiel

  1. Link: peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-GGUF-MTP (https://huggingface.co/peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-GGUF-MTP)
  2. Just like Tiel , agentic coding tasks, but stronger, less thinking(in my opinion) and uncensored
  3. Miniplus like what you did on Tiel

The MiniPlus V2.1 and NanoPlus quantizations are already underway; they will be available within an hour at most—follow me to get notified.

Wow
Thanks for your hard works and very fast reply, really impressive, this is the fastest model request i've seen on hf (or maybe you had already worked on this model from before :))

I will definitely check and test for if the Apex Miniplus CyberTiel work good too

Request for next model.
Reason: I've personally used k2 and it's the closest thing to match or exceed Qwen intelligence:
IFM/K2-Horizon-MoVA-36B-A4B

I personally tried the K2-Horizon, the smaller versions and that MoVA, and they didn't convince me because of the cost of the KV Cache.

I could barely get it to 8K context in Q4_0, not enough to properly test it, is it really any good?

For me, the fact that it had such an absurdly expensive KV cache was enough to completely rule it out. Did you guys test it in real-world use?

You were using Q4_0 KV?
How much VRAM on the rig?
I assume your models were tested on 24gb ram?

I tried it on an RTX 3090 and I simply didn't find that series of models appealing; the KV cache was too expensive.

Gotcha, no worries. Thanks for the responses, love your work.

Gotcha, no worries. Thanks for the responses, love your work.

If you find another model that you consider good, don't hesitate to let me know so I can quote it immediately.

are you interested in some thing called "Ternary-Bonsai-2-27B-gguf" (*edit:oh i tried result really bad, rather choose ornith 1.5)

Yeah, despite Prism-ML claiming over 90% retention, in practice they perform terribly. Ornith 1.5 is just way better.

among all 35b model, which provide you the best performance in agentic coding tasks ?
atm i'm using Tiel/CyberTiel

Tiel and CyberTiel are definitely great for that, but Occamy fits right in there as well.

While Tiel is very focused on direct syntax and fast tool execution, Occamy handles agentic coding equally well while offering a bit more balance on broader architectural context, complex debugging, and multi-file reasoning.

If you want to give it a spin to compare against your Tiel setup, here is the MiniPlus V2.1 release:
https://huggingface.co/IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF

Between Tiel/CyberTiel and Occamy, you're pretty much looking at the best options in the 35B tier for agentic workflows right now.

Yeah i see
But maybe i still go with tiel until qwen 4 35b a3b(just guess lol)
The author optimize this model kv cache really good so it take less vram/ram for large context amount
And it have mtp too
After apex miniplus exist so i change to mini plus when need quality, and comeback to iq3 xxs when need 262k context

Totally fair point. That DeltaNet KV cache efficiency plus MTP makes it a great daily driver for long context. Enjoy Tiel!

If you're dropping to flat IQ3_XXS just to fit the full 262K context, check out the NanoPlus release of Tiel:
https://huggingface.co/IsValorum/Tiel-Coder-35B-A3B-APEX-I-NanoPlus-GGUF

That's actually why I built NanoPlus: to give you that lower memory footprint for massive context while still maintaining Q4_K_M/L-class practical quality instead of the flat degradation of standard IQ3_XXS, keeping routers in F32 and output in Q6_K.

Yeah, but i think miniplus with 131k still work good, even can upto 180-190k if i reduce some prefill boost parameter and disable mtp

Thoughts on the new European Model? Aleph-Alpha/Kolibri-1 https://huggingface.co/Aleph-Alpha/Kolibri-1

78B with ~4B active, gets decent benchmark

Hi @userr99 ,

I looked into it as soon as it dropped today, but unfortunately it is physically impossible to make GGUFs for it right now.

The main blocker is that Aleph Alpha built a completely custom architecture from scratch (Kolibri1ForCausalLM). llama.cpp has zero support for it: the GGUF converter will simply crash because it cannot map the tensors, and the C++ engine has no kernels to execute its specific 4:1 sliding-window attention (4 SWA layers for every 1 GQA layer) or route across 384 micro-experts per layer (that is 19,200 expert FFN blocks across its 50 layers). It does not even run on standard vLLM without installing Aleph Alpha's own proprietary plugin package (aleph-alpha-inference).

On top of the architecture gap, the base weights are 145.5 GB in BF16. Even if someone in the community spends weeks writing a custom C++ PR to get kolibri1 into llama.cpp, a 78B MoE would still land around 32 to 35 GB in a MiniPlus quant, meaning it would not fit into a single 24 GB VRAM card anyway without spilling into RAM.

So for now, until upstream llama.cpp actually writes and merges full native support for the architecture, it is completely out of reach for GGUF.

i see something called "Xing4.0-29B-A4B" ,does it have any advantage when comparing to 35b models (coding, speed, weight,...)

Benchmark / Metric Xing4.0-29B-A4B Tiel-Coder-35B-A3B (Ornith-1.5) Qwen3.6-35B-A3B (Base Reference) Direct Comparison
Terminal-Bench 2.1 (Terminus-2) 57.50 67.80 52.50 Tiel +10.3 points ahead in interactive bash terminal execution
Terminal-Bench 2.1 (Claude Code) (Not reported) 68.50 49.20 Tiel specialized for CLI coding agent harnesses
SWE-bench Verified (GitHub issue solves) 75.00% 79.00% 73.40% Tiel +4.0% higher resolution rate on real production codebases
SWE-bench Multilingual (Polyglot coding) 66.00% 71.40% 67.20% Tiel +5.4% stronger multi-language support
SWE-bench Pro (Professional repo issues) (Not reported) 59.60% 49.50% Tiel +10.1% over stock Qwen3.6
DeepSWE (Complex architectural tasks) (Not reported) 22.00% 0.00% Tiel solves 22% where baseline 35Bs score 0%
SWE-bench-Live (25 live problems, Q4 GGUF) (Not tested) 12 / 25 (48.0%) 8 / 25 (32.0%) Tiel matches Claude Opus 4.6 medium (8.6 min median solve time)
Claw-Eval (Agentic workflows) 76.55 67.20 74.54 Xing4.0 leads in multi-step planning and tool-calling flows
MiniPlus V2.1 Model Weight 13.76 GB (3.42 BPW) 14.75 GB (3.45 BPW) 14.75 GB Xing4.0 saves approx. 1 GB VRAM on base model weights
NanoPlus Model Weight 11.74 GB (2.92 BPW) 12.60 GB (2.95 BPW) 12.60 GB Xing4.0 saves approx. 0.86 GB VRAM on ultra-compact weights
Attention Architecture MLA + mHC (Latent compression) GQA (Grouped-Query) GQA Xing4.0 compresses KV representation by approx. 44%
KV Cache Scaling (Large Context) Ultra-low footprint (MLA compressed) Standard footprint (GQA) Standard footprint Xing4.0 consumes far less KV cache VRAM at deep 128K to 256K context

Tiel seem still very op

And Cyber-Tiel is practically the same but Abliterated; I also have it under the MiniPlus and NanoPlus recipes if you want to try them.

Sign up or log in to comment