Non-uniform GGUF quantizations via GSQ + RCO: per-tensor mixed precision in standard GGUF form
AI & ML interests
None defined yet.
Recent Activity
View all activity
Papers
GPTQ-2D: Cubic-Time Two-Sided Adaptive Rounding
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling, https://huggingface.co/papers/2604.18556
-
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
Paper • 2604.18556 • Published • 11 -
ISTA-DASLab/Kimi-K2.6-2Bit-GSQ
Image-Text-to-Text • 84B • Updated • 23 • 1 -
ISTA-DASLab/Kimi-K2.5-2Bit-GSQ
Image-Text-to-Text • 84B • Updated • 68 • 1 -
ISTA-DASLab/Llama-3.1-70B-Instruct-2Bit-GSQ
Text Generation • 7B • Updated • 22
Non-uniform GGUF quantizations via GSQ + RCO: per-tensor mixed precision in standard GGUF form
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling, https://huggingface.co/papers/2604.18556
-
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
Paper • 2604.18556 • Published • 11 -
ISTA-DASLab/Kimi-K2.6-2Bit-GSQ
Image-Text-to-Text • 84B • Updated • 23 • 1 -
ISTA-DASLab/Kimi-K2.5-2Bit-GSQ
Image-Text-to-Text • 84B • Updated • 68 • 1 -
ISTA-DASLab/Llama-3.1-70B-Instruct-2Bit-GSQ
Text Generation • 7B • Updated • 22
models 164
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF
Text Generation • 27B • Updated • 6.8k • 62
ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ
Image-Text-to-Text • 27B • Updated • 1.71k • 6
ISTA-DASLab/Qwen3.6-35B-A3B-2Bit-GSQ
Image-Text-to-Text • 36B • Updated • 257 • 3
ISTA-DASLab/Kimi-K2.5-2Bit-GSQ
Image-Text-to-Text • 84B • Updated • 68 • 1
ISTA-DASLab/Kimi-K2.6-2Bit-GSQ
Image-Text-to-Text • 84B • Updated • 23 • 1
ISTA-DASLab/Qwen3.5-4B-GGUF-GSQ
Text Generation • 4B • Updated • 58 • 1
ISTA-DASLab/Kimi-K2.5-P48-NVFP4-W4A4-Preview
Image-Text-to-Text • Updated • 43 • 1
ISTA-DASLab/Llama-3.1-70B-Instruct-2Bit-GSQ
Text Generation • 7B • Updated • 22
ISTA-DASLab/Llama-3.1-70B-Instruct-3Bit-GSQ
Text Generation • 9B • Updated • 13
ISTA-DASLab/Qwen3-8B-GGUF-GSQ
Text Generation • 8B • Updated • 56 • 3