Paired-4:8 + NVFP4 W4A4 expert compressed MoE models for NVIDIA Blackwell Sparse Tensor Cores.
AI & ML interests
None defined yet.
Recent Activity
View all activity
Papers
Disaggregated Quantization: Specializing LLM Prefill and Decode
GPTQ-2D: Cubic-Time Two-Sided Adaptive Rounding
models 169
ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF
Image-Text-to-Text • 177B • Updated • 613k • 370
ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF
Image-Text-to-Text • 117B • Updated • 19.5k • 116
ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ
Text Generation • 646B • Updated
ISTA-DASLab/Qwen3.5-397B-A17B-P48NVFP4-MoESQ
Image-Text-to-Text • 258B • Updated
ISTA-DASLab/Qwen3-30B-A3B-P48NVFP4-MoESQ
Text Generation • 20B • Updated
ISTA-DASLab/Qwen3.8-27B-NVFP4-prefiller
Text Generation • 24B • Updated • 364 • 18
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF
Image-Text-to-Text • 27B • Updated • 1.68M • 1.82k
ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ
Image-Text-to-Text • 27B • Updated • 7.74k • 20
ISTA-DASLab/Qwen3.6-35B-A3B-2Bit-GSQ
Image-Text-to-Text • 36B • Updated • 3.56k • 16
ISTA-DASLab/Kimi-K2.5-2Bit-GSQ
Image-Text-to-Text • 84B • Updated • 461 • 1