how Qwen 3.8 adjusts the reasoning_effort
how Qwen 3.8 adjusts the reasoning_effort
how Qwen 3.8 adjusts the reasoning_effort
llama.cpp in progress pull request for smart caching of MoE experts, 16% to 35% TPS boost for my RTX 2080
preserve_thinking fix for Gemma 4 templates
Qwen 3.8 Max (2.4T-a95b) and 27B open weights being released next week
unsloth/DeepSeek-V4-Flash-0731-GGUF · Hugging Face
How to run Qwen 35b-a3b on 4GB to 8GB of VRAM (with 24+GB system RAM)
Sounds like you could just replace the template to fix it, there's a popular Qwen fixed template on hugging face, try that
https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates/blob/main/chat_template.jinja