Skip Navigation

InitialsDiceBearhttps://github.com/dicebear/dicebearhttps://creativecommons.org/publicdomain/zero/1.0/„Initials” (https://github.com/dicebear/dicebear) by „DiceBear”, licensed under „CC0 1.0” (https://creativecommons.org/publicdomain/zero/1.0/)B
Posts
5
Comments
10
Joined
1 wk. ago

  • In a few minutes a significant performance improvement incoming

    👀

  • this has been a crazy few weeks! lol

  • true, it's not perfectly clear

    also I just saw this

  • LocalLLaMA @sh.itjust.works

    preserve_thinking fix for Gemma 4 templates

  • Gemma is probably good for that, as long as it's consistently succeeding at the tool calls.

  • LocalLLaMA @sh.itjust.works

    Qwen 3.8 Max (2.4T-a95b) and 27B open weights being released next week

  • make sure that holds up with large context, you might need to step down to Q3 (which I've heard is still good for this model, many people are even using IQ2)

  • the original model has a lot of parts that were natively trained in 4 bit, so those layers can't go higher

  • LocalLLaMA @sh.itjust.works

    unsloth/DeepSeek-V4-Flash-0731-GGUF · Hugging Face

    huggingface.co /unsloth/DeepSeek-V4-Flash-0731-GGUF
  • --n-cpu-moe 36 --spec-type draft-mtp --spec-draft-n-max 3 does seem to speed up token generation for me

    Can't use llama-bench for MTP. In a basic tests it seems to improve from about 26 to 30 tokens per second output. But it seems to hurt my input speed from about 1300 pp down to 800.

  • LocalLLaMA @sh.itjust.works

    How to run Qwen 35b-a3b on 4GB to 8GB of VRAM (with 24+GB system RAM)