Skip Navigation

InitialsDiceBearhttps://github.com/dicebear/dicebearhttps://creativecommons.org/publicdomain/zero/1.0/„Initials” (https://github.com/dicebear/dicebear) by „DiceBear”, licensed under „CC0 1.0” (https://creativecommons.org/publicdomain/zero/1.0/)S
Posts
1
Comments
13
Joined
4 mo. ago

  • Do you care to explain ?

  • With which quantization ? I have a 16GB GPU, running Qwen3.5 9b UD_Q8_K_XL basically max out the VRAM usage. Maybe a MOE model fits you. Gemma 4 e4b is one with a total 8b.

    And what is the token speed you are getting ?

  • I guess you could wait until it is cracked. But by then, I think it is best to use a old pc.

    You say also if RAM can be upgraded. Consoles are not like PCs in terms of customizations. Only the old ones that have been cracked. You can't.

    But your enthusiasm is there. That's cool.

  • Yesterday I needed this. Will install this. Thanks.

    May I ask: have you noticed if the prompt processing speeds shown in llama-bench are vastly different from llama-server ? I have hundreds of tokens of difference.

  • You mean Gemma 4 ? You read in his discord ?

  • I confirm the same and it works now. I set it to maximum because fewer reasoning effort tokens cuts it directly.

    Thanks

  • LocalLLaMA @sh.itjust.works

    My models don't have reasoning ability in llama-b9543 server but have in llama-cli

  • I know it, seeing it in models titles.

  • Is uncensoring oneself a LLM difficult ?

  • Did you try any ? Because, I tried iglors and mradermacher, I got refusal to make a pipe bomb. Their answer are funny because they say to study academic engineering instead, lol. Still a refusal. I will try this one.

  • The Arch Linux forum make people do that.

  • Someone may have the same question in the future and there will be answers. You not responding is not that bad but it is even better that you do and provide an update to your situation, if you wish.

  • No. In fact, that is nice. I should try.