Skip Navigation

Posts
11
Comments
99
Joined
3 yr. ago

  • This isn't actually using a vision LLM, it's using a CLIP model. This image comes from an OpenAI blog from 2019 I think

  • 4/6 bait

  • gemini logo in the bottom right corner???

    i checked the twitter account and it's in russian. i guess they used nano banana to translate the text in an image of text. beautiful

  • won't care? yeah probably. aren't intelligent enough? that's an insane generalization, knowledgeable about technology ≠ smart

  • As a memory-poor user (hence the 8gb vram card), I consider Q8+ to be is higher precision, Q4-Q5 is mid-low precision (what i typically use), and below that is low precision

  • It's a webp animation. Maybe your client doesn't display it right, i'll replace it with a gif

    Regarding your other question, I tend to see better results with higher params + lower precision, versus low params + higher precision. That's just based on "vibes" though, I haven't done any real testing. Based on what I've seen, Q4 is the lowest safe quantization, and beyond that, the performance really starts to drop off. unfortunately even at 1 bit quantization I can't run GLM 4.6 on my system

  • this made me mad so i made a single, ultra minimal html page in 5 minutes that you can just paste in your url box

     
        
    data:text/html;base64,PCFkb2N0eXBlaHRtbD48Ym9keSBzdHlsZT10ZXh0LWFsaWduOmNlbnRlcjtmb250LWZhbWlseTpzYW5zLXNlcmlmO2JhY2tncm91bmQ6IzAwMDtjb2xvcjojMmYyPjxoMT5JcyBpdCBETlM/PC9oMT48cCBzdHlsZT1mb250LXNpemU6MTJyZW0+WWVz
    
      

    source code:

     html
        
    <!doctypehtml><body style=text-align:center;font-family:sans-serif;background:#000;color:#2f2><h1>Is it DNS?</h1><p style=font-size:12rem>Yes
    
      
  • You can run Qwen3 4b thinking at q4 quantization at 2.5GB, which is probably a better model too

  • there's also a "small" and "micro" variant, which are 32b a6b MoE and 3b dense models respectively

  • Sept

    Jump
  • I am not using an llm but holy bait

    Hop off the reddit voice

  • Sept

    Jump
  • 4/10 bait

  • Sept

    Jump
  • That is a thing, and it's called quantization aware training. Some open weight models like Gemma do it.

    The problem is that you need to re-train the whole model for that, and if you also want a full-quality version you need to train a lot more.

    It is still less precise, so it'll still be worse quality than full precision, but it does reduce the effect.

  • Sept

    Jump
  • rsync for backups? I guess it depends on what kind of backup

    for redundant backups of my data and configs that I still have a live copy of, I use restic, it compresses extremely well

    I have used rsync to permanently move something to another drive though

  • I guess it's like a CAPTCHA. It doesn't completely solve the problem the hoster wishes to solve, but it deters a lot of people from trying.

  • I tried this again, with a gguf quantized to Q4KM. It works quite well and can generate in ~7 minutes! Thanks!

  • My guy stop following me around