Skip Navigation

InitialsDiceBearhttps://github.com/dicebear/dicebearhttps://creativecommons.org/publicdomain/zero/1.0/„Initials” (https://github.com/dicebear/dicebear) by „DiceBear”, licensed under „CC0 1.0” (https://creativecommons.org/publicdomain/zero/1.0/)S
Posts
19
Comments
326
Joined
3 mo. ago

  • Agreed. It will be ironic if 1.58B models (Microsoft) turns out to be the great white hope.

    I looked at the recent Steam stats (which is a GPU sample of convenience); the most common GPU size was 6GB. Meanwhile you probably need what...64GB unified memory or a 5090 to drive a decent model at a decent speed/context?

    There's a real gap between the haves and the have nots and it's widening.

  • No idea because they failed to mention it.

    Spam isn't just automated spam bots posting. It's unsolicited, unexplained content. A "hey, I found this interesting because X" goes a long way towards humanising things like this but YMMV.

  • They're not. Call them via API on Open Router and see for yourself.

    There's a reason OAI and Anthropic are considered best in class and it's not just hype.

  • Myself - I've self hosted LLMs before, but with only 4-8GB vram (depending which card is in place), I can't run the good stuff at acceptable enough speeds.

    (Don't @ me - I know all the tricks with turbo quants, spec decoding, MoE etc. 192GB/s is 192GB/s)

    I do use Handy (STT) which is amazing (my fingers are arthritic and typing hurts after a while).

    My personal use case for LLM is quite simple - a trumped up super google and / or self reflection / journalling / sound board. Despite being glib about it, that's actually very useful to me.

    Work wise, I use the big winking orange asshole (Claude) when I have to. I have moral tension with with it, so am seriously looking at other options. I hear good things about GLM 5.2, but if I can't run Qwen 35B at any kind of decent speed, well....self hosted GLM is a pipe dream.

  • I'll make sure to send you flowers, Algernon lol

  • That Qwen 35B model is going to remain the people's champ for a long time I think. Surprisingly capable, even for code. I hear it loops badly at Q4 quant?

  • Nice bit of kit that. Very nice. Planning on serious AI shenanigans?

    I'll check it out.

    Cool. It's not ready any time soon but when it is, I'll announce it and make sure it's callable via SSH / terminal / OpenAI style chat end point.

    That way you don't need anything fancier than a nice terminal to call it.

  • Did you use OWUIs native "call simultaneous models to answer" feature for that or one of the AI debate harnesses?

  • You can get a P40 for much less than that, if your case can hold full height card. It's an old card but its 24GB, 400GB/s.

    Else yeah...$3-4,000 is about table stakes, which doesn't amortise for just AI (not for my use cases anyway). I'd love a Strix but Santa is stingy.

    Me - I have a fetish for tiny, low power computers. 1L lenovos, raspberry pis etc. That limits what I can run but with constraint comes inginuity. So I'm making an expert system for myself.

    https://codeberg.org/BobbyLLM/picoGURU

    It's not cooked yet (this is actually the first time I'm sharing it in public; it's not in installable state and the repo is new) but once it's done, I can have an always on local brain in a 2W envelope that runs fast. Might even port it to C64...I need an excuse to purchase the new Commodore ultimate.

  • Terrible posting etiquette though. Not a peep out of OP or any rationale.

    Is this a bot? Someone promoting their blog? Spam? Click bait?

  • Tags don't protect against that tho , if applied honestly. And humans aren't immune from making human slop all on their own.

    Spelunking the repo is 100% the answer if that's the threat model.

    The tag / no tag thing can only be part of the due diligence. I argue it (at best) is neutral to that end and at worst, completely flattens the reality of code gen in 2026. Nearly 100% of code gen now touches AI somewhere.

    Turbo encabulator style announcement is a much louder and more useful signal, and we already get that for free. Tag may actually end up blunting that.

  • Why's that?

  • Exactly. So if it sounds like a turbo encabulator, why so we need a tag?

  • It won't work. You need a text classifier to do sentiment analysis, because "ai" is a concept, not just "ai". TinyBERT or MiniLM I reckon could do it or if you really want to cut off your nose to spite your face, code the equivalent in python from scratch.

    Say what you want about M$, but TinyBERT / MiniLM are awesome.

    Smart play would be for the RSS reader to have that as optional plug in module, IMHO.

  • Sir, this is Lemmy. If you use AI in any way, you are clearly in league with the devil and deserve to burn.

    I agree with all your points, BTW.

    I posted this discussion because I wanted to explore both guard rails AND nuance around that sort of work flow, particularly for our new mod (and in light of several other scattered convos).

    A lot of the diffuse FuckAI Lemmy crowd have poor understanding of code workflow. "AI bad" knee jerks so hard it's going to dislocate something.

    I've tried to argue this point, because roughly... ooh...100% of code gen touches AI something. So, do we tag everything?

    What people really want is a [SLOP] tag, which is both lazy / not doing your own due diligence and impossible to implement.

    In hindsight, I think the pragmatic approach is ultimately the workable (albeit blunted) one. Have the ai tag. It flattens everything but if stops brigading and slop, that's the least amount of moderation work.

    I appreciate you posting btw.

  • "Are you now, or have you ever been..."

  • I think you might need both for the tag bot to work properly but dunno.

  • So state of play / preview of coming attractions - yes to tags, once tagged, cannot complain about said tagged content. Formal sticky etc next week.

    PS: I took a look at r/selfhosted - their bot seems to delete ALL new project posts and requires user to appeal / resubmit / verify directly. I think that's problematic (and ironic, if you think about it - you're trying to litigate ai/no ai with a bot) but not my circus, not my monkeys.