Skip Navigation

Posts
2
Comments
97
Joined
2 yr. ago

  • AMD Ryzen 9 9950X CPU and AMD radeon pro w7900 (48GB vram). I get 55tps output pretty consistently, but ingesting context starts around 1500tps and if context size reaches, say, 50K, tps drops to around 200tps. I often have to wait a bit, but it's a price I'm happy to pay for local-only AI

  • I use opencode with locally-hosted llama.cpp - usually with qwen3.6-35b-a3b.

    I tried opencode go for a couple month, and its definitely nice to have an lln runner with more gram and more GPUs, but I prefer to have all my stuff local whenever it's possible. Also, I'd use up my token allotments fairly quickly on opencode go.

    I also tried opentouter and it, too, was great - many more models. But I exhausted by credits even quicker than opencode go, and its also not local.

  • Very good points.

  • That's hot, and climate change exists. I get it. But 48C (118F) shutting things down? I used to live in Palm Springs, CA where it's over 37F for 5-6 months every year, and the hottest I personally experienced was 51C (124F) and I experienced a lot of 47C - and while that was indeed unpleasantly hot (but it's a dry heat! 🙂), life continued as normal. I remember even seeing tons of laborers building new houses mid day as a regular occurrence.

  • It does seem like what I'm seeing would be someone's idea/joke of a drunk mode, but it's how the page loads for me - I didn't click anything. And I can't read the buttons to see if one days "turn off drunk mode".

  • Holy shit how can you read that? For a minute I thought the text was encoded until maybe I clicked a button, then I thought it might be Arabic, now I see that's it's English but I can't even make out all the letters. No thanks.

  • Not my GPU. I have an AMD.

    Requirements

    • NVIDIA Ampere or newer GPU (A100, H100, H200, B200, RTX 3000+)
  • I'm my experience, running Ollama locally works great. I do have a beefy GPU, but even on affordable consumer grade GPUs you can get good results with smaller models.

    So it technically works to run an AI agent locally, but my experience has been that coding agents don't work well. I haven't tried using general AI agents.

    I think the amount of VRAM affordable/available to consumers is nowhere near enough to support a context length that's necessary for a coding agent to remain coherent. There are tools like Get Shit Done which are supposed to help with this, but I didn't have much luck.

    So I'm using OpenCode via OpenRouter to use LLMs in the cloud. Sad that I can't get local-only to work well enough to use for coding agents, but this arrangement works for me (for now).

  • Interesting! I've seen the sidebar but not thought of much advantage from it. I'll take another look. Model/provider change is a breeze, I just assumed Claude would do the same but maybe they want to make it harder to leave their models? Workspaces? Sounds interesting - gotta check it out.

    I just discovered how easy it is to view/switch sessions in opencode.

  • I use opencode, have seen Claude but never used it. What are some examples of things opencode does that Claude doesn't?

  • I i think it's great that you're spreading the word about the fediverse and helping get people off of big data.

    You might consider using a youtube alternative for your channel, perhaps a PeerTube instance.

  • Deleted

    Permanently Deleted

    Jump
  • Ah.

  • Deleted

    Permanently Deleted

    Jump
  • I don't care if AI was used in its creation. I do care if it's FOSS/libre.

    And also, it's a bit weird to me that copying YouTube's UI is considered good. I havent used YouTube in a long time, but I recall there being some good aspects and some bad. Why not create your own vesion of a UI?

  • Deleted

    Permanently Deleted

    Jump
  • I agree that more options is a good thing, and that activitypub would be a plus. But FYI, I wont be using it because of the license. I use only FOSS whenevr possible.

  • I don't think Lemmy's previous UI was bad but I don't think it was great. I've been using Tesseract on the desktop - not only is it's UI "better", but it has great additional features.

  • Server01: 64 Server02: 19 Plus a bunch of sidecar containers solely for configs that aren't running.

  • That's not the only scenario. I'm not a coder but there are many aspects of coding i enjoy, and AI is the tool that lets me enjoy doing it.

  • Thanks for explaining. 🙂

  • I'd assume that many services were created to send an activation link via email and don't know how to talk to a Matrix server. In those cases, do they email their activation link to a service or proxy that then communicates it to the appropriate matrix server/account/room?