AMD Ryzen 9 9950X CPU and AMD radeon pro w7900 (48GB vram). I get 55tps output pretty consistently, but ingesting context starts around 1500tps and if context size reaches, say, 50K, tps drops to around 200tps. I often have to wait a bit, but it's a price I'm happy to pay for local-only AI
- Posts
- 2
- Comments
- 97
- Joined
- 2 yr. ago
- Posts
- 2
- Comments
- 97
- Joined
- 2 yr. ago
- JumpDeleted
Permanently Deleted
- JumpDeleted
Permanently Deleted
- JumpDeleted
Permanently Deleted
FYI, I think opencode go is kind of a subscription model, not a direct credits-to-tokens model. In terms of value it's nowhere near as good as Claude's subscriptions, but it seems way more valuable that paying for tokens directly. However they only offer a few models - decent ones, but not many and a little behind the times.