Skip Navigation
Hacker News @lemmy.bestiver.se

Pool spare GPU capacity to run LLMs at larger scale

GitHub - michaelneale/mesh-llm: reference impl with llama.cpp compiled to distributed inference across machines, with real end to end demo

Comments

1