Pool spare GPU capacity to run LLMs at larger scale
Pool spare GPU capacity to run LLMs at larger scale
github.com
GitHub - michaelneale/mesh-llm: reference impl with llama.cpp compiled to distributed inference across machines, with real end to end demo
