compare
GPUStack vs llama.cpp
The same facts for both, read from GitHub every night, and the relation a person reviewed.
| Fact | GPUStack | llama.cpp |
|---|---|---|
| Language | Python | C++ |
| Licence | Apache-2.0 | MIT |
| Stars | 5.8k | 130k |
| Latest | v2.2.3 | v0.6.0 |
| Last push | 2026-09-30 | 2026-10-06 |
| Release cadence | too few releases | too few releases |
| Active contributors | 27 commit authors on the default branch in the last 90 days | 173+ commit authors on the default branch in the last 90 days |
| Flags | none | none |
How they relate
Both replace ChatGPT. Alternatives to ChatGPT →
PartialGPUStackManages a cluster of GPUs and serves open-weight models across it behind an OpenAI-compatible API.
Partialllama.cppRuns open-weight models on CPU or GPU, with a built-in server and a minimal web chat.
GPUStack
- v2.3.0rc22026-09-30pre-release
- v2.3.0rc12026-09-01pre-release
- v2.2.32026-07-31Fixed an issue where vLLM services routed through GPUStack were incompatible with Claude Code. (Issue #5934)
- v2.2.3rc12026-07-31pre-release
- v2.2.22026-07-24This release fixes two security vulnerabilities. Users are strongly advised to upgrade immediately.
llama.cpp
- b114452026-10-06vulkan : check for null vkEnumerateInstanceVersion (#29872)pre-release
- b114432026-10-06models : consolidate nextn row cropping into shared helpers (#30017)pre-release
- b114402026-10-06llama : re-reserve the sched when the nextn extraction flags change (#30020)pre-release
- b114392026-10-06ggml: refactor selective expert copying to user code (#29943)pre-release
- b114382026-10-06test-llama-archs : initialize backends before generating models (#30034)pre-release
GPUStack
unsigned The latest release, v2.2.3, carries no signature GitHub could verify.
Loading the security report
llama.cpp
unsigned The latest release, v0.6.0, carries no signature GitHub could verify.
Loading the security report