comparar
GPUStack vs llama.cpp
Os mesmos fatos para os dois, lidos do GitHub toda noite, e a relação que uma pessoa revisou.
| Fato | GPUStack | llama.cpp |
|---|---|---|
| Linguagem | Python | C++ |
| Licença | Apache-2.0 | MIT |
| Estrelas | 5.8k | 130k |
| Última release | v2.2.3 | v0.6.0 |
| Último push | 2026-09-30 | 2026-10-06 |
| Ritmo de releases | poucas releases | poucas releases |
| Contribuidores ativos | 27 autores de commits na branch padrão nos últimos 90 dias | 173+ autores de commits na branch padrão nos últimos 90 dias |
| Alertas | nenhum | nenhum |
Como se relacionam
Os dois substituem ChatGPT. Alternativas a ChatGPT →
ParcialGPUStackManages a cluster of GPUs and serves open-weight models across it behind an OpenAI-compatible API.
Parcialllama.cppRuns open-weight models on CPU or GPU, with a built-in server and a minimal web chat.
GPUStack
- v2.3.0rc22026-09-30pré-release
- v2.3.0rc12026-09-01pré-release
- v2.2.32026-07-31Fixed an issue where vLLM services routed through GPUStack were incompatible with Claude Code. (Issue #5934)
- v2.2.3rc12026-07-31pré-release
- v2.2.22026-07-24This release fixes two security vulnerabilities. Users are strongly advised to upgrade immediately.
llama.cpp
- b114452026-10-06vulkan : check for null vkEnumerateInstanceVersion (#29872)pré-release
- b114432026-10-06models : consolidate nextn row cropping into shared helpers (#30017)pré-release
- b114402026-10-06llama : re-reserve the sched when the nextn extraction flags change (#30020)pré-release
- b114392026-10-06ggml: refactor selective expert copying to user code (#29943)pré-release
- b114382026-10-06test-llama-archs : initialize backends before generating models (#30034)pré-release
GPUStack
não assinada A última release, v2.2.3, não tem nenhuma assinatura que o GitHub consiga verificar.
Carregando o relatório de segurança
llama.cpp
não assinada A última release, v0.6.0, não tem nenhuma assinatura que o GitHub consiga verificar.
Carregando o relatório de segurança