comparar
llama.cpp vs Xinference
Os mesmos fatos para os dois, lidos do GitHub toda noite, e a relação que uma pessoa revisou.
| Fato | llama.cpp | Xinference |
|---|---|---|
| Linguagem | C++ | Python |
| Licença | MIT | Apache-2.0 |
| Estrelas | 130k | 9.6k |
| Última release | v0.6.0 | v3.5.0 |
| Último push | 2026-10-06 | 2026-10-06 |
| Ritmo de releases | poucas releases | cerca de 15 dias entre releases |
| Contribuidores ativos | 173+ autores de commits na branch padrão nos últimos 90 dias | 38 autores de commits na branch padrão nos últimos 90 dias |
| Alertas | nenhum | nenhum |
Como se relacionam
Os dois substituem ChatGPT. Alternativas a ChatGPT →
Parcialllama.cppRuns open-weight models on CPU or GPU, with a built-in server and a minimal web chat.
ParcialXinferenceServes open-weight language, embedding and speech models behind an OpenAI-compatible API.
llama.cpp
- b114452026-10-06vulkan : check for null vkEnumerateInstanceVersion (#29872)pré-release
- b114432026-10-06models : consolidate nextn row cropping into shared helpers (#30017)pré-release
- b114402026-10-06llama : re-reserve the sched when the nextn extraction flags change (#30020)pré-release
- b114392026-10-06ggml: refactor selective expert copying to user code (#29943)pré-release
- b114382026-10-06test-llama-archs : initialize backends before generating models (#30034)pré-release
Xinference
llama.cpp
não assinada A última release, v0.6.0, não tem nenhuma assinatura que o GitHub consiga verificar.
Carregando o relatório de segurança
Xinference
✓ assinada A última release, v3.5.0, tem uma assinatura verificada pelo GitHub.
Carregando o relatório de segurança