comparar

llama.cpp vs Xinference

Os mesmos fatos para os dois, lidos do GitHub toda noite, e a relação que uma pessoa revisou.

llama.cpp vs Xinference
Fatollama.cppXinference
LinguagemC++Python
LicençaMITApache-2.0
Estrelas130k9.6k
Última releasev0.6.0v3.5.0
Último push2026-10-062026-10-06
Ritmo de releasespoucas releasescerca de 15 dias entre releases
Contribuidores ativos173+ autores de commits na branch padrão nos últimos 90 dias38 autores de commits na branch padrão nos últimos 90 dias
Alertasnenhumnenhum
  • Os dois substituem ChatGPT. Alternativas a ChatGPT →

    Parcialllama.cppRuns open-weight models on CPU or GPU, with a built-in server and a minimal web chat.

    ParcialXinferenceServes open-weight language, embedding and speech models behind an OpenAI-compatible API.