← voltar

categoria

Local model runtimes

Run open-weight language models on your own hardware, behind a local API.

Lado a lado

Local model runtimes
FerramentaEstrelasLicençaTermosAuto-hospedadoLinguagemÚltima releaseÚltimo push
Ollama182kMITCódigo abertoSimGov0.35.12026-10-06
llama.cpp130kMITCódigo abertoSimC++v0.6.02026-10-06
vLLM93kApache-2.0Código abertoSimPythonv0.31.02026-10-06
LocalAI49kMITCódigo abertoSimGov4.11.02026-10-06
exo48kApache-2.0Código abertoSimPythonv1.0.712026-10-06
SGLang37kApache-2.0Código abertoSimPythonv0.5.212026-10-06
llamafile26kOtherCódigo abertoSimC++0.10.62026-10-06
MLC LLM23kApache-2.0Código abertoSimPythonv0.20.02026-10-04
KTransformers20kApache-2.0Código abertoSimPythonv0.7.12026-10-01
TensorRT-LLM15kOtherCódigo abertoSimPythonv1.2.12026-10-06
Xinference9.6kApache-2.0Código abertoSimPythonv3.5.02026-10-06
LMDeploy8.1kApache-2.0Código abertoSimPythonv0.18.02026-09-28
mistral.rs7.7kMITCódigo abertoSimRustv0.9.42026-10-01
llama-swap5.9kMITCódigo abertoSimGov2622026-10-05
Lemonade5.8kApache-2.0Código abertoSimC++v2026.40.02026-10-06
GPUStack5.8kApache-2.0Código abertoSimPythonv2.2.32026-09-30
RamaLama3.1kMITCódigo abertoSimPythonv0.25.02026-10-05
TabbyAPI1.5kAGPL-3.0Código abertoSimPythonnenhuma2026-10-06

18 ferramentas

Ollama

Get up and running with Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.

Linguagem
Go
Licença
MIT
Estrelas
182k
Última release
v0.35.1
Último push
2026-10-06
llama.cpp

LLM inference in C/C++

Linguagem
C++
Licença
MIT
Estrelas
130k
Última release
v0.6.0
Último push
2026-10-06
vLLM

A high-throughput and memory-efficient inference and serving engine for LLMs

Linguagem
Python
Licença
Apache-2.0
Estrelas
93k
Última release
v0.31.0
Último push
2026-10-06
LocalAI

LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.

Linguagem
Go
Licença
MIT
Estrelas
49k
Última release
v4.11.0
Último push
2026-10-06
SGLang

SGLang is a high-performance serving framework for large language models and multimodal models.

Linguagem
Python
Licença
Apache-2.0
Estrelas
37k
Última release
v0.5.21
Último push
2026-10-06
llamafile

Distribute and run LLMs with a single file.

Linguagem
C++
Licença
Other
Estrelas
26k
Última release
0.10.6
Último push
2026-10-06
KTransformers

A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

Linguagem
Python
Licença
Apache-2.0
Estrelas
20k
Última release
v0.7.1
Último push
2026-10-01
TensorRT-LLM

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

Linguagem
Python
Licença
Other
Estrelas
15k
Última release
v1.2.1
Último push
2026-10-06
Xinference

Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.

Linguagem
Python
Licença
Apache-2.0
Estrelas
9.6k
Última release
v3.5.0
Último push
2026-10-06
LMDeploy

LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

Linguagem
Python
Licença
Apache-2.0
Estrelas
8.1k
Última release
v0.18.0
Último push
2026-09-28
mistral.rs

Fast, flexible LLM inference

Linguagem
Rust
Licença
MIT
Estrelas
7.7k
Última release
v0.9.4
Último push
2026-10-01
llama-swap

Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc

Linguagem
Go
Licença
MIT
Estrelas
5.9k
Última release
v262
Último push
2026-10-05
Lemonade

Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk

Linguagem
C++
Licença
Apache-2.0
Estrelas
5.8k
Última release
v2026.40.0
Último push
2026-10-06
GPUStack

A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.

Linguagem
Python
Licença
Apache-2.0
Estrelas
5.8k
Última release
v2.2.3
Último push
2026-09-30
RamaLama

RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.

Linguagem
Python
Licença
MIT
Estrelas
3.1k
Última release
v0.25.0
Último push
2026-10-05

O que essas ferramentas substituem