← back

category

Local model runtimes

Run open-weight language models on your own hardware, behind a local API.

Side by side

Local model runtimes
ToolStarsLicenceTermsSelf-hostedLanguageLatestLast push
Ollama182kMITOpen sourceYesGov0.35.12026-10-06
llama.cpp130kMITOpen sourceYesC++v0.6.02026-10-06
vLLM93kApache-2.0Open sourceYesPythonv0.31.02026-10-06
LocalAI49kMITOpen sourceYesGov4.11.02026-10-06
exo48kApache-2.0Open sourceYesPythonv1.0.712026-10-06
SGLang37kApache-2.0Open sourceYesPythonv0.5.212026-10-06
llamafile26kOtherOpen sourceYesC++0.10.62026-10-06
MLC LLM23kApache-2.0Open sourceYesPythonv0.20.02026-10-04
KTransformers20kApache-2.0Open sourceYesPythonv0.7.12026-10-01
TensorRT-LLM15kOtherOpen sourceYesPythonv1.2.12026-10-06
Xinference9.6kApache-2.0Open sourceYesPythonv3.5.02026-10-06
LMDeploy8.1kApache-2.0Open sourceYesPythonv0.18.02026-09-28
mistral.rs7.7kMITOpen sourceYesRustv0.9.42026-10-01
llama-swap5.9kMITOpen sourceYesGov2622026-10-05
Lemonade5.8kApache-2.0Open sourceYesC++v2026.40.02026-10-06
GPUStack5.8kApache-2.0Open sourceYesPythonv2.2.32026-09-30
RamaLama3.1kMITOpen sourceYesPythonv0.25.02026-10-05
TabbyAPI1.5kAGPL-3.0Open sourceYesPythonnone2026-10-06

18 tools

Ollama

Get up and running with Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.

Language
Go
Licence
MIT
Stars
182k
Latest
v0.35.1
Last push
2026-10-06
vLLM

A high-throughput and memory-efficient inference and serving engine for LLMs

Language
Python
Licence
Apache-2.0
Stars
93k
Latest
v0.31.0
Last push
2026-10-06
LocalAI

LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.

Language
Go
Licence
MIT
Stars
49k
Latest
v4.11.0
Last push
2026-10-06
SGLang

SGLang is a high-performance serving framework for large language models and multimodal models.

Language
Python
Licence
Apache-2.0
Stars
37k
Latest
v0.5.21
Last push
2026-10-06
TensorRT-LLM

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

Language
Python
Licence
Other
Stars
15k
Latest
v1.2.1
Last push
2026-10-06
Xinference

Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.

Language
Python
Licence
Apache-2.0
Stars
9.6k
Latest
v3.5.0
Last push
2026-10-06
llama-swap

Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc

Language
Go
Licence
MIT
Stars
5.9k
Latest
v262
Last push
2026-10-05
Lemonade

Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk

Language
C++
Licence
Apache-2.0
Stars
5.8k
Latest
v2026.40.0
Last push
2026-10-06
GPUStack

A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.

Language
Python
Licence
Apache-2.0
Stars
5.8k
Latest
v2.2.3
Last push
2026-09-30
RamaLama

RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.

Language
Python
Licence
MIT
Stars
3.1k
Latest
v0.25.0
Last push
2026-10-05

What these tools replace