← retour

catégorie

Local model runtimes

Run open-weight language models on your own hardware, behind a local API.

Côte à côte

Local model runtimes
OutilÉtoilesLicenceConditionsAuto-hébergeableLangageDernière releaseDernier push
Ollama182kMITOpen sourceOuiGov0.35.12026-10-06
llama.cpp130kMITOpen sourceOuiC++v0.6.02026-10-06
vLLM93kApache-2.0Open sourceOuiPythonv0.31.02026-10-06
LocalAI49kMITOpen sourceOuiGov4.11.02026-10-06
exo48kApache-2.0Open sourceOuiPythonv1.0.712026-10-06
SGLang37kApache-2.0Open sourceOuiPythonv0.5.212026-10-06
llamafile26kOtherOpen sourceOuiC++0.10.62026-10-06
MLC LLM23kApache-2.0Open sourceOuiPythonv0.20.02026-10-04
KTransformers20kApache-2.0Open sourceOuiPythonv0.7.12026-10-01
TensorRT-LLM15kOtherOpen sourceOuiPythonv1.2.12026-10-06
Xinference9.6kApache-2.0Open sourceOuiPythonv3.5.02026-10-06
LMDeploy8.1kApache-2.0Open sourceOuiPythonv0.18.02026-09-28
mistral.rs7.7kMITOpen sourceOuiRustv0.9.42026-10-01
llama-swap5.9kMITOpen sourceOuiGov2622026-10-05
Lemonade5.8kApache-2.0Open sourceOuiC++v2026.40.02026-10-06
GPUStack5.8kApache-2.0Open sourceOuiPythonv2.2.32026-09-30
RamaLama3.1kMITOpen sourceOuiPythonv0.25.02026-10-05
TabbyAPI1.5kAGPL-3.0Open sourceOuiPythonaucune2026-10-06

18 outils

Ollama

Get up and running with Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.

Langage
Go
Licence
MIT
Étoiles
182k
Dernière release
v0.35.1
Dernier push
2026-10-06
llama.cpp

LLM inference in C/C++

Langage
C++
Licence
MIT
Étoiles
130k
Dernière release
v0.6.0
Dernier push
2026-10-06
vLLM

A high-throughput and memory-efficient inference and serving engine for LLMs

Langage
Python
Licence
Apache-2.0
Étoiles
93k
Dernière release
v0.31.0
Dernier push
2026-10-06
LocalAI

LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.

Langage
Go
Licence
MIT
Étoiles
49k
Dernière release
v4.11.0
Dernier push
2026-10-06
SGLang

SGLang is a high-performance serving framework for large language models and multimodal models.

Langage
Python
Licence
Apache-2.0
Étoiles
37k
Dernière release
v0.5.21
Dernier push
2026-10-06
llamafile

Distribute and run LLMs with a single file.

Langage
C++
Licence
Other
Étoiles
26k
Dernière release
0.10.6
Dernier push
2026-10-06
KTransformers

A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

Langage
Python
Licence
Apache-2.0
Étoiles
20k
Dernière release
v0.7.1
Dernier push
2026-10-01
TensorRT-LLM

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

Langage
Python
Licence
Other
Étoiles
15k
Dernière release
v1.2.1
Dernier push
2026-10-06
Xinference

Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.

Langage
Python
Licence
Apache-2.0
Étoiles
9.6k
Dernière release
v3.5.0
Dernier push
2026-10-06
LMDeploy

LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

Langage
Python
Licence
Apache-2.0
Étoiles
8.1k
Dernière release
v0.18.0
Dernier push
2026-09-28
mistral.rs

Fast, flexible LLM inference

Langage
Rust
Licence
MIT
Étoiles
7.7k
Dernière release
v0.9.4
Dernier push
2026-10-01
llama-swap

Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc

Langage
Go
Licence
MIT
Étoiles
5.9k
Dernière release
v262
Dernier push
2026-10-05
Lemonade

Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk

Langage
C++
Licence
Apache-2.0
Étoiles
5.8k
Dernière release
v2026.40.0
Dernier push
2026-10-06
GPUStack

A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.

Langage
Python
Licence
Apache-2.0
Étoiles
5.8k
Dernière release
v2.2.3
Dernier push
2026-09-30
RamaLama

RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.

Langage
Python
Licence
MIT
Étoiles
3.1k
Dernière release
v0.25.0
Dernier push
2026-10-05

Ce que ces outils remplacent