← volver

categoría

Local model runtimes

Run open-weight language models on your own hardware, behind a local API.

Comparativa

Local model runtimes
HerramientaEstrellasLicenciaCondicionesAutoalojableLenguajeÚltima versiónÚltimo push
Ollama182kMITCódigo abiertoSíGov0.35.12026-10-06
llama.cpp130kMITCódigo abiertoSíC++v0.6.02026-10-06
vLLM93kApache-2.0Código abiertoSíPythonv0.31.02026-10-06
LocalAI49kMITCódigo abiertoSíGov4.11.02026-10-06
exo48kApache-2.0Código abiertoSíPythonv1.0.712026-10-06
SGLang37kApache-2.0Código abiertoSíPythonv0.5.212026-10-06
llamafile26kOtherCódigo abiertoSíC++0.10.62026-10-06
MLC LLM23kApache-2.0Código abiertoSíPythonv0.20.02026-10-04
KTransformers20kApache-2.0Código abiertoSíPythonv0.7.12026-10-01
TensorRT-LLM15kOtherCódigo abiertoSíPythonv1.2.12026-10-06
Xinference9.6kApache-2.0Código abiertoSíPythonv3.5.02026-10-06
LMDeploy8.1kApache-2.0Código abiertoSíPythonv0.18.02026-09-28
mistral.rs7.7kMITCódigo abiertoSíRustv0.9.42026-10-01
llama-swap5.9kMITCódigo abiertoSíGov2622026-10-05
Lemonade5.8kApache-2.0Código abiertoSíC++v2026.40.02026-10-06
GPUStack5.8kApache-2.0Código abiertoSíPythonv2.2.32026-09-30
RamaLama3.1kMITCódigo abiertoSíPythonv0.25.02026-10-05
TabbyAPI1.5kAGPL-3.0Código abiertoSíPythonninguna2026-10-06

18 herramientas

Ollama

Get up and running with Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.

Lenguaje
Go
Licencia
MIT
Estrellas
182k
Última versión
v0.35.1
Último push
2026-10-06
llama.cpp

LLM inference in C/C++

Lenguaje
C++
Licencia
MIT
Estrellas
130k
Última versión
v0.6.0
Último push
2026-10-06
vLLM

A high-throughput and memory-efficient inference and serving engine for LLMs

Lenguaje
Python
Licencia
Apache-2.0
Estrellas
93k
Última versión
v0.31.0
Último push
2026-10-06
LocalAI

LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.

Lenguaje
Go
Licencia
MIT
Estrellas
49k
Última versión
v4.11.0
Último push
2026-10-06
SGLang

SGLang is a high-performance serving framework for large language models and multimodal models.

Lenguaje
Python
Licencia
Apache-2.0
Estrellas
37k
Última versión
v0.5.21
Último push
2026-10-06
llamafile

Distribute and run LLMs with a single file.

Lenguaje
C++
Licencia
Other
Estrellas
26k
Última versión
0.10.6
Último push
2026-10-06
KTransformers

A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

Lenguaje
Python
Licencia
Apache-2.0
Estrellas
20k
Última versión
v0.7.1
Último push
2026-10-01
TensorRT-LLM

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

Lenguaje
Python
Licencia
Other
Estrellas
15k
Última versión
v1.2.1
Último push
2026-10-06
Xinference

Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.

Lenguaje
Python
Licencia
Apache-2.0
Estrellas
9.6k
Última versión
v3.5.0
Último push
2026-10-06
LMDeploy

LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

Lenguaje
Python
Licencia
Apache-2.0
Estrellas
8.1k
Última versión
v0.18.0
Último push
2026-09-28
mistral.rs

Fast, flexible LLM inference

Lenguaje
Rust
Licencia
MIT
Estrellas
7.7k
Última versión
v0.9.4
Último push
2026-10-01
llama-swap

Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc

Lenguaje
Go
Licencia
MIT
Estrellas
5.9k
Última versión
v262
Último push
2026-10-05
Lemonade

Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk

Lenguaje
C++
Licencia
Apache-2.0
Estrellas
5.8k
Última versión
v2026.40.0
Último push
2026-10-06
GPUStack

A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.

Lenguaje
Python
Licencia
Apache-2.0
Estrellas
5.8k
Última versión
v2.2.3
Último push
2026-09-30
RamaLama

RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.

Lenguaje
Python
Licencia
MIT
Estrellas
3.1k
Última versión
v0.25.0
Último push
2026-10-05

Lo que sustituyen estas herramientas