Get up and running with Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
Local model runtimes
Run open-weight language models on your own hardware, behind a local API.
Mudanças nesta categoria (RSS)
Lado a lado
| Ferramenta | Estrelas | Licença | Termos | Auto-hospedado | Linguagem | Última release | Último push |
|---|---|---|---|---|---|---|---|
| Ollama | 182k | MIT | Código aberto | Sim | Go | v0.35.1 | 2026-10-06 |
| llama.cpp | 130k | MIT | Código aberto | Sim | C++ | v0.6.0 | 2026-10-06 |
| vLLM | 93k | Apache-2.0 | Código aberto | Sim | Python | v0.31.0 | 2026-10-06 |
| LocalAI | 49k | MIT | Código aberto | Sim | Go | v4.11.0 | 2026-10-06 |
| exo | 48k | Apache-2.0 | Código aberto | Sim | Python | v1.0.71 | 2026-10-06 |
| SGLang | 37k | Apache-2.0 | Código aberto | Sim | Python | v0.5.21 | 2026-10-06 |
| llamafile | 26k | Other | Código aberto | Sim | C++ | 0.10.6 | 2026-10-06 |
| MLC LLM | 23k | Apache-2.0 | Código aberto | Sim | Python | v0.20.0 | 2026-10-04 |
| KTransformers | 20k | Apache-2.0 | Código aberto | Sim | Python | v0.7.1 | 2026-10-01 |
| TensorRT-LLM | 15k | Other | Código aberto | Sim | Python | v1.2.1 | 2026-10-06 |
| Xinference | 9.6k | Apache-2.0 | Código aberto | Sim | Python | v3.5.0 | 2026-10-06 |
| LMDeploy | 8.1k | Apache-2.0 | Código aberto | Sim | Python | v0.18.0 | 2026-09-28 |
| mistral.rs | 7.7k | MIT | Código aberto | Sim | Rust | v0.9.4 | 2026-10-01 |
| llama-swap | 5.9k | MIT | Código aberto | Sim | Go | v262 | 2026-10-05 |
| Lemonade | 5.8k | Apache-2.0 | Código aberto | Sim | C++ | v2026.40.0 | 2026-10-06 |
| GPUStack | 5.8k | Apache-2.0 | Código aberto | Sim | Python | v2.2.3 | 2026-09-30 |
| RamaLama | 3.1k | MIT | Código aberto | Sim | Python | v0.25.0 | 2026-10-05 |
| TabbyAPI | 1.5k | AGPL-3.0 | Código aberto | Sim | Python | nenhuma | 2026-10-06 |
18 ferramentas
LLM inference in C/C++
A high-throughput and memory-efficient inference and serving engine for LLMs
- Linguagem
- Python
- Licença
- Apache-2.0
- Estrelas
- 93k
- Última release
- v0.31.0
- Último push
- 2026-10-06
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
Run frontier AI locally.
- Linguagem
- Python
- Licença
- Apache-2.0
- Estrelas
- 48k
- Última release
- v1.0.71
- Último push
- 2026-10-06
SGLang is a high-performance serving framework for large language models and multimodal models.
- Linguagem
- Python
- Licença
- Apache-2.0
- Estrelas
- 37k
- Última release
- v0.5.21
- Último push
- 2026-10-06
Distribute and run LLMs with a single file.
Universal LLM Deployment Engine with ML Compilation
- Linguagem
- Python
- Licença
- Apache-2.0
- Estrelas
- 23k
- Última release
- v0.20.0
- Último push
- 2026-10-04
A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
- Linguagem
- Python
- Licença
- Apache-2.0
- Estrelas
- 20k
- Última release
- v0.7.1
- Último push
- 2026-10-01
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
- Linguagem
- Python
- Licença
- Apache-2.0
- Estrelas
- 9.6k
- Última release
- v3.5.0
- Último push
- 2026-10-06
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
- Linguagem
- Python
- Licença
- Apache-2.0
- Estrelas
- 8.1k
- Última release
- v0.18.0
- Último push
- 2026-09-28
Fast, flexible LLM inference
Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
- Linguagem
- C++
- Licença
- Apache-2.0
- Estrelas
- 5.8k
- Última release
- v2026.40.0
- Último push
- 2026-10-06
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
- Linguagem
- Python
- Licença
- Apache-2.0
- Estrelas
- 5.8k
- Última release
- v2.2.3
- Último push
- 2026-09-30
RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.
The official API server for Exllama. OAI compatible, lightweight, and fast.