← 戻る

カテゴリ

Local model runtimes

Run open-weight language models on your own hardware, behind a local API.

比較表

Local model runtimes
ツールスターライセンス利用条件セルフホスト言語最新リリース最終プッシュ
Ollama182kMITオープンソース可Gov0.35.12026-10-06
llama.cpp130kMITオープンソース可C++v0.6.02026-10-06
vLLM93kApache-2.0オープンソース可Pythonv0.31.02026-10-06
LocalAI49kMITオープンソース可Gov4.11.02026-10-06
exo48kApache-2.0オープンソース可Pythonv1.0.712026-10-06
SGLang37kApache-2.0オープンソース可Pythonv0.5.212026-10-06
llamafile26kOtherオープンソース可C++0.10.62026-10-06
MLC LLM23kApache-2.0オープンソース可Pythonv0.20.02026-10-04
KTransformers20kApache-2.0オープンソース可Pythonv0.7.12026-10-01
TensorRT-LLM15kOtherオープンソース可Pythonv1.2.12026-10-06
Xinference9.6kApache-2.0オープンソース可Pythonv3.5.02026-10-06
LMDeploy8.1kApache-2.0オープンソース可Pythonv0.18.02026-09-28
mistral.rs7.7kMITオープンソース可Rustv0.9.42026-10-01
llama-swap5.9kMITオープンソース可Gov2622026-10-05
Lemonade5.8kApache-2.0オープンソース可C++v2026.40.02026-10-06
GPUStack5.8kApache-2.0オープンソース可Pythonv2.2.32026-09-30
RamaLama3.1kMITオープンソース可Pythonv0.25.02026-10-05
TabbyAPI1.5kAGPL-3.0オープンソース可Pythonなし2026-10-06

18件のツール

Ollama

Get up and running with Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.

言語
Go
ライセンス
MIT
スター
182k
最新リリース
v0.35.1
最終プッシュ
2026-10-06
llama.cpp

LLM inference in C/C++

言語
C++
ライセンス
MIT
スター
130k
最新リリース
v0.6.0
最終プッシュ
2026-10-06
vLLM

A high-throughput and memory-efficient inference and serving engine for LLMs

言語
Python
ライセンス
Apache-2.0
スター
93k
最新リリース
v0.31.0
最終プッシュ
2026-10-06
LocalAI

LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.

言語
Go
ライセンス
MIT
スター
49k
最新リリース
v4.11.0
最終プッシュ
2026-10-06
SGLang

SGLang is a high-performance serving framework for large language models and multimodal models.

言語
Python
ライセンス
Apache-2.0
スター
37k
最新リリース
v0.5.21
最終プッシュ
2026-10-06
llamafile

Distribute and run LLMs with a single file.

言語
C++
ライセンス
Other
スター
26k
最新リリース
0.10.6
最終プッシュ
2026-10-06
MLC LLM

Universal LLM Deployment Engine with ML Compilation

言語
Python
ライセンス
Apache-2.0
スター
23k
最新リリース
v0.20.0
最終プッシュ
2026-10-04
KTransformers

A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

言語
Python
ライセンス
Apache-2.0
スター
20k
最新リリース
v0.7.1
最終プッシュ
2026-10-01
TensorRT-LLM

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

言語
Python
ライセンス
Other
スター
15k
最新リリース
v1.2.1
最終プッシュ
2026-10-06
Xinference

Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.

言語
Python
ライセンス
Apache-2.0
スター
9.6k
最新リリース
v3.5.0
最終プッシュ
2026-10-06
LMDeploy

LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

言語
Python
ライセンス
Apache-2.0
スター
8.1k
最新リリース
v0.18.0
最終プッシュ
2026-09-28
mistral.rs

Fast, flexible LLM inference

言語
Rust
ライセンス
MIT
スター
7.7k
最新リリース
v0.9.4
最終プッシュ
2026-10-01
llama-swap

Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc

言語
Go
ライセンス
MIT
スター
5.9k
最新リリース
v262
最終プッシュ
2026-10-05
Lemonade

Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk

言語
C++
ライセンス
Apache-2.0
スター
5.8k
最新リリース
v2026.40.0
最終プッシュ
2026-10-06
GPUStack

A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.

言語
Python
ライセンス
Apache-2.0
スター
5.8k
最新リリース
v2.2.3
最終プッシュ
2026-09-30
RamaLama

RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.

言語
Python
ライセンス
MIT
スター
3.1k
最新リリース
v0.25.0
最終プッシュ
2026-10-05

これらのツールが置き換えるもの