The Triton Inference Server provides an optimized cloud and edge inferencing solution.
- Language
- Python
- Licence
- BSD-3-Clause
- Stars
- 11k
- Latest
- v2.73.0
- Last push
- 2026-10-06
Serve trained machine learning models behind an API, with batching, scaling and versioning.
Changes in this category (RSS)
| Tool | Stars | Licence | Terms | Self-hosted | Language | Latest | Last push |
|---|---|---|---|---|---|---|---|
| Triton Inference Server | 11k | BSD-3-Clause | Open source | Yes | Python | v2.73.0 | 2026-10-06 |
| BentoML | 8.9k | Apache-2.0 | Open source | Yes | Python | v1.4.39 | 2026-10-05 |
| KServe | 6.1k | Apache-2.0 | Open source | Yes | Go | v0.21.0 | 2026-10-06 |
| MLServer | 900 | Apache-2.0 | Open source | Yes | Python | 1.7.1 | 2026-10-05 |
4 tools
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
An inference server for your machine learning models, including support for multiple frameworks, multi-model serving and more