Fast, flexible LLM inference
Runs text, vision and speech models locally behind an OpenAI-compatible API, without a desktop app.
alternatives to
Desktop app to download and run open-weight language models locally, with a chat window and a local API server. A closed product from Element Labs.
3 alternatives
Fast, flexible LLM inference
Runs text, vision and speech models locally behind an OpenAI-compatible API, without a desktop app.
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
A local server and app for running models, tuned for AMD GPUs and NPUs.
RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.
Pulls and runs models in containers from the command line, behind an OpenAI-compatible API.