Skip to content
#

ollama-alternative

Here are 21 public repositories matching this topic...

Rapid-MLX

The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.

  • Updated Jul 31, 2026
  • Python

⚡️ The fastest way to run local LLMs on Apple Silicon — sub-second model loads, beats Ollama on throughput, tail latency, and full-response time. OpenAI/Ollama-compatible. No cloud, no API keys.

  • Updated Jul 29, 2026
  • Python

OllamoMUI — #1 FREE Ollama alternative: local LLM proxy to 26 free models (OpenRouter/OpenAI/Anthropic/Groq/DeepSeek/Gemini) with RAG, memory & dashboard. Built by Rakibul Hasan (Rhasan@dev), a Dhaka, Bangladesh full-stack dev open to remote/WFH/freelance. Hire me for full-stack, backend & AI/LLM work.

  • Updated Jul 27, 2026
  • TypeScript
gemma-4-12b

One-click installer for Google Gemma 4 12B — optimized to run on 8GB laptops, 1M context, native multimodal. No terminal. Win/Mac, MIT.

  • Updated Jul 14, 2026
  • Python

Improve this page

Add a description, image, and links to the ollama-alternative topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the ollama-alternative topic, visit your repo's landing page and select "manage topics."

Learn more