llm-server
Here are 20 public repositories matching this topic...
Please see the newer: https://github.com/Quazmoz/openvino-windows-llm
-
Updated
May 20, 2026 - Python
Windows-first OpenAI-compatible local LLM server powered by OpenVINO GenAI for Intel CPU/GPU/NPU, with chat UI, model conversion, and setup scripts.
-
Updated
Aug 5, 2026 - Python
A lightweight LiteLLM server boilerplate pre-configured with uv and Docker for hosting your own OpenAI- and Anthropic-compatible endpoints. Includes LibreChat as an optional web UI.
-
Updated
Dec 8, 2025 - Python
single-executable / library which combines llama.cpp, whisper.cpp, and stable-diffusion.cpp
-
Updated
Aug 6, 2026 - C++
Function-calling API for LLM from multiple providers
-
Updated
Aug 10, 2024 - Go
-
Updated
May 27, 2026 - TypeScript
macOS GUI for managing pure mlx_lm.server on Apple Silicon in Direct Mode.
-
Updated
Jul 26, 2026 - Swift
A complete, menu-driven AI model interface for Windows that simplifies running local GGUF language models with llama.cpp. This tool automatically manages dependencies, provides multiple interaction modes, and prioritizes user privacy through fully offline operation.
-
Updated
Jan 30, 2026 - PowerShell
API server for `llm` CLI tool
-
Updated
Aug 12, 2025 - Python
PHP Frontend for Hosting local LLM's (run via VSCode or basic php execution methods/ add to project)
-
Updated
Jul 13, 2025 - PHP
Headless CLI for managing local MLX language-model HTTP servers on Apple Silicon Macs. Supports model discovery, server lifecycle management, performance benchmarking, and provider integration with OpenCode, Claude Code, and LiteLLM.
-
Updated
Jun 29, 2026 - Python
OpenAI-compatible local inference server for Apple Silicon using MLX. FastAPI server with Chat Completions and Responses APIs, multi-turn conversations, and streaming support.
-
Updated
Mar 7, 2026 - Python
A unified Monolithic API Server for Ollama that provides the architecture for System Mapping, Task Farming, Local Consultation, and Tool Orchestration.
-
Updated
Jul 13, 2026 - Python
A flexible FastAPI-based framework for handling AI tasks using Large Language Models (LLMs). Supports multiple providers, extensible tasks and routers, Redis caching, and OpenAI integration. Easily scalable for various LLM-based applications.
-
Updated
Sep 3, 2024 - Python
Turn a fresh Apple Silicon Mac into a tuned local-LLM dev server in an evening (MIT)
-
Updated
Aug 5, 2026 - TypeScript
Host an LLM and make it accessible on a network via API.
-
Updated
May 12, 2026 - Python
Unified simple LLM server wrapper with intelligent routing based on model ID
-
Updated
Aug 6, 2026 - Python
A lightweight, zero-cost FastAPI server for deploying GGUF models via llama.cpp, featuring streaming support and an OpenAI-compatible API.
-
Updated
Jul 7, 2026 - Python
Run local AI models in VS Code with automatic model detection, server start, and built-in MCP endpoint—no cloud or manual setup required.
-
Updated
Aug 6, 2026 - PHP
Improve this page
Add a description, image, and links to the llm-server topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the llm-server topic, visit your repo's landing page and select "manage topics."