You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Run Claude Code 100% on-device with local AI on Apple Silicon. MLX-native Anthropic-API server, 65 tok/s Qwen 3.5 122B, Llama 3.3 70B, Gemma 4 31B. Private, offline, airgap-ready. Built for NDA / legal / healthcare workflows.
Plug-and-play local AI studio: uncensored chat, image & video generation, coding agent. Runs abliterated LLMs + ComfyUI 100% offline. One installer, no Docker, no cloud.
Docker Compose for llama.cpp GGUF servers on AMD Strix Halo: Qwen, Gemma, and Laguna packages (abliterated and quantized), stock Vulkan plus ROCmFP4/MTP and ROCmFPX, parallel slots, with prefill/decode and quality metrics measured on this rig.
Abliterated GLM-5.2 with a true 4-bit NVFP4 KV cache on 4x DGX Spark: 316K context, 317,279-token pool (+58.6% vs fp8), 41.4 tok/s peak — with honest content-dependent decode measurements.
Reproducible vLLM recipe for shawnw3i/Huihui-Qwen3.6-27B-abliterated-AWQ-MTP on 2× RTX 3090 in a Proxmox LXC. MTP n=3, 256K context, full vision+tool-calling+reasoning. Silent 24/7 operation at 250W per card. Companion to the base-model recipe.