gitaskhub

A lean, fully-tested LLM inference server for the hardware you already have — free tiers, shared VPS, 2-core ARM boxes. OpenAI-compatible API on llama.cpp. On a CPU, never compute the same thing twice: it caches prompts, prefixes and past generations to disk, so request #100 costs a fraction of request #1.

Stars · 94
Language · C++
License · MIT
Ask anything about this repo to start.

By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).