Fully client-side VRAM planner for local LLMs — will your GGUF fit? Parses GGUF headers in-browser and models llama.cpp weight/KV/compute allocation. Zero runtime deps.
By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).