gitaskhub

CPU/RAM layerwise LLM inference in Rust — stream transformer blocks under a hard memory budget with background prefetch (AirLLM for CPU/RAM)

Stars · 4
Language · Rust
License · MIT
Ask anything about this repo to start.
Full explanation on explaingit →

By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).