CPU/RAM layerwise LLM inference in Rust — stream transformer blocks under a hard memory budget with background prefetch (AirLLM for CPU/RAM)
By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).