Implementation of a memory efficient multi-head attention as proposed in the paper, "Self-attention Does Not Need O(n²) Memory"
By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).