Deep MoE is an all-new mixture-of-experts architecture a transformer layer that does its thinking in a compressed thought space rather than on the wide residual stream.
By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).