VILA - a multi-image visual language model with training, inference and evaluation recipe, deployable from cloud to edge (Jetson Orin and laptops)
By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).