A minimal Vision-Language-Action model you can read: frozen CLIP + a tiny head on ManiSkill PickCube. LeRobot integration. Runs on a Mac, no GPU.
By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).