Mechanistic interpretability as reward signal for RL training of LLMs — SAE features + GRPO + anti-Goodhart framework
By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).