gitaskhub

Mechanistic interpretability as reward signal for RL training of LLMs — SAE features + GRPO + anti-Goodhart framework

Stars · 6
Language · Jupyter Notebook
License · Apache-2.0
Ask anything about this repo to start.

By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).