gitaskhub

Implementation of the paper: Beyond Correlation: The impact of human uncertainty in measuring the effectiveness of automatic evaluation and LLM-as-a-judge

Stars · 13
Language · Python
License · Apache-2.0
Ask anything about this repo to start.

By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).