gitaskhub

Multi-dataset evaluation of Anthropic's J-Space on Qwen3-4B. Demonstrates that internal 'workspace noise' can reliably route high-confidence confabulations, while mapping the structural limitations of the metric across adversarial, reasoning, and multiple-choice tasks.

Stars · 13
Language · Jupyter Notebook
License · MIT
Ask anything about this repo to start.
Full explanation on explaingit →

By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).