Change the repository type filter
All
Repositories list
13 repositories
- Instrumented agent harness for capturing SAE feature trajectories during SWE-bench Pro traces on Qwen3.6-27B (mech anatomy of agent reasoning failure)
web
PublicNext.js site for OpenInterpretability — the umbrella org for mechreward and public hybrid-architecture SAEsopeninterp-mcp
Publicprobebench-registry
Publicagentguard
Publicopeninterp-lab
Publicdecision-locator
PublicFind the layer where a language model commits a decision — and steer it. Any open-weight HF model. (WANDERING arc paper #6).github
Publicregistry
Publiccli
Publicopeninterp — Python SDK + CLI. FabricationGuard hallucination probe + ProbeBench leaderboard + Atlas search + Trace generation. pip install openinterpmechreward
PublicMechanistic interpretability as reward signal for RL training of LLMs — SAE features + GRPO + anti-Goodhart frameworknotebooks
PublicTrain your first SAE in 30 min → paper-grade at 27B. Free Colab · free Kaggle · cloud ladders. Every scale covered.
ProTip! When viewing an organization's repositories, you can use the
props. filter to filter by custom property.