Jacobian-Brainwash : A manual alignment tool for large language models built on Anthropic's Jacobian Lens. Results are exportable.
-
Updated
Jul 20, 2026 - Python
Jacobian-Brainwash : A manual alignment tool for large language models built on Anthropic's Jacobian Lens. Results are exportable.
Preregistered Jacobian-lens reliability research on Gemma 4: 25k prompts, frozen probes, public traces, a prospective transfer miss, and an answer-identity confound.
Watch your model think while your app talks to it: an OpenAI-compatible server with a live layer-by-layer view of the residual stream (logit lens, J-lens, steering, traces).
Agents that talk through model internals — activations & J-space — instead of text. No words pass between them. The lab on top of brainscope + hidden-directions.
LLM decision boundary; LLM reasoning trajectory; Geometric Lens; Laguerre Geometry
EVict and recOver KV cache Entries. Selective KV cache eviction and recovery for long-context LLM inference.
Anthropic's Jacobian Lens, ported to Mac. The official reference implementation targets CUDA GPUs — this fork runs it on Apple silicon (PyTorch MPS) and adds a local interactive explorer for Qwen with plain-language interpretability demos and a live cause-and-effect intervention.
A Jacobian-Lens (J-Lens) observer for vision-language models — read what a VLM is poised to say, before it says it. Multimodal J-Lens on Qwen3.5, concept-race, and a forward-only prompt helper.
Watch a language model's thoughts form before it speaks — interactive workbench + findings for Anthropic's Jacobian lens
Anthropic's jlens for single or multi-model chat replay, with a web interface.
Probing LLM J-spaces through the Jacobian lens on one RTX 3090 — 421-record research data dump (Units 0–15, incl. workspace-span batteries) with live dashboard, films, and per-record commentary
J-lens for audio-input LLMs with reproducible mixed text and audio fitting
Autonomous J-space search: causal concept-swap interventions inside a language model, built on the Jacobian lens (Anthropic, 2026).
Jacobian Lens experiments for language model interpretability.
Local-first macOS desktop application for fitting, applying, and exploring Jacobian Lenses on open-weight decoder language models.
Browser visualizer for Neuronpedia Jacobian-lens (J-space) chat exports: watch the concepts a model holds across its layers, a beat before each word.
See what your open model is really thinking — and whether it's making it up. J-space introspection + a drop-in OpenAI-compatible hallucination signal for any HuggingFace model.
Locating and editing refusal in the J-space workspace with the Jacobian lens: refusal is legible ~10 layers before the first token, and only ~1/3 lives in the verbalizable workspace.
Personal LLM interpretability research: how much of Anthropic's global-workspace/J-space findings transfer to small open-weights models (Jacobian lens experiments)
Jacobian Lens (J-Lens) and J-Space Toolkit for transformer mechanistic interpretability. Train linear lenses, decompose hidden states, and run causal interventions on decoder-only models.
Add a description, image, and links to the jacobian-lens topic page so that developers can more easily learn about it.
To associate your repository with the jacobian-lens topic, visit your repo's landing page and select "manage topics."