Content warning: this tool displays real-world text that was flagged as harmful. Expect shocking language.
Where the Lines Are is a visualization tool for studying how content moderation categories overlap, co-occur, and cluster in real prompt data. It treats a classification dataset not as a lookup table but as a structure worth seeing — a place where patterns in how humans produce harmful text become visible through careful graphic design.
The interface follows Edward Tufte's principles from The Visual Display of Quantitative Information: maximize data density, eliminate chartjunk, label everything directly, and let the data speak through its structure rather than through decoration. Every pixel either carries data or gets out of the way.
Eight labeled datasets ship with the tool, spanning 2018–2025, plus one taxonomy-only deployment-classifier layer (2025–26) — tracing the arc from comment-section toxicity detection through the current frontier paradigm where classifiers intercept model outputs on a CBRN/biosecurity axis:
| Dataset | Source | Year | Rows | Categories | License |
|---|---|---|---|---|---|
| Jigsaw Toxic Comments | Google Jigsaw | 2018 | 32,450 | 6 | CC0 |
| OpenAI Moderation | OpenAI moderation-api-release | 2022 | 1,680 | 8 | MIT |
| BeaverTails | PKU-Alignment | 2023 | 300,567 | 14 | CC-BY-NC-4.0 |
| PKU-SafeRLHF | PKU-Alignment | 2024 | 38,640 | 19 | CC-BY-NC-4.0 |
| NVIDIA Aegis v2 | NVIDIA | 2024 | 29,095 | 23 | CC-BY-4.0 |
| AIR-Bench 2024 | Stanford CRFM (arXiv 2407.17436) | 2024 | 5,694 | 16 | CC-BY-4.0 |
| HarmBench | CAIS (arXiv 2402.04249) | 2024 | 400 | 7 | MIT |
| MLCommons AILuminate v1.0 | MLCommons (arXiv 2503.05731) | 2025 | 1,200 | 12 | CC-BY-4.0 |
| Anthropic Constitutional Classifiers (taxonomy only) | Anthropic (arXiv 2501.18837) | 2025–26 | — | 4 (CBRN) | N/A |
The first eight share the same multi-label binary structure (each row has one or more flagged categories) but slice content moderation differently. Jigsaw uses 6 behavioral categories (toxic, obscene, insult, threat) from Wikipedia comments — the pre-AI-safety worldview. OpenAI introduced hierarchical severity (hate → hate/threatening). BeaverTails added financial crime, terrorism, and privacy. SafeRLHF expanded to 19 categories including cybercrime, mental manipulation, and environmental damage. Aegis reached 23 with profanity, malware, and unauthorized advice. The taxonomy evolution is the story — click a dataset panel at the top to switch; all visualizations rebuild from scratch.
Across the first five datasets — 70 categories in total — there is no biology or CBRN category at all. Yet that is precisely the axis frontier labs now draw lines on in production: real-time classifiers that read model outputs and block chemical, biological, radiological, and nuclear (CBRN) uplift. The 2024–26 layer closes that gap:
- AIR-Bench 2024 (Stanford CRFM) is a policy-derived taxonomy — 5,694 prompts distilled from 8 government regulations and 16 company policies into a 4-level tree (4 → 16 → 45 → 314). Rolled up to its 16 Level-2 categories, it is the first labeled dataset here whose taxonomy reaches Weapon Usage & Development (bioweapons, chemical, nuclear, radiological) and a dedicated cyber/Security Risks axis. It is single-label at the leaf, so its co-occurrence is near-diagonal by design — its contribution is taxonomic breadth, not co-activation structure.
- HarmBench (CAIS, Feb 2024) is an automated-red-teaming behavior set — 400 harmful instructions used to measure attack success and refusal robustness, each labeled with exactly one of seven semantic categories. Its chemical & biological category (56 behaviors) gives a third measured column on the CBRN/biosecurity axis. Single-label by construction (100% exclusivity, ~0% co-occurrence), and the behaviors are elicitation prompts — requests, not recipes — so the full text ships (MIT, already public on GitHub).
- MLCommons AILuminate v1.0 ships a public 1,200-prompt DEMO set (a 10% practice subset, CC-BY-4.0) labeled across the 12-hazard AIRR taxonomy — including a dedicated, measured Indiscriminate Weapons (CBRNE) hazard (100 prompts).
- Anthropic Constitutional Classifiers is a taxonomy-only layer: a real frontier deployment classifier whose CBRN categories (chemical, biological, radiological, nuclear) are published, but whose row-level corpus is not public. It is rendered as a labeled column in the Rosetta crosswalk and Drift timeline, badged "taxonomy only," appears greyed-out (non-selectable) in the dataset picker, and carries no counts, co-occurrence, or statistics — the honest representation of a category list with no measurable data behind it (see Measured vs. taxonomy-only, below).
Moderation categories are not independent. The visualizations expose their hidden geometry — differently for each dataset.
Some categories never travel alone. In Jigsaw, "severe toxic" is never flagged in isolation — it always co-occurs with other categories. In OpenAI, sexual/minors, hate/threatening, and violence/graphic behave identically — they appear only when a parent category is also flagged. Self-harm, by contrast, is 92% exclusive. Switch to BeaverTails and the profile changes: animal abuse is highly exclusive while discrimination is almost always shared.
Violence is the connective tissue of harm. The co-occurrence matrix shows that violence co-occurs with nearly every other category. Click the violence row and the word frequencies shift to "kill," "destroy," "war." Sexual content barely touches violence. These categories live in different neighborhoods — visible across all three datasets.
Word distributions reveal category boundaries. Click a word in the frequency strip and the breakdown panel shows how that word distributes across categories. The proportional bars make cross-category signatures immediately comparable.
Rare combinations are the most informative. The surprise metric sorts prompts by the rarity of their category combination. Edge cases reveal where category boundaries blur and where annotators were forced to make judgment calls across multiple dimensions simultaneously.
The binary matrix shows population structure. Each row of data becomes a thin strip of dark and light cells. Vertical dark bands show which categories dominate; horizontal patterns reveal clusters; scattered dark cells mark outliers. BeaverTails renders 300K rows at full density.
Cross-dataset comparison reveals taxonomy design choices. The Rosetta Stone table maps ~20 harm concepts across all nine taxonomies, showing how "privacy" becomes "PII/privacy" in Aegis, or how "minors" is split from "sexual" in some taxonomies but merged in others. The new CBRN / biosecurity row stays empty (—) across all five legacy datasets and only fills in for AIR-Bench, HarmBench, AILuminate, and the Anthropic deployment classifier — making the arrival of the bio axis visible at a glance. The Drift timeline shows this evolution chronologically — bold entries mark concepts appearing for the first time.
Measured vs. taxonomy-only. The tool draws a hard line between datasets with a real labeled corpus (which get counts, co-occurrence, exclusivity, and every other statistic) and taxonomy-only layers — published category lists with no public row data. Taxonomy-only entries appear only in the Rosetta crosswalk and Drift timeline, are badged in amber, and have no statistics computed or shown. Absence of row data is rendered as absence, never inflated into false density.
Beyond labeling — WMDP. A standalone panel below the Drift timeline contrasts the labeling taxonomies above with a capabilities benchmark: WMDP (Weapons of Mass Destruction Proxy, CAIS 2024, MIT), 3,668 multiple-choice questions measuring whether a model knows hazardous facts across biosecurity (1,273), chemical security (408), and cybersecurity (1,987). It is not a moderation taxonomy, so it never appears as a Rosetta column; and because its items are questions with answers — recipe-shaped, not request-shaped — the panel renders only domain counts, never the question content. It is the eval-side counterpoint to the deployment-classifier taxonomies: "what a model knows" versus "how prompts are labeled."
Annotators disagree more than you'd expect. The Split Verdict chart (SafeRLHF) shows that two independently classified responses disagree 8% of the time on privacy, but only 1% on trafficking. The safer response is not the better one 24% of the time. For Aegis, human labels agree perfectly while LLM jury labels diverge 36% between prompt and response safety.
Same prompt, different labels. 6,640 prompts appear in two or more datasets. The Doppelganger feature marks these in the results table — click to see how each taxonomy classified the identical text. The Consensus chart summarizes concept-level agreement: hate and privacy get 60%+ agreement across datasets, while toxicity and harassment get 0% (concepts that only some datasets track).
Beyond labels: what embeddings see that categories don't. Every statistic above lives in label space — it describes how categories relate to each other, never what the underlying text actually says. Two panels below the Consensus chart look at content directly, via offline sentence-embedding clustering (UMAP + HDBSCAN over MiniLM, computed once per dataset, no ML at runtime):
- Category coherence shows, per concept, what share of its prompts land in a single dominant embedding cluster. Narrow categories (trafficking, ransomware-specific cybercrime) cluster tightly; broad umbrella categories (cybercrime, harassment, fraud) turn out to cover many topically unrelated request types that just happen to share a taxonomy bucket — something a co-occurrence matrix cannot show, since it only ever looks at how labels relate to each other.
- Annotation outliers ranks labeled prompts by how much their 15 nearest embedding neighbors disagree with their label — flagging likely mislabels and prompts that lost their context (e.g. a single conversational turn sliced out of a longer multi-turn exchange). It is an item-level, browsable signal that the aggregate Split-Verdict/Consensus charts structurally cannot produce, since those never look at content, only at label agreement rates.
Both are clustered per dataset (each corpus gets its own embedding neighborhood, not one combined across all eight). BeaverTails (300K+ rows) is clustered from a stratified 50,000-row sample instead of the full corpus — rare categories kept in full, common ones thinned, the same shape of tradeoff already used elsewhere in this tool to keep things tractable on a laptop — and both panels say so when a dataset was sampled this way.
The visualizations apply Tufte's principles throughout: high data-ink ratio, direct labeling, small multiples, data-text integration, grayscale palette, and zero external dependencies. Every chart is rendered in purpose-built canvas code with no frameworks. All visualizations adapt their sizing, font, and layout automatically for 6 to 23 categories. A single amber accent is the one departure from pure grayscale — reserved exclusively for marking taxonomy-only data (the absence of a labeled corpus), so the encoding itself signals "no measured data here."
See principles.md for the full design rationale with specific Tufte page citations.
Open index.html in a browser. Everything updates reactively — click a category, click a matrix cell, click a word, type a search, toggle a pill. No server required for the smaller datasets (works via file://), though python3 -m http.server is recommended to load BeaverTails (51MB). A service worker caches datasets after first load for offline use.
Cache busting is automatic. The service worker's CACHE_NAME is a content hash of every file in URLS_TO_CACHE, derived by scripts/preprocess.py — every preprocess command ends by re-deriving it, so regenerating a dataset bumps the cache and a no-op regen doesn't. After hand-editing a precached UI file (index.html, static/*, dataset-loader.js), run python3 scripts/preprocess.py sw; tests/data/sw-cache.test.js (part of npm test) fails if CACHE_NAME is ever stale.
index.html Main page
dataset-loader.js Registry + dataset + xref loading (file:// and HTTP)
static/vis.js All canvas visualizations (~1000 lines)
static/styles.css Tufte-inspired stylesheet
datasets/
registry.json Dataset manifest with schemas, concepts, stats
xref.json Cross-dataset prompt matching index (6,640 entries, 1.2 MB)
xref-fuzzy.json Fuzzy (near-duplicate) cross-dataset matches
jigsaw.json 32,450 rows (13 MB)
openai.json 1,680 rows (1.2 MB)
beavertails.json 300,567 rows (51 MB)
saferlhf.json 38,640 rows (26 MB, includes divergence fields)
aegis.json 29,095 rows (14 MB, includes divergence fields)
airbench.json 5,694 rows (5 MB, AIR-Bench 2024 L2 rollup)
ailuminate.json 1,200 rows (0.3 MB, AILuminate v1.0 DEMO)
harmbench.json 400 rows (0.1 MB, HarmBench behaviors, single-label)
xref-semantic.json Embedding (semantic) cross-dataset matches
<id>-coherence.json Per concept: top-cluster concentration + example prompts
<id>-outliers.json Top ~200 least-homogeneous labeled prompts, ranked
*.js JS wrappers for file:// protocol
(Anthropic Constitutional Classifiers is taxonomy-only —
it lives in registry.json with rows:null, no data file.)
scripts/
preprocess.py One-time HuggingFace/GitHub → JSON pipeline + stats + xref +
embedding-clustering precompute (embed/cluster/noise-lens)
concepts.py Rosetta concept-crosswalk loader shared by the clustering steps
tests/
unit/ Pure-function unit tests (node --test)
data/ Independent data-validation for AIR-Bench, AILuminate,
HarmBench, and the taxonomy-only integrity invariants
browser/ Playwright smoke tests