Where the Lines Are
Rosetta stone
How eight taxonomies slice the same space — five legacy corpora, two frontier labeled datasets (AIR-Bench, AILuminate), and one taxonomy-only deployment classifier. Each row is a concept; columns are datasets.
Drift
The taxonomy grew from 6 categories (2018) through 23-category safety suites (2024) to frontier CBRN deployment classifiers (2025–26). Bold = first appearance. The amber-badged column is taxonomy-only — a published category list with no public corpus.
Exclusivity trend
Dark bar = avg. category exclusivity. Light bar = multi-label rate. Higher exclusivity = categories more often flagged alone. The jump to 100% exclusivity in the 2024–26 columns is a genre shift, not a labeling trend: the field moved from labeling found content (multi-label corpora) to authoring benchmark suites that assign exactly one category per prompt.
Beyond labeling — WMDP
Every column above is a labeling taxonomy — it sorts prompts into harm categories. WMDP (Weapons of Mass Destruction Proxy) measures something different: whether a model knows hazardous facts, via 3,668 multiple-choice questions across three security domains. It is a capabilities benchmark, not a moderation taxonomy, so it sits on its own here rather than as a Rosetta column. No question content is shown — these are hazardous-knowledge items (questions with correct answers), so the panel renders only structure and counts.
Source: CAIS, 2024 (arXiv 2403.03218) · MIT license · 3,668 questions — biosecurity 1,273 · chemical 408 · cybersecurity 1,987.
Concept comparison
How the same concept manifests across datasets. Dark bar = % flagged. Light bar = exclusivity ratio. Single-label benchmarks (AIR-Bench, HarmBench, AILuminate) always show 100% exclusivity by construction.
Contains potentially disturbing text.
Category definitions
| Category | Label | Definition |
|---|
Category counts
Category exclusivity
Dark = flagged for this category alone. Light = shared with other categories.
Click a category to focus on it. Click the matrix below to explore overlaps.
Pairwise co-occurrences
Darker cells = more prompts flagged for both categories. Click any cell to filter.
Network view
Force-directed layout. Node size = frequency. Edge darkness = co-occurrence strength. Only above-median edges shown.
Active filters order: schema
Word frequencies
Click a word to see which categories it appears in.
Prompt binary matrix
Each row is one prompt. Dark cells = category flagged.
0 Matching prompts
Sorted by surprise when unfiltered. Unique combinations appear first.
| Prompt | Surprise | Categories |
|---|