Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
36 commits
Select commit Hold shift + click to select a range
c649d9b
Harden G-protein and ligand annotation fidelity and unify curation ga…
iskoldt-X Jun 23, 2026
2544d33
Clean up lint in the generic-numbering table build script; add site_r…
iskoldt-X Jun 23, 2026
4f5c518
Make site-ref builder portable and guard against developer-local paths
iskoldt-X Jun 23, 2026
78bb707
Apply the G protein terminology convention to comments and docs
iskoldt-X Jun 29, 2026
f074d34
Normalize G protein terminology in user-visible output and review rou…
iskoldt-X Jun 29, 2026
ff40d22
Normalize G protein terminology in model-facing prompt and schema text
iskoldt-X Jun 29, 2026
32e0f64
Rename g_proteins.csv alpha columns to Alpha_identity / Alpha_alpha5_…
iskoldt-X Jun 30, 2026
6c0dea3
Fix six curator-flagged CSV output defects across ligands, partners, …
iskoldt-X Jul 1, 2026
21ea1b6
Refine G-protein subunit recovery and binder renaming: merge same-sub…
iskoldt-X Jul 1, 2026
34340d6
Update oligomer.py
iskoldt-X Jul 1, 2026
2f2574b
Make ligands.csv Residue_seq_id self-describing: prefix each residue …
iskoldt-X Jul 1, 2026
2be90eb
Exclude polymer residues from ligand copy selection
iskoldt-X Jul 2, 2026
67134d3
detect: carry each ligand copy's author identifier (auth_asym_id:auth…
iskoldt-X Jul 3, 2026
68f9945
annotate: add per-copy ligand site/role schema and prompt (ligand_cop…
iskoldt-X Jul 3, 2026
701db8d
annotate: validate per-copy ligand coverage each run (retry then degr…
iskoldt-X Jul 4, 2026
21751b7
aggregate: vote per-copy ligand site/role keyed by copy identifier
iskoldt-X Jul 4, 2026
91be41f
csv: fill each binding-site row's residues from per-copy votes
iskoldt-X Jul 4, 2026
6fef06e
annotate: add sequential batch submission mode and batch robustness f…
iskoldt-X Jul 5, 2026
e453b95
Harden annotation output and review gating
iskoldt-X Jul 7, 2026
1857543
curate: suppress pre-selected default at genuine identity disagreemen…
iskoldt-X Jul 8, 2026
3fb45c7
aggregate: rebuild small-molecule ligand rows from per-copy site votes
iskoldt-X Jul 9, 2026
ebd185a
aggregate: type-annotate the per-copy voted site as optional
iskoldt-X Jul 9, 2026
f1116d4
config: exclude FMN, NI, NH4, SCN, UNX from ligand list
iskoldt-X Jul 9, 2026
ca5af6a
aggregate: give site-undetermined ligand copies an explicit unknown-s…
iskoldt-X Jul 9, 2026
d1c43e0
config: correct ligand exclude-list membership (drop PLM, add UNL)
iskoldt-X Jul 9, 2026
661aab4
curate: don't conflate a null field value with the review quit signal
iskoldt-X Jul 11, 2026
eb5bf5b
Soft-fail audit directory creation so curation never crashes on it
iskoldt-X Jul 11, 2026
886f92d
Remove dead final-data null guard in the curation main loop
iskoldt-X Jul 11, 2026
28840e9
Make the per-copy ligand table reviewable during curation
iskoldt-X Jul 11, 2026
836cad5
Treat per-copy ligand role disagreements as advisory, not gating
iskoldt-X Jul 11, 2026
e42a757
Pass a clean per-copy ligand block through review without a prompt
iskoldt-X Jul 11, 2026
1b85b37
feat: catalogue auxiliary small molecules on a dedicated output lane
iskoldt-X Jul 20, 2026
61eb96c
prompt: drop the Mg co-agonist worked example from metal-ligand guidance
iskoldt-X Jul 20, 2026
f332403
docs(readme): document annotate --sequential, read-only decision brie…
iskoldt-X Jul 21, 2026
bd5a896
Potential fix for pull request finding
iskoldt-X Jul 21, 2026
3f820cc
Normalize remaining 'G-protein' to 'G protein' across strings, docstr…
iskoldt-X Jul 21, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,9 @@ coverage.xml
# Ruff
.ruff_cache/

# Per-checkout extra denylist for the no-local-paths hook (machine-specific tokens)
scripts/.local-denylist

# Docker
.docker/

Expand All @@ -54,3 +57,4 @@ agent_reviews/
benchmark*
CLAUDE.md
GEMINI.md
CONVENTIONS.md
9 changes: 9 additions & 0 deletions .pre-commit-config.yaml
Original file line number Diff line number Diff line change
@@ -1,4 +1,13 @@
repos:
- repo: local
hooks:
- id: no-local-paths
name: reject developer-local paths
entry: python scripts/check_no_local_paths.py
language: system
types: [text]
pass_filenames: true

- repo: https://github.com/astral-sh/ruff-pre-commit
rev: v0.11.6
hooks:
Expand Down
25 changes: 14 additions & 11 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,7 +61,7 @@ Each step is **resumable** and **idempotent** — re-running any command skips a

A coordinate-driven detect stage (built on [gemmi](https://gemmi.readthedocs.io/)) runs before annotation and supplies the model with objective structural **facts** — not computed verdicts — leaving the final judgment to the AI:

- **G-protein coupling subtype** — Identifies the G-alpha subtype by matching the structure's alpha5 C-terminal window against reference sequences; subtypes that share an identical alpha5 helix (an inseparable set) are routed to family-level review instead of guessing one confident subtype.
- **G protein coupling subtype** — Identifies the G-alpha subtype by matching the structure's alpha5 C-terminal window against reference sequences; subtypes that share an identical alpha5 helix (an inseparable set) are routed to family-level review instead of guessing one confident subtype.
- **Binding-site geometry** — Each ligand's contacted residues are mapped to GPCRdb generic numbers and segments, with an ANVIL-style membrane-frame fit (oriented from the receptor's own intracellular landmarks: DRY, NPxxY, H8) that reports lipid-facing vs pocket-facing fraction and signed membrane depth. The model infers `site_ref` from these facts plus the paper, with `unknown` a first-class answer.
- **Dimer coupling protomer** — For an obligate Class C dimer, the protomer the G-alpha actually engages is detected from coordinates and used to pick the dimer's primary chain (e.g. GABA-B's GABBR2 couples while GABBR1 binds the agonist); the partner protomer is recorded.
- **Incidental-candidate ligands** — Dual-use molecules (cholesterol, palmitate, etc.) are surfaced to the model for a functional-vs-structural judgment rather than being silently dropped.
Expand All @@ -74,7 +74,7 @@ A coordinate-driven detect stage (built on [gemmi](https://gemmi.readthedocs.io/
- **Context-rich prompts** — The AI receives not just the paper PDF but also pre-enriched PDB metadata, the detect stage's structural evidence, a per-chain polymer table carrying each chain's 7TM status and residue length (to tell a true 7TM receptor from a non-receptor partner), a chain inventory reminder, and sibling structure warnings — reducing hallucination by grounding the model in API-verified and coordinate-derived facts.
- **Model-judged oligomeric state** — The model annotates the receptor's `oligomeric_state` (monomer / homo-/hetero-dimer / etc.) from neutral facts, counting only GPCR protomers (not transducer or ligand partners).
- **Flexible model selection** — Switch models at runtime via `--model` flag or `GPCR_GEMINI_MODEL` environment variable without code changes; sampling depth is tunable via `--temperature` and `--thinking-level` (threaded through both single and batch paths).
- **Batch API support** — Large-scale annotation via Gemini Batch API with JSONL submission, polling, and automatic result recovery; submissions are sharded into jobs (never splitting a structure's runs) and tracked in a registry for idempotent recovery.
- **Batch API support** — Large-scale annotation via Gemini Batch API with JSONL submission, polling, and automatic result recovery; submissions are sharded into jobs (never splitting a structure's runs) and tracked in a registry for idempotent recovery. A `--sequential` mode submits one shard per invocation and refuses to overlap an in-flight job, so scheduled or repeated runs advance the corpus without overrunning the provider's enqueued-token limit.
- **Rate-limited client** — Sliding-window rate limiting (1000 RPM) with exponential backoff on 429 responses.

### Post-Annotation Validation
Expand All @@ -90,12 +90,13 @@ A coordinate-driven detect stage (built on [gemmi](https://gemmi.readthedocs.io/

#### Warn-only safety cross-checks (surface, don't rewrite)

A family of checks routes likely mistakes to the review channel that disables one-click accept-all, while leaving the model's answer untouched: role-vs-site contradictions (e.g. an allosteric role at the orthosteric site), mis-filed GPCR protomers evicted from auxiliary proteins (sparing crystallization fusions and soluble partners), co-agonist reminders when multiple agonists are present, BRIL / T4-lysozyme fusion advisories, unannotated non-GPCR polymer chains, hallucinated ligands in ligand-free structures, and unrecognised G-alpha subtypes or G-protein-derived peptides mis-filed as ligands. Assembly-vs-oligomer mismatches are informational, not alerts.
A family of checks routes likely mistakes to the review channel that disables one-click accept-all, while leaving the model's answer untouched: role-vs-site contradictions (e.g. an allosteric role at the orthosteric site), mis-filed GPCR protomers evicted from auxiliary proteins (sparing crystallization fusions and soluble partners), co-agonist reminders when multiple agonists are present, BRIL / T4-lysozyme fusion advisories, unannotated non-GPCR polymer chains, hallucinated ligands in ligand-free structures, and unrecognised G-alpha subtypes or G protein-derived peptides mis-filed as ligands. Assembly-vs-oligomer mismatches are informational, not alerts.

### Expert Curation

- **Rich terminal dashboard** — An ergonomic review interface built with [Rich](https://github.com/Textualize/rich) for rapid, informed decision-making.
- **Context-aware validation alerts** — Real-time display of ghost chains, hallucinated ligands, UniProt identity clashes, and chimera warnings alongside the data being reviewed.
- **Read-only decision brief** — Each gated structure opens with a ranked list of every signal that needs a decision (validation findings, oligomer findings, and each AI-run voting fork), so the reviewer sees the full picture before drilling in.
- **Recursive review engine** — Navigate field-by-field through the annotation tree, with controversy highlights guiding attention to disputed values.
- **Append-only audit trail** — Every human decision (accept / edit / reject) is logged to `audit_trail.jsonl` with timestamps, providing full reproducibility.
- **Resumable sessions** — Curation progress is persisted; interrupted sessions resume exactly where they left off.
Expand Down Expand Up @@ -230,7 +231,7 @@ docker run --rm -it \

### `gpcr-tools detect`

Pre-annotation structural detection: compute coordinate-driven evidence (G-protein coupling, binding-site geometry, oligomeric state, chimera provenance) for the AI and flag hard cases for review.
Pre-annotation structural detection: compute coordinate-driven evidence (G protein coupling, binding-site geometry, oligomeric state, chimera provenance) for the AI and flag hard cases for review.

```bash
gpcr-tools detect # All enriched PDBs (tops up missing/degraded)
Expand All @@ -251,6 +252,7 @@ gpcr-tools annotate --prompt prompts/custom.md # Custom prompt template
gpcr-tools annotate --temperature 0.7 # Sampling temperature (default: model's own)
gpcr-tools annotate --thinking-level low # Reasoning depth: minimal|low|medium|high
gpcr-tools annotate --batch # Submit via Batch API
gpcr-tools annotate --batch --sequential # One shard/run; won't overlap an in-flight batch (cron-safe)
gpcr-tools annotate --check-batch # Poll batch status
gpcr-tools annotate --recover # Re-process raw batch output
```
Expand Down Expand Up @@ -285,7 +287,7 @@ Print an operational report over pipeline outputs.
```bash
gpcr-tools report pdf-coverage # Paper-PDF outcomes
gpcr-tools report full-audit # Validation warnings + chimera conflicts across PDBs
gpcr-tools report tail-analysis # G-protein chimera score distribution
gpcr-tools report tail-analysis # G protein chimera score distribution
gpcr-tools report run-manifest # Per-target accounting (no-PDF / incomplete /
# acceptable / gated, with provenance);
# writes output/run_manifest.{json,md}
Expand Down Expand Up @@ -379,8 +381,9 @@ Tab-separated, normalized files ready for database ingestion:
| File | Contents |
|------|----------|
| `structures.csv` | PDB ID, receptor UniProt, method, resolution, state, chain, date, and (for a heterodimer) the partner protomer's UniProt + chain |
| `ligands.csv` | Ligand names, PubChem IDs, roles, binding-site type (`Site`, from the geometry-informed `site_ref`), entity types, SMILES, InChIKey, sequences, and whether the bound compound is an endogenous ligand (`is_endogenous`, GtoPdb). Incidental molecules the model judged non-functional are omitted. |
| `g_proteins.csv` | G-protein subunit UniProt IDs and chain assignments |
| `ligands.csv` | Ligand identity (`Name`, the PDBe chemical-component code such as `RET` or `U0G`, falling back to the descriptive name when no component code exists), the full descriptive name (`Title`), PubChem IDs, roles, binding-site type (`Site`, from the geometry-informed `site_ref`), entity types, SMILES, InChIKey, sequences, the residue numbers of each modelled copy (`Residue_seq_id`, comma-joined and aligned copy-for-copy with `label_asym_id`), and whether the bound compound is an endogenous ligand (`is_endogenous`, GtoPdb). Molecules that are not functional ligands (ions, cofactors, glycans, detergents, matrix lipids, and molecules the model judged non-functional or structural) are catalogued in `auxiliary_small_molecules.csv` instead. |
| `auxiliary_small_molecules.csv` | Small molecules present but not functional receptor ligands: `Name` (component code), `Type` (`Ion` / `Lipid` / `Detergent` / `Other`), `Function` (`Cofactor` or blank), located like `ligands.csv` (`ChainID` + `label_asym_id` + `Residue_seq_id`) |
| `g_proteins.csv` | G protein subunit identities and chain assignments: `Alpha_identity` (alpha subunit UniProt entry name), `Alpha_alpha5_identity` (alpha5-helix functional coupling), `Alpha_backbone` (modelled scaffold), plus beta/gamma UniProt IDs and chains |
| `arrestins.csv` | Arrestin UniProt IDs and chains |
| `fusion_proteins.csv` | Fusion protein names |
| `nanobodies.csv`, `antibodies.csv`, `scfv.csv` | Binding partner names |
Expand Down Expand Up @@ -426,8 +429,8 @@ src/gpcr_tools/
├── detector/ # Pre-annotation detect stage (runs before annotate)
│ ├── signals.py # DetectSignal contract (advisory→prompt, review→curator)
│ ├── gprotein.py # G-protein alpha5 identity detector
│ ├── coupling.py # G-protein-coupling protomer of a dimer (geometry)
│ ├── gprotein.py # G protein alpha5 identity detector
│ ├── coupling.py # G protein-coupling protomer of a dimer (geometry)
│ ├── site_ref.py # Ligand binding-site detector (geometry → generic numbers)
│ ├── geometry.py # Dual-role ligand detector (multi-pocket burial)
│ ├── ligands.py # Incidental-candidate ligand detector (cholesterol, palmitate)
Expand All @@ -447,7 +450,7 @@ src/gpcr_tools/
│ └── runner.py # 12-step orchestration with error isolation
├── validator/ # Cross-validation + enrichment modules
│ ├── chimera.py # G-protein alpha5 identity (sequence matching)
│ ├── chimera.py # G protein alpha5 identity (sequence matching)
│ ├── receptor_validator.py # UniProt identity verification
│ ├── ligand_validator.py # PDB-CCD existence check + endogenous tagging
│ ├── endogenous.py # Endogenous-ligand classifier (GtoPdb table)
Expand Down Expand Up @@ -503,7 +506,7 @@ pytest tests/ -v

### Test Suite

The test suite includes 1,100+ tests:
The test suite includes 1,700+ tests:

- **Unit tests** for every module across all five pipeline stages
- **Integration tests** for the full aggregation pipeline, error isolation, and atomic write safety
Expand Down
8 changes: 8 additions & 0 deletions scripts/.local-denylist.example
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
# Per-checkout extra tokens for the no-local-paths pre-commit hook.
# Copy this file to scripts/.local-denylist (git-ignored) and list any
# machine-specific strings you want blocked from commits — one per line,
# matched as plain substrings. Lines starting with '#' are comments.
#
# Example (replace with your own):
# my-username
# my-conda-env-name
99 changes: 68 additions & 31 deletions scripts/build_site_ref_table.py
Original file line number Diff line number Diff line change
Expand Up @@ -4,30 +4,63 @@
non-human UniProt accession present in the local corpus (current orthologs).
Output: {accession: {"c": class, "e": entry_name, "r": {seqnum: [x_label, segment, aa]}}}
gzipped JSON. Keyed by UniProt accession (what enriched/RCSB-align provides).
Requires the local GPCRdb DB (gpcrdb-db). Run once; ship the artifact.
Requires the local GPCRdb DB (gpcrdb-db) and GPCR_CORPUS_DIR set to the corpus
root. Run once from the repo root; ships to src/gpcr_tools/data/.
"""
import json, gzip, glob, subprocess, os

import glob
import gzip
import json
import os
import subprocess


def psql(sql):
return subprocess.run(["docker","exec","gpcrdb-db","psql","-U","protwis","-d","protwis","-tA","-F","\t","-c",sql],
capture_output=True, text=True).stdout
return subprocess.run(
[
"docker",
"exec",
"gpcrdb-db",
"psql",
"-U",
"protwis",
"-d",
"protwis",
"-tA",
"-F",
"\t",
"-c",
sql,
],
capture_output=True,
text=True,
).stdout


# 1. corpus non-human accessions (from enriched GPCR uniprots)
base="${GPCR_CORPUS_DIR}"
corpus_accs=set()
base = os.environ.get("GPCR_CORPUS_DIR")
if not base:
raise SystemExit(
"Set GPCR_CORPUS_DIR to the corpus root (the directory holding */enriched/*.json)."
)
corpus_accs = set()
for ep in glob.glob(f"{base}/*/enriched/*.json"):
try: raw=json.load(open(ep))
except: continue
e=(raw.get("data") or {}).get("entry") or raw
for ent in (e.get("polymer_entities") or []):
for u in (ent.get("uniprots") or []):
try:
with open(ep) as fh:
raw = json.load(fh)
except Exception:
continue
e = (raw.get("data") or {}).get("entry") or raw
for ent in e.get("polymer_entities") or []:
for u in ent.get("uniprots") or []:
if u.get("gpcrdb_entry_name_slug"):
acc=(u.get("rcsb_id") or "").strip()
if acc: corpus_accs.add(acc)
acc = (u.get("rcsb_id") or "").strip()
if acc:
corpus_accs.add(acc)
print(f"corpus GPCR accessions: {len(corpus_accs)}")

in_list = ",".join("'%s'" % a.replace("'","") for a in corpus_accs) or "''"
sql=f"""SELECT p.accession, p.entry_name, split_part(pf.slug,'_',1) AS cls,
in_list = ",".join("'" + a.replace("'", "") + "'" for a in corpus_accs) or "''"
sql = f"""SELECT p.accession, p.entry_name, split_part(pf.slug,'_',1) AS cls,
r.sequence_number, COALESCE(gn.label,''), COALESCE(ps.slug,''), r.amino_acid
FROM residue r
JOIN protein_conformation pc ON pc.id=r.protein_conformation_id
Expand All @@ -42,23 +75,27 @@ def psql(sql):
AND (sp.latin_name='Homo sapiens' OR p.accession IN ({in_list}))
ORDER BY p.accession, r.sequence_number;"""

table={}
n=0
table = {}
n = 0
for line in psql(sql).strip().split("\n"):
p=line.split("\t")
if len(p)<7 or not p[0]: continue
acc,entry,cls,seqn,xlab,seg,aa=p[0],p[1],p[2],p[3],p[4],p[5],p[6]
rec=table.setdefault(acc,{"c":cls,"e":entry,"r":{}})
rec["r"][seqn]=[xlab or None, seg or None, aa]
n+=1
p = line.split("\t")
if len(p) < 7 or not p[0]:
continue
acc, entry, cls, seqn, xlab, seg, aa = p[0], p[1], p[2], p[3], p[4], p[5], p[6]
rec = table.setdefault(acc, {"c": cls, "e": entry, "r": {}})
rec["r"][seqn] = [xlab or None, seg or None, aa]
n += 1

out="/tmp/gpcrdb_generic_numbers.json.gz"
with gzip.open(out,"wt",encoding="utf-8") as f:
json.dump(table,f,separators=(",",":"))
print(f"receptors: {len(table)} | residue rows: {n} | gz size: {os.path.getsize(out)/1e6:.2f} MB")
out = os.environ.get("GPCR_SITE_REF_OUT", "src/gpcr_tools/data/gpcrdb_generic_numbers.json.gz")
with gzip.open(out, "wt", encoding="utf-8") as f:
json.dump(table, f, separators=(",", ":"))
print(f"receptors: {len(table)} | residue rows: {n} | gz size: {os.path.getsize(out) / 1e6:.2f} MB")
# spot check
for acc in ("P07550","Q9NYV8","P41180"):
for acc in ("P07550", "Q9NYV8", "P41180"):
if acc in table:
r=table[acc]["r"]
landmarks={k:v for k,v in r.items() if v[0] in ("3x32","2x50","6x48","7x39")}
print(f" {acc} {table[acc]['e']} class {table[acc]['c']}: {len(r)} residues; landmarks {landmarks}")
r = table[acc]["r"]
landmarks = {k: v for k, v in r.items() if v[0] in ("3x32", "2x50", "6x48", "7x39")}
print(
f" {acc} {table[acc]['e']} class {table[acc]['c']}: "
f"{len(r)} residues; landmarks {landmarks}"
)
Loading