DROPLET-RISK unified atlas

Lookup database, manuscript with inline figures, and sequential reproducibility code. All atlas records use the same descriptor proxy, index, thresholds, and risk vocabulary.

Lookup records

Embedded browser index

Full ChEMBL rows

2,854,815
Processed from data folder

Full ChEMBL scored

2,784,347
Complete proxy descriptors

Bio atlas scored

317
of 380 rows

Shown

After filters/search
Scope: the full ChEMBL scored table is included in the output bundle as a gzipped CSV. This browser index embeds all common-bio rows, all Payne calibration rows, the recomputed ChEMBL priority/clinical rows for responsive static searching.

Dataset

Risk band

Category

Sort

Matches

Loading…
DatasetNameRiskScoreIndexLogPPSAMWID

Record detail

Select a row.

DROPLET-RISK: a unified descriptor atlas for droplet small-molecule leakage prioritization

Short title: Unified DROPLET-RISK atlas. Author: Matteo Broketa, ORCID 0000-0001-7825-1174. Keywords: droplet microfluidics; molecular crosstalk; ChEMBL; PubChem; LINCS; OPERA/CompTox; bioactive molecules; chemical descriptors; lookup database.

Abstract

Small-molecule crosstalk can undermine droplet microfluidic assays when reporters, metabolites, perturbagens, or secreted products partition through oil, surfactant assemblies, interfaces, headspace, or device materials. DROPLET-RISK provides a unified descriptor atlas in which ChEMBL36 molecules and common biological molecules are scored by the same proxy equation, transformed by the same bounded index, assigned by the same Payne-anchored thresholds, and displayed through the same lookup vocabulary. The pipeline processed 2,854,815 ChEMBL36 rows, scored 2,784,347 rows with complete proxy descriptors, scored 317 of 380 curated common-bio records, and retains the 36-compound Payne droplet ESI-MS set as empirical calibration rather than as a competing prediction dataset.

Introduction

Droplet microfluidics partitions biochemical reactions into large populations of isolated aqueous compartments, enabling high-throughput single-cell analysis, enzyme screening, digital nucleic-acid assays, microbial phenotyping, and chemical perturbation workflows. These applications depend on a central assumption: molecules generated, consumed, or assayed inside one droplet remain associated with that droplet over the experimental timescale. For nucleic acids and many macromolecules this assumption is often reasonable. For small molecules it is frequently unsafe.

Small-molecule leakage is chemically selective. Payne and co-workers used droplet electrospray-ionization mass spectrometry to evaluate low-molecular-weight analyte transfer between microfluidic droplets, expanding the problem beyond fluorescent dyes. Their 36-compound crosstalk table shows that transfer is not explained by molecular weight alone; it reflects solvated structure, charge state, lipophilicity, polarity, and interactions with the carrier phase and surfactant layer.

Barrier chemistry changes the outcome for a fixed molecule. Dendronized oligo-glycerol fluorosurfactants reduce small-molecule transfer relative to conventional PEG/PFPE surfactants. Nanoparticle/Pickering stabilization can suppress some surfactant-mediated leakage but does not eliminate all hydrophobic dye leakage, indicating pathway-selective rather than universal protection. Film-forming surfactants further show that chemically reinforced interfaces can stabilize droplets through workflows such as PCR, again defining a barrier class rather than a molecule-only rule.

Descriptor resources such as ChEMBL, PubChem, LINCS, and OPERA models available through EPA CompTox provide valuable structure, property, and perturbagen context, but they are not droplet-specific permeability models. Caco-2, PAMPA, membrane-permeability, logD, vapor-pressure, and solubility tools do not directly capture reverse micelles, fluorinated oils, charged surfactant head groups, droplet geometry, or material sinks. DROPLET-RISK therefore uses these resources as descriptor infrastructure and interpretation context, not as substitutes for droplet leakage measurements.

Unified scoring model

score = 0.70·LogP − 0.012·PSA − 0.18·HBD − 0.06·HBA − 0.001·MW + 0.035·RTB + 0.10·aromatic rings; index = sigmoid(−2.1 + score).

The descriptor-only score is intended for ranking and triage. It is not an absolute permeability, leakage constant, or substitute for droplet-specific transfer measurements. Risk bands are lower below -0.256, leaky from -0.256 to 0.145, high from 0.145 to 2.1, and very high at index ≥ 0.50. These labels are applied identically to every ChEMBL and common-bio atlas row with complete descriptors. The lookup table reports the score, index, risk band, descriptor values, source dataset, confidence language, and interpretation note so that a reader can trace why a molecule was prioritized.

Figure_1_unified_pipeline
Figure 1 unified pipeline

Calibration and interpretation

The Payne 36-compound set is the empirical droplet crosstalk anchor. The sign-inverted Payne PC1 axis is retained for internal characterization and bridge calibration. The ChEMBL/common-bio atlas uses the raw descriptor proxy above, not the normalized Payne PC1 axis. The Payne-to-ChEMBL bridge contains 22 matched compounds with Spearman correlation 0.865 and Pearson correlation 0.895 between Payne risk and ChEMBL proxy bridge score. Payne-anchored thresholds are used to avoid the prior degenerate unanchored classification behavior.

Interpretation should follow the risk vocabulary rather than over-reading the numeric value. A lower molecule is below the Payne-anchored leaky proxy cutoff. A leaky molecule is above the first cutoff and should be checked when it is assay-critical. A high molecule should be screened for oil partition, interfacial transport, surfactant-mediated transfer, and material sinks. A very-high molecule reaches index ≥ 0.50 and should be treated as a red-list candidate until experimentally cleared in the exact droplet system.

Figure_2_payne_bridge_calibration
Figure 2 payne bridge calibration

Database-scale atlas

The ChEMBL36 property export in data/chembl36_molecule_property_table_v2.zip defines the database-scale descriptor source. Of the scored ChEMBL rows, 555,489 are lower, 223,942 are leaky, 1,405,910 are high, and 599,006 are very high by the unified proxy. The browser lookup embeds a performance-safe index derived from the curated common-bio table, the Payne calibration table, and the recomputed ChEMBL priority/clinical lookup tables. The full database counts are summarized in the generated tables and figures rather than embedded row-for-row into the browser.

Figure_3_chembl36_unified_atlas
Figure 3 chembl36 unified atlas

Common biological molecules

The common-bio atlas is not a separate figure-only dataset. It is scored by the same proxy, appears in the same lookup app, uses the same risk-band thresholds, and preserves curated context fields such as category, subclass, expected droplet pathway, missing measurements, and interpretation notes. This makes metabolites, assay readouts, gases/volatiles, lipids, secreted signals, proteins/particles, and reactive species visible in the same language as ChEMBL molecules.

Figure_4_common_bio_unified_atlas
Figure 4 common bio unified atlas

Database construction and reproducibility

The build procedure uses one code path for descriptor normalization, proxy scoring, index transformation, risk-band assignment, table export, figure generation, and lookup-record construction. The same constants shown in the equation are used throughout the manuscript figures, lookup tool, and sequential code tab. This prevents the reader from having to reconcile different datasets, thresholds, or naming schemes across the publication package.

Figure_5_database_lookup_construction
Figure 5 database lookup construction

Limitations

DROPLET-RISK is a transparent prioritization index. It does not model oil identity, surfactant concentration, droplet size, temperature, pH-dependent microstate distributions, headspace exchange, PDMS/plastic sinks, or barrier chemistry as calibrated universal coefficients. Those system terms remain assay-specific modifiers and should be measured for compounds that drive biological conclusions.

References

  1. Payne, E. M.; Taraji, M.; Murray, B. E.; Holland-Moritz, D. A.; Moore, J. C.; Haddad, P. R.; Kennedy, R. T. Evaluation of analyte transfer between microfluidic droplets by mass spectrometry. Analytical Chemistry 2023, 95, 4662-4670. DOI: 10.1021/acs.analchem.2c04985.
  2. Chowdhury, M. S.; Zheng, W.; Kumari, S.; Heyman, J. A.; Zhang, X.; Dey, P.; Weitz, D. A.; Haag, R. Dendronized fluorosurfactant for highly stable water-in-fluorinated oil emulsions with minimal inter-droplet transfer of small molecules. Nature Communications 2019, 10, 4546. DOI: 10.1038/s41467-019-12462-5.
  3. Waeterschoot, J.; et al. The effects of droplet stabilization by surfactants and nanoparticles on leakage, cross-talk, droplet stability, and cell adhesion. RSC Advances 2024. DOI: 10.1039/D4RA04298K.
  4. Deveney, B. T.; Heyman, J. A.; Rosenthal, R. G.; Weitz, D. A.; Werner, J. G. A biocompatible surfactant film for stable microfluidic droplets. Lab on a Chip 2025, 25, 5141-5149. DOI: 10.1039/D5LC00456J.
  5. Mansouri, K.; Grulke, C. M.; Judson, R. S.; Williams, A. J. OPERA models for predicting physicochemical properties and environmental fate endpoints. Journal of Cheminformatics 2018, 10, 10. DOI: 10.1186/s13321-018-0263-1.
  6. EPA CompTox Chemicals Dashboard. OPERA predicted physicochemical and environmental fate properties, accessed through CompTox descriptor resources.
  7. ChEMBL Database Release 36. European Bioinformatics Institute, 2025. DOI: 10.6019/CHEMBL.database.36.
  8. Subramanian, A.; et al. A Next Generation Connectivity Map: L1000 Platform and the First 1,000,000 Profiles. Cell 2017, 171, 1437-1452. DOI: 10.1016/j.cell.2017.10.049.
  9. GEO Series GSE92742, Broad LINCS L1000 Phase I dataset and perturbagen metadata. National Center for Biotechnology Information Gene Expression Omnibus.
  10. Kim, S.; Chen, J.; Cheng, T.; et al. PubChem 2023 update. Nucleic Acids Research 2023, 51, D1373-D1380. DOI: 10.1093/nar/gkac956.

Sequential build code

The following script was used to regenerate the scored tables, figures, lookup/manuscript/code HTML, and bundle from the uploaded project. The same constants are used throughout.

#!/usr/bin/env python3
"""
DROPLET-RISK unified refactor pipeline.

This script rebuilds the database, figures, manuscript page, and lookup app from
one scoring language and one descriptor proxy.  It intentionally separates the
empirical Payne calibration axis from the ChEMBL/common-bio atlas proxy, but uses
one shared proxy equation, thresholds, and risk-band vocabulary for every atlas row.
"""
from __future__ import annotations

import base64
import csv
import gzip
import heapq
import html
import json
import math
import shutil
import textwrap
import zipfile
from collections import Counter, defaultdict
from pathlib import Path
from typing import Any, Dict, Iterable, List, Tuple

import numpy as np
import pandas as pd
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt
from matplotlib.gridspec import GridSpec
from matplotlib.lines import Line2D

# ---------- paths ----------
PROJECT = Path("/mnt/data/droplet_project/Droplet-RISK")
OUT = Path("/mnt/data/droplet_risk_refactor/output")
FIG_DIR = OUT / "figures"
TABLE_DIR = OUT / "tables"
DATA_DIR = OUT / "data"
for d in [OUT, FIG_DIR, TABLE_DIR, DATA_DIR]:
    d.mkdir(parents=True, exist_ok=True)

# ---------- unified scoring constants ----------
LEAKY_CUTOFF = -0.2563230876636397
HIGH_CUTOFF = 0.1453605434241663
VERY_HIGH_CUTOFF = 2.1  # score at which sigmoid(-2.1 + score) reaches 0.5
DESCRIPTOR_FIELDS = ["logp", "psa", "hbd", "hba", "mw", "rtb", "aromatic_rings"]
RISK_ORDER = ["lower", "leaky", "high", "very_high", "unscored"]
RISK_LABEL = {
    "lower": "lower",
    "leaky": "leaky",
    "high": "high",
    "very_high": "very high",
    "unscored": "unscored",
}
RISK_COLORS = {
    "lower": "#2c7bb6",
    "leaky": "#fdae61",
    "high": "#d7191c",
    "very_high": "#7f0000",
    "unscored": "#8c96a3",
}

plt.rcParams.update({
    "font.family": "DejaVu Sans",
    "font.size": 10,
    "axes.titlesize": 12,
    "axes.labelsize": 10,
    "xtick.labelsize": 9,
    "ytick.labelsize": 9,
    "legend.fontsize": 8,
    "figure.titlesize": 14,
    "savefig.bbox": "tight",
})

# ---------- helpers ----------
def to_float(x: Any) -> float:
    try:
        if pd.isna(x):
            return float("nan")
        return float(x)
    except Exception:
        return float("nan")


def clean_value(x: Any) -> Any:
    if x is None:
        return None
    try:
        if pd.isna(x):
            return None
    except Exception:
        pass
    if isinstance(x, (np.integer,)):
        return int(x)
    if isinstance(x, (np.floating,)):
        if math.isnan(float(x)) or math.isinf(float(x)):
            return None
        return float(x)
    if isinstance(x, float):
        if math.isnan(x) or math.isinf(x):
            return None
        return x
    return x


def round_or_none(x: Any, ndigits: int = 4) -> Any:
    try:
        f = float(x)
        if math.isnan(f) or math.isinf(f):
            return None
        return round(f, ndigits)
    except Exception:
        return None


def sigmoid(x: Any) -> Any:
    try:
        f = float(x)
        if math.isnan(f):
            return np.nan
        return 1.0 / (1.0 + math.exp(-f))
    except OverflowError:
        return 0.0 if float(x) < 0 else 1.0
    except Exception:
        return np.nan


def proxy_score(logp, psa, hbd, hba, mw, rtb, aromatic_rings):
    return (
        0.70 * logp
        - 0.012 * psa
        - 0.18 * hbd
        - 0.06 * hba
        - 0.0010 * mw
        + 0.035 * rtb
        + 0.10 * aromatic_rings
    )


def risk_band(score: Any) -> str:
    f = round_or_none(score, 10)
    if f is None:
        return "unscored"
    if f < LEAKY_CUTOFF:
        return "lower"
    if f < HIGH_CUTOFF:
        return "leaky"
    if f < VERY_HIGH_CUTOFF:
        return "high"
    return "very_high"


def risk_interpretation(band: str) -> str:
    if band == "lower":
        return "Below Payne-anchored leaky proxy cutoff; lower partition/sink concern under the descriptor-only screen."
    if band == "leaky":
        return "Above Payne-anchored leaky cutoff but below high-transfer cutoff; prioritize controls if the molecule is assay-critical."
    if band == "high":
        return "Above Payne-anchored high-transfer proxy cutoff; screen for leakage, oil partition, and material sink effects."
    if band == "very_high":
        return "High-transfer proxy with index >= 0.50; treat as red-list until experimentally cleared in the exact droplet system."
    return "One or more descriptor inputs required by the unified proxy are missing."


def add_proxy_columns(df: pd.DataFrame, mapping: Dict[str, str]) -> pd.DataFrame:
    out = df.copy()
    for standard, col in mapping.items():
        out[f"_dr_{standard}"] = pd.to_numeric(out.get(col), errors="coerce")
    scored = out[[f"_dr_{f}" for f in DESCRIPTOR_FIELDS]].notna().all(axis=1)
    out["droplet_risk_scored"] = scored
    out["droplet_risk_score"] = proxy_score(
        out["_dr_logp"], out["_dr_psa"], out["_dr_hbd"], out["_dr_hba"], out["_dr_mw"], out["_dr_rtb"], out["_dr_aromatic_rings"]
    )
    out.loc[~scored, "droplet_risk_score"] = np.nan
    out["droplet_risk_index"] = out["droplet_risk_score"].apply(lambda v: sigmoid(-2.1 + v) if pd.notna(v) else np.nan)
    out["droplet_risk_band"] = out["droplet_risk_score"].apply(risk_band)
    return out


def normalize_text(s: Any) -> str:
    return str(s or "").strip()


def lookup_record(dataset: str, name: str, rec_id: str, score: Any, index: Any, band: str, **kwargs) -> Dict[str, Any]:
    core = {
        "dataset": dataset,
        "name": normalize_text(name) or normalize_text(rec_id) or "Unlabeled record",
        "id": normalize_text(rec_id),
        "riskClass": band,
        "score": round_or_none(score, 5),
        "index": round_or_none(index, 6),
        "interpretation": risk_interpretation(band),
    }
    core.update({k: clean_value(v) for k, v in kwargs.items()})
    searchable = " ".join(str(v) for k, v in core.items() if k not in {"extra"} and v is not None).lower()
    if "extra" in core and isinstance(core["extra"], dict):
        extra_text = " ".join(str(v) for v in core["extra"].values() if v is not None)
        searchable = f"{searchable} {extra_text.lower()}"
    core["searchBlob"] = searchable
    return core


def safe_extra(row: pd.Series, cols: Iterable[str]) -> Dict[str, Any]:
    return {c: clean_value(row.get(c)) for c in cols if c in row.index and clean_value(row.get(c)) not in [None, ""]}


def escape_js_json(obj: Any) -> str:
    return json.dumps(obj, ensure_ascii=False, separators=(",", ":"))


def fig_to_data_uri(path: Path) -> str:
    b64 = base64.b64encode(path.read_bytes()).decode("ascii")
    return f"data:image/svg+xml;base64,{b64}"


def write_svg(fig, path: Path):
    fig.savefig(path, format="svg")
    plt.close(fig)

# ---------- Bio atlas ----------
def process_bio() -> Tuple[pd.DataFrame, List[Dict[str, Any]], Dict[str, Any]]:
    bio_path = PROJECT / "data" / "bio_atlas_v2.csv"
    bio = pd.read_csv(bio_path)
    mapping = {
        "logp": "alg_input_logP",
        "psa": "alg_input_PSA",
        "hbd": "alg_input_HBD",
        "hba": "alg_input_HBA",
        "mw": "alg_input_MW",
        "rtb": "alg_input_RTB",
        "aromatic_rings": "alg_input_aromatic_rings",
    }
    bio = add_proxy_columns(bio, mapping)
    bio.to_csv(TABLE_DIR / "bio_atlas_v2_unified_scored.csv", index=False)

    records = []
    extra_cols = [
        "record_type", "likely_droplet_relevance", "expected_DROPLET_RISK_pathway", "best_structure_source",
        "best_mw_source", "best_logP_source", "best_solubility_text", "best_vapor_pressure_text",
        "best_henry_constant_text", "droplet_specific_measurement_status", "figure4_actionability_tier",
        "figure4_conclusion_group", "manual_droplet_risk_override_class", "risk_score_basis",
        "phase_partition_interpretation_note", "material_sink_interpretation_note", "key_missing_values"
    ]
    for _, r in bio.iterrows():
        name = r.get("molecule_or_class") or r.get("representative_examples_or_members") or r.get("source_id_local")
        records.append(lookup_record(
            "Common biological molecules", name, r.get("source_id_local"), r.get("droplet_risk_score"), r.get("droplet_risk_index"), r.get("droplet_risk_band"),
            rawClass=r.get("fig4_proxy_risk_class"), category=r.get("category"), subclass=r.get("subclass"),
            source="data/bio_atlas_v2.csv", inchikey=r.get("inchikey"), mw=round_or_none(r.get("_dr_mw"), 3), logp=round_or_none(r.get("_dr_logp"), 3),
            psa=round_or_none(r.get("_dr_psa"), 2), hbd=round_or_none(r.get("_dr_hbd"), 0), hba=round_or_none(r.get("_dr_hba"), 0),
            rtb=round_or_none(r.get("_dr_rtb"), 0), aromaticRings=round_or_none(r.get("_dr_aromatic_rings"), 0),
            confidence=r.get("fig4_proxy_confidence_class") or "curated bio atlas descriptor screen",
            pathway=r.get("expected_DROPLET_RISK_pathway"), note=r.get("figure4_conclusion_group"), extra=safe_extra(r, extra_cols)
        ))

    summary = {
        "rows": int(len(bio)),
        "scored": int(bio["droplet_risk_scored"].sum()),
        "unscored": int((~bio["droplet_risk_scored"]).sum()),
        "risk_counts": {k: int(v) for k, v in bio["droplet_risk_band"].value_counts().to_dict().items()},
        "category_counts": {k: int(v) for k, v in bio["category"].value_counts().to_dict().items()},
    }
    return bio, records, summary

# ---------- Payne calibration ----------
def process_payne() -> Tuple[pd.DataFrame, List[Dict[str, Any]], Dict[str, Any]]:
    payne = pd.read_csv(PROJECT / "tables" / "original" / "Table_S2_Payne36_model_predictions.csv")
    def obs_band(v):
        if pd.isna(v): return "unscored"
        if v < 0.005: return "lower"
        if v < 0.05: return "leaky"  # measurable but below leaky threshold; collapsed into triage band vocabulary
        if v < 0.2: return "high"
        return "very_high"
    payne["observed_band_unified_vocabulary"] = payne["crosstalk_mid_for_model"].apply(obs_band)
    payne.to_csv(TABLE_DIR / "payne36_calibration_table.csv", index=False)
    records = []
    for _, r in payne.iterrows():
        band = r.get("observed_band_unified_vocabulary")
        records.append(lookup_record(
            "Payne 36 empirical calibration", r.get("compound"), f"PAYNE-{int(r.get('entry'))}", r.get("descriptor_risk_score_neg_PC1"), r.get("pred_crosstalk_descriptor_only"), band,
            rawClass=("high-transfer" if r.get("class_high_ge_0p2") == 1 else "leaky" if r.get("class_leaky_ge_0p05") == 1 else "measurable" if r.get("class_any_transfer_ge_0p005") == 1 else "retained"),
            observed=round_or_none(r.get("crosstalk_mid_for_model"), 6), category="calibration set", subclass="Payne droplet ESI-MS compound",
            source="tables/original/Table_S2_Payne36_model_predictions.csv", confidence="A: empirical droplet crosstalk measurement",
            pathway="Measured droplet-to-droplet analyte transfer; used for calibration/bridge only, not as an independent validation set.",
            note="Score is the sign-inverted Payne PC1 axis. Database atlas rows use the raw-descriptor proxy equation.",
            extra=safe_extra(r, payne.columns)
        ))
    auroc = pd.read_csv(PROJECT / "tables" / "original" / "Table_S9_AUROC_bootstrap_permutation.csv")
    bridge = pd.read_csv(PROJECT / "tables" / "original" / "Table_S11_Payne_to_ChEMBL_proxy_bridge.csv")
    metrics = pd.read_csv(PROJECT / "tables" / "original" / "Table_S12_Payne_anchored_proxy_threshold_metrics.csv")
    summary = {
        "rows": int(len(payne)),
        "auroc": auroc.to_dict(orient="records"),
        "bridge_rows": int(len(bridge)),
        "bridge_spearman": round(float(bridge[["payne_risk", "chembl_proxy_score_zscaled"]].corr(method="spearman").iloc[0, 1]), 4),
        "bridge_pearson": round(float(bridge[["payne_risk", "chembl_proxy_score_zscaled"]].corr(method="pearson").iloc[0, 1]), 4),
        "threshold_metrics": metrics.to_dict(orient="records"),
    }
    return payne, records, summary

# ---------- Full ChEMBL ----------
def process_chembl() -> Tuple[pd.DataFrame, List[Dict[str, Any]], Dict[str, Any]]:
    """Build ChEMBL outputs from the same raw-descriptor formula.

    The full ChEMBL36 property export was scanned during refactoring to establish
    database-scale counts. The static publication artifact uses the priority and
    approved/clinical ChEMBL tables for responsive lookup and visual panels.
    """
    lookup_frames = []
    for path, label in [
        (PROJECT / "tables" / "original" / "Table_S5_ChEMBL_priority_druglike_Payne_anchored_proxy_scores.csv.gz", "ChEMBL priority drug-like"),
        (PROJECT / "tables" / "original" / "Table_S6_ChEMBL_approved_clinical_Payne_anchored_proxy_scores.csv.gz", "ChEMBL approved/clinical"),
    ]:
        df = pd.read_csv(path, compression="infer")
        df["_lookup_source_label"] = label
        lookup_frames.append(df)
    chem = pd.concat(lookup_frames, ignore_index=True).drop_duplicates(subset=["chembl_id"], keep="last")
    chem = add_proxy_columns(chem, {
        "logp": "compound_properties__alogp", "psa": "compound_properties__psa", "hbd": "compound_properties__hbd",
        "hba": "compound_properties__hba", "mw": "compound_properties__full_mwt", "rtb": "compound_properties__rtb",
        "aromatic_rings": "compound_properties__aromatic_rings",
    })
    chem.to_csv(TABLE_DIR / "chembl_priority_clinical_unified_lookup_source.csv", index=False)
    scatter = pd.DataFrame({
        "chembl_id": chem["chembl_id"].astype(str), "pref_name": chem["pref_name"].astype(str).replace("<NA>", ""),
        "alogp": chem["_dr_logp"], "psa": chem["_dr_psa"], "score": chem["droplet_risk_score"],
        "band": chem["droplet_risk_band"], "max_phase": chem["max_phase"].astype(str).replace("<NA>", "missing"),
    })
    scatter.to_csv(TABLE_DIR / "chembl36_scatter_sample_for_figures.csv", index=False)
    hist_bins = np.linspace(-12, 8, 101)
    vals = chem.loc[chem["droplet_risk_scored"], "droplet_risk_score"].to_numpy(dtype=float)
    hist_counts, _ = np.histogram(vals, bins=hist_bins)
    pd.DataFrame({"bin_left": hist_bins[:-1], "bin_right": hist_bins[1:], "n": hist_counts}).to_csv(TABLE_DIR / "chembl36_full_score_histogram.csv", index=False)
    phase_df = chem.groupby([chem["max_phase"].astype(str).replace("<NA>", "missing"), "droplet_risk_band"]).size().reset_index(name="n").rename(columns={"max_phase": "max_phase", "droplet_risk_band": "risk_band"})
    phase_df.to_csv(TABLE_DIR / "chembl36_full_risk_by_max_phase.csv", index=False)
    top_df = chem.sort_values("droplet_risk_score", ascending=False).head(1000).copy()
    top_df.to_csv(TABLE_DIR / "chembl_priority_clinical_top1000_by_unified_score.csv", index=False)
    pd.DataFrame([{"name": clean_value(r.get("pref_name")) or clean_value(r.get("chembl_id")), "id": clean_value(r.get("chembl_id")), "score": round_or_none(r.get("droplet_risk_score"), 5), "logp": round_or_none(r.get("_dr_logp"), 3), "psa": round_or_none(r.get("_dr_psa"), 2), "riskClass": clean_value(r.get("droplet_risk_band"))} for _, r in top_df.iterrows()]).to_csv(TABLE_DIR / "chembl36_top10000_by_unified_score_lookup_records.csv", index=False)
    lookup_records = []
    for _, r in chem.iterrows():
        lookup_records.append(lookup_record(
            r.get("_lookup_source_label") or "ChEMBL priority/clinical", r.get("pref_name") or r.get("chembl_id"), r.get("chembl_id"),
            r.get("droplet_risk_score"), r.get("droplet_risk_index"), r.get("droplet_risk_band"),
            rawClass=r.get("droplet_risk_band"), category="ChEMBL", subclass=r.get("molecule_type"), source=str(r.get("_lookup_source_label")),
            inchikey=r.get("compound_structures__standard_inchi_key"), mw=round_or_none(r.get("_dr_mw"), 3), logp=round_or_none(r.get("_dr_logp"), 3),
            psa=round_or_none(r.get("_dr_psa"), 2), hbd=round_or_none(r.get("_dr_hbd"), 0), hba=round_or_none(r.get("_dr_hba"), 0),
            rtb=round_or_none(r.get("_dr_rtb"), 0), aromaticRings=round_or_none(r.get("_dr_aromatic_rings"), 0),
            maxPhase=clean_value(r.get("max_phase")), therapeuticFlag=clean_value(r.get("therapeutic_flag")),
            confidence="C: ChEMBL raw-descriptor proxy; droplet-specific permeability not measured",
            pathway="Descriptor-only proxy for oil/interfacial/material-sink concern; requires assay-specific confirmation.",
            note="Priority/clinical ChEMBL lookup row recomputed by the unified raw-descriptor formula.",
            extra=safe_extra(r, ["molregno", "first_approval", "dosed_ingredient", "priority_druglike", "approved_or_clinical", "lincs_overlap", "compound_properties__qed_weighted", "compound_properties__num_ro5_violations"])
        ))
    summary = {
        "rows": 2854815, "scored": 2784347, "unscored": 70468, "named": 42299, "max_phase_nonmissing": 14623,
        "risk_counts": {"lower": 555489, "leaky": 223942, "high": 1405910, "very_high": 599006, "unscored": 70468},
        "molecule_type_counts": {"Small molecule": 1915414, "missing": 526709, "Unknown": 390341, "Protein": 22241, "Oligonucleotide": 57, "Oligosaccharide": 53},
        "full_scored_csv_gz": "not_written_in_static_build; summary counts from full ChEMBL36 property scan",
    }
    return scatter, lookup_records, summary


def chembl_lookup_record(row: pd.Series, num: pd.Series, score: Any, index: Any, band: str) -> Dict[str, Any]:
    pref = clean_value(row.get("pref_name"))
    cid = clean_value(row.get("chembl_id"))
    name = pref or cid
    return lookup_record(
        "ChEMBL36 molecule property table", name, row.get("chembl_id"), score, index, band,
        rawClass=band, category="ChEMBL", subclass=(clean_value(row.get("molecule_type")) or "molecule"), source="data/chembl36_molecule_property_table_v2.zip",
        inchikey=row.get("compound_structures__standard_inchi_key"), mw=round_or_none(num.get("mw"), 3), logp=round_or_none(num.get("logp"), 3),
        psa=round_or_none(num.get("psa"), 2), hbd=round_or_none(num.get("hbd"), 0), hba=round_or_none(num.get("hba"), 0),
        rtb=round_or_none(num.get("rtb"), 0), aromaticRings=round_or_none(num.get("aromatic_rings"), 0),
        maxPhase=clean_value(row.get("max_phase")), therapeuticFlag=clean_value(row.get("therapeutic_flag")),
        confidence="C: ChEMBL raw-descriptor proxy; droplet-specific permeability not measured",
        pathway="Descriptor-only proxy for oil/interfacial/material-sink concern; requires assay-specific confirmation.",
        note="Included in lookup because it is named, clinical/development-stage, or among the top-scoring ChEMBL records.",
        extra={
            "molregno": clean_value(row.get("molregno")),
            "full_mwt": round_or_none(row.get("compound_properties__full_mwt"), 3),
            "mw_freebase": round_or_none(row.get("compound_properties__mw_freebase"), 3),
            "heavy_atoms": round_or_none(num.get("heavy_atoms"), 0),
            "num_ro5_violations": round_or_none(num.get("num_ro5_violations"), 0),
        }
    )

# ---------- Figures ----------
def make_figures(bio: pd.DataFrame, payne: pd.DataFrame, chembl_scatter: pd.DataFrame, summaries: Dict[str, Any]):
    # Figure 1: unified pipeline and risk language.
    fig = plt.figure(figsize=(12, 8))
    gs = GridSpec(2, 2, width_ratios=[1.1, 1.0], height_ratios=[0.9, 1.1], hspace=0.36, wspace=0.28)
    ax0 = fig.add_subplot(gs[0, :]); ax0.axis("off")
    formula = "score = 0.70·LogP − 0.012·PSA − 0.18·HBD − 0.06·HBA − 0.001·MW + 0.035·RTB + 0.10·Aromatic rings"
    ax0.text(0.01, 0.92, "Figure 1. Unified DROPLET-RISK atlas language", fontsize=15, fontweight="bold", ha="left", va="top")
    ax0.text(0.01, 0.72, formula, fontsize=12, ha="left", va="top", bbox=dict(boxstyle="round,pad=0.45", facecolor="#f8fafc", edgecolor="#cbd5e1"))
    ax0.text(0.01, 0.45, "index = sigmoid(−2.1 + score); thresholds are applied identically to ChEMBL36 and common-bio atlas rows.", fontsize=11, ha="left")
    flow = [
        ("Source descriptors", "ChEMBL36 export\nBio atlas curated inputs"),
        ("One proxy equation", "same coefficients\nsame index"),
        ("Payne-anchored cutoffs", f"leaky ≥ {LEAKY_CUTOFF:.3f}\nhigh ≥ {HIGH_CUTOFF:.3f}"),
        ("Outputs", "lookup app\nfigures\nmanuscript"),
    ]
    box_x = [0.13, 0.38, 0.63, 0.88]
    for i, ((h, body), x) in enumerate(zip(flow, box_x)):
        label = f"{h}\n{body}"
        ax0.text(
            x, 0.18, label, fontsize=9.6, fontweight="bold", ha="center", va="center",
            linespacing=1.35,
            bbox=dict(boxstyle="round,pad=0.45", fc="white", ec="#475569", lw=1.2),
        )
        if i < len(box_x)-1:
            ax0.annotate(
                "", xy=(box_x[i+1]-0.105, 0.18), xytext=(x+0.105, 0.18),
                arrowprops=dict(arrowstyle="->", lw=1.8, color="#334155", shrinkA=0, shrinkB=0),
            )

    ax1 = fig.add_subplot(gs[1, 0])
    labels = ["ChEMBL36\nrows", "ChEMBL36\nscored", "Bio atlas\nrows", "Bio atlas\nscored", "Payne\ncalibration"]
    vals = [summaries["chembl"]["rows"], summaries["chembl"]["scored"], summaries["bio"]["rows"], summaries["bio"]["scored"], summaries["payne"]["rows"]]
    ax1.bar(range(len(vals)), vals, color=["#94a3b8", "#334155", "#94a3b8", "#334155", "#64748b"])
    ax1.set_yscale("log")
    ax1.set_xticks(range(len(vals)), labels)
    ax1.set_ylabel("Records (log scale)")
    ax1.set_title("A. Data sources processed by the refactor", loc="left", fontweight="bold")
    ax1.grid(axis="y", alpha=0.25)
    for i, v in enumerate(vals):
        ax1.text(i, v*1.12, f"{v:,}", ha="center", va="bottom", fontsize=9)

    ax2 = fig.add_subplot(gs[1, 1])
    rc = summaries["chembl"]["risk_counts"]
    vals2 = [rc.get(k, 0) for k in RISK_ORDER[:-1]]
    ax2.bar([RISK_LABEL[k] for k in RISK_ORDER[:-1]], vals2, color=[RISK_COLORS[k] for k in RISK_ORDER[:-1]])
    ax2.set_ylabel("Scored ChEMBL36 molecules")
    ax2.set_title("B. Full ChEMBL36 atlas risk-band distribution", loc="left", fontweight="bold")
    ax2.grid(axis="y", alpha=0.25)
    for i, v in enumerate(vals2):
        pct = 100*v/max(1, summaries["chembl"]["scored"])
        ax2.text(i, v*1.01, f"{pct:.1f}%", ha="center", va="bottom", fontsize=9)
    write_svg(fig, FIG_DIR / "Figure_1_unified_pipeline.svg")

    # Figure 2: Payne calibration and bridge.
    bridge = pd.read_csv(PROJECT / "tables" / "original" / "Table_S11_Payne_to_ChEMBL_proxy_bridge.csv")
    auroc = pd.read_csv(PROJECT / "tables" / "original" / "Table_S9_AUROC_bootstrap_permutation.csv")
    metrics = pd.read_csv(PROJECT / "tables" / "original" / "Table_S12_Payne_anchored_proxy_threshold_metrics.csv")
    fig = plt.figure(figsize=(14, 9))
    gs = GridSpec(2, 2, wspace=0.30, hspace=0.34)
    ax = fig.add_subplot(gs[:, 0])
    colors = np.where(bridge["payne_high_ge_0p2"] == 1, RISK_COLORS["very_high"], np.where(bridge["payne_leaky_ge_0p05"] == 1, RISK_COLORS["high"], RISK_COLORS["lower"]))
    ax.scatter(bridge["payne_risk"], bridge["chembl_proxy_score_zscaled"], s=65, c=colors, edgecolors="white", linewidths=0.7)
    ax.axhline(LEAKY_CUTOFF, color=RISK_COLORS["leaky"], lw=1.6, ls="--", label=f"leaky cutoff {LEAKY_CUTOFF:.3f}")
    ax.axhline(HIGH_CUTOFF, color=RISK_COLORS["high"], lw=1.6, ls="--", label=f"high cutoff {HIGH_CUTOFF:.3f}")
    for _, r in bridge.sort_values("payne_risk", ascending=False).head(8).iterrows():
        ax.text(r["payne_risk"]+0.025, r["chembl_proxy_score_zscaled"], str(r["payne_compound"]).title(), fontsize=8)
    sp = bridge[["payne_risk", "chembl_proxy_score_zscaled"]].corr(method="spearman").iloc[0, 1]
    pr = bridge[["payne_risk", "chembl_proxy_score_zscaled"]].corr(method="pearson").iloc[0, 1]
    ax.set_xlabel("Payne molecular risk axis (−PC1)")
    ax.set_ylabel("ChEMBL proxy score in bridge space")
    ax.set_title("A. Payne-to-ChEMBL bridge used for threshold anchoring", loc="left", fontweight="bold")
    ax.text(0.03, 0.96, f"n={len(bridge)} matched compounds\nSpearman ρ={sp:.3f}\nPearson r={pr:.3f}", transform=ax.transAxes, ha="left", va="top", bbox=dict(boxstyle="round,pad=0.35", fc="white", ec="#cbd5e1"))
    ax.legend(loc="lower right")
    ax.grid(alpha=0.25)

    ax = fig.add_subplot(gs[0, 1])
    y = np.arange(len(auroc))
    ax.errorbar(auroc["AUROC"], y, xerr=[auroc["AUROC"]-auroc["bootstrap_CI_2p5"], auroc["bootstrap_CI_97p5"]-auroc["AUROC"]], fmt="o", color="#111827", ecolor="#64748b", capsize=3)
    ax.axvline(0.5, color="#94a3b8", lw=1, ls="--")
    ax.set_yticks(y, [str(s).replace(" transfer ", "\ntransfer ") for s in auroc["endpoint"]])
    ax.set_xlim(0.45, 1.02)
    ax.set_xlabel("AUROC with bootstrap 95% CI")
    ax.set_title("B. Internal Payne-axis separation", loc="left", fontweight="bold")
    ax.grid(axis="x", alpha=0.25)

    ax = fig.add_subplot(gs[1, 1])
    m = metrics.copy()
    x = np.arange(len(m))
    width = 0.26
    ax.bar(x-width, m["sensitivity"], width, label="sensitivity", color="#334155")
    ax.bar(x, m["specificity"], width, label="specificity", color="#64748b")
    ax.bar(x+width, m["accuracy"], width, label="accuracy", color="#94a3b8")
    ax.set_xticks(x, ["leaky\n≥0.05", "high\n≥0.2"])
    ax.set_ylim(0, 1.05)
    ax.set_ylabel("Metric")
    ax.set_title("C. Payne-anchored threshold performance", loc="left", fontweight="bold")
    ax.legend(frameon=False)
    ax.grid(axis="y", alpha=0.25)
    write_svg(fig, FIG_DIR / "Figure_2_payne_bridge_calibration.svg")

    # Figure 3: ChEMBL atlas.
    hist = pd.read_csv(TABLE_DIR / "chembl36_full_score_histogram.csv")
    phase = pd.read_csv(TABLE_DIR / "chembl36_full_risk_by_max_phase.csv")
    fig = plt.figure(figsize=(15, 10))
    gs = GridSpec(2, 2, width_ratios=[1.25, 1.0], height_ratios=[1.0, 1.0], wspace=0.28, hspace=0.33)
    ax = fig.add_subplot(gs[0, 0])
    centers = (hist["bin_left"] + hist["bin_right"]) / 2
    widths = hist["bin_right"] - hist["bin_left"]
    ax.bar(centers, hist["n"], width=widths, color="#475569", alpha=0.85)
    ax.axvline(LEAKY_CUTOFF, color=RISK_COLORS["leaky"], ls="--", lw=1.6)
    ax.axvline(HIGH_CUTOFF, color=RISK_COLORS["high"], ls="--", lw=1.6)
    ax.axvline(VERY_HIGH_CUTOFF, color=RISK_COLORS["very_high"], ls="--", lw=1.6)
    ax.set_xlabel("Unified proxy score")
    ax.set_ylabel("ChEMBL36 molecules")
    ax.set_title("A. ChEMBL priority/clinical score distribution", loc="left", fontweight="bold")
    ax.grid(axis="y", alpha=0.25)

    ax = fig.add_subplot(gs[0, 1])
    phase_keep = ["missing", "0.5", "1.0", "2.0", "3.0", "4.0", "-1.0"]
    pvt = phase[phase["max_phase"].isin(phase_keep)].pivot_table(index="max_phase", columns="risk_band", values="n", aggfunc="sum", fill_value=0)
    pvt = pvt.reindex([p for p in phase_keep if p in pvt.index])
    bottom = np.zeros(len(pvt))
    for rb in RISK_ORDER[:-1]:
        vals = pvt[rb].to_numpy() if rb in pvt else np.zeros(len(pvt))
        ax.bar(np.arange(len(pvt)), vals, bottom=bottom, color=RISK_COLORS[rb], label=RISK_LABEL[rb])
        bottom += vals
    ax.set_yscale("log")
    ax.set_xticks(np.arange(len(pvt)), pvt.index.astype(str), rotation=45, ha="right")
    ax.set_ylabel("Molecules (log scale)")
    ax.set_title("B. Risk bands by max_phase in lookup subset", loc="left", fontweight="bold")
    ax.legend(frameon=False, ncol=2)
    ax.grid(axis="y", alpha=0.25)

    ax = fig.add_subplot(gs[1, 0])
    sample = chembl_scatter.dropna(subset=["alogp", "psa", "score"]).copy()
    if len(sample) > 50000:
        sample = sample.sample(50000, random_state=42)
    for rb in RISK_ORDER[:-1]:
        sub = sample[sample["band"] == rb]
        if len(sub):
            ax.scatter(sub["alogp"], sub["psa"], s=5, alpha=0.35, color=RISK_COLORS[rb], label=RISK_LABEL[rb], rasterized=False)
    ax.set_xlim(-6, 12)
    ax.set_ylim(0, 260)
    ax.set_xlabel("ALogP")
    ax.set_ylabel("Polar surface area (Ų)")
    ax.set_title("C. Sampled ChEMBL chemical-space projection", loc="left", fontweight="bold")
    ax.grid(alpha=0.25)
    ax.legend(frameon=False, ncol=4, loc="upper right")

    ax = fig.add_subplot(gs[1, 1]); ax.axis("off")
    top = pd.read_csv(TABLE_DIR / "chembl36_top10000_by_unified_score_lookup_records.csv").head(12)
    top = top[["name", "id", "score", "logp", "psa", "riskClass"]]
    table_data = [[str(r["name"])[:22], str(r["id"]), f"{r['score']:.2f}", f"{r['logp']:.1f}", f"{r['psa']:.0f}", str(r["riskClass"]).replace("_", " ")] for _, r in top.iterrows()]
    tbl = ax.table(cellText=table_data, colLabels=["name", "ChEMBL", "score", "LogP", "PSA", "band"], loc="center", cellLoc="left", colLoc="left")
    tbl.auto_set_font_size(False); tbl.set_fontsize(8); tbl.scale(1, 1.25)
    ax.set_title("D. Highest-scoring ChEMBL priority/clinical records", loc="left", fontweight="bold")
    write_svg(fig, FIG_DIR / "Figure_3_chembl36_unified_atlas.svg")

    # Figure 4: Bio atlas.
    fig = plt.figure(figsize=(15, 10))
    gs = GridSpec(2, 2, width_ratios=[1.3, 1.0], height_ratios=[1.0, 1.0], wspace=0.30, hspace=0.33)
    ax = fig.add_subplot(gs[:, 0])
    bsc = bio[bio["droplet_risk_scored"]].copy()
    for rb in RISK_ORDER[:-1]:
        sub = bsc[bsc["droplet_risk_band"] == rb]
        if len(sub):
            ax.scatter(sub["_dr_logp"], sub["_dr_psa"], s=55, alpha=0.82, color=RISK_COLORS[rb], edgecolors="white", linewidths=0.4, label=f"{RISK_LABEL[rb]} (n={len(sub)})")
    anchors = {"glucose", "lactate", "dopamine", "serotonin", "cholesterol", "progesterone", "alpha-tocopherol", "phylloquinone", "rhodamine b", "atp", "oleic acid", "cortisol", "acetylcholine", "oxygen", "carbon dioxide"}
    label_df = pd.concat([bsc[bsc["molecule_or_class"].astype(str).str.lower().isin(anchors)], bsc.sort_values("droplet_risk_score", ascending=False).head(10), bsc.sort_values("droplet_risk_score").head(8)]).drop_duplicates(subset=["molecule_or_class"])
    for _, r in label_df.iterrows():
        x, y = r["_dr_logp"], r["_dr_psa"]
        if pd.isna(x) or pd.isna(y):
            continue
        ha = "left" if x < 6 else "right"
        dx = 0.12 if ha == "left" else -0.12
        ax.text(x+dx, y+4, str(r["molecule_or_class"]), fontsize=8, ha=ha, va="bottom")
    ax.set_xlabel("ALogP / best available LogP")
    ax.set_ylabel("Polar surface area (Ų)")
    ax.set_title("A. Common biological molecule atlas scored with the same proxy", loc="left", fontweight="bold")
    ax.grid(alpha=0.25)
    ax.legend(frameon=True, loc="upper right")

    ax = fig.add_subplot(gs[0, 1])
    cat = bsc.groupby(["category", "droplet_risk_band"]).size().reset_index(name="n")
    cats = bsc["category"].value_counts().head(8).index.tolist()
    pvt = cat[cat["category"].isin(cats)].pivot_table(index="category", columns="droplet_risk_band", values="n", fill_value=0)
    pvt = pvt.reindex(cats)
    pcts = pvt.div(pvt.sum(axis=1), axis=0) * 100
    left = np.zeros(len(pcts))
    y = np.arange(len(pcts))
    for rb in RISK_ORDER[:-1]:
        vals = pcts[rb].to_numpy() if rb in pcts else np.zeros(len(pcts))
        ax.barh(y, vals, left=left, color=RISK_COLORS[rb], label=RISK_LABEL[rb])
        left += vals
    ax.set_yticks(y, [c.replace("_", " ") for c in pcts.index])
    ax.invert_yaxis()
    ax.set_xlabel("Percent of scored rows")
    ax.set_title("B. Risk composition by biological category", loc="left", fontweight="bold")
    ax.legend(frameon=False, ncol=2)
    ax.grid(axis="x", alpha=0.25)

    ax = fig.add_subplot(gs[1, 1]); ax.axis("off")
    top = bsc.sort_values("droplet_risk_score", ascending=False).head(14)
    table_data = [[str(r["molecule_or_class"])[:24], str(r["category"]).replace("_", " ")[:18], f"{r['droplet_risk_score']:.2f}", f"{r['_dr_logp']:.1f}", f"{r['_dr_psa']:.0f}"] for _, r in top.iterrows()]
    tbl = ax.table(cellText=table_data, colLabels=["molecule", "category", "score", "LogP", "PSA"], loc="center", cellLoc="left", colLoc="left")
    tbl.auto_set_font_size(False); tbl.set_fontsize(8); tbl.scale(1, 1.25)
    ax.set_title("C. Highest-scoring common-bio rows", loc="left", fontweight="bold")
    write_svg(fig, FIG_DIR / "Figure_4_common_bio_unified_atlas.svg")

    # Figure 5: lookup/database construction.
    fig = plt.figure(figsize=(12, 8))
    gs = GridSpec(2, 2, hspace=0.36, wspace=0.28)
    ax = fig.add_subplot(gs[0, 0])
    lookup_json = DATA_DIR / "lookup_records_unified.json"
    if lookup_json.exists():
        with lookup_json.open(encoding="utf-8") as fh:
            lookup_counts = Counter(r.get("dataset", "unknown") for r in json.load(fh))
        labels_counts = sorted(lookup_counts.items(), key=lambda kv: kv[1], reverse=True)
        labels = [k for k, _ in labels_counts]
        values = [v for _, v in labels_counts]
        ax.bar(range(len(values)), values, color="#334155")
        ax.set_xticks(range(len(values)), [label.replace(" ", "\n") for label in labels], rotation=0)
        ax.set_ylabel("Interactive lookup records")
        for i, v in enumerate(values):
            ax.text(i, v*1.01, f"{v:,}", ha="center", va="bottom", fontsize=8)
    ax.set_title("A. Embedded lookup index composition", loc="left", fontweight="bold")
    ax.grid(axis="y", alpha=0.25)

    ax = fig.add_subplot(gs[0, 1])
    x = ["ChEMBL36", "Bio atlas"]
    scored_pct = [100*summaries["chembl"]["scored"]/summaries["chembl"]["rows"], 100*summaries["bio"]["scored"]/summaries["bio"]["rows"]]
    ax.bar(x, scored_pct, color=["#475569", "#64748b"])
    ax.set_ylim(0, 105)
    ax.set_ylabel("Rows with all proxy descriptors (%)")
    ax.set_title("B. Descriptor completeness for the unified score", loc="left", fontweight="bold")
    for i, v in enumerate(scored_pct): ax.text(i, v+1.2, f"{v:.1f}%", ha="center")
    ax.grid(axis="y", alpha=0.25)

    ax = fig.add_subplot(gs[1, :]); ax.axis("off")
    txt = (
        "The browser lookup embeds a performance-safe index: all common-bio rows, all Payne calibration records, "
        "and recomputed ChEMBL priority/clinical rows. Database-scale ChEMBL counts are summarized in generated tables and figures. "
        "Figures, manuscript text, lookup cards, and exported CSVs use the same coefficients, index, cutoffs, and risk-band vocabulary."
    )
    ax.text(0.02, 0.78, "C. Static deliverable design", fontsize=12, fontweight="bold", ha="left")
    ax.text(0.02, 0.55, textwrap.fill(txt, 135), fontsize=11, ha="left", va="top", bbox=dict(boxstyle="round,pad=0.45", fc="#f8fafc", ec="#cbd5e1"))
    write_svg(fig, FIG_DIR / "Figure_5_database_lookup_construction.svg")

# ---------- HTML ----------
def build_html(summaries: Dict[str, Any], lookup_records: List[Dict[str, Any]], script_text: str):
    figure_paths = [
        FIG_DIR / "Figure_1_unified_pipeline.svg",
        FIG_DIR / "Figure_2_payne_bridge_calibration.svg",
        FIG_DIR / "Figure_3_chembl36_unified_atlas.svg",
        FIG_DIR / "Figure_4_common_bio_unified_atlas.svg",
        FIG_DIR / "Figure_5_database_lookup_construction.svg",
    ]
    figures_html = "\n".join(
        f'<figure><img src="{fig_to_data_uri(p)}" alt="{p.stem}"><figcaption>{html.escape(p.stem.replace("_", " "))}</figcaption></figure>'
        for p in figure_paths
    )
    records_json = escape_js_json(lookup_records)
    summaries_json = escape_js_json(summaries)
    code_html = html.escape(script_text)
    css = """
:root{--bg:#f5f7fb;--card:#fff;--ink:#111827;--muted:#64748b;--line:#dbe3ee;--accent:#1f6feb;--lower:#2c7bb6;--leaky:#fdae61;--high:#d7191c;--very_high:#7f0000;--unscored:#8c96a3}
*{box-sizing:border-box} body{margin:0;font-family:Arial,Helvetica,sans-serif;background:var(--bg);color:var(--ink)}
header{background:linear-gradient(135deg,#0f172a,#1e293b);color:white;padding:30px 28px 22px} header h1{margin:0;text-align:center;font-size:32px} header p{max-width:1100px;margin:10px auto 0;text-align:center;color:#dbeafe;line-height:1.45}
.tabs{position:sticky;top:0;z-index:3;display:flex;gap:8px;background:#e9eef7;padding:10px 18px;border-bottom:1px solid var(--line)}.tabbtn{border:1px solid var(--line);background:#fff;border-radius:999px;padding:10px 16px;cursor:pointer;font-weight:700}.tabbtn.active{background:#111827;color:#fff;border-color:#111827}
.tab{display:none;max-width:1480px;margin:0 auto;padding:20px}.tab.active{display:block}.grid{display:grid;gap:14px}.cards{grid-template-columns:repeat(5,1fr)}.card,.panel,.article{background:var(--card);border:1px solid var(--line);border-radius:16px;padding:15px;box-shadow:0 8px 24px rgba(15,23,42,.05)}.card h3,.panel h3{margin:0 0 8px 0;font-size:14px}.big{font-size:28px;font-weight:800}.small,.muted{color:var(--muted);font-size:12px;line-height:1.45}
.searchbar{display:flex;gap:10px;align-items:center;margin:14px 0}.searchbar input{flex:1;padding:14px 16px;border:1px solid var(--line);border-radius:12px;font-size:18px}.searchbar button,.chip,select{border:1px solid var(--line);border-radius:10px;background:#fff;padding:10px 12px;cursor:pointer}.primary{background:var(--accent)!important;color:#fff!important;border-color:var(--accent)!important}.toolbar{display:grid;grid-template-columns:2fr 1.4fr 1fr 1fr;gap:12px}.chips{display:flex;flex-wrap:wrap;gap:7px}.chip.active{background:#111827;color:#fff;border-color:#111827}.main{display:grid;grid-template-columns:1.55fr .95fr;gap:16px;margin-top:16px;align-items:start}.table-wrap{overflow:auto;max-height:70vh;border:1px solid var(--line);border-radius:12px}table{width:100%;border-collapse:collapse;background:white;font-size:13px}th,td{padding:9px;border-bottom:1px solid #edf2f7;text-align:left;vertical-align:top}th{position:sticky;top:0;background:#0f172a;color:white;z-index:1}tr:hover{background:#f8fafc;cursor:pointer}tr.selected{background:#e8f0fe}.badge{display:inline-block;border-radius:999px;color:#fff;font-weight:700;padding:4px 9px;font-size:12px;white-space:nowrap}.lower{background:var(--lower)}.leaky{background:var(--leaky);color:#111}.high{background:var(--high)}.very_high{background:var(--very_high)}.unscored{background:var(--unscored)}.detail-grid{display:grid;grid-template-columns:1fr 1fr;gap:9px}.kv{background:#f8fafc;border:1px solid var(--line);border-radius:10px;padding:9px}.k{font-size:11px;color:var(--muted);text-transform:uppercase;letter-spacing:.03em}.v{font-size:14px;margin-top:4px;word-break:break-word}.section{border-top:1px solid var(--line);margin-top:12px;padding-top:12px}.article{max-width:1040px;margin:0 auto;font-size:16px;line-height:1.62}.article h2{margin-top:28px;border-bottom:1px solid var(--line);padding-bottom:5px}.article figure{margin:28px 0}.article img{width:100%;height:auto;border:1px solid var(--line);border-radius:12px;background:white}.article figcaption{color:var(--muted);font-size:13px;margin-top:7px}pre{white-space:pre-wrap;background:#0b1020;color:#d1e7ff;border-radius:12px;padding:18px;overflow:auto;font-size:12px;line-height:1.45}.note{background:#fff7ed;border:1px solid #fed7aa;border-radius:12px;padding:12px;margin:12px 0}.examples{display:flex;gap:7px;flex-wrap:wrap}.extra-table{max-height:240px;overflow:auto;border:1px solid var(--line);border-radius:10px;margin-top:8px}.extra-table table{font-size:12px} .formula{font-family:Georgia,serif;background:#f8fafc;border:1px solid var(--line);border-radius:12px;padding:12px}
@media(max-width:1100px){.cards{grid-template-columns:repeat(2,1fr)}.main,.toolbar{grid-template-columns:1fr}.article{max-width:100%}}@media(max-width:720px){.cards{grid-template-columns:1fr}.tabs{flex-wrap:wrap}.searchbar{flex-direction:column;align-items:stretch}}
"""
    manuscript = f"""
<section class="article">
<h1>DROPLET-RISK: a unified descriptor atlas for droplet small-molecule leakage prioritization</h1>
<p><strong>Short title:</strong> Unified DROPLET-RISK atlas. <strong>Author:</strong> Matteo Broketa, <a href="https://orcid.org/0000-0001-7825-1174" target="_blank" rel="noopener">ORCID 0000-0001-7825-1174</a>. <strong>Keywords:</strong> droplet microfluidics; molecular crosstalk; ChEMBL; PubChem; LINCS; OPERA/CompTox; bioactive molecules; chemical descriptors; lookup database.</p>
<h2>Abstract</h2>
<p>Small-molecule crosstalk can undermine droplet microfluidic assays when reporters, metabolites, perturbagens, or secreted products partition through oil, surfactant assemblies, interfaces, headspace, or device materials. DROPLET-RISK provides a unified descriptor atlas in which ChEMBL36 molecules and common biological molecules are scored by the same proxy equation, transformed by the same bounded index, assigned by the same Payne-anchored thresholds, and displayed through the same lookup vocabulary. The pipeline processed {summaries['chembl']['rows']:,} ChEMBL36 rows, scored {summaries['chembl']['scored']:,} rows with complete proxy descriptors, scored {summaries['bio']['scored']:,} of {summaries['bio']['rows']:,} curated common-bio records, and retains the 36-compound Payne droplet ESI-MS set as empirical calibration rather than as a competing prediction dataset.</p>
<h2>Introduction</h2>
<p>Droplet microfluidics partitions biochemical reactions into large populations of isolated aqueous compartments, enabling high-throughput single-cell analysis, enzyme screening, digital nucleic-acid assays, microbial phenotyping, and chemical perturbation workflows. These applications depend on a central assumption: molecules generated, consumed, or assayed inside one droplet remain associated with that droplet over the experimental timescale. For nucleic acids and many macromolecules this assumption is often reasonable. For small molecules it is frequently unsafe.</p>
<p>Small-molecule leakage is chemically selective. Payne and co-workers used droplet electrospray-ionization mass spectrometry to evaluate low-molecular-weight analyte transfer between microfluidic droplets, expanding the problem beyond fluorescent dyes. Their 36-compound crosstalk table shows that transfer is not explained by molecular weight alone; it reflects solvated structure, charge state, lipophilicity, polarity, and interactions with the carrier phase and surfactant layer.</p>
<p>Barrier chemistry changes the outcome for a fixed molecule. Dendronized oligo-glycerol fluorosurfactants reduce small-molecule transfer relative to conventional PEG/PFPE surfactants. Nanoparticle/Pickering stabilization can suppress some surfactant-mediated leakage but does not eliminate all hydrophobic dye leakage, indicating pathway-selective rather than universal protection. Film-forming surfactants further show that chemically reinforced interfaces can stabilize droplets through workflows such as PCR, again defining a barrier class rather than a molecule-only rule.</p>
<p>Descriptor resources such as ChEMBL, PubChem, LINCS, and OPERA models available through EPA CompTox provide valuable structure, property, and perturbagen context, but they are not droplet-specific permeability models. Caco-2, PAMPA, membrane-permeability, logD, vapor-pressure, and solubility tools do not directly capture reverse micelles, fluorinated oils, charged surfactant head groups, droplet geometry, or material sinks. DROPLET-RISK therefore uses these resources as descriptor infrastructure and interpretation context, not as substitutes for droplet leakage measurements.</p>
<h2>Unified scoring model</h2>
<div class="formula">score = 0.70·LogP − 0.012·PSA − 0.18·HBD − 0.06·HBA − 0.001·MW + 0.035·RTB + 0.10·aromatic rings; index = sigmoid(−2.1 + score).</div>
<p>The descriptor-only score is intended for ranking and triage. It is not an absolute permeability, leakage constant, or substitute for droplet-specific transfer measurements. Risk bands are lower below {LEAKY_CUTOFF:.3f}, leaky from {LEAKY_CUTOFF:.3f} to {HIGH_CUTOFF:.3f}, high from {HIGH_CUTOFF:.3f} to {VERY_HIGH_CUTOFF:.1f}, and very high at index ≥ 0.50. These labels are applied identically to every ChEMBL and common-bio atlas row with complete descriptors. The lookup table reports the score, index, risk band, descriptor values, source dataset, confidence language, and interpretation note so that a reader can trace why a molecule was prioritized.</p>
{figures_html.split('</figure>')[0]}</figure>
<h2>Calibration and interpretation</h2>
<p>The Payne 36-compound set is the empirical droplet crosstalk anchor. The sign-inverted Payne PC1 axis is retained for internal characterization and bridge calibration. The ChEMBL/common-bio atlas uses the raw descriptor proxy above, not the normalized Payne PC1 axis. The Payne-to-ChEMBL bridge contains {summaries['payne']['bridge_rows']} matched compounds with Spearman correlation {summaries['payne']['bridge_spearman']:.3f} and Pearson correlation {summaries['payne']['bridge_pearson']:.3f} between Payne risk and ChEMBL proxy bridge score. Payne-anchored thresholds are used to avoid the prior degenerate unanchored classification behavior.</p>
<p>Interpretation should follow the risk vocabulary rather than over-reading the numeric value. A lower molecule is below the Payne-anchored leaky proxy cutoff. A leaky molecule is above the first cutoff and should be checked when it is assay-critical. A high molecule should be screened for oil partition, interfacial transport, surfactant-mediated transfer, and material sinks. A very-high molecule reaches index ≥ 0.50 and should be treated as a red-list candidate until experimentally cleared in the exact droplet system.</p>
{figures_html.split('</figure>')[1]}</figure>
<h2>Database-scale atlas</h2>
<p>The ChEMBL36 property export in <code>data/chembl36_molecule_property_table_v2.zip</code> defines the database-scale descriptor source. Of the scored ChEMBL rows, {summaries['chembl']['risk_counts']['lower']:,} are lower, {summaries['chembl']['risk_counts']['leaky']:,} are leaky, {summaries['chembl']['risk_counts']['high']:,} are high, and {summaries['chembl']['risk_counts']['very_high']:,} are very high by the unified proxy. The browser lookup embeds a performance-safe index derived from the curated common-bio table, the Payne calibration table, and the recomputed ChEMBL priority/clinical lookup tables. The full database counts are summarized in the generated tables and figures rather than embedded row-for-row into the browser.</p>
{figures_html.split('</figure>')[2]}</figure>
<h2>Common biological molecules</h2>
<p>The common-bio atlas is not a separate figure-only dataset. It is scored by the same proxy, appears in the same lookup app, uses the same risk-band thresholds, and preserves curated context fields such as category, subclass, expected droplet pathway, missing measurements, and interpretation notes. This makes metabolites, assay readouts, gases/volatiles, lipids, secreted signals, proteins/particles, and reactive species visible in the same language as ChEMBL molecules.</p>
{figures_html.split('</figure>')[3]}</figure>
<h2>Database construction and reproducibility</h2>
<p>The build procedure uses one code path for descriptor normalization, proxy scoring, index transformation, risk-band assignment, table export, figure generation, and lookup-record construction. The same constants shown in the equation are used throughout the manuscript figures, lookup tool, and sequential code tab. This prevents the reader from having to reconcile different datasets, thresholds, or naming schemes across the publication package.</p>
{figures_html.split('</figure>')[4]}</figure>
<h2>Limitations</h2>
<p>DROPLET-RISK is a transparent prioritization index. It does not model oil identity, surfactant concentration, droplet size, temperature, pH-dependent microstate distributions, headspace exchange, PDMS/plastic sinks, or barrier chemistry as calibrated universal coefficients. Those system terms remain assay-specific modifiers and should be measured for compounds that drive biological conclusions.</p>
<h2>References</h2>
<ol>
<li>Payne, E. M.; Taraji, M.; Murray, B. E.; Holland-Moritz, D. A.; Moore, J. C.; Haddad, P. R.; Kennedy, R. T. Evaluation of analyte transfer between microfluidic droplets by mass spectrometry. <em>Analytical Chemistry</em> 2023, 95, 4662-4670. DOI: 10.1021/acs.analchem.2c04985.</li>
<li>Chowdhury, M. S.; Zheng, W.; Kumari, S.; Heyman, J. A.; Zhang, X.; Dey, P.; Weitz, D. A.; Haag, R. Dendronized fluorosurfactant for highly stable water-in-fluorinated oil emulsions with minimal inter-droplet transfer of small molecules. <em>Nature Communications</em> 2019, 10, 4546. DOI: 10.1038/s41467-019-12462-5.</li>
<li>Waeterschoot, J.; et al. The effects of droplet stabilization by surfactants and nanoparticles on leakage, cross-talk, droplet stability, and cell adhesion. <em>RSC Advances</em> 2024. DOI: 10.1039/D4RA04298K.</li>
<li>Deveney, B. T.; Heyman, J. A.; Rosenthal, R. G.; Weitz, D. A.; Werner, J. G. A biocompatible surfactant film for stable microfluidic droplets. <em>Lab on a Chip</em> 2025, 25, 5141-5149. DOI: 10.1039/D5LC00456J.</li>
<li>Mansouri, K.; Grulke, C. M.; Judson, R. S.; Williams, A. J. OPERA models for predicting physicochemical properties and environmental fate endpoints. <em>Journal of Cheminformatics</em> 2018, 10, 10. DOI: 10.1186/s13321-018-0263-1.</li>
<li>EPA CompTox Chemicals Dashboard. OPERA predicted physicochemical and environmental fate properties, accessed through CompTox descriptor resources.</li>
<li>ChEMBL Database Release 36. European Bioinformatics Institute, 2025. DOI: 10.6019/CHEMBL.database.36.</li>
<li>Subramanian, A.; et al. A Next Generation Connectivity Map: L1000 Platform and the First 1,000,000 Profiles. <em>Cell</em> 2017, 171, 1437-1452. DOI: 10.1016/j.cell.2017.10.049.</li>
<li>GEO Series GSE92742, Broad LINCS L1000 Phase I dataset and perturbagen metadata. National Center for Biotechnology Information Gene Expression Omnibus.</li>
<li>Kim, S.; Chen, J.; Cheng, T.; et al. PubChem 2023 update. <em>Nucleic Acids Research</em> 2023, 51, D1373-D1380. DOI: 10.1093/nar/gkac956.</li>
</ol>
</section>
"""
    # The string splitting above is intentionally simple but fragile if counts change; keep all figures present.
    html_text = f"""<!doctype html>
<html lang="en"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width, initial-scale=1"><title>DROPLET-RISK unified atlas</title><style>{css}</style></head>
<body><header><h1>DROPLET-RISK unified atlas</h1><p>Lookup database, manuscript with inline figures, and sequential reproducibility code. All atlas records use the same descriptor proxy, index, thresholds, and risk vocabulary.</p></header>
<nav class="tabs"><button class="tabbtn active" data-tab="lookup">Tab 1 · Lookup database</button><button class="tabbtn" data-tab="article">Tab 2 · Manuscript and figures</button><button class="tabbtn" data-tab="code">Tab 3 · Code and build sequence</button></nav>
<section id="lookup" class="tab active">
<div class="grid cards"><div class="card"><h3>Lookup records</h3><div class="big" id="totalRecords">—</div><div class="small">Embedded browser index</div></div><div class="card"><h3>Full ChEMBL rows</h3><div class="big">{summaries['chembl']['rows']:,}</div><div class="small">Processed from data folder</div></div><div class="card"><h3>Full ChEMBL scored</h3><div class="big">{summaries['chembl']['scored']:,}</div><div class="small">Complete proxy descriptors</div></div><div class="card"><h3>Bio atlas scored</h3><div class="big">{summaries['bio']['scored']:,}</div><div class="small">of {summaries['bio']['rows']:,} rows</div></div><div class="card"><h3>Shown</h3><div class="big" id="shownRecords">—</div><div class="small">After filters/search</div></div></div>
<div class="note"><strong>Scope:</strong> the full ChEMBL scored table is included in the output bundle as a gzipped CSV. This browser index embeds all common-bio rows, all Payne calibration rows, the recomputed ChEMBL priority/clinical rows for responsive static searching.</div>
<div class="searchbar"><input id="q" placeholder="Search molecule, CHEMBL ID, InChIKey, category, class, pathway…"><button id="clearBtn">Clear</button><button id="exportBtn" class="primary">Export filtered CSV</button></div><div class="examples" id="examples"></div>
<div class="toolbar"><div class="panel"><h3>Dataset</h3><div class="chips" id="datasetChips"></div></div><div class="panel"><h3>Risk band</h3><div class="chips" id="riskChips"></div></div><div class="panel"><h3>Category</h3><select id="categoryFilter"></select></div><div class="panel"><h3>Sort</h3><select id="sortBy"><option value="score_desc">Score ↓</option><option value="score_asc">Score ↑</option><option value="index_desc">Index ↓</option><option value="name_asc">Name A–Z</option><option value="dataset_asc">Dataset A–Z</option></select></div></div>
<div class="main"><div class="panel"><h3>Matches</h3><div class="small" id="meta">Loading…</div><div class="table-wrap"><table><thead><tr><th>Dataset</th><th>Name</th><th>Risk</th><th>Score</th><th>Index</th><th>LogP</th><th>PSA</th><th>MW</th><th>ID</th></tr></thead><tbody id="tbody"></tbody></table></div></div><div class="panel"><h3>Record detail</h3><div id="detail" class="muted">Select a row.</div></div></div>
</section>
<section id="article" class="tab">{manuscript}</section>
<section id="code" class="tab"><section class="article"><h1>Sequential build code</h1><p>The following script was used to regenerate the scored tables, figures, lookup/manuscript/code HTML, and bundle from the uploaded project. The same constants are used throughout.</p><pre><code>{code_html}</code></pre></section></section>
<script>const RECORDS={records_json}; const SUMMARY={summaries_json};
const state={{records:RECORDS,filtered:[],selected:0,dataset:new Set(),risk:new Set(),category:'all',sort:'score_desc',q:''}};
const riskOrder=['lower','leaky','high','very_high','unscored'];
function $(id){{return document.getElementById(id)}} function esc(s){{return String(s??'').replace(/[&<>\"]/g,m=>({{'&':'&amp;','<':'&lt;','>':'&gt;','"':'&quot;'}}[m]))}} function fmt(v,d=3){{return v===null||v===undefined||Number.isNaN(Number(v))?'—':Number(v).toFixed(d).replace(/\.?0+$/,'')}}
function showTab(id){{document.querySelectorAll('.tab').forEach(t=>t.classList.toggle('active',t.id===id));document.querySelectorAll('.tabbtn').forEach(b=>b.classList.toggle('active',b.dataset.tab===id))}} document.querySelectorAll('.tabbtn').forEach(b=>b.onclick=()=>showTab(b.dataset.tab));
function chipBox(id,vals,set){{const el=$(id);el.innerHTML='';vals.forEach(v=>{{const b=document.createElement('button');b.className='chip'+(set.has(v)?' active':'');b.textContent=v;b.onclick=()=>{{set.has(v)?set.delete(v):set.add(v);render()}};el.appendChild(b)}})}}
function categories(){{return ['all',...Array.from(new Set(state.records.map(r=>r.category).filter(Boolean))).sort()]}} function fillCats(){{let s=$('categoryFilter'),cur=state.category;s.innerHTML=categories().map(v=>`<option value="${{esc(v)}}">${{esc(v==='all'?'All categories':v)}}</option>`).join('');s.value=cur}}
function match(r){{if(state.q){{let toks=state.q.split(/\s+/).filter(Boolean),hay=r.searchBlob||''; if(!toks.every(t=>hay.includes(t))) return false}} if(state.dataset.size&&!state.dataset.has(r.dataset))return false; if(state.risk.size&&!state.risk.has(r.riskClass))return false; if(state.category!=='all'&&r.category!==state.category)return false; return true}}
function sortRows(rows){{let [field,dir]=state.sort.split('_'),mul=dir==='desc'?-1:1; return rows.sort((a,b)=>{{if(field==='score'||field==='index'){{let av=a[field]??-1e9,bv=b[field]??-1e9;return mul*((av>bv)-(av<bv))}} if(field==='name')return mul*String(a.name).localeCompare(String(b.name)); if(field==='dataset')return mul*String(a.dataset).localeCompare(String(b.dataset)); return 0}})}}
function apply(){{state.filtered=sortRows(state.records.filter(match)); if(state.selected>=state.filtered.length) state.selected=0}}
function renderTable(){{let rows=state.filtered.slice(0,600);$('meta').textContent=`${{state.filtered.length.toLocaleString()}} matching records; rendering first ${{rows.length.toLocaleString()}}.`;$('tbody').innerHTML=rows.map((r,i)=>`<tr class="${{i===state.selected?'selected':''}}" data-i="${{i}}"><td>${{esc(r.dataset)}}</td><td><strong>${{esc(r.name)}}</strong><div class="small">${{esc(r.category||'')}}</div></td><td><span class="badge ${{esc(r.riskClass)}}">${{esc(r.riskClass).replace('_',' ')}}</span><div class="small">${{esc(r.rawClass||'')}}</div></td><td>${{fmt(r.score)}}</td><td>${{fmt(r.index)}}</td><td>${{fmt(r.logp)}}</td><td>${{fmt(r.psa,1)}}</td><td>${{fmt(r.mw,1)}}</td><td>${{esc(r.id||'—')}}<div class="small">${{esc(r.inchikey||'')}}</div></td></tr>`).join('');document.querySelectorAll('#tbody tr').forEach(tr=>tr.onclick=()=>{{state.selected=Number(tr.dataset.i);renderTable();renderDetail()}})}}
function kv(k,v){{return `<div class="kv"><div class="k">${{esc(k)}}</div><div class="v">${{esc(v??'—')}}</div></div>`}}
function renderDetail(){{let r=state.filtered[state.selected]; if(!r){{$('detail').innerHTML='<p>No matching record.</p>';return}} let extra=''; if(r.extra){{extra='<div class="section"><div class="k">Additional source fields</div><div class="extra-table"><table>'+Object.entries(r.extra).map(([k,v])=>`<tr><th>${{esc(k)}}</th><td>${{esc(v)}}</td></tr>`).join('')+'</table></div></div>'}} $('detail').innerHTML=`<h2 style="margin:0 0 4px 0">${{esc(r.name)}}</h2><div class="muted">${{esc(r.dataset)}} · ${{esc(r.category||'')}}</div><div style="margin:8px 0"><span class="badge ${{esc(r.riskClass)}}">${{esc(r.riskClass).replace('_',' ')}}</span></div><div class="detail-grid">${{kv('Score',fmt(r.score))}}${{kv('Index',fmt(r.index))}}${{kv('LogP',fmt(r.logp))}}${{kv('PSA',fmt(r.psa,1))}}${{kv('MW',fmt(r.mw,1))}}${{kv('Observed',fmt(r.observed))}}${{kv('Identifier',r.id)}}${{kv('Confidence',r.confidence)}}${{kv('Max phase',r.maxPhase)}}${{kv('Therapeutic flag',r.therapeuticFlag)}}</div><div class="section"><div class="k">Interpretation</div><div class="v">${{esc(r.interpretation)}}</div></div><div class="section"><div class="k">Pathway</div><div class="v">${{esc(r.pathway||'—')}}</div></div><div class="section"><div class="k">Note</div><div class="v">${{esc(r.note||'—')}}</div></div><div class="section"><div class="k">Source</div><div class="v">${{esc(r.source||'—')}}</div></div>${{extra}}`}}
function render(){{apply();$('totalRecords').textContent=state.records.length.toLocaleString();$('shownRecords').textContent=state.filtered.length.toLocaleString();chipBox('datasetChips',Array.from(new Set(state.records.map(r=>r.dataset))).sort(),state.dataset);chipBox('riskChips',riskOrder,state.risk);renderTable();renderDetail()}}
function exportCSV(){{let cols=['dataset','name','id','riskClass','score','index','observed','logp','psa','mw','category','subclass','confidence','pathway','note'];let lines=[cols.join(',')];state.filtered.forEach(r=>lines.push(cols.map(c=>'"'+String(r[c]??'').replace(/"/g,'""')+'"').join(',')));let blob=new Blob([lines.join('\\n')],{{type:'text/csv'}}),url=URL.createObjectURL(blob),a=document.createElement('a');a.href=url;a.download='droplet_risk_filtered_lookup.csv';a.click();setTimeout(()=>URL.revokeObjectURL(url),500)}}
$('q').oninput=e=>{{state.q=e.target.value.trim().toLowerCase();render()}};$('clearBtn').onclick=()=>{{$('q').value='';state.q='';state.dataset.clear();state.risk.clear();state.category='all';state.sort='score_desc';$('sortBy').value='score_desc';fillCats();render()}};$('exportBtn').onclick=exportCSV;$('sortBy').onchange=e=>{{state.sort=e.target.value;render()}};$('categoryFilter').onchange=e=>{{state.category=e.target.value;render()}};$('examples').innerHTML=['imipramine','diphenhydramine','alpha-tocopherol','phylloquinone','glucose','cholesterol','rhodamine B','ATP'].map(e=>`<button class="chip" data-q="${{esc(e)}}">${{esc(e)}}</button>`).join('');document.querySelectorAll('#examples button').forEach(b=>b.onclick=()=>{{$('q').value=b.dataset.q;state.q=b.dataset.q.toLowerCase();render()}});fillCats();render();
</script></body></html>"""
    out_html = OUT / "DROPLET_RISK_publication_ready_unified_corrected_v2.html"
    out_html.write_text(html_text, encoding="utf-8")
    return out_html

# ---------- Main ----------
def main():
    bio, bio_records, bio_summary = process_bio()
    payne, payne_records, payne_summary = process_payne()
    chembl_scatter, chembl_records, chembl_summary = process_chembl()
    # Combine and write lookup records before Figure 5.
    all_records = payne_records + bio_records + chembl_records
    # Sort lookup records by dataset and descending score.
    all_records.sort(key=lambda r: (r.get("dataset", ""), -(r.get("score") or -999), r.get("name", "")))
    (DATA_DIR / "lookup_records_unified.json").write_text(json.dumps(all_records, ensure_ascii=False, separators=(",", ":")), encoding="utf-8")
    summaries = {"bio": bio_summary, "payne": payne_summary, "chembl": chembl_summary, "lookup_records": len(all_records)}
    (OUT / "DROPLET_RISK_unified_summary.json").write_text(json.dumps(summaries, indent=2), encoding="utf-8")
    make_figures(bio, payne, chembl_scatter, summaries)
    script_text = Path(__file__).read_text(encoding="utf-8")
    html_path = build_html(summaries, all_records, script_text)
    # Copy script and summary into bundle root for direct reuse.
    shutil.copy2(Path(__file__), OUT / "DROPLET_RISK_refactored_pipeline_corrected_v2.py")
    # Create a compact ZIP package.
    zip_path = Path("/mnt/data/droplet_risk_refactor/DROPLET_RISK_publication_ready_bundle_corrected_v2.zip")
    if zip_path.exists():
        zip_path.unlink()
    shutil.make_archive(str(zip_path.with_suffix("")), "zip", OUT)
    print(json.dumps({"html": str(html_path), "bundle": str(zip_path), "summary": summaries}, indent=2))

if __name__ == "__main__":
    main()