← Back to Virtual Screening

Methods & validation

Method

For each target, a per-target QSAR model (or, for targets with few known ligands, a 2D interaction pharmacophore) ranks an 8.4M in-stock library; a diversity-selected ~10k shortlist is docked with Uni-Dock into one curated structure. Each target is then ranked by whichever of Vina or gnina CNNaffinity measurably ranks potency better for that target, decided per target from a censored concordance index over ChEMBL compounds rather than by protein class. The top 1000 per target are kept.

Post-filtering: physically clashing poses (Vina > 0) and PAINS are removed; zinc-enzyme and aminergic shortlists are restricted to the appropriate zinc-binding-group / basic-amine chemotype. Scores are relative ranking scores, not calibrated affinities.

Some targets are withdrawn from the catalog rather than served with a screen we do not believe.

Validation 1 — the ranking check

The question a ranked hit list has to answer is whether the score knows a strong binder from a weak one. For each target we take its known ChEMBL compounds with a measured potency, dock them into the same curated structure the campaign used, and count: given two compounds whose measured potency differs by at least 100×, how often does the score put the stronger one first? Guessing scores 50%. Perfect ordering scores 100%.

Only pairs whose true order is unambiguous are counted — a compound reported as “> 10 µM” is treated as a bound, never as an exact value, so censored measurements cannot manufacture agreement. The interval is a 1000-resample bootstrap over compounds; a target passes only when the lower end of that interval clears 50%. The same test is run for Vina and for gnina CNNaffinity on the identical compound set, and each target is then ranked by whichever won — that is where the per-target choice of score comes from. (The statistic is a censored concordance index, or C-index, if you want the name for it.)

A target that does not pass is not necessarily wrong; most often it simply has too few compounds with potencies far enough apart to resolve anything. But we do not know that its ordering is any better than chance, and the target page says so. A target marked no data could not be tested at all — ChEMBL holds too few compounds for it that are both library-eligible and separated enough in measured potency for any pair to count. That is a fact about the published pharmacology, not about the docking, and no further computing changes it.

Validation 2 — known-active recovery

Each target's top 1000 is matched by scaffold (InChIKey skeleton) against its ChEMBL actives (pChEMBL ≥ 7). A target “recovers” a known active when that active's scaffold appears in its top 1000. Because most known actives are research compounds not purchasable in the library, this is a conservative signal: recovery means the docking floated a genuine active to the top of the in-stock matter that was available.

Target Class ranking check # known # recovered best rank top-100 sim

“ranking check” = how often the score puts the more potent of two compounds first, over pairs at least 100× apart, with the 95% bootstrap interval in brackets; bold green means the interval clears 50%. “not tested” means the measurement has not been run for that target, which is not the same as failing it. “top-100 sim” = median Morgan(2,2048) Tanimoto of the top-100 hits to their nearest known active (how active-like the ranked matter is). Sort any column by clicking its header.

Models cards and validation Cite how to cite these results Models are updated in place, so a result reflects the models as they stood on the day it was produced. Record the date alongside anything you carry forward.