Project

project picture

Summer 2026 challenge: phase 2 contestant

Rational Design of Isoform-Selective CA II Inhibitors through Chemical Space Exploration

Ernestas Urniežius, Institute of Biotechnology, Vilnius University, Vilnius, Lithuania

During the first three months, the project advanced from data and structure curation to a tiered chemical-space and selectivity workflow. Two main obstacles emerged: high conservation of CA active sites and the scale of accessible sulfamoyl chemistry, which made simultaneous isozyme-wide modeling and exhaustive screening difficult to perform. We addressed these by focusing on CA II with CA I, CA VII and CA IX as initial comparators, defining objective scaffold and physicochemical reduction criteria, and combining ligand-based analysis with protein-side information from PCM and mixed-solvent MD. Chemical space mapping, clustering and scaffold analysis converted a very large search space into tractable scaffold neighborhoods. The project has therefore moved from broad exploration toward restrained virtual screening and CADD based bottom-up cycle in which promising scaffolds can be experimentally validated first and their local chemical neighborhoods analyzed in the next cycle.
After 3 months, Ernestas has achieved the following milestones:
  1. Milestone 1 was reached. A structural reference set spanning 12 catalytically active human CA isozymes was curated, with 115 experimental structures selected across 10 isozymes. CA VA and CA VB currently lack experimentally solved structures. CA II is represented by 57 unique protein–ligand structures, while PLBD provides binding affinity data for all isozymes and linked structural data for several isozymes. Protein structures were prepared with MODELLER, PROPKA and H++, and ligands with Open Babel and RDKit. Literature, sequence data and superimposed structures were analyzed in SeeSAR, PyMOL and ChimeraX to map conserved and isozyme-variable active-site residues. Because these differences are subtle, the initial comparative workflow focused on CA II with CA I, CA VII and CA IX as main comparators. Selection criteria include pose plausibility, isozyme-variable interactions, scaffold diversity, physicochemical and pharmacokinetic properties, binding thermodynamics and accessibility.
  2. Milestone 2 has substantially progressed. Primary sulfamoyl groups were screened with SpaceMACS as the essential Zn-binding motif. SpaceMACS mapping of Enamine REAL identified at least 27 million compounds spanning primary sulfonamides and related sulfamate and sulfamide isosteres across 10–61 heavy atoms bins. This is a lower-bound estimate because several heavy-atom bins reached the 1 million retrieval limit, preventing full enumeration of the space. The scale of the search was addressed by hierarchical reduction using clustering methods, infiniSee and RDKit. A pragmatic 10-26 HA window was selected for focused scaffold exploration to cover broad chemistry, from monocyclic to linked or fused ring systems. Direct ring systems and ring-containing chemotypes with ≤2 linker atoms between the sulfonamide sulfur and scaffold were explored. Current strategy uses similarity methods, local neighborhoods, chemotype diversity, physicochemical and pharmacokinetic properties, and accessibility.
  3. Proteochemometric models trained on curated CA–sulfonamide affinity data from PLBD database reach Pearson r 0.76-0.88 and RMSE = 3.7-6.3 kJ/mol across isozyme-specific test sets. ML SHAP analysis provides residue-level importance signals associated with recognition and isozyme selectivity, while mixed-solvent MD maps hydrophobic and polar hotspots and hydration patterns. Together, these data guide pharmacophoric restraints for tethered rDock screening. SeeSAR template docking with constraints, HYDE and free-energy calculations will support consensus ranking. Because PCM training data are benzenesulfonamide dominated, predictions will be weighted using defined applicability domain and confidence estimates. Approximately 20–50 scaffolds and their small-substituent derivatives will be selected for validation based on computational evaluation and accessibility. Later cycles will sample derivatives around validated scaffolds through established CADD workflows, with optional xREAL expansion.