An Explainable Machine Learning-Based QSAR Framework for Predicting Thrombin Inhibitory Activity


Kaya A. O., Emre M. C.

PHARMACEUTICALS, cilt.19, sa.9, ss.1411, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 19 Sayı: 9
  • Basım Tarihi: 2026
  • Doi Numarası: 10.3390/ph19091411
  • Dergi Adı: PHARMACEUTICALS
  • Derginin Tarandığı İndeksler: Academic Search Ultimate (EBSCO), Biomedical Reference Collection: Corporate Edition (EBSCO), Scopus, Science Citation Index Expanded (SCI-EXPANDED), EMBASE, Directory of Open Access Journals
  • Sayfa Sayıları: ss.1411
  • Yozgat Bozok Üniversitesi Adresli: Evet

Özet

Background/Objectives: Public thrombin bioactivity records contain heterogeneous endpoints, replicate measurements, related chemical series, and potentially reactive compounds that may bias quantitative structure–activity relationship models. In this study, we developed an explainable, assay-aware, and leakage-safe machine learning framework for predicting thrombin-inhibitory activity. Methods: Exact Ki and IC50 records for human thrombin (CHEMBL204) were standardized, converted to pActivity, and aggregated using predefined criteria. The final dataset comprised 5189 unique compounds represented by development-filtered Mordred descriptors and Morgan fingerprint. The models were optimized using development-only out-of-fold validation and evaluated using random and scaffold-disjoint held-out tests. Results: In the random-split analysis, ConsensusAll achieved an out-of-fold R2 of 0.7588 and a held-out test R2 of 0.7480, with an RMSE of 0.7360, MAE of 0.5307, and concordance correlation coefficient of 0.8540. In the scaffold-disjoint locked test, ConsensusTop5 achieved R2 = 0.5831, RMSE = 0.9484, and MAE = 0.7290, respectively. The applicability domain covered 92.68% of the locked test compounds and yielded R2 = 0.6041. One hundred Y-randomization runs produced a mean R2 of −0.1233 (empirical p = 0.0099). Conclusions: The framework provides useful predictions within the represented chemical space and measurable generalization for unseen scaffolds. This supports compound prioritization, although prospective biochemical validation remains necessary.