{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,12,30]],"date-time":"2024-12-30T18:31:09Z","timestamp":1735583469524},"reference-count":40,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2015,5,20]],"date-time":"2015-05-20T00:00:00Z","timestamp":1432080000000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Cheminform"],"published-print":{"date-parts":[[2015,12]]},"abstract":"Abstract<\/jats:title>\n \n Background<\/jats:title>\n Cheminformaticians are equipped with a very rich toolbox when carrying out molecular similarity calculations. A large number of molecular representations exist, and there are several methods (similarity and distance metrics) to quantify the similarity of molecular representations. In this work, eight well-known similarity\/distance metrics are compared on a large dataset of molecular fingerprints with sum of ranking differences (SRD) and ANOVA analysis. The effects of molecular size, selection methods and data pretreatment methods on the outcome of the comparison are also assessed.<\/jats:p>\n <\/jats:sec>\n \n Results<\/jats:title>\n A supplier database (https:\/\/mcule.com\/<\/jats:ext-link>) was used as the source of compounds for the similarity calculations in this study. A large number of datasets, each consisting of one hundred compounds, were compiled, molecular fingerprints were generated and similarity values between a randomly chosen reference compound and the rest were calculated for each dataset. Similarity metrics were compared based on their ranking of the compounds within one experiment (one dataset) using sum of ranking differences (SRD), while the results of the entire set of experiments were summarized on box and whisker plots. Finally, the effects of various factors (data pretreatment, molecule size, selection method) were evaluated with analysis of variance (ANOVA).<\/jats:p>\n <\/jats:sec>\n \n Conclusions<\/jats:title>\n This study complements previous efforts to examine and rank various metrics for molecular similarity calculations. Here, however, an entirely general approach was taken to neglect any a priori<\/jats:italic> knowledge on the compounds involved, as well as any bias introduced by examining only one or a few specific scenarios. The Tanimoto index, Dice index, Cosine coefficient and Soergel distance were identified to be the best (and in some sense equivalent) metrics for similarity calculations, i.e<\/jats:italic>. these metrics could produce the rankings closest to the composite (average) ranking of the eight metrics. The similarity metrics derived from Euclidean and Manhattan distances are not recommended on their own, although their variability and diversity from other similarity metrics might be advantageous in certain cases (e.g.<\/jats:italic> for data fusion). Conclusions are also drawn regarding the effects of molecule size, selection method and data pretreatment on the ranking behavior of the studied metrics.<\/jats:p>\n <\/jats:sec>","DOI":"10.1186\/s13321-015-0069-3","type":"journal-article","created":{"date-parts":[[2015,5,19]],"date-time":"2015-05-19T15:35:55Z","timestamp":1432049755000},"update-policy":"http:\/\/dx.doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":914,"title":["Why is Tanimoto index an appropriate choice for fingerprint-based similarity calculations?"],"prefix":"10.1186","volume":"7","author":[{"given":"D\u00e1vid","family":"Bajusz","sequence":"first","affiliation":[]},{"given":"Anita","family":"R\u00e1cz","sequence":"additional","affiliation":[]},{"given":"K\u00e1roly","family":"H\u00e9berger","sequence":"additional","affiliation":[]}],"member":"297","published-online":{"date-parts":[[2015,5,20]]},"reference":[{"key":"69_CR1","doi-asserted-by":"publisher","first-page":"3204","DOI":"10.1039\/b409813g","volume":"2","author":"A Bender","year":"2004","unstructured":"Bender A, Glen RC. Molecular similarity: a key technique in molecular informatics. Org Biomol Chem. 2004;2:3204\u201318.","journal-title":"Org Biomol Chem"},{"key":"69_CR2","doi-asserted-by":"publisher","first-page":"3186","DOI":"10.1021\/jm401411z","volume":"57","author":"G Maggiora","year":"2014","unstructured":"Maggiora G, Vogt M, Stumpfe D, Bajorath J. Molecular similarity in medicinal chemistry. J Med Chem. 2014;57:3186\u2013204.","journal-title":"J Med Chem"},{"key":"69_CR3","doi-asserted-by":"publisher","first-page":"225","DOI":"10.1023\/A:1027221424359","volume":"9-11","author":"H Kubinyi","year":"1998","unstructured":"Kubinyi H. Similarity and dissimilarity: a medicinal chemist\u2019s view. Perspect Drug Discov Des. 1998;9-11:225\u201352.","journal-title":"Perspect Drug Discov Des"},{"key":"69_CR4","doi-asserted-by":"publisher","first-page":"225","DOI":"10.1016\/j.drudis.2007.01.011","volume":"12","author":"H Eckert","year":"2007","unstructured":"Eckert H, Bajorath J. Molecular similarity analysis in virtual screening: foundations, limitations and novel approaches. Drug Discov Today. 2007;12:225\u201333.","journal-title":"Drug Discov Today"},{"key":"69_CR5","doi-asserted-by":"publisher","first-page":"108","DOI":"10.1021\/ci800249s","volume":"49","author":"A Bender","year":"2009","unstructured":"Bender A, Jenkins JL, Scheiber J, Sukuru SCK, Glick M, Davies JW. How similar are similarity searching methods?: a principal component analysis of molecular descriptor space. J Chem Inf Model. 2009;49:108\u201319.","journal-title":"J Chem Inf Model"},{"key":"69_CR6","doi-asserted-by":"publisher","first-page":"742","DOI":"10.1021\/ci100050t","volume":"50","author":"D Rogers","year":"2010","unstructured":"Rogers D, Hahn M. Extended-connectivity fingerprints. J Chem Inf Model. 2010;50:742\u201354.","journal-title":"J Chem Inf Model"},{"key":"69_CR7","doi-asserted-by":"publisher","first-page":"58","DOI":"10.1016\/j.ymeth.2014.08.005","volume":"71","author":"A Cereto-Massagu\u00e9","year":"2015","unstructured":"Cereto-Massagu\u00e9 A, Ojeda MJ, Valls C, Mulero M, Garcia-Vallv\u00e9 S, Pujadas G. Molecular fingerprint similarity search in virtual screening. Methods. 2015;71:58\u201363.","journal-title":"Methods"},{"key":"69_CR8","doi-asserted-by":"publisher","first-page":"571","DOI":"10.1109\/TENCON.2003.1273228","volume-title":"Proceedings of TENCON 2003 Conference on Convergent Technologies for the Asia-Pacific Region","author":"M Kokare","year":"2003","unstructured":"Kokare M, Chatterji BN, Biswas PK. Comparison of similarity metrics for texture image retrieval. In: Proceedings of TENCON 2003 Conference on Convergent Technologies for the Asia-Pacific Region, vol. 2. Edited by IEEE. 2003. p. 571\u20135."},{"key":"69_CR9","first-page":"58","volume-title":"Proceedings of the Workshop on Artificial Intelligence for Web Search (AAAI 2000)","author":"A Strehl","year":"2000","unstructured":"Strehl A, Strehl E, Ghosh J, Mooney R. Impact of similarity measures on web-page clustering. In: Proceedings of the Workshop on Artificial Intelligence for Web Search (AAAI 2000). Edited by AAAI. 2000. p. 58\u201364."},{"key":"69_CR10","doi-asserted-by":"publisher","first-page":"291","DOI":"10.1145\/1718487.1718524","volume-title":"Proceedings of the third ACM international conference on Web search and data mining","author":"H Becker","year":"2010","unstructured":"Becker H, Naaman M, Gravano L. Learning similarity metrics for event identification in social media. In: Proceedings of the third ACM international conference on Web search and data mining. New York, NY, USA: ACM; 2010. p. 291\u2013300. WSDM\u201910."},{"key":"69_CR11","doi-asserted-by":"publisher","first-page":"1284","DOI":"10.1177\/1087057113501390","volume":"18","author":"F Reisen","year":"2013","unstructured":"Reisen F, Zhang X, Gabriel D, Selzer P. Benchmarking of multivariate similarity measures for high-content screening fingerprints in phenotypic drug discovery. J Biomol Screen. 2013;18:1284\u201397.","journal-title":"J Biomol Screen"},{"key":"69_CR12","doi-asserted-by":"publisher","first-page":"155","DOI":"10.2174\/1386207024607338","volume":"5","author":"JD Holliday","year":"2002","unstructured":"Holliday JD, Hu C-Y, Willett P. Grouping of coefficients for the calculation of inter-molecular similarity and dissimilarity using 2D fragment Bit-strings. Comb Chem High Throughput Screen. 2002;5:155\u201366.","journal-title":"Comb Chem High Throughput Screen"},{"key":"69_CR13","doi-asserted-by":"publisher","first-page":"1407","DOI":"10.1021\/ci025531g","volume":"42","author":"X Chen","year":"2002","unstructured":"Chen X, Reynolds CH. Performance of similarity measures in 2D fragment-based similarity searching: comparison of structural descriptors and similarity coefficients. J Chem Inf Comput Sci. 2002;42:1407\u201314.","journal-title":"J Chem Inf Comput Sci"},{"key":"69_CR14","doi-asserted-by":"publisher","first-page":"435","DOI":"10.1021\/ci025596j","volume":"43","author":"N Salim","year":"2003","unstructured":"Salim N, Holliday J, Willett P. Combination of fingerprint-based similarity coefficients using data fusion. J Chem Inf Comput Sci. 2003;43:435\u201342.","journal-title":"J Chem Inf Comput Sci"},{"key":"69_CR15","doi-asserted-by":"publisher","first-page":"1046","DOI":"10.1016\/j.drudis.2006.10.005","volume":"11","author":"P Willett","year":"2006","unstructured":"Willett P. Similarity-based virtual screening using 2D fingerprints. Drug Discov Today. 2006;11:1046\u201353.","journal-title":"Drug Discov Today"},{"key":"69_CR16","doi-asserted-by":"publisher","first-page":"2884","DOI":"10.1021\/ci300261r","volume":"52","author":"R Todeschini","year":"2012","unstructured":"Todeschini R, Consonni V, Xiang H, Holliday J, Buscema M, Willett P. Similarity coefficients for binary chemoinformatics data: overview and extended comparison using simulated and real data sets. J Chem Inf Model. 2012;52:2884\u2013901.","journal-title":"J Chem Inf Model"},{"key":"69_CR17","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1021\/ci300547g","volume":"53","author":"P Willett","year":"2013","unstructured":"Willett P. Combination of similarity rankings using data fusion. J Chem Inf Model. 2013;53:1\u201310.","journal-title":"J Chem Inf Model"},{"key":"69_CR18","doi-asserted-by":"publisher","first-page":"1840","DOI":"10.1021\/ci049867x","volume":"44","author":"M Whittle","year":"2004","unstructured":"Whittle M, Gillet VJ, Willett P, Alex A, Loesel J. Enhancing the effectiveness of virtual screening by fusing nearest neighbor lists: a comparison of similarity coefficients. J Chem Inf Comput Sci. 2004;44:1840\u20138.","journal-title":"J Chem Inf Comput Sci"},{"key":"69_CR19","doi-asserted-by":"publisher","first-page":"379","DOI":"10.1021\/ci970437z","volume":"38","author":"DR Flower","year":"1998","unstructured":"Flower DR. On the properties of Bit string-based measures of chemical similarity. J Chem Inf Comput Sci. 1998;38:379\u201386.","journal-title":"J Chem Inf Comput Sci"},{"key":"69_CR20","doi-asserted-by":"crossref","first-page":"65","DOI":"10.1007\/BF03380182","volume":"7\/8","author":"MS Lajiness","year":"1997","unstructured":"Lajiness MS. Dissimilarity-based compound selection techniques. Perspect Drug Discov Des. 1997;7\/8:65\u201384.","journal-title":"Perspect Drug Discov Des"},{"key":"69_CR21","doi-asserted-by":"publisher","first-page":"2887","DOI":"10.1021\/jm980708c","volume":"42","author":"SL Dixon","year":"1999","unstructured":"Dixon SL, Koehler RT. The hidden component of size in two-dimensional fragment descriptors: side effects on sampling in bioactive libraries. J Med Chem. 1999;42:2887\u2013900.","journal-title":"J Med Chem"},{"key":"69_CR22","doi-asserted-by":"publisher","first-page":"819","DOI":"10.1021\/ci034001x","volume":"43","author":"JD Holliday","year":"2003","unstructured":"Holliday JD, Salim N, Whittle M, Willett P. Analysis and display of the size dependence of chemical similarity coefficients. J Chem Inf Comput Sci. 2003;43:819\u201328.","journal-title":"J Chem Inf Comput Sci"},{"key":"69_CR23","doi-asserted-by":"publisher","first-page":"163","DOI":"10.1021\/ci990316u","volume":"40","author":"JW Godden","year":"2000","unstructured":"Godden JW, Xue L, Bajorath J. Combinatorial preferences affect molecular similarity\/diversity calculations using binary fingerprints and Tanimoto coefficients. J Chem Inf Comput Sci. 2000;40:163\u20136.","journal-title":"J Chem Inf Comput Sci"},{"key":"69_CR24","doi-asserted-by":"publisher","first-page":"766","DOI":"10.1145\/1066157.1066244","volume-title":"Proceedings of the 2005 ACM SIGMOD international conference on Management of data","author":"X Yan","year":"2005","unstructured":"Yan X, Yu P, Han J. Substructure similarity search in graph databases. In: Proceedings of the 2005 ACM SIGMOD international conference on Management of data. Edited by ACM 2005. p. 766\u201377."},{"key":"69_CR25","volume-title":"Proceedings of the 9th Joint Conference on Information Sciences","author":"S Klinger","year":"2006","unstructured":"Klinger S, Austin J. Weighted superstructures for chemical similarity searching. In: Proceedings of the 9th Joint Conference on Information Sciences. 2006."},{"key":"69_CR26","doi-asserted-by":"publisher","first-page":"151","DOI":"10.1002\/cem.1320","volume":"25","author":"K H\u00e9berger","year":"2011","unstructured":"H\u00e9berger K, Koll\u00e1r-Hunek K. Sum of ranking differences for method discrimination and its validation: comparison of ranks with random numbers. J Chemom. 2011;25:151\u20138.","journal-title":"J Chemom"},{"key":"69_CR27","doi-asserted-by":"publisher","first-page":"101","DOI":"10.1016\/j.trac.2009.09.009","volume":"29","author":"K H\u00e9berger","year":"2010","unstructured":"H\u00e9berger K. Sum of ranking differences compares methods or models fairly. TrAC Trends Anal Chem. 2010;29:101\u20139.","journal-title":"TrAC Trends Anal Chem"},{"key":"69_CR28","doi-asserted-by":"publisher","first-page":"17","DOI":"10.1186\/1758-2946-4-17","volume":"4","author":"R Kiss","year":"2012","unstructured":"Kiss R, S\u00e1ndor M, Szalai FA. http:\/\/Mcule.com: a public web service for drug discovery. J Cheminform. 2012;4:17.","journal-title":"J Cheminform"},{"key":"69_CR29","unstructured":"KNIME | Konstanz Information Miner, University of Konstanz, Germany. 2014. [https:\/\/www.knime.org\/]"},{"key":"69_CR30","unstructured":"JChem 2.8.2, ChemAxon LLC, Budapest, Hungary. 2014 [http:\/\/www.chemaxon.com]"},{"key":"69_CR31","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1016\/j.chemolab.2005.11.001","volume":"87","author":"R Todeschini","year":"2007","unstructured":"Todeschini R, Ballabio D, Consonni V, Mauri A, Pavan M. CAIMAN (classification and influence matrix analysis): a new approach to the classification based on leverage-scaled functions. Chemom Intell Lab Syst. 2007;87:3\u201317.","journal-title":"Chemom Intell Lab Syst"},{"key":"69_CR32","doi-asserted-by":"publisher","first-page":"983","DOI":"10.1021\/ci9800211","volume":"38","author":"P Willett","year":"1998","unstructured":"Willett P, Barnard J, Downs G. Chemical similarity searching. J Chem Inf Comput Sci. 1998;38:983\u201396.","journal-title":"J Chem Inf Comput Sci"},{"key":"69_CR33","doi-asserted-by":"publisher","first-page":"139","DOI":"10.1016\/j.chemolab.2013.06.007","volume":"127","author":"K Koll\u00e1r-Hunek","year":"2013","unstructured":"Koll\u00e1r-Hunek K, H\u00e9berger K. Method and model comparison by sum of ranking differences in cases of repeated observations (ties). Chemom Intell Lab Syst. 2013;127:139\u201346.","journal-title":"Chemom Intell Lab Syst"},{"key":"69_CR34","unstructured":"Chemical Hashed Fingerprint [https:\/\/docs.chemaxon.com\/display\/CD\/Chemical+Hashed+Fingerprint]."},{"key":"69_CR35","unstructured":"RDKit: Cheminformatics and Machine Learning Software, Open-source. 2014. [http:\/\/www.rdkit.org\/]"},{"key":"69_CR36","series-title":"31","doi-asserted-by":"publisher","DOI":"10.1007\/978-0-387-21606-5","volume-title":"Elements of Statistical Learning: Data Mining, Inference, and Prediction","author":"T Hastie","year":"2001","unstructured":"Hastie T, Tibshirani R, Friedman J. Overview of supervised learning. In: Elements of Statistical Learning: Data Mining, Inference, and Prediction, 31. New York: Springer; 2001."},{"key":"69_CR37","unstructured":"STATISTICA 12.5, StatSoft, Inc., Tulsa, OK 74104, USA, 2014. [http:\/\/www.statsoft.com\/Products\/STATISTICA-Features\/Version-12]."},{"key":"69_CR38","doi-asserted-by":"publisher","first-page":"987","DOI":"10.1016\/S1359-6446(05)03511-7","volume":"10","author":"RAE Carr","year":"2005","unstructured":"Carr RAE, Congreve M, Murray CW, Rees DC. Fragment-based lead discovery: leads by design. Drug Discov Today. 2005;10:987\u201392.","journal-title":"Drug Discov Today"},{"key":"69_CR39","doi-asserted-by":"publisher","first-page":"3743","DOI":"10.1002\/(SICI)1521-3773(19991216)38:24<3743::AID-ANIE3743>3.0.CO;2-U","volume":"38","author":"SJ Teague","year":"1999","unstructured":"Teague SJ, Davis AM, Leeson PD, Oprea T. The design of leadlike combinatorial libraries. Angew Chemie Int Ed. 1999;38:3743\u20138.","journal-title":"Angew Chemie Int Ed"},{"key":"69_CR40","doi-asserted-by":"publisher","first-page":"235","DOI":"10.1016\/S1056-8719(00)00107-6","volume":"44","author":"CA Lipinski","year":"2000","unstructured":"Lipinski CA. Drug-like properties and the causes of poor solubility and poor permeability. J Pharmacol Toxicol Methods. 2000;44:235\u201349.","journal-title":"J Pharmacol Toxicol Methods"}],"container-title":["Journal of Cheminformatics"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s13321-015-0069-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1186\/s13321-015-0069-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s13321-015-0069-3","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13321-015-0069-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,2]],"date-time":"2021-09-02T13:46:47Z","timestamp":1630590407000},"score":1,"resource":{"primary":{"URL":"https:\/\/jcheminf.biomedcentral.com\/articles\/10.1186\/s13321-015-0069-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,5,20]]},"references-count":40,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2015,12]]}},"alternative-id":["69"],"URL":"https:\/\/doi.org\/10.1186\/s13321-015-0069-3","relation":{"has-review":[{"id-type":"doi","id":"10.3410\/f.725546403.793541622","asserted-by":"object"}]},"ISSN":["1758-2946"],"issn-type":[{"value":"1758-2946","type":"electronic"}],"subject":[],"published":{"date-parts":[[2015,5,20]]},"assertion":[{"value":"2 December 2014","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 April 2015","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 May 2015","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"20"}}