Ensemble Model with BERT, RoBERTa and XLNet for Molecular Property Prediction
Article Information
Abstract
Molecular property prediction is a fundamental task in drug discovery and materials science, yet most high-performing approaches depend on large-scale pretraining that demands substantial computational resources. This work proposes a pretraining-free ensemble framework that trains multiple Transformer-based architectures—BERT, RoBERTa, and XLNet—from random initialization using the Atom-in-SMILES (AIS) molecular representation, which provides richer atomic-level semantics than conventional SMILES. The three Transformer encoders are coupled with BiLSTM prediction heads and integrated via a BaggingRegressor to reduce variance and improve generalization. Experiments on the ZINC250k and ZINC310k benchmarks demonstrate that the proposed framework achieves competitive performance against pretrained baselines including GROVER, CHEM-BERT, and D-MPNN, while requiring only task-specific end-to-end training with adaptive early stopping. These results establish that carefully designed molecular representations combined with heterogeneous ensemble learning can serve as a practical and resource-efficient alternative to pretraining-based paradigms in molecular modeling.
Graphical Abstract
Keywords
Data Availability Statement
Funding
Conflicts of Interest
AI Use Statement
Ethical Approval and Consent to Participate
References
- Goh, G. B., Hodas, N. O., & Vishnu, A. (2017). Deep Learning for Computational Chemistry. Journal of Computational Chemistry, 38(16), 1291-1307.
[CrossRef] [Google Scholar] - Lavecchia, A. (2019). Deep learning in drug discovery: opportunities, challenges and future prospects. Drug discovery today, 24(10), 2017-2032.
[CrossRef] [Google Scholar] - Wu, Z., Ramsundar, B., Feinberg, E. N., Gomes, J., Geniesse, C., Pappu, A. S., ... & Pande, V. (2018). MoleculeNet: a benchmark for molecular machine learning. Chemical science, 9(2), 513-530.
[CrossRef] [Google Scholar] - Li, Z., Jiang, M., Wang, S., & Zhang, S. (2022). Deep Learning Methods for Molecular Representation and Property Prediction. Drug Discovery Today, 27(12), 103373.
[CrossRef] [Google Scholar] - Walters, W. P., & Barzilay, R. (2020). Applications of deep learning in molecule generation and molecular property prediction. Accounts of chemical research, 54(2), 263-270.
[CrossRef] [Google Scholar] - Rogers, D., & Hahn, M. (2010). Extended-connectivity fingerprints. Journal of chemical information and modeling, 50(5), 742-754.
[CrossRef] [Google Scholar] - Chen, H., Engkvist, O., Wang, Y., Olivecrona, M., & Blaschke, T. (2018). The rise of deep learning in drug discovery. Drug discovery today, 23(6), 1241-1250.
[CrossRef] [Google Scholar] - Goh, G. B., Hodas, N. O., Siegel, C., & Vishnu, A. (2017). Smiles2vec: An Interpretable General-Purpose Deep Neural Network for Predicting Chemical Properties. arXiv preprint arXiv:1712.02034.
[CrossRef] [Google Scholar] - Pinheiro, G. A., Mucelini, J., Soares, M. D., Prati, R. C., da Silva, J. L. F., & Quiles, M. G. (2020). Machine Learning Prediction of Nine Molecular Properties Based on the SMILES Representation of the QM9 Quantum-Chemistry Dataset. The Journal of Physical Chemistry A, 124(47), 9854-9866.
[CrossRef] [Google Scholar] - Jo, J., Kwak, B., Choi, H. S., & Yoon, S. (2020). The message passing neural networks for chemical property prediction on SMILES. Methods, 179, 65-72.
[CrossRef] [Google Scholar] - Wieder, O., Kohlbacher, S., Kuenemann, M., Garon, A., Ducrot, P., Seidel, T., & Langer, T. (2020). A compact review of molecular property prediction with graph neural networks. Drug Discovery Today: Technologies, 37, 1-12.
[CrossRef] [Google Scholar] - Zhang, Z., Liu, Q., Wang, H., Lu, C., & Lee, C. K. (2021). Motif-based graph self-supervised learning for molecular property prediction. Advances in Neural Information Processing Systems, 34, 15870-15882.
[Google Scholar] - Jiang, D., Wu, Z., Hsieh, C. Y., Chen, G., Liao, B., Wang, Z., ... & Hou, T. (2021). Could graph neural networks learn better molecular representation for drug discovery? A comparison study of descriptor-based and graph-based models. Journal of cheminformatics, 13(1), 12.
[CrossRef] [Google Scholar] - Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, {\L., & Polosukhin, I. (2017). Attention Is All You Need. Advances in Neural Information Processing Systems, 30.
[Google Scholar] - Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019, June). Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers) (pp. 4171-4186).
[CrossRef] [Google Scholar] - Sun, M., Zhao, S., Gilvary, C., Elemento, O., Zhou, J., & Wang, F. (2020). Graph convolutional networks for computational drug development and discovery. Briefings in bioinformatics, 21(3), 919-935.
[CrossRef] [Google Scholar] - Xiong, Z., Wang, D., Liu, X., Zhong, F., Wan, X., Li, X., ... & Zheng, M. (2019). Pushing the boundaries of molecular representation for drug discovery with the graph attention mechanism. Journal of medicinal chemistry, 63(16), 8749-8760.
[CrossRef] [Google Scholar] - Chithrananda, S., Grand, G., & Ramsundar, B. (2020). ChemBERTa: large-scale self-supervised pretraining for molecular property prediction. arXiv preprint arXiv:2010.09885.
[CrossRef] [Google Scholar] - Irwin, R., Dimitriadis, S., He, J., & Bjerrum, E. J. (2022). Chemformer: a pre-trained transformer for computational chemistry. Machine Learning: Science and Technology, 3(1), 015022.
[CrossRef] [Google Scholar] - Altae-Tran, H., Ramsundar, B., Pappu, A. S., & Pande, V. (2017). Low data drug discovery with one-shot learning. ACS central science, 3(4), 283.
[CrossRef] [Google Scholar] - Sagi, O., & Rokach, L. (2018). Ensemble learning: A survey. Wiley interdisciplinary reviews: data mining and knowledge discovery, 8(4), e1249.
[CrossRef] [Google Scholar] - Zhou, Z. H. (2021). Ensemble learning. In Machine learning (pp. 181-210). Singapore: Springer Singapore.
[CrossRef] [Google Scholar] - Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., ... & Stoyanov, V. (2019). Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
[CrossRef] [Google Scholar] - Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R. R., & Le, Q. V. (2019). Xlnet: Generalized autoregressive pretraining for language understanding. Advances in neural information processing systems, 32.
[Google Scholar] - Fabian, B., Edlich, T., Gaspar, H., Segler, M., Meyers, J., Fiscato, M., & Ahmed, M. (2020). Molecular representation learning with language models and domain-relevant auxiliary tasks. arXiv preprint arXiv:2011.13230.
[CrossRef] [Google Scholar] - Zheng, S., Yan, X., Yang, Y., & Xu, J. (2019). Identifying structure–property relationships through SMILES syntax analysis with self-attention mechanism. Journal of chemical information and modeling, 59(2), 914-923.
[CrossRef] [Google Scholar] - Ying, C., Cai, T., Luo, S., Zheng, S., Ke, G., He, D., ... & Liu, T. Y. (2021). Do transformers really perform badly for graph representation?. Advances in neural information processing systems, 34, 28877-28888.
[Google Scholar] - Karpov, P., Godin, G., & Tetko, I. V. (2020). Transformer-CNN: Swiss knife for QSAR modeling and interpretation. Journal of cheminformatics, 12(1), 17.
[CrossRef] [Google Scholar] - Breiman, L. (1996). Bagging predictors. Machine learning, 24(2), 123-140.
[CrossRef] [Google Scholar] - Weininger, D. (1988). SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules. Journal of chemical information and computer sciences, 28(1), 31-36.
[CrossRef] [Google Scholar] - David, L., Thakkar, A., Mercado, R., & Engkvist, O. (2020). Molecular representations in AI-driven drug discovery: a review and practical guide. Journal of cheminformatics, 12(1), 56.
[CrossRef] [Google Scholar] - O'Boyle, N., & Dalke, A. (2018). DeepSMILES: An Adaptation of SMILES for Use in Machine Learning of Chemical Structures. ChemRxiv preprint.
[CrossRef] [Google Scholar] - Krenn, M., Häse, F., Nigam, A., Friederich, P., & Aspuru-Guzik, A. (2020). Self-referencing embedded strings (SELFIES): A 100\% robust molecular string representation. Machine Learning: Science and Technology, 1(4), 045024.
[CrossRef] [Google Scholar] - Ucak, U. V., Ashyrmamatov, I., & Lee, J. (2023). Improving the Quality of Chemical Language Model Outcomes with Atom-in-SMILES Tokenization. Journal of Cheminformatics, 15(1), 55.
[CrossRef] [Google Scholar] - Friedman, J. H. (2001). Greedy function approximation: a gradient boosting machine. Annals of Statistics, 29(5), 1189-1232.
[CrossRef] [Google Scholar] - Smola, A. J., & Sch\"{olkopf, B. (2004). A tutorial on support vector regression. Statistics and Computing, 14(3), 199-222.
[CrossRef] [Google Scholar] - Loh, W. Y. (2011). Classification and regression trees. Wiley interdisciplinary reviews: data mining and knowledge discovery, 1(1), 14-23.
[CrossRef] [Google Scholar] - Seeger, M. (2004). Gaussian processes for machine learning. International journal of neural systems, 14(02), 69-106.
[CrossRef] [Google Scholar] - Ross, J., Belgodere, B., Chenthamarakshan, V., Padhi, I., Mroueh, Y., & Das, P. (2022). Large-scale chemical language representations capture molecular structure and properties. Nature Machine Intelligence, 4(12), 1256-1264.
[CrossRef] [Google Scholar] - Born, J., & Manica, M. (2023). Regression Transformer Enables Concurrent Sequence Regression and Generation for Molecular Language Modelling. Nature Machine Intelligence, 5(4), 432-444.
[CrossRef] [Google Scholar] - Wang, S., Guo, Y., Wang, Y., Sun, H., & Huang, J. (2019, September). Smiles-bert: large scale unsupervised pre-training for molecular property prediction. In Proceedings of the 10th ACM international conference on bioinformatics, computational biology and health informatics (pp. 429-436).
[CrossRef] [Google Scholar] - Yu, J., Zhang, C., Cheng, Y., Yang, Y.-F., She, Y.-B., Liu, F., Su, W., & Su, A. (2023). SolvBERT for Solvation Free Energy and Solubility Prediction: A Demonstration of an NLP Model for Predicting the Properties of Molecular Complexes. Digital Discovery, 2(2), 409-421.
[CrossRef] [Google Scholar] - Li, J., & Jiang, X. (2021). Mol‐BERT: an effective molecular representation with BERT for molecular property prediction. Wireless Communications and Mobile Computing, 2021(1), 7181815.
[CrossRef] [Google Scholar] - Liu, Y., Zhang, R., Li, T., Jiang, J., Ma, J., & Wang, P. (2023). MolRoPE-BERT: An Enhanced Molecular Representation with Rotary Position Embedding for Molecular Property Prediction. Journal of Molecular Graphics and Modelling, 118, 108344.
[CrossRef] [Google Scholar] - Irwin, J. J., Sterling, T., Mysinger, M. M., Bolstad, E. S., & Coleman, R. G. (2012). ZINC: a free tool to discover chemistry for biology. Journal of chemical information and modeling, 52(7), 1757-1768.
[CrossRef] [Google Scholar] - Sterling, T., & Irwin, J. J. (2015). ZINC 15--Ligand Discovery for Everyone. Journal of Chemical Information and Modeling, 55(11), 2324-2337.
[CrossRef] [Google Scholar] - Breiman, L. (2001). Random forests. Machine learning, 45(1), 5-32.
[CrossRef] [Google Scholar] - Alperstein, Z., Cherkasov, A., & Rolfe, J. T. (2019). All smiles variational autoencoder. arXiv preprint arXiv:1905.13343.
[CrossRef] [Google Scholar] - Gómez-Bombarelli, R., Wei, J. N., Duvenaud, D., Hernández-Lobato, J. M., Sánchez-Lengeling, B., Sheberla, D., ... & Aspuru-Guzik, A. (2018). Automatic chemical design using a data-driven continuous representation of molecules. ACS central science, 4(2), 268.
[CrossRef] [Google Scholar] - Winter, R., Montanari, F., Noé, F., & Clevert, D. A. (2019). Learning continuous and data-driven molecular descriptors by translating equivalent chemical representations. Chemical science, 10(6), 1692-1701.
[CrossRef] [Google Scholar] - Rong, Y., Bian, Y., Xu, T., Xie, W., Wei, Y., Huang, W., & Huang, J. (2020). Self-Supervised Graph Transformer on Large-Scale Molecular Data. Advances in Neural Information Processing Systems, 33, 12559-12571.
[Google Scholar] - Kim, H., Lee, J., Ahn, S., & Lee, J. R. (2021). A merged molecular representation learning for molecular properties prediction with a web-based service. Scientific Reports, 11(1), 11028.
[CrossRef] [Google Scholar] - Gilmer, J., Schoenholz, S. S., Riley, P. F., Vinyals, O., & Dahl, G. E. (2017, July). Neural message passing for quantum chemistry. In International conference on machine learning (pp. 1263-1272). Pmlr.
[Google Scholar] - Yang, K., Swanson, K., Jin, W., Coley, C., Eiden, P., Gao, H., ... & Barzilay, R. (2019). Analyzing learned molecular representations for property prediction. Journal of chemical information and modeling, 59(8), 3370-3388.
[CrossRef] [Google Scholar]
Cite This Article
TY - JOUR AU - Hu, Junling PY - 2026 DA - 2026/08/06 TI - Ensemble Model with BERT, RoBERTa and XLNet for Molecular Property Prediction JO - ICCK Transactions on Emerging Topics in Artificial Intelligence T2 - ICCK Transactions on Emerging Topics in Artificial Intelligence JF - ICCK Transactions on Emerging Topics in Artificial Intelligence VL - 3 IS - 3 SP - 170 EP - 187 DO - 10.62762/TETAI.2026.604672 UR - https://www.icck.org/article/abs/TETAI.2026.604672 KW - ensemble learning KW - BERT KW - RoBERTa KW - XLNet KW - molecular property prediction AB - Molecular property prediction is a fundamental task in drug discovery and materials science, yet most high-performing approaches depend on large-scale pretraining that demands substantial computational resources. This work proposes a pretraining-free ensemble framework that trains multiple Transformer-based architectures—BERT, RoBERTa, and XLNet—from random initialization using the Atom-in-SMILES (AIS) molecular representation, which provides richer atomic-level semantics than conventional SMILES. The three Transformer encoders are coupled with BiLSTM prediction heads and integrated via a BaggingRegressor to reduce variance and improve generalization. Experiments on the ZINC250k and ZINC310k benchmarks demonstrate that the proposed framework achieves competitive performance against pretrained baselines including GROVER, CHEM-BERT, and D-MPNN, while requiring only task-specific end-to-end training with adaptive early stopping. These results establish that carefully designed molecular representations combined with heterogeneous ensemble learning can serve as a practical and resource-efficient alternative to pretraining-based paradigms in molecular modeling. SN - 3068-6652 PB - Institute of Central Computation and Knowledge LA - English ER -
@article{Hu2026Ensemble,
author = {Junling Hu},
title = {Ensemble Model with BERT, RoBERTa and XLNet for Molecular Property Prediction},
journal = {ICCK Transactions on Emerging Topics in Artificial Intelligence},
year = {2026},
volume = {3},
number = {3},
pages = {170-187},
doi = {10.62762/TETAI.2026.604672},
url = {https://www.icck.org/article/abs/TETAI.2026.604672},
abstract = {Molecular property prediction is a fundamental task in drug discovery and materials science, yet most high-performing approaches depend on large-scale pretraining that demands substantial computational resources. This work proposes a pretraining-free ensemble framework that trains multiple Transformer-based architectures—BERT, RoBERTa, and XLNet—from random initialization using the Atom-in-SMILES (AIS) molecular representation, which provides richer atomic-level semantics than conventional SMILES. The three Transformer encoders are coupled with BiLSTM prediction heads and integrated via a BaggingRegressor to reduce variance and improve generalization. Experiments on the ZINC250k and ZINC310k benchmarks demonstrate that the proposed framework achieves competitive performance against pretrained baselines including GROVER, CHEM-BERT, and D-MPNN, while requiring only task-specific end-to-end training with adaptive early stopping. These results establish that carefully designed molecular representations combined with heterogeneous ensemble learning can serve as a practical and resource-efficient alternative to pretraining-based paradigms in molecular modeling.},
keywords = {ensemble learning, BERT, RoBERTa, XLNet, molecular property prediction},
issn = {3068-6652},
publisher = {Institute of Central Computation and Knowledge}
}
Article Metrics
Publisher's Note
ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and Permissions
Copyright © 2026 by the Author(s). Published by Institute of Central Computation and Knowledge. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
Portico