ICCK Transactions on Emerging Topics in Artificial Intelligence | Volume 3, Issue 3: 170-187, 2026 | DOI: 10.62762/TETAI.2026.604672
Abstract
Molecular property prediction is a fundamental task in drug discovery and materials science, yet most high-performing approaches depend on large-scale pretraining that demands substantial computational resources. This work proposes a pretraining-free ensemble framework that trains multiple Transformer-based architectures—BERT, RoBERTa, and XLNet—from random initialization using the Atom-in-SMILES (AIS) molecular representation, which provides richer atomic-level semantics than conventional SMILES. The three Transformer encoders are coupled with BiLSTM prediction heads and integrated via a BaggingRegressor to reduce variance and improve generalization. Experiments on the ZINC250k and ZINC... More >
Graphical Abstract