Acute graft-versus-host disease (aGvHD) remains a major cause of morbidity and mortality after allogeneic hematopoietic stem cell transplantation (HSCT). We performed RNA analysis of 1408 candidate genes in bone marrow samples obtained from 167 patients undergoing HSCT. RNA expression data were used in a machine learning algorithm to predict the presence or absence of aGvHD using either random forest or extreme gradient boosting algorithms. Patients were randomly divided into training (2/3 of patients) and validation (1/3 of patients) sets. Using post-HSCT RNA data, the machine learning algorithm selected 92 genes for predicting aGvHD that appear to play a role in PI3/AKT, MAPK, and FOXO signaling, as well as microRNA. The algorithm selected 20 genes for predicting survival included genes involved in MAPK and chemokine signaling. Using pre-HSCT RNA data, the machine learning algorithm selected 400 genes and 700 genes predicting aGvHD and overall survival, but candidate signaling pathways could not be specified in this analysis. These data show that NGS analyses of RNA expression using machine learning algorithms may be useful biomarkers of aGvHD and overall survival for patients undergoing HSCT, allowing for the identification of major signaling pathways associated with HSCT outcomes and helping to dissect the complex steps involved in the development of aGvHD. The analysis of pre-HSCT bone marrow samples may lead to pre-HSCT interventions including choice of remission induction regimens and modifications in patient health before HSCT.
Keywords: acute GvHD; allogeneic transplantation; machine learning; post-transplant survival; signaling pathways; transcriptomics.