보건 및 환경
Predicting birth weight by multivariate functional principal component regressions
Functional data analysis (FDA) provides a powerful statistical framework for analyzing complex data, such as curves or functions, over high-dimensional domains. In this paper, we focus on functional predictor regression (scalar-on-function) models applied to the prediction of birth weight using maternal dietary patterns during pregnancy. Specifically, we analyze trajectories of nine components of the Alternative Healthy Eating Index (AHEI) as multivariate functional predictors. We adopt Multivariate Functional Principal Component Analysis (MVFPCA) to obtain low-dimensional, uncorrelated representations of the multivariate functional predictors through their FPC scores. Building on MVFPCA, we develop a novel multivariate functional principal component regression (MVFPCR) model to predict birth weight effectively. Our method accommodates various regression approaches, including linear and quantile regression, depending on the distribution of birth weights. Through simulation studies and the application to a fetal growth study dataset, we demonstrate the effectiveness of our proposed model in functional predictor regression.
2026
보건 및 환경
A community effort to optimize sequence-based deep learning models of gene regulation
A systematic evaluation of how model architectures and training strategies impact genomics model performance is needed. To address this gap, we held a DREAM Challenge where competitors trained models on a dataset of millions of random promoter DNA sequences and corresponding expression levels, experimentally determined in yeast. For a robust evaluation of the models, we designed a comprehensive suite of benchmarks encompassing various sequence types. All top-performing models used neural networks but diverged in architectures and training strategies. To dissect how architectural and training choices impact performance, we developed the Prix Fixe framework to divide models into modular building blocks. We tested all possible combinations for the top three models, further improving their performance. The DREAM Challenge models not only achieved state-of-the-art results on our comprehensive yeast dataset but also consistently surpassed existing benchmarks on Drosophila and human genomic datasets, demonstrating the progress that can be driven by gold-standard genomics datasets.Co-authors: Abdul Muntakim Rafi, Daria Nogina, Dmitry Penzar, Dohoon Lee, Danyeong Lee, Nayeon Kim, Sangyeup Kim, Dohyeon Kim, Yeojin Shin, Georgy Meshcheryakov, Andrey Lando, Arsenii Zinkevich, Byeong-Chan Kim, Juhyun Lee, Taein Kang, Eeshit Dhaval Vaishnav, Payman Yadollahpour, Random Promoter DREAM Challenge Consortium, Sun Kim, Jake Albrecht, Aviv Regev, Wuming Gong, Ivan V. Kulakovskiy, Pablo Meyer & Carl G. de Boer
2024