Sequence-based modeling of plant epigenomes reveals cell-type-specific cis-regulatory grammar
Sequence-based modeling of plant epigenomes reveals cell-type-specific cis-regulatory grammar
Yao, J.; Li, J.; Zhang, X.; Li, X.; Marand, A. P.; Pickering, E.; Schmitz, R. J.
AbstractHow cis-regulatory sequences and their genetic variation govern chromatin accessibility, regulate gene expression, and the establishment of plant cell identities, making them fundamental to development, environmental responses, and phenotypic diversity. Here we present PEAgent, a framework for training, evaluating and interpreting deep-learning models that predict single-cell chromatin accessibility directly from DNA sequence, packaged in an interactive web portal and toolkit. Models were trained on single-cell chromatin-accessibility atlases of soybean, maize and rice, together spanning over 355,000 cells and 320 cell types and ~150 million years of evolution. We unraveled a lexicon of 243 cell-type-resolved regulatory patterns, half of them composite, with TCP and bHLH showing the greatest influence and strongest conservation across species. Co-occurrence and in silico synergy analyses, explicitly modeling motif orientation and spacing, revealed two distinct cooperative modes acting at short and nucleosome-scale distances. We further showed that model predictions distinguish grass-conserved from rice-specific regulatory sequences far more accurately than sequence conservation scores alone, and validated the model's predicted variant effects against cell-type-level chromatin accessible QTLs. PEAgent provides a foundational resource for decoding cell-type-specific plant cis-regulatory logic and interpreting noncoding variation in plants.