Research Refining sequence-to-activity models by increasing model resolution

A base-pair resolution model reveals finer regulatory detail in immune cell chromatin

Understanding how DNA sequence controls gene expression is key to understanding cell identity and disease. One way to probe this is ATAC-seq, an experimental technique that measures which regions of the genome are physically accessible to the cell's machinery — accessibility that is a prerequisite for transcription factors to bind DNA and switch genes on or off. Building on the BPNet framework, the study presents bpAI-TAC, a deep learning model that predicts chromatin accessibility across immune cell types directly from DNA sequence, at base-pair resolution. The key idea is to model two complementary signals separately: the total ATAC-seq counts in a region, which reflect how accessible it is, and the shape of the profile within it, which reveals where exactly transcription factors bind. Learning to predict this fine-scale shape as a secondary task helps the model learn a better regulatory grammar, which in turn improves its predictions of total-count differences across cell types — the metric that actually captures cell-type-specific regulatory control. The work shows that this higher-resolution training consistently improves predictions of differential accessibility between cell types, that training jointly across many related immune cell types outperforms training separate models per cell type, and that comparing sequence attributions with and without profile information reveals new regulatory motifs with strong effects that only emerge at base-pair detail. The approach mirrors a trend also seen recently in AlphaGenome, which likewise adopted base-pair resolution ATAC-seq modeling to improve regulatory syntax learning — though the higher resolution comes at a computational cost, since the model must now make 250 predictions per region instead of one, requiring larger, wider architectures. The study was led by Nuria Alina Chandra, an undergraduate researcher supervised by Alexander Sasse during his postdoctoral work in Sara Mostafavi's lab at the University of Washington, and builds on the group's earlier AI-TAC model.

overview