AI Predicts Complex DNA Binding by Transcription Factors

Summary: Researchers at North Carolina State University have developed BINND (Binding and Interaction Neural Network for DNA), a novel deep learning model trained on an unprecedented dataset of 144 million sequence pairs. BINND predicts complex DNA–DNA binding affinity and interaction patterns with high accuracy, outperforming prior methods and providing a practical tool to advance DNA computing, molecular data storage, and diagnostic applications.

Key Facts

  • Modeling hyperconnected DNA networks: Unlike traditional approaches that treat binding as a binary, isolated event, BINND is designed to predict how many different DNA strands interact simultaneously in crowded, biologically realistic environments.
  • Empirical dataset at scale: The team built a physical library of 144 million DNA sequence pairs and used those measured binding events to train BINND, avoiding reliance on simplified biophysical extrapolations that miss non-linear behaviors.
  • Strong predictive performance: In proof-of-concept tests BINND achieved 83.5% accuracy in predicting binding behavior, improving on state-of-the-art models by at least 10% and generalizing across diverse sequences.
  • Conservative failure mode: When BINND errs, it more often predicts a non-binding outcome when binding actually occurs, a bias that reduces the risk of false positive cross-binding and helps prevent harmful crosstalk in molecular diagnostics.
  • Practical demonstration: The team used BINND to construct a cross-binding matrix mapping interactions among ninety-six 20-nucleotide sequences and twenty-six other 20-nucleotide sequences, creating a reliable “address book” for molecular data storage and retrieval.
  • Enabling scalable DNA computing and storage: By accurately mapping which strands will bind, BINND helps solve a central scaling challenge for DNA-based data storage and DNA computing, enabling more confident design and rapid, error-minimized retrieval from molecular archives.

Source: North Carolina State University

Overview: The BINND model brings deep learning to the problem of predicting non-complementary and partial DNA–DNA binding in hyperconnected systems. This capability is useful for designing sensitive molecular diagnostics, building DNA-based computing systems, and engineering large-scale DNA data storage where unintended molecular interactions can corrupt information.

“We often think about binding as a very simple relationship – Molecule A binds to Molecule B,” says Albert Keung, co-corresponding author and associate professor of chemical and biomolecular engineering at NC State. “But in biological systems, one strand may partially bind dozens of others. Capturing that hypercomplexity is a significant challenge, and it’s essential for both understanding natural genetic systems and building robust biomolecular technologies.”

This shows DNA.
The deep learning model BINND can accurately map hypercomplex, non-complementary DNA–DNA binding behaviors, overcoming a primary scaling bottleneck for molecular data storage. Credit: Neuroscience News

The research team recognized that existing predictive tools were limited by small datasets and by reliance on biophysical models that fail to capture many non-linear interactions. To address this, they generated a massive empirical dataset of actual binding measurements—144 million sequence pairs—and trained BINND directly on this ground-truth data.

Co-lead author Gunavaran Brihadiswaran, a Ph.D. student at NC State, notes: “Deep learning models can capture complex, hidden patterns, but they require a robust training set. By creating an expansive experimental library, we empowered BINND to learn real molecular behaviors rather than relying on approximations.”

In evaluations the BINND model predicted binding with 83.5% accuracy and demonstrated a useful safety bias: false negatives were more common than false positives. “That conservative failure mode helps reduce catastrophic cross-binding, which is crucial for molecular diagnostics and data retrieval,” says co-lead author Karishma Matange.

To show practical value, the investigators used BINND to compile a searchable cross-binding database in matrix form, illustrating interactions among selected 20-nucleotide sequences. This resource serves as a guide for selecting sequences that minimize unintended interactions and for designing safe probe-based retrieval in DNA storage systems. The team has made BINND and its resources publicly available as a repository for researchers to use and build upon.

“A major question for DNA data storage and DNA computing has been scalability,” says James Tuck, co-corresponding author and professor of electrical and computer engineering. “BINND gives engineers and biologists a tool to predict complex binding networks quickly and reliably, helping move molecular memory and computation toward practical scale.”

The peer-reviewed study, titled “Deep Learning Predicts Dissimilar DNA–DNA Binding and Engineers Hyperconnected Networks,” is published open access in the journal Nature Communications. The paper’s co-authors include Karishma Matange, Gunavaran Brihadiswaran, Kyle J. Tomek, Kevin Volkel, Doug Townsend, James M. Tuck, and Albert J. Keung. DOI: 10.1038/s41467-026-75395-w

Funding: This work was supported by the National Science Foundation (grants 2027655, 1901324, 2403352), the National Institutes of Health (R41HG013877), the U.S. Department of Education Graduate Assistance in Areas of Need fellowship (P200A160061), and the Simons Foundation (grant 990252).

Key Questions Answered:

Q: Why is DNA–DNA binding more complicated than the simple “A–T” and “C–G” pairing taught in school?

A: Classroom rules describe ideal, fully complementary pairing, but in real systems DNA strands can partially bind to many imperfect matches with varying strengths. These partial interactions form a dense network of possible bindings—“hypercomplexity”—that can cause cross-talk and errors in diagnostics and data storage if not properly predicted.

Q: How does BINND help build practical DNA computers and storage systems?

A: DNA storage encodes digital data as sequences of A, C, T, and G. To retrieve data, a fluorescent probe must bind specifically to the intended target. BINND maps likely interactions among sequences so designers can choose probes and data strands that avoid unwanted binding, reducing corruption and improving retrieval accuracy.

Q: Why did the large training dataset matter?

A: Machine learning models require extensive, high-quality examples to learn complex patterns. Prior models relied on limited datasets and physics-based approximations. The 144 million empirically measured sequence pairs provided BINND with enough real-world interactions to learn subtle, non-linear binding behaviors that are difficult to predict analytically.

Editorial Notes:

  • This article was edited by a Neuroscience News editor.
  • The journal paper was reviewed in full by the editorial team.
  • Additional contextual information was added by staff to clarify technical implications.

About this AI and genetics research news

Author: Matt Shipman
Source: North Carolina State University
Contact: Matt Shipman – North Carolina State University
Image credit: Neuroscience News


Abstract

Deep Learning Predicts Dissimilar DNA–DNA Binding and Engineers Hyperconnected Networks

Traditional design frameworks in molecular bioengineering often focus on orthogonality, treating weak or non-specific interactions as problems to avoid. This restricts usable sequence space and limits scalability, especially when synthetic systems must operate in natural backgrounds with high sequence diversity. Accurate, fast prediction of non-orthogonal interactions—validated against ground truth data—is essential to harness the full sequence space.

The authors present BINND, combining an ultra-high-throughput experimental platform with a deep learning model trained on millions of measured interactions. BINND achieves above 80% accuracy, generalizes across diverse sequences, and operates far faster than prior models. The study illustrates BINND’s value with a searchable DNA interaction network and discusses applications in diagnostics, bioengineering, DNA origami, and large-scale DNA data storage.