How AI Is Designing the Human Body’s Most Elusive Proteins

Summary: Researchers have developed a new machine learning approach capable of designing intrinsically disordered proteins (IDPs) — flexible, shape-shifting biomolecules that constitute almost 30% of human proteins and have so far resisted accurate prediction and design by tools such as AlphaFold. By combining physics-based molecular simulations with automatic differentiation, the team can tune amino-acid sequences to produce specific ensemble behaviors, opening new directions for synthetic biology, precision therapeutics, and the study of diseases like Parkinson’s and cancer.

Where conventional AI systems predict single, static protein structures, this method optimizes sequences against dynamic, physics-grounded simulations. That lets scientists design sequences that purposely populate particular structural ensembles or respond predictably to environmental cues — capabilities essential for engineering sensors, linkers, or therapeutic molecules based on IDPs.

Key Facts:

  • New AI-driven design: Uses automatic differentiation to invert physics-based molecular simulations and identify sequences that produce desired disordered ensembles.
  • Addresses a major blind spot: Intrinsically disordered proteins, which do not adopt a single stable fold, are prevalent in human biology and were not well handled by structure-prediction models like AlphaFold.
  • Wide applications: Enables design of synthetic IDPs for biomedical uses, molecular sensing, and materials engineering, and provides a tool for probing disease-associated mutations.
This shows DNA and a person's head.
With automatic differentiation, the researchers were able to train a computer to recognize how small changes in protein sequences – even single amino acid changes – affect the final desired properties of proteins. Credit: Neuroscience News

Intrinsically disordered proteins continually sample many conformations rather than settling into one fixed shape. That conformational plasticity underlies essential functions — molecular recognition, signaling, and the formation of dynamic assemblies — but it also complicates design and prediction. Mutations within IDPs are implicated in neurodegenerative diseases and cancer, so a method that can rationally design and probe such sequences is valuable both for basic science and therapeutic development.

The new computational framework was developed by researchers at the Harvard John A. Paulson School of Engineering and Applied Sciences (SEAS) in collaboration with Northwestern University. The study, published in Nature Computational Science, was co-led by SEAS graduate student Ryan Krueger and Krishna Shrinivas (NSF-Simons QuantBio Fellow, now assistant professor at Northwestern), with Michael Brenner at SEAS.

Unlike approaches that rely on large datasets to train surrogate models, this method directly couples differentiable optimization to molecular dynamics simulations. Automatic differentiation — the mathematical technique used to compute exact derivatives of simulation outputs with respect to sequence inputs — enables gradient-based search through sequence space. The result is an efficient, principled way to identify sequences that yield target ensemble properties, such as specific dimensions, loop formation, linker behavior, or sensitivity to physicochemical stimuli.

The authors describe the technique as a search engine that evaluates how even single amino-acid substitutions change ensemble features predicted by physics-informed simulations. By leveraging extant molecular models rather than replacing them, the approach preserves the underlying physics and produces designs grounded in realistic dynamic behavior. This makes the designed proteins “differentiable” in the computational sense: their ensemble statistics change smoothly and predictably as sequence parameters are adjusted.

Demonstrated applications include engineering IDPs with predefined ensemble dimensions, constructing loops and linkers with controlled flexibility, designing sensors that respond sharply to environmental changes, and creating binders tailored to disordered targets with particular conformational biases. These examples illustrate the method’s versatility and its potential to generate bespoke IDPs for research and applied use.

Funding: The project received support from the National Science Foundation AI Institute of Dynamic Systems, the Office of Naval Research, the Harvard Materials Research Science and Engineering Center, and the NSF-Simons Center for Mathematical and Statistical Analysis of Biology at Harvard.

Key Questions Answered:

Q: Why can’t current AI systems like AlphaFold predict all protein structures?

A: Many proteins—called intrinsically disordered proteins—do not adopt a single stable three-dimensional fold. They instead exist as heterogeneous ensembles, which makes static structure prediction inadequate for capturing their functional behavior.

Q: How does this new method differ from standard protein prediction tools?

A: Rather than predicting a single structure, it combines molecular dynamics simulations with automatic differentiation to directly optimize sequences for ensemble-level properties, using real physics to guide design.

Q: Why are disordered proteins important?

A: IDPs mediate critical biological processes like signaling and molecular assembly and are associated with diseases such as Parkinson’s, Alzheimer’s, and certain cancers, making them important targets for both basic research and therapeutic strategies.

About this AI and genetics research news

Author: Anne Manning
Source: Harvard
Contact: Anne Manning – Harvard
Image: The image is credited to Neuroscience News

Original Research: Closed access.
“Generalized design of sequence–ensemble–function relationships for intrinsically disordered proteins” by Ryan Krueger et al. Nature Computational Science


Abstract

Generalized design of sequence–ensemble–function relationships for intrinsically disordered proteins

Design methods for folded proteins have advanced rapidly, but many biologically important sequences are intrinsically disordered and encode a wide ensemble of conformations rather than a single fold. That heterogeneity makes rational design difficult. This work introduces a computational framework to design IDPs by inverting molecular simulations: using differentiable optimization to identify sequences that produce desired ensemble properties. The approach supports diverse design goals and sequence constraints, yielding IDPs with target dimensions, loop and linker behavior, sensitive physicochemical sensors, and binders that recognize disordered substrates with specific conformational biases. Overall, the method offers a general strategy for engineering sequence–ensemble–function relationships in biological macromolecules.