Dauparas, J. et al. Robust deep learning–based protein sequence design using ProteinMPNN. Science 378, 49–56 (2022).
Google Scholar
Watson, J. et al. De novo design of protein structure and function with RFdiffusion. Nature 620, 1089–1100 (2023).
Google Scholar
Ingraham, J. et al. Illuminating protein space with a programmable generative model. Nature 623, 1070–1078 (2023).
Google Scholar
Hsien-Wei Yeh, A. et al. De novo design of luciferases using deep learning. Nature 614, 774–780 (2023).
Google Scholar
Bennett, N. et al. Atomically accurate de novo design of antibodies with RFdiffusion. Nature 649, 183–193 (2026).
Google Scholar
Lisanza, S. et al. Multistate and functional protein design using RoseTTAFold sequence space diffusion. Nat. Biotechnol. 43, 1288–1298 (2025).
Google Scholar
Hong, L. & Kortemme, T. An integrative approach to protein sequence design through multiobjective optimization. PLoS Comput. Biol. 20, e1011953 (2024).
Google Scholar
Ziegler, D. et al. Fine-tuning language models from human preferences. Preprint at https://doi.org/10.48550/arXiv.1909.08593 (2019).
Ouyang, L. et al. Training language models to follow instructions with human feedback. In Proc. 35th Conference on Advances in Neural Information Processing Systems (eds Oh, A. et al.) (ACM, 2022).
Ruffolo, J. A. et al. Design of highly functional genome editors by modeling CRISPR–Cas sequences. Nature 645, 518–525 (2025).
Google Scholar
Nijkamp, E., Ruffolo, J., Weinstein, E., Naik, N. & Madani, A. ProGen2: exploring the boundaries of protein language models. Cell Syst. 14, 968–978 (2023).
Google Scholar
Ivančić, D. et al. Discovery and protein language model-guided design of hyperactive transposases. Nat. Biotechnol. https://doi.org/10.1038/s41587-025-02816-4 (2025).
Munsamy, G. et al. Conditional language models enable the efficient design of proficient enzymes. Preprint at bioRxiv https://doi.org/10.1101/2024.05.03.592223 (2024).
Kotha, S., Springer, J. & Raghunathan, A. Understanding catastrophic forgetting in language models via implicit inference. In Proc. 12th International Conference on Learning Representations (ed. Kim, B.) (ICLR, 2024).
Luo, Y. et al. An empirical study of catastrophic forgetting in large language models during continual fine-tuning. IEEE Trans. Audio Speech Lang. Process. 33, 3776–3786 (2025).
Google Scholar
ESM Team. ESM Cambrian: Revealing the mysteries of proteins with unsupervised learning. EvolutionaryScale https://evolutionaryscale.ai/blog/esm-cambrian (2024).
Nisonoff, H., Xiong, J., Allenspach, S. & Listgarten, J. Unlocking guidance for discrete state-space diffusion and flow models. In Proc. 13th International Conference on Learning Representations (ed. Yue, Y.) (ICLR, 2025).
Dickstein, J., Weiss, E., Maheswaranathan, N. & Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In Proc. 37th International Conference on Machine Learning (eds Bach, F. & Blei, D.) (PMLR, 2015).
Song, Y. et al. Score-based generative modeling through stochastic differential equations. In Proc. 9th International Conference on Learning Representations (ed. Mohamed, S.) (ICLR, 2021).
Dhariwal, P. & Nichol, A. Diffusion models beat GANs on image synthesis. In Proc. 34th Conference on Advances in Neural Information Processing Systems (eds Ranzato, M. et al.) (ACM, 2021).
Ho, J. & Salimans, T. Classifier-free diffusion guidance. In Proc. Neural Information Processing Systems Workshop on Deep Generative Models and Downstream Applications (eds Zhang, C. et al.) (NeurIPS, 2021).
Campbell, A. et al. A continuous time framework for discrete denoising models. In Proc. 35th Conference on Advances in Neural Information Processing Systems (eds Oh, A. et al.) (ACM, 2022).
Sun, H., Yu, L., Dai, B., Schuurmans, D. & Dai, H. Score-based continuous-time discrete diffusion models. In Proc. 11th International Conference on Learning Representations (ed. Liu, Y.) (ICLR, 2023).
Lou, A., Meng, C. & Ermon, S. Discrete diffusion modeling by estimating the ratios of the data distribution. In Proc. 41st International Conference on Machine Learning (eds Salakhutdinov, R. et al.) (PMLR, 2024).
Campbell, A., Yim, J., Barzilay, R., Rainforth, T. & Jaakkola, T. Generative flows on discrete state-spaces: enabling multimodal flows with applications to protein co-design. In Proc. 12th International Conference on Machine Learning (ed. Kim, B.) (ICLR, 2024).
Gat, I. et al. Discrete flow matching. In Proc. 37th Conference on Advances in Neural Information Processing Systems (eds Zhang, C. et al.) (ACM, 2024).
Austin, J., Johnson, D., Ho, J., Tarlow, D. & Berg, R. Structured denoising diffusion models in discrete state-spaces. In Proc. 34th Conference on Advances in Neural Information Processing Systems (eds Ranzato, M. et al.) (ACM, 2021).
Hoogeboom, E. et al. Autoregressive diffusion models. In Proc. 9th International Conference on Learning Representations (ed. Mohamed, S.) (ICLR, 2021).
Shi, J., Han, K., Wang, Z., Doucet, A. & Titsias, M. Simplified and generalized masked diffusion for discrete data. In Proc. 37th Conference on Advances in Neural Information Processing Systems (eds Zhang, C. et al.) (ACM, 2024).
Sahoo, S. et al. Simple and effective masked diffusion language models. In Proc. 37th Conference on Advances in Neural Information Processing Systems (eds Zhang, C. et al.) (ACM, 2024).
Ou, J. et al. Your absorbing discrete diffusion secretly models the conditional distributions of clean data. In Proc. 13th International Conference on Learning Representations (ed. Yue, Y.) (ICLR, 2025).
Gordon, C., Lu, A. & Abbeel, P. Protein language model fitness is a matter of preference. In Proc. 13th International Conference on Learning Representations (ed. Yue, Y.) (ICLR, 2025).
Chen, L. et al. Target sequence-conditioned design of peptide binders using masked language modeling. Nat. Biotechnol. 44, 1002–1010 (2026).
Google Scholar
Blalock, N. et al. Functional alignment of protein language models via reinforcement learning. Preprint at bioRxiv https://doi.org/10.1101/2025.05.02.651993 (2025).
Stocco, F. et al. Guiding generative protein language models with reinforcement learning. Preprint at https://doi.org/10.48550/arXiv.2412.12979 (2024).
Rafailov, R. et al. Direct preference optimization: your language model is secretly a reward model. In Proc. 36th Conference on Advances in Neural Information Processing Systems (eds Oh, A. et al.) (ACM, 2023).
Widatalla, T., Rafailov, R. & Hie, B. Aligning protein generative models with experimental fitness via direct preference optimization. Preprint at bioRxiv https://doi.org/10.1101/2024.05.20.595026 (2024).
Chennakesavalu, S., Hu, F., Ibarraran, S. & Rotskoff, G. Aligning chemical and protein language models with continuous feedback using energy rank alignment. In Proc. ICLR Workshop on Generative and Experimental Perspectives for Biomolecular Design (eds Liu, C. et al.) (ICLR, 2025).
Wang, C. et al. Fine-tuning discrete diffusion models via reward optimization with applications to dna, and protein design. In Proc. 13th International Conference on Learning Representations (ed. Yue, Y.) (ICLR, 2025).
Brooks, J. et al. Steering masked discrete diffusion models via discrete denoising posterior prediction. In Proc. 13th International Conference on Learning Representations (ed. Yue, Y.) (ICLR, 2025).
Wallace, B. et al. Diffusion model alignment using direct preference optimization. In Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (ed. Ceballos, C.) (IEEE, 2024).
Yang, J. et al. Steering generative models with experimental data for protein fitness optimization. In Proc. 38th Conference on Advances in Neural Information Processing Systems (eds Belgrave, D. et al.) (ACM, 2025).
Emami, P., Perreault, A., Law, J., Biagioni, D. & John, P. Plug & play directed evolution of proteins with gradient-based discrete MCMC. In Proc. Neural Information Processing Systems Machine Learning in Structural Biology Workshop (eds Rao, R. et al.) (NeurIPS, 2022).
Wu, L., Trippe, B., Naesseth, C., Blei, D. & Cunningham, J. Practical and asymptotically exact conditional sampling in diffusion models. In Proc. 36th Conference on Advances in Neural Information Processing Systems (eds Oh, A. et al.) (ACM, 2023).
Li, X. et al. Derivative-free guidance in continuous and discrete diffusion models with soft value-based decoding. In Proc. 38th Conference on Advances in Neural Information Processing Systems (eds Belgrave, D. et al.) (ACM, 2025).
Singhal, R. et al. A general framework for inference-time scaling and steering of diffusion models. In Proc. 42nd International Conference on Machine Learning (eds Singh, A. et al.) (PMLR, 2025).
Hayes, T. et al. Simulating 500 million years of evolution with a language model. Science 387, 850–858 (2025).
Google Scholar
Hsu, C. et al. Learning inverse folding from millions of predicted structures. In Proc. 39th International Conference on Machine Learning (eds Chaudhuri, K. et al.) (PMLR, 2022).
Alamdari, S. et al. Protein generation with evolutionary diffusion: sequence is all you need. Preprint at bioRxiv https://doi.org/10.1101/2023.09.11.556673 (2023).
Wang, X. et al. Diffusion language models are versatile protein learners. In Proc. 41st International Conference on Machine Learning (eds Salakhutdinov, R. et al.) (PMLR, 2024).
Madani, A. et al. Large language models generate functional protein sequences across diverse families. Nat. Biotechnol. 41, 1099–1106 (2023).
Google Scholar
Hu, E. J. et al. LoRA: low-rank adaptation of large language models. In Proc. 10th International Conference on Learning Representations (eds Hofmann, K. & Rush, A.) (ICLR, 2022).
Tsuboyama, K. et al. Mega-scale experimental analysis of protein folding stability in biology and design. Nature 620, 434–444 (2023).
Google Scholar
Chen, Y. et al. Deep mutational scanning of an oxygen-independent fluorescent protein CreiLOV for comprehensive profiling of mutational and epistatic effects. ACS Synth. Biol. 12, 1461–1473 (2023).
Google Scholar
Starita, L. et al. Activity-enhancing mutations in an E3 ubiquitin ligase identified by high-throughput mutagenesis. Proc. Natl Acad. Sci. USA 110, E1263–E1272 (2013).
Google Scholar
Melamed, D., Young, D., Gamble, C., Miller, C. & Fields, S. Deep mutational scanning of an RRM domain of the Saccharomyces cerevisiae poly(A)-binding protein. RNA 19, 1537–1551 (2013).
Google Scholar
Wang, B. et al. Active learning-guided optimization of cell-free biosensors for lead testing in drinking water. Nat. Commun. 17, 261 (2026).
Google Scholar
Rees, H. & Liu, D. Base editing: precision chemistry on the genome and transcriptome of living cells. Nat. Rev. Genet. 19, 770–788 (2018).
Google Scholar
Gaudelli, N. et al. Programmable base editing of A•T to G•C in genomic DNA without DNA cleavage. Nature 551, 464–471 (2017).
Google Scholar
Gaudelli, N. et al. Directed evolution of adenine base editors with increased activity and therapeutic application. Nat. Biotechnol. 38, 892–900 (2020).
Google Scholar
Richter, M. et al. Phage-assisted evolution of an adenine base editor with improved Cas domain compatibility and activity. Nat. Biotechnol. 38, 883–891 (2020).
Google Scholar
Xiao, Y., Wu, Y. & Tang, W. An adenine base editor variant expands context compatibility. Nat. Biotechnol. 42, 1442–1453 (2024).
Google Scholar
Ranzau, B., Rallapalli, K., Evanoff, M., Paesani, F. & Komor, A. The wild-type tRNA adenosine deaminase enzyme TadA is capable of sequence-specific DNA base editing. ChemBioChem 24, e202200788 (2023).
Google Scholar
Ghazvininejad, M., Levy, O., Liu, Y. & Zettlemoyer, L. Mask-Predict: parallel decoding of conditional masked language models. In Proc. 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (eds Inui, K. et al.) (ACL, 2019).
Wang, A. & Cho, K. BERT has a mouth, and it must speak: BERT as a Markov random field language model. In Proc. Workshop on Methods for Optimizing and Evaluating Neural Language Generation (eds Bosselut, A. et al.) (ACL, 2019).
Chang, H., Zhang, H., Jiang, L., Liu, C. & Freeman, W. MaskGIT: masked generative image transformer. In Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (ed. O’Conner, L.) (IEEE, 2022).
Li, Y. et al. Promises and pitfalls of generative masked language modeling: theoretical framework and practical guidelines. In Proc. 41st International Conference on Machine Learning (eds Salakhutdinov, R. et al.) (PMLR, 2024).
Uria, B., Murray, I. & Larochelle, H. A deep and tractable density estimator. In Proc. 31st International Conference on Machine Learning (eds Xing, E. P. & Jebara, T.) (PMLR, 2014).
Yang, Z. et al. Xlnet: generalized autoregressive pretraining for language understanding. In Proc. 32nd Conference on Advances in Neural Information Processing Systems (eds Wallach, H.) (ACM, 2019).
Jing, B. et al. Generating functional and multistate proteins with a multimodal diffusion transformer. Preprint at bioRxiv https://doi.org/10.1101/2025.09.03.672144 (2025).
Lin, Z. et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130 (2023).
Google Scholar
Ho, J., Jain, A. & Abbeel, P. Denoising diffusion probabilistic models. In Proc. 33rd Conference on Advances in Neural Information Processing Systems (eds Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M. F. & Lin, H.) (ACM, 2020).
Gillespie, D. Exact stochastic simulation of coupled chemical reactions. J. Phys. Chem. 81, 2340–2361 (1977).
Google Scholar
Gillespie, D. Approximate accelerated stochastic simulation of chemically reacting systems. J. Phys. Chem. 115, 1716–1733 (2001).
Google Scholar
Hoogeboom, E., Nielsen, D., Jaini, P., Forré, P. & Welling, M. Argmax flows and multinomial diffusion: learning categorical distributions. In Proc. 34th Conference on Advances in Neural Information Processing Systems (eds Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P. S. & Wortman Vaughan, J.) (ACM, 2021).
Albergo, M. & Eijnden, E. Building normalizing flows with stochastic interpolants. In Proc. 11th International Conference on Learning Representations (ed. Liu, Y.) (ICLR, 2023).
Lipman, Y., Chen, R., Hamu, H., Nickel, M. & Le, M. Flow matching for generative modeling. In Proc. 11th International Conference on Learning Representations (ed. Liu, Y.) (ICLR, 2023).
Liu, X., Gong, C. & Liu, Q. Flow straight and fast: learning to generate and transfer data with rectified flow. In Proc. 11th International Conference on Learning Representations (ed. Liu, Y.) (ICLR, 2023).
Kingma, D. & Gao, R. Understanding diffusion objectives as the ELBO with simple data augmentation. In Proc. 36th Conference on Advances in Neural Information Processing Systems (eds Oh, A. et al.) (ACM, 2023).
Peng, F. et al. Path planning for masked diffusion model sampling. Preprint at https://doi.org/10.48550/arXiv.2502.03540 (2025).
Devlin, J., Chang, M., Lee, K. & Toutanova, K. BERT: pre-training of deep bidirectional transformers for language understanding. In Proc. 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (eds Burstein, J., Doran, C. & Solorio, T.) (ACL, 2019).
Strauss, R. & Oliva, J. Arbitrary conditional distributions with energy. In Proc. 34th Conference on Advances in Neural Information Processing Systems (eds Ranzato, M. et al.) (ACM, 2021).
Karras, T. et al. Guiding a diffusion model with a bad version of itself. In Proc. 37th Conference on Advances in Neural Information Processing Systems (eds Zhang, C. et al.) (ACM, 2024).
Grathwohl, W., Swersky, K., Hashemi, M., Duvenaud, D. & Maddison, C. Oops I took a gradient: scalable sampling for discrete distributions. In Proc. 38th International Conference on Machine Learning (eds Meila, M. & Zhang, T.) (PMLR, 2021).
Schiff, Y. et al. Simple Guidance Mechanisms for Discrete Diffusion Models. In Proc. 13th International Conference on Learning Representations (ed. Yue, Y.) (ICLR, 2025).
Chen, Z. et al. Fast sampling via discrete non-Markov diffusion models with predetermined transition time. In Proc. 37th Conference on Advances in Neural Information Processing Systems (eds Zhang, C. et al.) (ACM, 2024).
Zheng, K. et al. Masked diffusion models are secretly time-agnostic masked models and exploit inaccurate categorical sampling. In Proc. 13th International Conference on Learning Representations (ed. Yue, Y.) (ICLR, 2025).
Del Moral, P. & Penev, S. Stochastic Processes: From Applications to Theory 1st edn (Chapman and Hall/CRC, 2017).
Stocco, F., Garibbo, M. & Ferruz, N. Steering generative models for protein design: aligning and conditioning strategies. Curr. Opin. Struct. Biol. 98, 103250 (2026).
Google Scholar
Dettmers, T., Pagnoni, A., Holtzman, A. & Zettlemoyer, L. QLORA: efficient finetuning of quantized LLMs. In Proc. 36th Conference on Advances in Neural Information Processing Systems (eds Oh, A. et al.) (ACM, 2023).
Jiang, H. et al. SMART: robust and efficient fine-tuning for pre-trained natural language models through principled regularized optimization. In Proc. 58th Annual Meeting of the Association for Computational Linguistics (eds Jurafsky, D. et al.) (ACL, 2020).
Dodge, J. et al. Fine-tuning pretrained language models: weight initializations, data orders, and early stopping. Preprint at https://doi.org/10.48550/arXiv.2002.06305 (2020).
Mosbach, M., Andriushchenko., M. & Klakow, D. On the stability of fine-tuning BERT: misconceptions, explanations, and strong baselines. In Proc. 9th International Conference on Learning Representations (ed. Mohamed, S.) (ICLR, 2021).
Loshchilov, I., and Hutter, F. Decoupled Weight Decay Regularization. In Proc. 7th International Conference on Learning Representations (ed. Sainath, T.) (ICLR, 2019).
Hsu, C., Nisonoff, H., Fannjiang, C. & Listgarten, J. Learning protein fitness models from evolutionary and assay-labeled data. Nat. Biotechnol. 40, 1114–1122 (2022).
Google Scholar
Brookes, D., Park, H. & Listgarten, J. Conditioning by adaptive sampling for robust design. In Proc. 36th International Conference on Machine Learning (eds Chaudhuri, K. & Salakhutdinov, R.) (PMLR, 2019).
UniProt Consortium UniProt: the universal protein knowledgebase in 2021. Nucleic Acids Res. 49, D480–D489 (2021).
Google Scholar
Sanderson, T., Bileschi, M., Belanger, D. & Colwell, L. ProteInfer, deep neural networks for protein functional inference. eLife 12, e80942 (2023).
Google Scholar
Yu, T. et al. Enzyme function prediction using contrastive learning. Science 379, 1358–1363 (2023).
Google Scholar
Dawson, N. et al. CATH: an expanded resource to predict protein function through structure and sequence. Nucleic Acids Res. 45, D289–D295 (2017).
Google Scholar
Geffner, T. et al. Proteina: scaling flow-based protein structure generative models. In Proc. 13th International Conference on Learning Representations (ed. Yue, Y.) (ICLR, 2025).
Kingma, D. & Ba, J. Adam: a method for stochastic optimization. In Proc. 3rd International Conference on Learning Representations (eds Bengio, Y. & LeCun, Y.) (ICLR, 2015).
Grandini, M., Bagli, E. & Visani, G. Metrics for multi-class classification: an overview. Preprint at https://doi.org/10.48550/arXiv.2008.05756 (2020).
Yim, J. et al. SE(3) diffusion model with application to protein backbone generation. In Proc. 40th International Conference on Machine Learning (eds Krause, A. et al.) (PMLR, 2023).
Budzko, L., Sobiech, K., Jackowiak, P. & Figlerowicz, M. Engineered deaminases as a key component of DNA and RNA editing tools. Mol. Ther. Nucleic Acids 34, 102062 (2023).
Google Scholar
Del Arco, J., Acosta, J. & Lucas, J. Biotechnological applications of purine and pyrimidine deaminases. Biotechnol. Adv. 77, 108473 (2024).
Google Scholar
Hopf, T. et al. Mutation effects predicted from sequence co-variation. Nat. Biotechnol. 35, 128–135 (2017).
Google Scholar
Paszke, A. et al. Pytorch: an imperative style, high-performance deep learning library. In Proc. 32nd Conference on Advances in Neural Information Processing Systems (eds Wallach, H.) (ACM, 2019).
Ansel, J. et al. PyTorch 2: faster machine learning through dynamic Python bytecode transformation and graph compilation. In Proc. 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (eds Abu-Ghazaleh, N., Gupta, R., Musuvathi, M. & Tsafrir, D.) (ACM, 2024).
Xiong, J., Nisonoff, H. & Gaur, I. ProteinGuide data. Zenodo https://doi.org/10.5281/zenodo.15635060 (2026).