Work overview

Section 01 of 06

1 Introduction

seqme: a Python library for evaluating biological sequence design from generative models

Rasmus Møller-Larsen, Adam Izdebski, Jan Olszewski, Pankhil Gawade, Michal Kmicikiewicz, Wojciech Zarzecki, and Ewa Szczurek · 2026

Contents

Section 01 of 06

  1. 011 Introduction
  2. 022 Metrics
  3. 033 Additional functionalities of seqme
  4. 044 Implementation
  5. 055 Case studies
  6. 066 Discussion
Text size
Work overview

Section 1 of 6

1 Introduction

Rasmus Møller-Larsen, Adam Izdebski, Jan Olszewski, Pankhil Gawade, Michal Kmicikiewicz, Wojciech Zarzecki, and Ewa Szczurek · about 1 minutes

Biological sequences are composed of ordered chains—of nucleotides in DNA and RNA, amino acids in peptides and proteins, and atoms in small molecules. These sequences carry the information on structure, interactions, and function of the encoded molecules, making sequence modeling central to understanding and engineering biology (Anfinsen 1973, Kozak 1986). Recent years have witnessed a surge of generative AI models and algorithms for biological sequence design, including the development of novel compounds (Tang et al. 2024, Izdebski et al. 2025), peptides (Szymczak et al. 2025), and proteins (Kortemme 2024, Kmicikiewicz et al. 2025). Evaluating biological designs requires multiple complementary criteria, such as fidelity to a reference dataset’s distribution, novelty, diversity, and optimization of desired properties. Overlooking these aspects when evaluating computational methods, in particular generative AI models, can result in failures to generate sequences that are both high quality and functionally useful (Theis et al. 2016). Therefore, a plethora of evaluation metrics have been proposed to measure performance and identify failure modes of generative models (Manduchi et al. 2025). However, to date, no software library has been developed that unifies these efforts and is dedicated specifically to evaluating biological sequence design.

To address this gap, we introduce seqme—the first library for end-to-end evaluation of generative AI and computational algorithms for de novo biological sequence design. The library provides access to a collection of metrics, embedding models and property predictors, while remaining highly extensible. It supports major types of biological sequences, including small molecules, DNA, ncRNA, mRNA, peptides, and proteins. In addition, seqme enables assessment of single-shot design and iterative sequence discovery pipelines. By offering a unified and reproducible framework, the library establishes a practical foundation for fair comparison and accelerated progress in biological sequence design.