Academic paper
Assembly Theory and the Smallest Grammar Problem
Abstract
Assembly theory (AT) quantifies complexity through the assembly index (ASI) -- a metric proven equivalent to the size of the smallest straight-line program (SLP). While this equivalence links AT to data compression -- a field shaped by decades of research -- the practical efficacy of specific algorithms as ASI approximators remains largely unexplored. This study provides an empirical evaluation of eight compression algorithms, spanning grammar-based and dictionary schemes (CAs) across a dataset of 408 strings -- comprising 368 synthetic strings (including max-complexity strings and strings with varying levels of symbol distribution balance, quantified by Shannon entropy) and 40 natural biological sequences (genomic and proteomic). We introduce Re-Pair T-NDR, a branch-and-bound tie-resolving variant of Re-Pair, providing a tighter upper bound on the ASI than the other CAs researched in this study. Increasing alphabet size extends the regime in which CAs closely track the ASI, delaying a ``divergence point'', where a CA deviates from the optimal assembly path. The results establish compression algorithms as practical tools for ASI estimation.
This public page contains bibliographic metadata and the author abstract. Use the reader for licensed document access.
Open licensed paper reader