arXiv++ Combinatorics

Browse math.CO papers from arXiv

phylogenetic

427 papers tagged with this keyword
Inferring DAGs and Phylogenetic Networks from Least Common Ancestors
A least common ancestor (LCA) of two leaves in a directed acyclic graph (DAG) is a vertex that is an ancestor of both leaves and has no proper descendant that is also their common ancestor. LCAs capture hierarchical relationships in rooted trees and, more generally, in DAGs. In 1981, Aho et al. introduced the problem of determining whether a set of pairwise LCA constraints on a set $X$, of the form $(i,j)<(k,l)$ with $i,j,k,l\in X$, can be realized by a rooted tree whose leaf set is $X$, such that whenever $(i,j)<(k,l)$, the LCA of $i,j$ is a descendant of that of $k,l$. They also presented a polynomial-time algorithm, BUILD, to solve this problem. However, many such constraint systems cannot be realized by any tree, prompting the question of whether they can be realized by a more general DAG. We extend Aho et al.'s framework from trees to DAGs, providing both theoretical and algorithmic foundations for reasoning about LCA constraints in this broader setting. Given a collection $R$ of LCA constraints, we define its $+$-closure $R^+$, capturing additional LCA relations implied by $R$. Using $R^+$, we construct a canonical DAG $G_R$ and prove that $R$ is DAG-realizable if and only if it is realized by $G_R$. We further adapt this construction to phylogenetic networks, defining a canonical network $N_R$ and prove that it is regular, i.e., it coincides with the Hasse diagram of its underlying set system. Finally, we show that for any DAG-realizable $R$, its classical closure - comprising all LCA constraints that hold in every DAG realizing $R$ - coincides with its $+$-closure. All constructions are computable in polynomial time, and we provide explicit algorithms for each.
Characterizations of undirected 2-quasi best match graphs
Bipartite best match graphs (BMG) and their generalizations arise in mathematical phylogenetics as combinatorial models describing evolutionary relationships among related genes in a pair of species. In this work, we characterize the class of \emph{undirected 2-quasi-BMGs} (un2qBMGs), which form a proper subclass of the $P_6$-free chordal bipartite graphs. We show that un2qBMGs are exactly the class of bipartite graphs free of $P_6$, $C_6$, and the eight-vertex Sunlet$_4$ graph. Equivalently, a bipartite graph $G$ is un2qBMG if and only if every connected induced subgraph contains a ``heart-vertex'' which is adjacent to all the vertices of the opposite color. We further provide a $O(|V(G)|^3)$ algorithm for the recognition of un2qBMGs that, in the affirmative case, constructs a labeled rooted tree that ``explains'' $G$. Finally, since un2qBMGs coincide with the $(P_6,C_6)$-free bi-cographs, they can also be recognized in linear time.
Generalizing matrix representations to fully heterochronous ranked tree shapes
Phylogenetic tree shapes capture fundamental signatures of evolution. We consider ``ranked'' tree shapes, which are equipped with a total order on the internal nodes compatible with the tree graph. Recent work has established an elegant bijection of ranked tree shapes and a class of integer matrices, called \textbf{F}-matrices, defined by simple inequalities. This formulation is for isochronous ranked tree shapes, where all leaves share the same sampling time, such as in the study of ancient human demography from present-day individuals. Another important style of phylogenetics concerns trees where the ``timing'' of events is by branch length rather than calendar time. This style of tree, called a rooted phylogram, is output by popular maximum-likelihood methods. These trees are broadly relevant, such as to study the affinity maturation of B cells in the immune system. Discretizing time in a rooted phylogram gives a fully heterochronous ranked tree shape, where leaves are part of the total order. Here we extend the \textbf{F}-matrix framework to such fully heterochronous ranked tree shapes. We establish an explicit bijection between a class of \textbf{F}-matrices and the space of such tree shapes. The matrix representation has the key feature that values at any entry are highly constrained via four previous entries, enabling straightforward enumeration of all valid tree shapes. We also use this framework to develop probabilistic models on ranked tree shapes. Our work extends understanding of combinatorial objects that have a rich history in the literature.
2025-10-23 v2
Labeling and folding multi-labeled trees
In 1989 Erdős and Székely showed that there is a bijection between (i) the set of rooted trees with $n+1$ vertices whose leaves are bijectively labeled with the elements of $[\ell]=\{1,2,\dots,\ell\}$ for some $\ell \leq n$, and (ii) the set of partitions of $[n]=\{1,2,\dots,n\}$. They established this via a labeling algorithm based on the anti-lexicographic ordering of non-empty subsets of $[n]$ which extends the labeling of the leaves of a given tree to a labeling of all of the vertices of that tree. In this paper, we generalize their approach by developing a labeling algorithm for multi-labeled trees, that is, rooted trees whose leaves are labeled by positive integers but in which distinct leaves may have the same label. In particular, we show that certain orderings of the set of all finite, non-empty multisets of positive integers can be used to characterize partitions of a multiset that arise from labelings of multi-labeled trees. As an application, we show that the recently introduced class of labelable phylogenetic networks is precisely the class of phylogenetic networks that are stable relative to the so-called folding process on multi-labeled trees. We also give a bijection between the labelable phylogenetic networks with leaf-set $[n]$ and certain partitions of multisets.
Recognizing Leaf Powers and Pairwise Compatibility Graphs is NP-Complete
Leaf powers and pairwise compatibility graphs were introduced over twenty years ago as simplified graph models for phylogenetic trees. Despite significant research, several properties of these graph classes remain poorly understood. In this paper, we establish that the recognition problem for both classes is NP-complete. We extend this hardness result to a broader hierarchy of graph classes, including pairwise compatibility graphs and their generalizations, multi interval pairwise compatibility graphs.
2025-10-10 v2
Parameterized Algorithms for Diversity of Networks with Ecological Dependencies
For a phylogenetic tree, the phylogenetic diversity of a set A of taxa is the total weight of edges on paths to A. Finding small sets of maximal diversity is crucial for conservation planning, as it indicates where limited resources can be invested most efficiently. In recent years, efficient algorithms have been developed to find sets of taxa that maximize phylogenetic diversity either in a phylogenetic network or in a phylogenetic tree subject to ecological constraints, such as a food web. However, these aspects have mostly been studied independently. Since both factors are biologically important, it seems natural to consider them together. In this paper, we introduce decision problems where, given a phylogenetic network, a food web, and integers k, and D, the task is to find a set of k taxa with phylogenetic diversity of at least D under the maximize all paths measure, while also satisfying viability conditions within the food web. Here, we consider different definitions of viability, which all demand that a "sufficient" number of prey species survive to support surviving predators. We investigate the parameterized complexity of these problems and present several fixed-parameter tractable (FPT) algorithms. Specifically, we provide a complete complexity dichotomy characterizing which combinations of parameters - out of the size constraint k, the acceptable diversity loss D, the scanwidth of the food web, the maximum in-degree in the network, and the network height h - lead to W[1]-hardness and which admit FPT algorithms. Our primary methodological contribution is a novel algorithmic framework for solving phylogenetic diversity problems in networks where dependencies (such as those from a food web) impose an order, using a color coding approach.
2025-09-19
Ordered Leaf Attachment (OLA) Vectors can Identify Reticulation Events even in Multifurcated Trees
Recently, a new vector encoding, Ordered Leaf Attachment (OLA), was introduced that represents $n$-leaf phylogenetic trees as $n-1$ length integer vectors by recording the placement location of each leaf. Both encoding and decoding of trees run in linear time and depend on a fixed ordering of the leaves. Here, we investigate the connection between OLA vectors and the maximum acyclic agreement forest (MAAF) problem. A MAAF represents an optimal breakdown of $k$ trees into reticulation-free subtrees, with the roots of these subtrees representing reticulation events. We introduce a corrected OLA distance index over OLA vectors of $k$ trees, which is easily computable in linear time. We prove that the corrected OLA distance corresponds to the size of a MAAF, given an optimal leaf ordering that minimizes that distance. Additionally, a MAAF can be easily reconstructed from optimal OLA vectors. We expand these results to multifurcated trees: we introduce an $O(kn \cdot m\log m)$ algorithm that optimally resolves a set of multifurcated trees given a leaf-ordering, where $m$ is the size of a largest multifurcation, and show that trees resolved via this algorithm also minimize the size of a MAAF. These results suggest a new approach to fast computation of phylogenetic networks and identification of reticulation events via random permutations of leaves. Additionally, in the case of microbial evolution, a natural ordering of leaves is often given by the sample collection date, which means that under mild assumptions, reticulation events can be identified in polynomial time on such datasets.
2025-09-05
Revealing the building blocks of tree balance: fundamental units of the Sackin and Colless Indices
Over the past decades, more than 25 phylogenetic tree balance indices and several families of such indices have been proposed in the literature -- some of which even contain infinitely many members. It is well established that different indices have different strengths and perform unequally across application scenarios. For example, power analyses have shown that the ability of an index to detect the generative model of a given phylogenetic tree varies significantly between indices. This variation in performance motivates the ongoing search for new and possibly \enquote{better} (im)balance indices. An easy way to generate a new index is to construct a compound index, e.g., a linear combination of established indices. Two of the most prominent and widely used imbalance indices in this context are the Sackin index and the Colless index. In this study, we show that these classic indices are themselves compound in nature: they can be decomposed into more elementary components that independently satisfy the defining properties of a tree (im)balance index. We further show that the difference Colless minus Sackin results in another imbalance index that is minimized (amongst others) by all Colless minimal trees. Conversely, the difference Sackin minus Colless forms a balance index. Finally, we compare the building blocks of which the Sackin and the Colless indices consist to these indices as well as to the stairs2 index, which is another index from the literature. Our results suggest that the elementary building blocks we identify are not only foundational to established indices but also valuable tools for analyzing disagreement among indices when comparing the balance of different trees.
2025-08-29 v2
When Many Trees Go to War: On Sets of Phylogenetic Trees With Almost No Common Structure
It is known that any two trees on the same $n$ leaves can be displayed by a network with $n-2$ reticulations, and there are two trees that cannot be displayed by a network with fewer reticulations. But how many reticulations are needed to display multiple trees? For any set of $t$ trees on $n$ leaves, there is a trivial network with $(t - 1)n$ reticulations that displays them. To do better, we have to exploit common structure of the trees to embed non-trivial subtrees of different trees into the same part of the network. In this paper, we show that for $t \in o(\sqrt{\lg n})$, there is a set of $t$ trees with virtually no common structure that could be exploited. More precisely, we show for any $t\in o(\sqrt{\lg n})$, there are $t$ trees such that any network displaying them has $(t-1)n - o(n)$ reticulations. For $t \in o(\lg n)$, we obtain a slightly weaker bound. We also prove that already for $t = c\lg n$, for any constant $c > 0$, there is a set of $t$ trees that cannot be displayed by a network with $o(n \lg n)$ reticulations, matching up to constant factors the known upper bound of $O(n \lg n)$ reticulations sufficient to display \emph{all} trees with $n$ leaves. These results are based on simple counting arguments and extend to unrooted networks and trees.
2025-08-21
Defining a phylogenetic tree with the minimum number of small-state characters
Phylogenetic trees represent evolutionary relationships and can be uniquely defined by sets of finite-state biological characteristics. Despite prior work showing that sufficiently large trees can be determined by $r$-state character sets, the minimal leaf thresholds $n_r$ remain largely unknown. In this work, we establish the 3-state case as $n_3 = 8$, providing a concrete base for higher-state analyses. We then resolve the 5-state problem by constructing a counterexample for $n=15$ and proving that for $n \geq 16$, $\lceil (n-3)/4 \rceil$ 5-state characters suffice to uniquely define any tree. Our approach relies on rigorous mathematical induction with complete verification of base cases and logically consistent inductive steps, offering new insights into the minimal conditions for character-based tree identification.
2025-08-19
A sharp lower bound for the number of phylogenetic trees displayed by a tree-child network
A normal (phylogenetic) network with $k$ reticulations displays $2^k$ phylogenetic trees. In this paper, we establish an analogous result for tree-child (phylogenetic) networks with no underlying $3$-cycles. In particular, we show that a tree-child network with $k\ge 2$ reticulations and no underlying $3$-cycles displays at least $2^{k/2}$ phylogenetic trees if $k$ is even and at least $\frac{3}{2\sqrt{2}}2^{k/2}$ if $k$ is odd. Moreover, we show that these bounds are sharp and characterise the tree-child networks that attain these bounds.
2025-08-07
Identifiability of Large Phylogenetic Mixtures for Many Phylogenetic Model Structures
Identifiability of phylogenetic models is a necessary condition to ensure that the model parameters can be uniquely determined from data. Mixture models are phylogenetic models where the probability distributions in the model are convex combinations of distributions in simpler phylogenetic models. Mixture models are used to model heterogeneity in the substitution process in DNA sequences. While many basic phylogenetic models are known to be identifiable, mixture models in generality have only been shown to be identifiable in certain cases. We expand the main theorem of [Rhodes, Sullivant 2012] to prove identifiability of mixture models in equivariant phylogenetic models, specifically the Jukes-Cantor, Kimura 2-parameter model, Kimura 3-parameter model and the Strand Symmetric model.
2025-08-04 v2
Tropical cluster varieties of type C
We explicitly describe the tropicalization of a cluster variety of finite type C, realizing it as the space of axially symmetric phylogenetic trees. We also find all occurring sign patterns of coordinates, for both the cluster variety and the cluster configuration space. We show that each of the corresponding signed tropicalizations is, combinatorially, dual to either a cyclohedron or an associahedron. As additional results, we construct Gröbner and tropical bases for the defining ideals of both varieties, and classify the arising toric degenerations.
2025-07-30
Phylogenetic network models as graphical models
The displayed tree phylogenetic network model is shown to sit as a natural submodel of the graphical model associated to a directed acyclic graph (DAG). This representation allows to derive a number of results about the displayed tree model. In particular, the concept of a local modification to a DAG model is developed and applied to the displayed tree model. As an application, some nonidentifiability issues related to the displayed tree models are highlighted as they relate to reticulation edges and stacked reticulations in the networks. We also derive rank conditions on flattenings of probability tensors for the displayed tree model, generalizing classic results for phylogenetic tree models.
Characterizing semi-directed phylogenetic networks and their multi-rootable variants
In evolutionary biology, phylogenetic networks are graphs that provide a flexible framework for representing complex evolutionary histories that involve reticulate evolutionary events. Recently phylogenetic studies have started to focus on a special class of such networks called semi-directed networks. These graphs are defined as mixed graphs that can be obtained by de-orienting some of the arcs in some rooted phylogenetic network, that is, a directed acyclic graph whose leaves correspond to a collection of species and that has a single source or root vertex. However, this definition of semi-directed networks is implicit in nature since it is not clear when a mixed-graph enjoys this property or not. In this paper, we introduce novel, explicit mathematical characterizations of semi-directed networks, and also multi-semi-directed networks, that is, mixed graphs that can be obtained from directed phylogenetic networks that may have more than one root. In addition, through extending foundational tools from the theory of rooted networks into the semi-directed setting - such as cherry picking sequences, omnians, and path partitions - we characterize when a (multi-)semi-directed network can be obtained by de-orienting some rooted network that is contained in one of the well-known classes of tree-child, orchard, tree-based or forest-based networks. These results address structural aspects of (multi-)semi-directed networks and pave the way to improved theoretical and computational analyses of such networks, for example, within the development of algebraic evolutionary models that are based on such networks.
Distinguishing Phylogenetic Level-2 Networks with Quartets and Inter-Taxon Quartet Distances
The inference of phylogenetic networks, which model complex evolutionary processes including hybridization and gene flow, remains a central challenge in evolutionary biology. Until now, statistically consistent inference methods have been limited to phylogenetic level-1 networks, which allow no interdependence between reticulate events. In this work, we establish the theoretical foundations for a statistically consistent inference method for a much broader class: semi-directed level-2 networks that are outer-labeled planar and galled. We precisely characterize the features of these networks that are distinguishable from the topologies of their displayed quartet trees. Moreover, we prove that an inter-taxon distance derived from these quartets is circular decomposable, enabling future robust inference of these networks from quartet data, such as concordance factors obtained from gene tree distributions under the Network Multispecies Coalescent model. Our results also have novel identifiability implications across different data types and evolutionary models, applying to any setting in which displayed quartets can be distinguished.
2025-07-21
On Ward Numbers and Increasing Schröder Trees
The Ward numbers $W(n,k)$ combinatorially enumerate set partitions with block sizes $\geq 2$ and phylogenetic trees (total partition trees). We prove that $W(n,k)$ also counts \emph{increasing Schröder trees} by verifying they satisfy Ward's recurrence. We construct a direct type-preserving bijection between total partition trees and increasing Schröder trees, complementing known type-preserving bijections to set partitions (including Chen's decomposition for increasing Schröder trees). Weighted generalizations extend these bijections to enriched increasing Schröder trees trees and Schröder trees trees, yielding new links to labeled rooted trees. Finally, we deduce a functional equation for weighted increasing Schröder trees, whose solution using Chen's decomposition leads to a combinatorial interpretation of a Lagrange inversion variant.
Order-Dependent Dissimilarity Measures on Phylogenetic Trees
Ordered leaf attachment, Phylo2Vec, and HOP are three recently introduced vector representations for rooted phylogenetic trees where the representation is determined by an ordering of the underlying leaf set X. Comparing the vectors of two rooted phylogenetic X-trees T and T' for a fixed ordering on X leads to polynomial-time computable measure for the dissimilarity of T and T', albeit dependent on the choice of the leaf ordering. For each of ordered leaf attachment, Phylo2Vec, and HOP, we compare this measure with the rooted subtree prune and regraft distance (rSPR), the hybrid number, and the temporal tree-child hybrid number of T and T'. Although there is no direct relationship between rSPR and any of the three vector-based measures, we show that, when minimized over all orderings, the hybrid number is equivalent to HOP, and an upper bound on the other two. Moreover, when minimized over all orderings induced by common cherry-picking sequences of T and T', the temporal tree-child hybrid number of T and T' is equivalent to each of the three vector-based measures.
2025-06-29 v3
A dichotomy law for certain classes of phylogenetic networks
Many classes of phylogenetic networks have been proposed in the literature. A feature of several of these classes is that if one restricts a network in the class to a subset of its leaves, then the resulting network may no longer lie within this class. This has implications for their biological applicability, since some species -- which are the leaves of an underlying evolutionary network -- may be missing (e.g., they may have become extinct, or there are no data available for them) or we may simply wish to focus attention on a subset of the species. On the other hand, certain classes of networks are `closed' when we restrict to subsets of leaves, such as (i) the classes of all phylogenetic networks or all phylogenetic trees; (ii) the classes of galled networks, simplicial networks, galled trees; and (iii) the classes of networks that have some parameter that is monotone-under-leaf-subsampling (e.g., the number of reticulations, height, etc.) bounded by some fixed value. It is easily shown that a closed subclass of phylogenetic trees is either all trees or a vanishingly small proportion of them (as the number of leaves grows). In this short paper, we explore whether this dichotomy phenomenon holds for other classes of phylogenetic networks, and their subclasses.
A strengthened bound on the number of states required to characterize maximum parsimony distance
In this article we prove that the distance $d_{\mathrm{MP}}(T_1,T_2) = k$ between two unrooted binary phylogenetic trees $T_1, T_2$ on the same set of taxa can be defined by a character that is convex on one of $T_1, T_2$ and which has at most $2k$ states. This significantly improves upon the previous bound of $7k-5$ states. We also show that for every $k \geq 1$ there exist two trees $T_1, T_2$ with $d_{\mathrm{MP}}(T_1,T_2) = k$ such that at least $k+1$ states are necessary in any character that achieves this distance and which is convex on one of $T_1, T_2$. We augment these lower and upper bounds with an empirical analysis which shows that in practice significantly fewer than $k+1$ states are usually required.