arXiv++ Combinatorics

Browse math.CO papers from arXiv

phylogenetic

427 papers tagged with this keyword
Bounds on the Treewidth of Level-k Rooted Phylogenetic Networks
Phylogenetic networks are directed acyclic graphs that depict the genomic evolution of related taxa. Reticulation nodes in such networks (nodes with more than one parent) represent reticulate evolutionary events, such as recombination, reassortment, hybridization, or horizontal gene transfer. Typically, the complexity of a phylogenetic network is expressed in terms of its level, i.e., the maximum number of edges that are required to be removed from each biconnected component of the phylogenetic network to turn it into a tree. Here, we study the relationship between the level of a phylogenetic network and another popular graph complexity parameter - treewidth. We show a $\frac{k+3}{2}$ upper bound on the treewidth of level-$k$ phylogenetic networks and an improved $(1/3 + δ) k$ upper bound for large $k$. These bounds imply that many computational problems on phylogenetic networks, such as the small parsimony problem or some variants of phylogenetic diversity maximization, are polynomial-time solvable on level-$k$ networks with constant $k$. Our first bound is applicable to any $k$, and it allows us to construct an explicit tree decomposition of width $\frac{k+3}{2}$ that can be used to analyze phylogenetic networks generated by tools like SNAQ that guarantee bounded network level. Finally, we show a $k/13$ lower bound on the maximum treewidth among level-$k$ phylogenetic networks for large enough $k$ based on expander graphs.
Closing the complexity gap of the double distance problem
Genome rearrangement has been an active area of research in computational comparative genomics for the last three decades. While initially mostly an interesting algorithmic endeavor, now the practical application by applying rearrangement distance methods and more advanced phylogenetic tasks is becoming common practice, given the availability of many completely sequenced genomes. Several genome rearrangement models have been developed over time, sometimes with surprising computational properties. A prominent example is the fact that computing the reversal distance of two signed permutations is possible in linear time, while for two unsigned permutations it is NP-hard. Therefore one has always to be careful about the precise problem formulation and complexity analysis of rearrangement problems in order not to be fooled. The double distance is the minimum number of genomic rearrangements between a singular and a duplicated genome that, in addition to rearrangements, are separated by a whole genome duplication. At the same time it allows to assign the genes of the duplicated genome to the two paralogous chromosome copies that existed right after the duplication event. Computing the double distance is another example of a tricky hardness landscape: If the distance measure underlying the double distance is the simple breakpoint distance, the problem can be solved in linear time, while with the more elaborate DCJ distance it is NP-hard. Indeed, there is a family of distance measures, parameterized by an even number k, between the breakpoint distance (k=2) and the DCJ distance (k=\infty). Little was known about the hardness border between these extremes; the problem complexity was known only for k=4 and k=6. In this paper, we close the gap, providing a full picture of the hardness landscape when computing the double distance.
2024-11-01 v3
Simplifying and Characterizing DAGs and Phylogenetic Networks via Least Common Ancestor Constraints
Rooted phylogenetic networks, or more generally, directed acyclic graphs (DAGs), are widely used to model species or gene relationships that traditional rooted trees cannot fully capture, especially in the presence of reticulate processes or horizontal gene transfers. Such networks or DAGs are typically inferred from observable data (e.g. genomic sequences of extant species), providing only an estimate of the true evolutionary history. However, these inferred DAGs are often complex and difficult to interpret. In particular, many contain vertices that do not serve as least common ancestors (LCAs) for any subset of the underlying genes or species, thus may lack direct support from the observable data. In contrast, LCA vertices are witnessed by historical traces justifying their existence and thus represent ancestral states substantiated by the data. To reduce unnecessary complexity and eliminate unsupported vertices, we aim to simplify a DAG to retain only LCA vertices while preserving essential evolutionary information. In this paper, we characterize $\mathrm{LCA}$-relevant and $\mathrm{lca}$-relevant DAGs, defined as those in which every vertex serves as an LCA (or unique LCA) for some subset of taxa. We introduce methods to identify LCAs in DAGs and efficiently transform any DAG into an $\mathrm{LCA}$-relevant or $\mathrm{lca}$-relevant one while preserving key structural properties of the original DAG or network. This transformation is achieved using a simple operator ``$\ominus$'' that mimics vertex suppression.
On the lumpability of tree-valued Markov chains
Phylogenetic trees constitute an interesting class of objects for stochastic processes due to the non-standard nature of the space they inhabit. In particular, many statistical applications require the construction of Markov processes on the space of trees, whose cardinality grows superexponentially with the number of leaves considered. We investigate whether certain lower-dimensional projections of tree space preserve the Markov property in tree-valued Markov processes. We study exact lumpability of tree shapes and $\varepsilon$-lumpability of clades, exploiting the combinatorial structure of the SPR graph to obtain bounds on the lumping error under the random walk and Metropolis-Hastings processes. Finally, we show how to use these results in empirical investigation, leveraging exact and $\varepsilon$-lumpability to improve Monte Carlo estimation of tree-related quantities.
2024-10-17 v2
Mathematically tractable models of random phylogenetic networks: an overview of some recent developments
Models of random phylogenetic networks have been used since the inception of the field, but the introduction and rigorous study of mathematically tractable models is a much more recent topic that has gained momentum in the last 5 years. This manuscript discusses some recent developments in the field through a selection of examples. The emphasis is on the techniques rather than on the results themselves, and on probabilistic tools rather than on combinatorial ones.
2024-10-17 v2
Tropicalizing binary geometries
The type A cluster configuration space, commonly known as $\mathcal M_{0,n}$, is the very affine part of the binary geometry associated with the associahedron. The tropicalization of $\mathcal M_{0,n}$ can be realized as the space of phylogenetic trees and its signed tropicalizations as the dual-associahedron subfans. We give a concise overview of this construction and propose an extension to type C. The type C cluster configuration space $\mathcal M_{\mathrm C_l}$ arises from the binary geometry associated with the cyclohedron. We define a space of axially symmetric phylogenetic trees containing many dual-associahedron and dual-cyclohedron subfans. We conjecturally realize the tropicalization of $\mathcal M_{\mathrm C_l}$ as the defined space and its signed tropicalizations as the aforementioned subfans.
Tree height and the asymptotic mean of the Colijn-Plazzotta rank of unlabeled binary rooted trees
The Colijn--Plazzotta ranking is a bijective encoding of the unlabeled binary rooted trees with positive integers. We show that the rank $f(t)$ of a tree $t$ is closely related to its height $h$, the length of the longest path from a leaf to the root. We consider the rank $f(τ_n)$ of a random $n$-leaf tree $τ_n$ under each of three models: (i) uniformly random unlabeled unordered binary rooted trees, or unlabeled topologies; (ii) uniformly random leaf-labeled binary trees, or labeled topologies under the uniform model; and (iii) random binary search trees, or labeled topologies under the Yule--Harding model. Relying on the close relationship between tree rank and tree height, we obtain results concerning the asymptotic properties of $\log \log f(τ_n)$. In particular, we find $\mathbb{E} \{\log_2 \log f(τ_n)\} \sim 2 \sqrt{πn}$ for uniformly random unlabeled ordered binary rooted trees and uniformly random leaf-labeled binary trees, and for a constant $α\approx 4.31107$, $\mathbb{E}\{\log_2 \log f(τ_n)\} \sim α\log n $ for leaf-labeled binary trees under the Yule--Harding model. We show that the mean of $f(τ_n)$ itself under the three models is largely determined by the rank $c_{n-1}$ of the highest-ranked tree -- the caterpillar -- obtaining an asymptotic relationship with $π_n c_{n-1}$, where $π_n$ is a model-specific function of $n$. The results resolve open problems, providing a new class of results on an encoding useful in mathematical phylogenetics.
2024-08-30
Characterising rooted and unrooted tree-child networks
Rooted phylogenetic networks are used by biologists to infer and represent complex evolutionary relationships between species that cannot be accurately explained by a phylogenetic tree. Tree-child networks are a particular class of rooted phylogenetic networks that has been extensively investigated in recent years. In this paper, we give a novel characterisation of a tree-child network $\mathcal{R}$ in terms of cherry-picking sequences that are sequences on the leaves of $\mathcal{R}$ and reduce it to a single vertex by repeatedly applying one of two reductions to its leaves. We show that our characterisation extends to unrooted tree-child networks which are mostly unexplored in the literature and, in turn, also offers a new approach to settling the computational complexity of deciding if an unrooted phylogenetic network can be oriented as a rooted tree-child network.
2024-08-28 v3
Combinatorial invariants of finite metric spaces and the Wasserstein arrangement
In 2010, Vershik proposed a new combinatorial invariant of metric spaces given by a class of polytopes that arise in the theory of optimal transport and are called ``Wasserstein polytopes'' or ``Kantorovich-Rubinstein polytopes'' in the literature. Answering a question posed by Vershik, we describe the stratification of the metric cone induced by the combinatorial type of these polytopes through a hyperplane arrangement. Moreover, we study its relationships with the stratification by combinatorial type of the injective hull (i.e., the tight span) and, in particular, with certain types of metrics arising in phylogenetic analysis. We also compute enumerative invariants in the case of metrics on up to six points.
2024-08-26
Bounding the number of reticulation events for displaying multiple trees in a phylogenetic network
Reconstructing a parsimonious phylogenetic network that displays multiple phylogenetic trees is an important problem in theory of phylogenetics, where the complexity of the inferred networks is measured by reticulation numbers. The reticulation number for a set of trees is defined as the minimum number of reticulations in a phylogenetic network that displays those trees. A mathematical problem is bounding the reticulation number for multiple trees over a fixed number of taxa. While this problem has been extensively studied for two trees, much less is known about the upper bounds on the reticulation numbers for three or more arbitrary trees. In this paper, we present a few non-trivial upper bounds on reticulation numbers for three or more trees.
When are quarnets sufficient to reconstruct semi-directed phylogenetic networks?
Phylogenetic networks are graphs that are used to represent evolutionary relationships between different taxa. They generalize phylogenetic trees since for example, unlike trees, they permit lineages to combine. Recently, there has been rising interest in semi-directed phylogenetic networks, which are mixed graphs in which certain lineage combination events are represented by directed edges coming together, whereas the remaining edges are left undirected. One reason to consider such networks is that it can be difficult to root a network using real data. In this paper, we consider the problem of when a semi-directed phylogenetic network is defined or encoded by the smaller networks that it induces on the 4-leaf subsets of its leaf set. These smaller networks are called quarnets. We prove that semi-directed binary level-2 phylogenetic networks are encoded by their quarnets, but that this is not the case for level-3. In addition, we prove that the so-called blob tree of a semi-directed binary network, a tree that give the coarse-grained structure of the network, is always encoded by the quarnets of the network. These results are relevant for proving the statistical consistency of programs that are currently being developed for reconstructing phylogenetic networks from practical data, such as the recently developed Squirrel software tool.
A Wild Sheep Chase Through an Orchard
Orchards are a biologically relevant class of phylogenetic networks as they can describe treelike evolutionary histories augmented with horizontal transfer events. Moreover, the class has attractive mathematical characterizations that can be exploited algorithmically. On the other hand, undirected orchard networks have hardly been studied yet. Here, we prove that deciding whether an undirected, binary phylogenetic network is an orchard -- or equivalently, whether it has an orientation that makes it a rooted orchard -- is NP-hard. For this, we introduce a new characterization of undirected orchards which could be useful for proving positive results.
2024-08-13 v2
A complete characterization of pairs of binary phylogenetic trees with identical $A_k$-alignments
Phylogenetic trees play a key role in the reconstruction of evolutionary relationships. Typically, they are derived from aligned sequence data (like DNA, RNA, or proteins) by using optimization criteria like, e.g., maximum parsimony (MP). It is believed that the latter is able to reconstruct the \enquote{true} tree, i.e., the tree that generated the data, whenever the number of substitutions required to explain the data with that tree is relatively small compared to the size of the tree (measured in the number $n$ of leaves of the tree, which represent the species under investigation). However, reconstructing the correct tree from any alignment first and foremost requires the given alignment to perform differently on the \enquote{correct} tree than on others. A special type of alignments, namely so-called $A_k$-alignments, has gained considerable interest in recent literature. These alignments consist of all binary characters (\enquote{sites}) which require precisely $k$ substitutions on a given tree. It has been found that whenever $k$ is small enough (in comparison to $n$), $A_k$-alignments uniquely characterize the trees that generated them. However, recent literature has left a significant gap between $n\leq 2k+2$ -- namely the cases in which no such characterization is possible -- and $n\geq 4k$ -- namely the cases in which this characterization works. It is the main aim of the present manuscript to close this gap, i.e., to present a full characterization of all pairs of trees that share the same $A_k$-alignment. In particular, we show that indeed every binary phylogenetic tree with $n$ leaves is uniquely defined by its $A_k$-alignments if $n\geq 2k+3$. By closing said gap, we also ensure that our result is optimal.
2024-08-06 v4
Compactifications of phylogenetic systems and species of electrical networks
We describe new spaces and maps. Our graphical map is a visual and numerical correspondence between spaces of circular electrical networks and circular planar split systems. When restricted to the planar circular electrical case, this graphical map finds the split system uniquely associated with the Kalmanson resistance distance of the dual network, matching the induced split system familiar from phylogenetics. This correspondence is extended to compactifications of the respective spaces, taking cactus networks to the cactus split systems defined herein. The graphical map preserves both network components and cactus structure, allowing an elegant enumeration of induced phylogenetic split systems via combinatorial species. We introduce the global spaces of circular planar electrical networks and circular split systems. These new spaces are also CW complexes, but the 0-cells of each are counted by the Bell numbers as opposed to the Catalan numbers. As species, the two sorts of global cacti are seen to be compositions in complementary ways.
The $B_2$ index of galled trees
In recent years, there has been an effort to extend the classical notion of phylogenetic balance, originally defined in the context of trees, to networks. One of the most natural ways to do this is with the so-called $B_2$ index. In this paper, we study the $B_2$ index for a prominent class of phylogenetic networks: galled trees. We show that the $B_2$ index of a uniform leaf-labeled galled tree converges in distribution as the network becomes large. We characterize the corresponding limiting distribution, and show that its expected value is 2.707911858984... This is the first time that a balance index has been studied to this level of detail for a random phylogenetic network. One specificity of this work is that we use two different and independent approaches, each with its advantages: analytic combinatorics, and local limits. The analytic combinatorics approach is more direct, as it relies on standard tools; but it involves slightly more complex calculations. Because it has not previously been used to study such questions, the local limit approach requires developing an extensive framework beforehand; however, this framework is interesting in itself and can be used to tackle other similar problems.
2024-07-18 v2
Sackin Indices for Labeled and Unlabeled Classes of Galled Trees
The Sackin index is an important measure for the balance of phylogenetic trees. We investigate two extensions of the Sackin index to the class of galled trees and two of its subclasses (simplex galled trees and normal galled trees) where we consider both labeled and unlabeled galled trees. In all cases, we show that the mean of the Sackin index for a network which is uniformly sampled from its class is asymptotic to $μn^{3/2}$ for an explicit constant $μ$. In addition, we show that the scaled Sackin index convergences weakly and with all its moments to the Airy distribution.
2024-07-09
When is a set of phylogenetic trees displayed by a normal network?
A normal network is uniquely determined by the set of phylogenetic trees that it displays. Given a set $\mathcal{P}$ of rooted binary phylogenetic trees, this paper presents a polynomial-time algorithm that reconstructs the unique binary normal network whose set of displayed binary trees is $\mathcal{P}$, if such a network exists. Additionally, we show that any two rooted phylogenetic trees can be displayed by a normal network and show that this result does not extend to more than two trees. This is in contrast to tree-child networks where it has been previously shown that any collection of rooted phylogenetic trees can be displayed by a tree-child network. Lastly, we introduce a type of cherry-picking sequence that characterises when a collection $\mathcal{P}$ of rooted phylogenetic trees can be displayed by a normal network and, further, characterise the minimum number of reticulations needed over all normal networks that display $\mathcal{P}$. We then exploit these sequences to show that, for all $n\ge 3$, there exist two rooted binary phylogenetic trees on $n$ leaves that can be displayed by a tree-child network with a single reticulation, but cannot be displayed by a normal network with less than $n-2$ reticulations.
Network Representation and Modular Decomposition of Combinatorial Structures: A Galled-Tree Perspective
In phylogenetics, reconstructing rooted trees from distances between taxa is a common task. Böcker and Dress generalized this concept by introducing symbolic dated maps $δ:X \times X \to Υ$, where distances are replaced by symbols, and showed that there is a one-to-one correspondence between symbolic ultrametrics and labeled rooted phylogenetic trees. Many combinatorial structures fall under the umbrella of symbolic dated maps, such as 2-dissimilarities, symmetric labeled 2-structures, or edge-colored complete graphs, and are here referred to as strudigrams. Strudigrams have a unique decomposition into non-overlapping modules, which can be represented by a modular decomposition tree (MDT). In the absence of prime modules, strudigrams are equivalent to symbolic ultrametrics, and the MDT fully captures the relationships $δ(x,y)$ between pairs of vertices $x,y$ in $X$ through the label of their least common ancestor in the MDT. However, in the presence of prime vertices, this information is generally hidden. To provide this missing structural information, we aim to locally replace the prime vertices in the MDT to obtain networks that capture full information about the strudigrams. While starting with the general framework of prime-vertex replacement networks, we then focus on a specific type of such networks obtained by replacing prime vertices with so-called galls, resulting in labeled galled-trees. We introduce the concept of galled-tree explainable (GATEX) strudigrams, provide their characterization, and demonstrate that recognizing these structures and reconstructing the labeled networks that explain them can be achieved in polynomial time.
2024-05-30
Bounding the softwired parsimony score of a phylogenetic network
In comparison to phylogenetic trees, phylogenetic networks are more suitable to represent complex evolutionary histories of species whose past includes reticulation such as hybridisation or lateral gene transfer. However, the reconstruction of phylogenetic networks remains challenging and computationally expensive due to their intricate structural properties. For example, the small parsimony problem that is solvable in polynomial time for phylogenetic trees, becomes NP-hard on phylogenetic networks under softwired and parental parsimony, even for a single binary character and structurally constrained networks. To calculate the parsimony score of a phylogenetic network $N$, these two parsimony notions consider different exponential-size sets of phylogenetic trees that can be extracted from $N$ and infer the minimum parsimony score over all trees in the set. In this paper, we ask: What is the maximum difference between the parsimony score of any phylogenetic tree that is contained in the set of considered trees and a phylogenetic tree whose parsimony score equates to the parsimony score of $N$? Given a gap-free sequence alignment of multi-state characters and a rooted binary level-$k$ phylogenetic network, we use the novel concept of an informative blob to show that this difference is bounded by $k+1$ times the softwired parsimony score of $N$. In particular, the difference is independent of the alignment length and the number of character states. We show that an analogous bound can be obtained for the softwired parsimony score of semi-directed networks, while under parental parsimony on the other hand, such a bound does not hold.
2024-05-30
Phylogenetic degrees for Jukes-Cantor model
Jukes-Cantor model is one of the most meaningful statistical models from a biological perspective. We are interested in computing the algebraic degrees for phylogenetic varieties, which we call phylogenetic degrees, associated to the Jukes-Cantor model and any tree. As these varieties are toric, their geometry is hidden in the associated polytopes. For this reason, we provide two different combinatorial approaches to compute the volume for these polytopes.