phylogenetic
427 papers tagged with this keyword
Representing Partitions on Trees
Published
• View Publication
• BIB
In evolutionary biology, biologists often face the problem of constructing a phylogenetic tree on a set $X$ of species from a multiset $Π$ of partitions corresponding to various attributes of these species. One approach that is used to solve this problem is to try instead to associate a tree (or even a network) to the multiset $Σ_Π$ consisting of all those bipartitions $\{A,X-A\}$ with $A$ a part of some partition in $Π$. The rational behind this approach is that a phylogenetic tree with leaf set $X$ can be uniquely represented by the set of bipartitions of $X$ induced by its edges. Motivated by these considerations, given a multiset $Σ$ of bipartitions corresponding to a phylogenetic tree on $X$, in this paper we introduce and study the set $P(Σ)$ consisting of those multisets of partitions $Π$ of $X$ with $Σ_Π=Σ$. More specifically, we characterize when $P(Σ)$ is non-empty, and also identify some partitions in $P(Σ)$ that are of maximum and minimum size. We also show that it is NP-complete to decide when $P(Σ)$ is non-empty in case $Σ$ is an arbitrary multiset of bipartitions of $X$. Ultimately, we hope that by gaining a better understanding of the mapping that takes an arbitrary partition system $Π$ to the multiset $Σ_Π$, we will obtain new insights into the use of median networks and, more generally, split-networks to visualize sets of partitions.
On the Maximum Parsimony distance between phylogenetic trees
Published
• View Publication
• BIB
Within the field of phylogenetics there is great interest in distance measures to quantify the dissimilarity of two trees. Here, based on an idea of Bruen and Bryant, we propose and analyze a new distance measure: the Maximum Parsimony (MP) distance. This is based on the difference of the parsimony scores of a single character on both trees under consideration, and the goal is to find the character which maximizes this difference. In this article we show that this new distance is a metric and provides a lower bound to the well-known Subtree Prune and Regraft (SPR) distance. We also show that to compute the MP distance it is sufficient to consider only characters that are convex on one of the trees, and prove several additional structural properties of the distance. On the complexity side, we prove that calculating the MP distance is in general NP-hard, and identify an interesting island of tractability in which the distance can be calculated in polynomial time.
A short note on exponential-time algorithms for hybridization number
In this short note we prove that, given two (not necessarily binary) rooted phylogenetic trees T_1, T_2 on the same set of taxa X, where |X|=n, the hybridization number of T_1 and T_2 can be computed in time O^{*}(2^n) i.e. O(2^{n} poly(n)). The result also means that a Maximum Acyclic Agreement Forest (MAAF) can be computed within the same time bound.
Tropical Grassmannian and Tropical Linear Varieties from phylogenetic trees
In this paper we study tropicalization of Grassmannian and linear varieties. In particular, we study the tropical linear spaces cor- responding to the phylogenetic trees. We prove that corresponding to each subtree of the phylogenetic tree there is a point on the tropical grassmannian. We deduce a necessary and sufficient condition for it to be on the facet of the tropical linear space.
Polyhedral Covers of Tree Space
Published in SIAM Journal of Discrete Mathematics 28 (2014) 1508 - 1514
• View Publication
• BIB
The phylogenetic tree space, introduced by Billera, Holmes, and Vogtmann, is a cone over a simplicial complex. In this short article, we construct this complex from local gluings of classical polytopes, the associahedron and the permutohedron. Its homotopy is also reinterpreted and calculated based on polytope data.
Reconstructing a phylogenetic level-1 network from quartets
Published
• View Publication
• BIB
We describe a method that will reconstruct an unrooted binary phylogenetic level-1 network on n taxa from the set of all quartets containing a certain fixed taxon, in O(n^3) time. We also present a more general method which can handle more diverse quartet data, but which takes O(n^6) time. Both methods proceed by solving a certain system of linear equations over GF(2).
For a general dense quartet set (containing at least one quartet on every four taxa) our O(n^6) algorithm constructs a phylogenetic level-1 network consistent with the quartet set if such a network exists and returns an (O(n^2) sized) certificate of inconsistency otherwise. This answers a question raised by Gambette, Berry and Paul regarding the complexity of reconstructing a level-1 network from a dense quartet set.
The asymptotic enhanced negative type of finite ultrametric spaces
Published
• View Publication
• BIB
Negative type inequalities arise in the study of embedding properties of metric spaces, but they often reduce to intractable combinatorial problems. In this paper we study more quantitative versions of these inequalities involving the so-called $p$-negative type gap. In particular, we focus our attention on the class of finite ultrametric spaces which are important in areas such as phylogenetics and data mining.
Let $(X,d)$ be a given finite ultrametric space with minimum non-zero distance $α$. Then the $p$-negative type gap $Γ_{X}(p)$ of $(X,d)$ is positive for all $p \geq 0$. In this paper we compute the value of the limit \begin{eqnarray*} Γ_{X}(\infty) & = & \lim\limits_{p \rightarrow \infty} \frac{Γ_{X}(p)}{α^{p}}. \end{eqnarray*} It turns out that this value is positive and it may be given explicitly by an elegant combinatorial formula. On the basis of our calculations we are then able to characterize when $Γ_{X}(p)/ α^{p}$ is constant on $[0, \infty)$.
The determination of $Γ_{X}(\infty)$ also leads to new, asymptotically sharp, families of enhanced $p$-negative type inequalities for $(X,d)$. Indeed, suppose that $G \in (0, Γ_{X}(\infty))$. Then, for all sufficiently large $p$, we have \begin{eqnarray*} \frac{G \cdot α^{p}}{2} \left( \sum\limits_{k=1}^{n} |ζ_{k}| \right)^{2} + \sum\limits_{j,i =1}^{n} d(z_{j},z_{i})^{p} ζ_{j} ζ_{i} & \leq & 0 \end{eqnarray*} for each finite subset $\{ z_{1}, \ldots, z_{n} \} \subseteq X$ and each choice of real numbers $ζ_{1}, \ldots, ζ_{n}$ with $ζ_{1} + \cdots + ζ_{n} = 0$. We note that these results do not extend to general finite metric spaces.
A matroid associated with a phylogenetic tree
Published
• View Publication
• BIB
A (pseudo-)metric $D$ on a finite set $X$ is said to be a `tree metric' if there is a finite tree with leaf set $X$ and non-negative edge weights so that, for all $x,y \in X$, $D(x,y)$ is the path distance in the tree between $x$ and $y$. It is well known that not every metric is a tree metric. However, when some such tree exists, one can always find one whose interior edges have strictly positive edge weights and that has no vertices of degree 2, any such tree is -- up to canonical isomorphism -- uniquely determined by $D$, and one does not even need all of the distances in order to fully (re-)construct the tree's edge weights in this case. Thus, it seems of some interest to investigate which subsets of $\binom{X}{2}$ suffice to determine (`lasso') these edge weights. In this paper, we use the results of a previous paper to discuss the structure of a matroid that can be associated with an (unweighted) $X-$tree $T$ defined by the requirement that its bases are exactly the `tight edge-weight lassos' for $T$, i.e, the minimal subsets $\cl$ of $\ch$ that lasso the edge weights of $T$.
Distance-based phylogenetic methods around a polytomy
Published
• View Publication
• BIB
Distance-based phylogenetic algorithms attempt to solve the NP-hard least squares phylogeny problem by mapping an arbitrary dissimilarity map representing biological data to a tree metric. The set of all dissimilarity maps is a Euclidean space properly containing the space of all tree metrics as a polyhedral fan. Outputs of distance-based tree reconstruction algorithms such as UPGMA and Neighbor-Joining are points in the maximal cones in the fan. Tree metrics with polytomies lie at the intersections of maximal cones.
A phylogenetic algorithm divides the space of all dissimilarity maps into regions based upon which combinatorial tree is reconstructed by the algorithm. Comparison of phylogenetic methods can be done by comparing the geometry of these regions. We use polyhedral geometry to compare the local nature of the subdivisions induced by least squares phylogeny, UPGMA, and Neighbor-Joining. Our results suggest that in some circumstances, UPGMA and Neighbor-Joining poorly match least squares phylogeny when the true tree has a polytomy.
Low degree minimal generators of phylogenetic semigroups
Published
• View Publication
• BIB
The phylogenetic semigroup on a graph generalizes the Jukes-Cantor binary model on a tree. Minimal generating sets of phylogenetic semigroups have been described for trivalent trees by Buczyńska and Wiśniewski, and for trivalent graphs with first Betti number 1 by Buczyńska. We characterize degree two minimal generators of the phylogenetic semigroup on any trivalent graph. Moreover, for any graph with first Betti number 1 and for any trivalent graph with first Betti number 2 we describe the minimal generating set of its phylogenetic semigroup.
On Computing the Maximum Parsimony Score of a Phylogenetic Network
Published
• View Publication
• BIB
Phylogenetic networks are used to display the relationship of different species whose evolution is not treelike, which is the case, for instance, in the presence of hybridization events or horizontal gene transfers. Tree inference methods such as Maximum Parsimony need to be modified in order to be applicable to networks. In this paper, we discuss two different definitions of Maximum Parsimony on networks, "hardwired" and "softwired", and examine the complexity of computing them given a network topology and a character. By exploiting a link with the problem Multicut, we show that computing the hardwired parsimony score for 2-state characters is polynomial-time solvable, while for characters with more states this problem becomes NP-hard but is still approximable and fixed parameter tractable in the parsimony score. On the other hand we show that, for the softwired definition, obtaining even weak approximation guarantees is already difficult for binary characters and restricted network topologies, and fixed-parameter tractable algorithms in the parsimony score are unlikely. On the positive side we show that computing the softwired parsimony score is fixed-parameter tractable in the level of the network, a natural parameter describing how tangled reticulate activity is in the network. Finally, we show that both the hardwired and softwired parsimony score can be computed efficiently using Integer Linear Programming. The software has been made freely available.
Polyhedral computational geometry for averaging metric phylogenetic trees
Published
• View Publication
• BIB
This paper investigates the computational geometry relevant to calculations of the Frechet mean and variance for probability distributions on the phylogenetic tree space of Billera, Holmes and Vogtmann, using the theory of probability measures on spaces of nonpositive curvature developed by Sturm. We show that the combinatorics of geodesics with a specified fixed endpoint in tree space are determined by the location of the varying endpoint in a certain polyhedral subdivision of tree space. The variance function associated to a finite subset of tree space has a fixed $C^\infty$ algebraic formula within each cell of the corresponding subdivision, and is continuously differentiable in the interior of each orthant of tree space. We use this subdivision to establish two iterative methods for producing sequences that converge to the Frechet mean: one based on Sturm's Law of Large Numbers, and another based on descent algorithms for finding optima of smooth functions on convex polyhedra. We present properties and biological applications of Frechet means and extend our main results to more general globally nonpositively curved spaces composed of Euclidean orthants.
Approximation algorithms for nonbinary agreement forests
Published
• View Publication
• BIB
Given two rooted phylogenetic trees on the same set of taxa X, the Maximum Agreement Forest problem (MAF) asks to find a forest that is, in a certain sense, common to both trees and has a minimum number of components. The Maximum Acyclic Agreement Forest problem (MAAF) has the additional restriction that the components of the forest cannot have conflicting ancestral relations in the input trees. There has been considerable interest in the special cases of these problems in which the input trees are required to be binary. However, in practice, phylogenetic trees are rarely binary, due to uncertainty about the precise order of speciation events. Here, we show that the general, nonbinary version of MAF has a polynomial-time 4-approximation and a fixed-parameter tractable (exact) algorithm that runs in O(4^k poly(n)) time, where n = |X| and k is the number of components of the agreement forest minus one. Moreover, we show that a c-approximation algorithm for nonbinary MAF and a d-approximation algorithm for the classical problem Directed Feedback Vertex Set (DFVS) can be combined to yield a d(c+3)-approximation for nonbinary MAAF. The algorithms for MAF have been implemented and made publicly available.
Trinets encode tree-child and level-2 phylogenetic networks
Published
• View Publication
• BIB
Phylogenetic networks generalize evolutionary trees, and are commonly used to represent evolutionary histories of species that undergo reticulate evolutionary processes such as hybridization, recombination and lateral gene transfer. Recently, there has been great interest in trying to develop methods to construct rooted phylogenetic networks from triplets, that is rooted trees on three species. However, although triplets determine or encode rooted phylogenetic trees, they do not in general encode rooted phylogenetic networks, which is a potential issue for any such method. Motivated by this fact, Huber and Moulton recently introduced trinets as a natural extension of rooted triplets to networks. In particular, they showed that level-1 phylogenetic networks are encoded by their trinets, and also conjectured that all "recoverable" rooted phylogenetic networks are encoded by their trinets. Here we prove that recoverable binary level-2 networks and binary tree-child networks are also encoded by their trinets. To do this we prove two decomposition theorems based on trinets which hold for all recoverable binary rooted phylogenetic networks. Our results provide some additional evidence in support of the conjecture that trinets encode all recoverable rooted phylogenetic networks, and could also lead to new approaches to construct phylogenetic networks from trinets.
Betti numbers of cut ideals of trees
Published in J. Alg. Stat., 4(1):108-117, 2013
• View Publication
• BIB
Cut ideals, introduced by Sturmfels and Sullivant, are used in phylogenetics and algebraic statistics. We study the minimal free resolutions of cut ideals of tree graphs. By employing basic methods from topological combinatorics, we obtain upper bounds for the Betti numbers of this type of ideals. These take the form of simple formulas on the number of vertices, which arise from the enumeration of induced subgraphs of certain incomparability graphs associated to the edge sets of trees.
Invariant polynomial functions on tensors under the action of a product of orthogonal groups
Published
• View Publication
• BIB
Let K be the product O(n_1) x O(n_2) x ... x O(n_r) of orthogonal groups. Let V the r-fold tensor product of defining representations of each orthogonal factor. We compute a stable formula for the dimension of the K-invariant algebra of degree d homogeneous polynomial functions on V. To accomplish this, we compute a formula for the number of matchings which commute with a fixed permutation. Finally, we provide formulas for the invariants and describe a bijection between a basis for the space of invariants and the isomorphism classes of certain r-regular graphs on d vertices, as well as a method of associating each invariant to other combinatorial settings such as phylogenetic trees.
Lassoing and corraling rooted phylogenetic trees
Published in Bulletin of Mathematical Biology: Volume 75, Issue 3 (2013), Page 444-465
• View Publication
• BIB
The construction of a dendogram on a set of individuals is a key component of a genomewide association study. However even with modern sequencing technologies the distances on the individuals required for the construction of such a structure may not always be reliable making it tempting to exclude them from an analysis. This, in turn, results in an input set for dendogram construction that consists of only partial distance information which raises the following fundamental question. For what subset of its leaf set can we reconstruct uniquely the dendogram from the distances that it induces on that subset. By formalizing a dendogram in terms of an edge-weighted, rooted phylogenetic tree on a pre-given finite set X with |X|>2 whose edge-weighting is equidistant and a set of partial distances on X in terms of a set L of 2-subsets of X, we investigate this problem in terms of when such a tree is lassoed, that is, uniquely determined by the elements in L. For this we consider four different formalizations of the idea of "uniquely determining" giving rise to four distinct types of lassos. We present characterizations for all of them in terms of the child-edge graphs of the interior vertices of such a tree. Our characterizations imply in particular that in case the tree in question is binary then all four types of lasso must coincide.
A simple fixed parameter tractable algorithm for computing the hybridization number of two (not necessarily binary) trees
Published
• View Publication
• BIB
Here we present a new fixed parameter tractable algorithm to compute the hybridization number r of two rooted, not necessarily binary phylogenetic trees on taxon set X in time (6^r.r!).poly(n)$, where n=|X|. The novelty of this approach is its use of terminals, which are maximal elements of a natural partial order on X, and several insights from the softwired clusters literature. This yields a surprisingly simple and practical bounded-search algorithm and offers an alternative perspective on the underlying combinatorial structure of the hybridization number problem.
Perfect taxon sampling and fixing taxon traceability: Introducing a class of phylogenetically decisive collections of taxon sets
Published
• View Publication
• BIB
Phylogenetically decisive collections of taxon sets have the property that if trees are chosen for each of their elements, as long as these trees are compatible, the resulting supertree is unique. This means that as long as the trees describing the phylogenetic relationships of the (input) species sets are compatible, they can only be combined into a common supertree in precisely one way. This setting is sometimes also referred to as \enquote{perfect taxon sampling}. While for rooted trees, the decision if a given set of input taxon sets is phylogenetically decisive can be made in polynomial time, the decision problem to determine whether a collection of taxon sets is phylogenetically decisive concerning \emph{unrooted} trees is unfortunately coNP-complete and therefore in practice hard to solve for large instances. This shows that recognizing such sets is often difficult. In this paper, we explain phylogenetic decisiveness and introduce a class of input taxon sets, namely so-called \emph{fixing taxon traceable} sets, which are guaranteed to be phylogenetically decisive and which can be recognized in polynomial time. Using both combinatorial approaches as well as simulations, we compare properties of fixing taxon traceability and phylogenetic decisiveness, e.g., by deriving lower and upper bounds for the number of quadruple sets (i.e., sets of 4-tuples) needed in the input set for each of these properties. In particular, we correct an erroneous lower bound concerning phylogenetic decisiveness from the literature.
We have implemented the algorithm to determine if a given collection of taxon sets is fixing taxon traceable in \textsf{R} and made our software package \verb+FixingTaxonTraceR+ publicly available.
Searching for Realizations of Finite Metric Spaces in Tight Spans
Published in Discrete Optimization 10 (2013), no. 4, 310-319
• View Publication
• BIB
An important problem that commonly arises in areas such as internet traffic-flow analysis, phylogenetics and electrical circuit design, is to find a representation of any given metric $D$ on a finite set by an edge-weighted graph, such that the total edge length of the graph is minimum over all such graphs. Such a graph is called an optimal realization and finding such realizations is known to be NP-hard. Recently Varone presented a heuristic greedy algorithm for computing optimal realizations. Here we present an alternative heuristic that exploits the relationship between realizations of the metric $D$ and its so-called tight span $T_D$. The tight span $T_D$ is a canonical polytopal complex that can be associated to $D$, and our approach explores parts of $T_D$ for realizations in a way that is similar to the classical simplex algorithm. We also provide computational results illustrating the performance of our approach for different types of metrics, including $l_1$-distances and two-decomposable metrics for which it is provably possible to find optimal realizations in their tight spans.