Papers by Michael Hendriksen
14 paper(s) by this author
· All BibTeX
Tree Buckets and the Reconstruction of Pairs of Phylogenetic Trees
Phylogenetic trees are used in evolutionary biology to represent the evolutionary history of a collection of taxa. As we have incomplete information about any evolutionary history, recovering trees from partial information is a focus of phylogenetic combinatorics. However, in some cases the available data does not describe a single phylogenetic tree. We consider recovery of pairs of phylogenetic trees from their combined subtrees with $k$ leaves, which we call a $k$-bucket. We establish the exact cases in which these pairs of trees are recoverable from their subtrees with a single leaf removed, both when just considering the structure of the trees, and when additionally considering the set of taxa on the leaves. We also consider recovery of pairs of trees with labelled leaves from their rooted triples, and establish that they are recoverable up to a sequence of subtree swaps.
Enumerating monophyletic characters in mathematical phylogenetics
Grouping species according to their phylogenetic relationships often results in different groups than grouping them according to their shared traits. Monophyletic groups play an important role in this regard, as they are groups of species sharing the same trait and being uniquely defined by a joint phylogenetic subtree. This immediately leads to the question of how to identify possible monophyletic groups in characters, which assign each present-day species a certain trait and which are typically used for phylogenetic tree reconstruction.
In our manuscript, we provide a general formula to quantify how many different characters are monophyletic on any given tree and provide simple formulae for binary characters and for certain tree shapes. We also investigate relations between monophyly and the well-known phylogenetic tree reconstruction criterion maximum parsimony by providing a linear-time algorithm which determines the parsimony score together with the monophyly type of a character on a tree.
Classes of phylogenetic networks that are robust to root placement
Standard phylogenetic reconstruction techniques often yield unrooted phylogenetic networks; these are subsequently rooted to infer evolutionary history. A common problem in this process is to determine the structural classes to which the resulting network will belong. In this paper, we investigate unrooted networks in which the choice of any root results in a valid rooted phylogenetic network, a property we define as {\em robustly orientable}. We then establish a strict structural condition for this class, specifically, that an unrooted network is robustly orientable if and only if it contains no sink components. We also show that if an unrooted network is level-$2$ or less, or if it is tree-based, then it is robustly orientable. Furthermore, we define an unrooted network to be {\em robustly class $\mathcal C$} if the choice of any root results in a network belonging to class $\mathcal C$. We demonstrate that an unrooted network is robustly tree-child or robustly stack-free if and only if it is level-$1$ or less. Finally, we show that a phylogenetic network is robustly normal if and only if it is a phylogenetic tree.
Counting spinal phylogenetic networks
Phylogenetic networks are an important way to represent evolutionary histories that involve reticulations such as hybridization or horizontal gene transfer, yet fundamental questions such as how many networks there are that satisfy certain properties are very difficult. A new way to encode a large class of networks, using expanding covers, may provide a way to approach such problems. Expanding covers encode a large class of phylogenetic networks, called labellable networks. This class does not include all networks, but does include many familiar classes, including orchard, normal, tree-child and tree-sibling networks. As expanding covers are a combinatorial structure, it is possible that they can be used as a tool for counting such classes for a fixed number of leaves and reticulations, for which, in many cases, a closed formula has not yet been found. More recently, a new class of networks was introduced, called spinal networks, which are analogous to caterpillar trees for phylogenetic trees and can be fully described using covers. In the present article, we describe a method for counting networks that are both spinal and belong to some more familiar class, with the hope that these form a base case from which to attack the more general classes.
A polynomial invariant for a new class of phylogenetic networks
Published
• View Publication
• BIB
Invariants for complicated objects such as those arising in phylogenetics, whether they are invariants as matrices, polynomials, or other mathematical structures, are important tools for distinguishing and working with such objects. In this paper, we generalize a complete polynomial invariant on trees to a class of phylogenetic networks called separable networks, which will include orchard networks. Networks are becoming increasingly important for their ability to represent reticulation events, such as hybridization, in evolutionary history. We provide a function from the space of internally multi-labelled phylogenetic networks, a more generic graph structure than phylogenetic networks where the reticulations are also labelled, to a polynomial ring. We prove that the separability condition allows us to characterize, via the polynomial, the phylogenetic networks with the same number of leaves and same number of reticulations by considering their internally labelled versions. While the invariant for trees is a polynomial in Z[x_1,..., x_n,y] where n is the number of leaves, the invariant for internally multi-labelled phylogenetic networks is an element of Z[x_1,..., x_n,lambda_1,...,lambda_r,y], where r is the number of reticulations in the network. When the networks are considered without leaf labels the number of variables reduces to r+2.
On the enumeration of integer tetrahedra
Published
• View Publication
• BIB
We consider the problem of enumerating integer tetrahedra of fixed perimeter (sum of side-lengths) and/or diameter (maximum side-length), up to congruence. As we will see, this problem is considerably more difficult than the corresponding problem for triangles, which has long been solved. We expect there are no closed-form solutions to the tetrahedron enumeration problems, but we explore the extent to which they can be approached via classical methods, such as orbit enumeration. We also discuss algorithms for computing the numbers, and present several tables and figures that can be used to visualise the data. Several intriguing patterns seem to emerge, leading to a number of natural conjectures. The central conjecture is that the number of integer tetrahedra of perimeter $n$, up to congruence, is asymptotic to $n^5/C$ for some constant $C\approx 229000$.
A survey of the monotonicity and non-contradiction of consensus methods and supertree methods
Published
• View Publication
• BIB
In a recent study, Bryant, Francis and Steel investigated the concept of \enquote{future-proofing} consensus methods in phylogenetics. That is, they investigated if such methods can be robust against the introduction of additional data like added trees or new species. In the present manuscript, we analyze consensus methods under a different aspect of introducing new data, namely concerning the discovery of new clades. In evolutionary biology, often formerly unresolved clades get resolved by refined reconstruction methods or new genetic data analyses. In our manuscript we investigate which properties of consensus methods can guarantee that such new insights do not disagree with previously found consensus trees, but merely refine them, a property termed \emph{monotonicity}. Along the lines of analyzing monotonicity, we also study two {established} supertree methods, namely Matrix Representation with Parsimony (MRP) and Matrix Representation with Compatibility (MRC), which have also been suggested as consensus methods in the literature. While we (just like Bryant, Francis and Steel in their recent study) unfortunately have to conclude some negative answers concerning general consensus methods, we also state some relevant and positive results concerning the majority rule ($\mathtt{MR}$) and strict consensus methods, which are amongst the most frequently used consensus methods. Moreover, we show that there exist infinitely many consensus methods which are monotonic and have some other desirable properties.
\textbf{Keywords:} consensus tree, phylogenetics, majority rule, tree refinement, matrix representation with parsimony
\textbf{MSC:} C92B05, 05C05
Phylosymmetric algebras: mathematical properties of a new tool in phylogenetics
Published
• View Publication
• BIB
In phylogenetics it is of interest for rate matrix sets to satisfy closure under matrix multiplication as this makes finding the set of corresponding transition matrices possible without having to compute matrix exponentials. It is also advantageous to have a small number of free parameters as this, in applications, will result in a reduction of computation time. We explore a method of building a rate matrix set from a rooted tree structure by assigning rates to internal tree nodes and states to the leaves, then defining the rate of change between two states as the rate assigned to the most recent common ancestor of those two states. We investigate the properties of these matrix sets from both a linear algebra and a graph theory perspective and show that any rate matrix set generated this way is closed under matrix multiplication. The consequences of setting two rates assigned to internal tree nodes to be equal are then considered. This methodology could be used to develop parameterised models of amino acid substitution which have a small number of parameters but convey biological meaning.
On the comparison of incompatibility of split systems across different numbers of taxa
Published
• View Publication
• BIB
The concept of $k$-compatibility measures how many phylogenetic trees it would take to display all splits in a given set. A set of trees that display every single possible split is termed a \textit{universal tree set}. In this note, we find $A(n)$, the minimal size of a universal tree set for $n$ taxa. By normalising the $k$-compatibility using $A(n)$, one can then compare incompatibility of split systems across different taxa sizes. We demonstrate this application by comparing two SplitsTree networks of different sizes derived from archaeal genomes.
Tree-metrizable HGT networks
Published
• View Publication
• BIB
Phylogenetic trees are often constructed by using a metric on the set of taxa that label the leaves of the tree. While there are a number of methods for constructing a tree using a given metric, such trees will only display the metric if it satisfies the so-called "four point condition", established by Buneman in 1971. While this condition guarantees that a unique tree will display the metric, meaning that the distance between any two leaves can be found by adding the distances on arcs in the path between the leaves, it doesn't exclude the possibility that a phylogenetic network might also display the metric. This possibility was recently pointed out and "tree-metrized" networks --- that display a tree metric --- with a single reticulation were characterized. In this paper, we show that in the case of HGT (horizontal gene transfer) networks, in fact there are tree-metrized networks containing many reticulations.
A partial order and cluster-similarity metric on rooted phylogenetic trees
Metrics on rooted phylogenetic trees are integral to a number of areas of phylogenetic analysis. Cluster-similarity metrics have recently been introduced in order to limit skew in the distribution of distances, and to ensure that trees in the neighbourhood of each other have similar hierarchies. In the present paper we introduce a new cluster-similarity metric on rooted phylogenetic tree space that has an associated local operation, allowing for easy calculation of neighbourhoods, a trait that is desirable for MCMC calculations. The metric is defined by the distance on the Hasse diagram induced by a partial order on the set of rooted phylogenetic trees, itself based on the notion of a hierarchy-preserving map between trees. The partial order we introduce is a refinement of the well-known refinement order on hierarchies. Both the partial order and the hierarchy-preserving maps may also be of independent interest.
Lattice consensus: A partial order on phylogenetic trees that induces an associatively stable consensus method
There is a long tradition of the axiomatic study of consensus methods in phylogenetics that satisfy certain desirable properties. One recently-introduced property is associative stability, which is desirable because it confers a computational advantage, in that the consensus method only needs to be computed "pairwise". In this paper, we introduce a phylogenetic consensus method that satisfies this property, in addition to being "regular". The method is based on the introduction of a partial order on the set of rooted phylogenetic trees, itself based on the notion of a hierarchy-preserving map between trees. This partial order may be of independent interest. We call the method "lattice consensus", because it takes the unique maximal element in a lattice of trees defined by the partial order. Aside from being associatively stable, lattice consensus also satisfies the property of being Pareto on rooted triples, answering in the affirmative a question of Bryant et al (2017). We conclude the paper with an answer to another question of Bryant et al, showing that there is no regular extension stable consensus method for binary trees.
Tree-Based Unrooted Nonbinary Phylogenetic Networks
Published
• View Publication
• BIB
Phylogenetic networks are a generalisation of phylogenetic trees that allow for more complex evolutionary histories that include hybridisation-like processes. It is of considerable interest whether a network can be considered `tree-like' or not, which lead to the introduction of \textit{tree-based} networks in the rooted, binary context. Tree-based networks are those networks which can be constructed by adding additional edges into a given phylogenetic tree, called the \textit{base tree}. Previous extensions have considered extending to the binary, unrooted case and the nonbinary, rooted case. We extend tree-based networks to the context of unrooted, nonbinary networks in three ways, depending on the types of additional edges that are permitted. A phylogenetic network in which every embedded tree is a base tree is termed a \textit{fully tree-based} network. We also extend this concept to unrooted, nonbinary phylogenetic networks and classify the resulting networks. We also derive some results on the colourability of tree-based networks, which can be useful to determine whether a network is tree-based.
Behind Every Great Tree is a Great (Phylogenetic) Network
In Francis and Steel (2015), it was shown that there exists non-trivial networks on $4$ leaves upon which the distance metric affords a metric on a tree which is not the base tree of the network. In this paper we extend this result in two directions. We show that for any tree $T$ there exists a family of non-trivial HGT networks $N$ for which the distance metric $d_N$ affords a metric on $T$. We additionally provide a class of networks on any number of leaves upon which the distance metric affords a metric on a tree which is not the base tree of the network.
The family of networks are all "floating" networks, a subclass of a novel family of networks introduced in this paper, and referred to as "versatile" networks. Versatile networks are then characterised.
Additionally, we find a lower bound for the number of `useful' HGT arcs in such networks, in a sense explained in the paper. This lower bound is equal to the number of HGT arcs required for each floating network in the main results, and thus our networks are minimal in this sense.