Papers by K. T. Huber
7 paper(s) by this author
· All BibTeX
Injective split systems
Published
• View Publication
• BIB
A split system $\mathcal S$ on a finite set $X$, $|X|\ge3$, is a set of bipartitions or splits of $X$ which contains all splits of the form $\{x,X-\{x\}\}$, $x \in X$. To any such split system $\mathcal S$ we can associate the Buneman graph $\mathcal B(\mathcal S)$ which is essentially a median graph with leaf-set $X$ that displays the splits in $\mathcal S$. In this paper, we consider properties of injective split systems, that is, split systems $\mathcal S$ with the property that $\mathrm{med}_{\mathcal B(\mathcal S)}(Y) \neq \mathrm{med}_{\mathrm B(\mathcal S)}(Y')$ for any 3-subsets $Y,Y'$ in $X$, where $\mathrm {med}_{\mathcal B(\mathcal S)}(Y)$ denotes the median in $\mathcal B(\mathcal S)$ of the three elements in $Y$ considered as leaves in $\mathcal B(\mathcal S)$. In particular, we show that for any set $X$ there always exists an injective split system on $X$, and we also give a characterization for when a split system is injective. We also consider how complex the Buneman graph $\mathcal B(\mathcal S)$ needs to become in order for a split system $\mathcal S$ on $X$ to be injective. We do this by introducing a quantity for $|X|$ which we call the injective dimension for $|X|$, as well as two related quantities, called the injective 2-split and the rooted-injective dimension. We derive some upper and lower bounds for all three of these dimensions and also prove that some of these bounds are tight. An underlying motivation for studying injective split systems is that they can be used to obtain a natural generalization of symbolic tree maps. An important consequence of our results is that any three-way symbolic map on $X$ can be represented using Buneman graphs.
Overlaid species forests
Published
• View Publication
• BIB
Introgression is an evolutionary process in which genes or other types of genetic material are introduced into a genome. It is an important evolutionary process that can, for example, play a fundamental role in speciation. Recently the concept of an overlaid species forest was introduced to represent introgression histories. Basically this approach takes a putative gene history in the form of a phylogenetic gene tree and tries to overlay this onto a forest which usually consists of a collection of lineage trees for the species of interest. The result is a network called an overlaid species forest in which genes jump or introgress between lineages. In this paper we study properties of overlaid species forests, showing that they have various connections with models for lateral gene transfer, maximum parsimony, and unfolding of phylogenetic networks. In particular, we show that a certain algorithm called OSF-B UILDER for constructing overlaid species forests is guaranteed to a produce a special type of overlaid species forest with a minimum number introgressions, as well as providing some characterizations for networks that can arise from overlaid species forests. We expect that these results will be useful in developing new methods for representing introgression histories, a growing area of interest in phylogenetics.
The polytopal structure of the tight-span of a totally split-decomposable metric
Published
• View Publication
• BIB
The tight-span of a finite metric space is a polytopal complex that has appeared in several areas of mathematics. In this paper we determine the polytopal structure of the tight-span of a totally split decomposable (finite) metric. Totally split-decomposable metrics are a generalization of tree-metrics and have importance within phylogenetics. In previous work, we showed that the cells of the tight-span of such a metric are zonotopes that are polytope isomorphic to either hypercubes or rhombic dodecahedra. Here, we extend these results and show that the tight-spanof a totally split-decomposable metric can be broken up into a canonical collection of polytopal complexes whose polytopal structures can be directly determined from the metric. This allows us to also completely determine the polytopal structure of the tight-span of a totally split-decomposable metric in a very direct way.We anticipate that our improved understanding of this structure may ultimately lead to improved techniques for phylogenetic inference.
On the challenge of reconstructing level-1 phylogenetic networks from triplets and clusters
Published
• View Publication
• BIB
Phylogenetic networks have gained prominence over the years due to their ability to represent complex non-treelike evolutionary events such as recombination or hybridization. Popular combinatorial objects used to construct them are triplet systems and cluster systems, the motivation being that any network $N$ induces a triplet system $\mathcal R(N)$ and a softwired cluster system $\mathcal S(N)$. Since in real-world studies it cannot be guaranteed that all triplets/softwired clusters induced by a network are available it is of particular interest to understand whether subsets of $\mathcal R(N)$ or $\mathcal S(N)$ allow one to uniquely reconstruct the underlying network $N$. Here we show that even within the highly restricted yet biologically interesting space of level-1 phylogenetic networks it is not always possible to uniquely reconstruct a level-1 network $N$ even when all triplets in $\mathcal R(N)$ or all clusters in $\mathcal S(N)$ are available. On the positive side, we introduce a reasonably large subclass of level-1 networks the members of which are uniquely determined by their induced triplet/softwired cluster systems. Along the way, we also establish various enumerative results, both positive and negative, including results which show that certain special subclasses of level-1 networks $N$ can be uniquely reconstructed from proper subsets of $\mathcal R(N)$ and $\mathcal S(N)$. We anticipate these results to be of use in the design of, for example, algorithms for phylogenetic network inference.
Characterizing Block Graphs in Terms of their Vertex-Induced Partitions
Given a finite connected simple graph $G=(V,E)$ with vertex set $V$ and edge set $E\subseteq \binom{V}{2}$, we will show that
$1.$ the (necessarily unique) smallest block graph with vertex set $V$ whose edge set contains $E$ is uniquely determined by the $V$-indexed family ${\bf P}_G:=\big(π_0(G^{(v)})\big)_{v \in V}$ of the various partitions $π_0(G^{(v)})$ of the set $V$ into the set of connected components of the graph $G^{(v)}:=(V,\{e\in E: v\notin e\})$,
$2.$ the edge set of this block graph coincides with set of all $2$-subsets $\{u,v\}$ of $V$ for which $u$ and $v$ are, for all $w\in V-\{u,v\}$, contained in the same connected component of $G^{(w)}$,
$3.$ and an arbitrary $V$-indexed family ${\bf P}p=({\bf p}_v)_{v \in V}$ of partitions $π_v$ of the set $V$ is of the form ${\bf P}p={\bf P}p_G$ for some connected simple graph $G=(V,E)$ with vertex set $V$ as above if and only if, for any two distinct elements $u,v\in V$, the union of the set in ${\bf p}_v$ that contains $u$ and the set in ${\bf p}_u$ that contains $v$ coincides with the set $V$, and $\{v\}\in {\bf p}_v$ holds for all $v \in V$.
As well as being of inherent interest to the theory of block graphs, these facts are also useful in the analysis of compatible decompositions and block realizations of finite metric spaces.
Blocks and Cut Vertices of the Buneman Graph
Published
• View Publication
• BIB
Given a set $\Sg$ of bipartitions of some finite set $X$ of cardinality at least 2, one can associate to $\Sg$ a canonical $X$-labeled graph $\B(\Sg)$, called the Buneman graph. This graph has several interesting mathematical properties - for example, it is a median network and therefore an isometric subgraph of a hypercube. It is commonly used as a tool in studies of DNA sequences gathered from populations. In this paper, we present some results concerning the {\em cut vertices} of $\B(\Sg)$, i.e., vertices whose removal disconnect the graph, as well as its {\em blocks} or 2-{\em connected components} - results that yield, in particular, an intriguing generalization of the well-known fact that $\B(\Sg)$ is a tree if and only if any two splits in $\Sg$ are compatible.
Phylogenetic networks form partial trees
A contemporary and fundamental problem faced by many evolutionary biologists is how to puzzle together a collection $\mathcal P$ of partial trees (leaf-labelled trees whose leaves are bijectively labelled by species or, more generally, taxa, each supported by e. g. a gene) into an overall parental structure that displays all trees in $\mathcal P$. This already difficult problem is complicated by the fact that the trees in $\mathcal P$ regularly support conflicting phylogenetic relationships and are not on the same but only overlapping taxa sets. A desirable requirement on the sought after parental structure therefore is that it can accommodate the observed conflicts. Phylogenetic networks are a popular tool capable of doing precisely this. However, not much is known about how to construct such networks from partial trees, a notable exception being the $Z$-closure super-network approach and the recently introduced $Q$-imputation approach. Here, we propose the usage of closure rules to obtain such a network. In particular, we introduce the novel $Y$-closure rule and show that this rule on its own or in combination with one of Meacham's closure rules (which we call the $M$-rule) has some very desirable theoretical properties. In addition, we use the $M$- and $Y$-rule to explore the dependency of Rivera et al.'s ``ring of life'' on the fact that the underpinning phylogenetic trees are all on the same data set. Our analysis culminates in the presentation of a collection of induced subtrees from which this ring can be reconstructed.