arXiv++ Combinatorics

Browse math.CO papers from arXiv

Papers by Mareike Fischer

43 paper(s) by this author · All BibTeX
Enumerating monophyletic characters in mathematical phylogenetics
Grouping species according to their phylogenetic relationships often results in different groups than grouping them according to their shared traits. Monophyletic groups play an important role in this regard, as they are groups of species sharing the same trait and being uniquely defined by a joint phylogenetic subtree. This immediately leads to the question of how to identify possible monophyletic groups in characters, which assign each present-day species a certain trait and which are typically used for phylogenetic tree reconstruction. In our manuscript, we provide a general formula to quantify how many different characters are monophyletic on any given tree and provide simple formulae for binary characters and for certain tree shapes. We also investigate relations between monophyly and the well-known phylogenetic tree reconstruction criterion maximum parsimony by providing a linear-time algorithm which determines the parsimony score together with the monophyly type of a character on a tree.
On Agreement Subtrees in Multiple Pylogenetic Trees
Snir and Yuster [Discrete Appl. Math. 347 (2026) 160--171] asked for the least number $h(k)$ such that $k$ unrooted binary phylogenetic trees on the same $h(k)$ leaves always share a common quartet. We give a new upper bound for the $k$-tree version of the Maximum Agreement Subtree problem, namely an upper bound for the number of leaves, on which $k$ unrooted binary phylogenetic trees always share a common induced binary subtree on $n$ leaves, which is a four-times iterated exponential function. For $h(k)$, this implies a four-times iterated exponential upper bound. We also set an exponential lower bound for $h(k)$.
2026-01-21 v2
A height-based metaconcept for rooted tree balance and its implications for the $B_1$ index
Tree balance has received considerable attention in recent years, both in phylogenetics and in other areas. Numerous (im)balance indices have been proposed to quantify the (im)balance of rooted trees. A recent comprehensive survey summarized this literature and showed that many existing indices are based on similar underlying principles. To unify these approaches, three general metaconcepts were introduced, providing a framework to classify, analyze, and extend imbalance indices. In this context, a metaconcept is a function $Φ_f$ that depends on another function $f$ capturing some aspect of tree shape. In this manuscript, we extend this line of research by introducing a new metaconcept based on the heights of the pending subtrees of all inner vertices. We provide a thorough analysis of this metaconcept and use it to answer open questions concerning the well-known $B_1$ balance index. In particular, we characterize the tree shapes that maximize the $B_1$ index in two cases: (i) arbitrary rooted trees and (ii) binary rooted trees. For both cases, we also determine the corresponding maximum values of the index. Finally, while the $B_1$ index is induced by a so-called third-order metaconcept, we explicitly introduce three new (im)balance indices derived from the first- and second-order height metaconcepts, respectively, thereby demonstrating that pending subtree heights give rise to a variety of novel (im)balance indices.
2025-09-05
Revealing the building blocks of tree balance: fundamental units of the Sackin and Colless Indices
Over the past decades, more than 25 phylogenetic tree balance indices and several families of such indices have been proposed in the literature -- some of which even contain infinitely many members. It is well established that different indices have different strengths and perform unequally across application scenarios. For example, power analyses have shown that the ability of an index to detect the generative model of a given phylogenetic tree varies significantly between indices. This variation in performance motivates the ongoing search for new and possibly \enquote{better} (im)balance indices. An easy way to generate a new index is to construct a compound index, e.g., a linear combination of established indices. Two of the most prominent and widely used imbalance indices in this context are the Sackin index and the Colless index. In this study, we show that these classic indices are themselves compound in nature: they can be decomposed into more elementary components that independently satisfy the defining properties of a tree (im)balance index. We further show that the difference Colless minus Sackin results in another imbalance index that is minimized (amongst others) by all Colless minimal trees. Conversely, the difference Sackin minus Colless forms a balance index. Finally, we compare the building blocks of which the Sackin and the Colless indices consist to these indices as well as to the stairs2 index, which is another index from the literature. Our results suggest that the elementary building blocks we identify are not only foundational to established indices but also valuable tools for analyzing disagreement among indices when comparing the balance of different trees.
A strengthened bound on the number of states required to characterize maximum parsimony distance
In this article we prove that the distance $d_{\mathrm{MP}}(T_1,T_2) = k$ between two unrooted binary phylogenetic trees $T_1, T_2$ on the same set of taxa can be defined by a character that is convex on one of $T_1, T_2$ and which has at most $2k$ states. This significantly improves upon the previous bound of $7k-5$ states. We also show that for every $k \geq 1$ there exist two trees $T_1, T_2$ with $d_{\mathrm{MP}}(T_1,T_2) = k$ such that at least $k+1$ states are necessary in any character that achieves this distance and which is convex on one of $T_1, T_2$. We augment these lower and upper bounds with an empirical analysis which shows that in practice significantly fewer than $k+1$ states are usually required.
2025-06-10 v3
Metaconcepts of rooted tree balance
Measures of tree balance play an important role in many different research areas such as mathematical phylogenetics or theoretical computer science. Typically, tree balance is quantified by a single number which is assigned to the tree by a balance or imbalance index, of which several exist in the literature. Most of these indices are based on structural aspects of tree shape, such as clade sizes or leaf depths. For instance, indices like the Sackin index, total cophenetic index, and $\widehat{s}$-shape statistic all quantify tree balance through clade sizes, albeit with different definitions and properties. In this paper, we formalize the idea that many tree (im)balance indices are functions of similar underlying tree shape characteristics by introducing metaconcepts of tree balance. A metaconcept is a function $Φ_f$ that depends on a function $f$ capturing some aspect of tree shape, such as balance values, clade sizes, or leaf depths. These metaconcepts encompass existing indices but also provide new means of measuring tree balance. The versatility and generality of metaconcepts allow for the systematic study of entire families of (im)balance indices, providing deeper insights that extend beyond index-by-index analysis.
2025-02-18 v3
The GFB Tree and Tree Imbalance Indices
Tree balance plays an important role in various research areas in phylogenetics and computer science. Typically, it is measured with the help of a balance index or imbalance index. There are more than 25 such indices available, recently surveyed in a book by Fischer et al. They are used to rank rooted binary trees on a scale from the most balanced to the least balanced. We show that a wide range of subtree-size based measures satisfying concavity and monotonicity conditions are minimized by the complete or greedy-from-the-bottom (GFB) tree and maximized by the caterpillar tree, yielding an infinitely large family of distinct new imbalance indices. Answering an open question from the literature, we show that one such established measure, the $\widehat{s}$-shape statistic, has the GFB tree as its unique minimizer. We also provide an alternative characterization of GFB trees, showing that they are equivalent to complete trees, which arise in different contexts. We give asymptotic bounds on the expected $\widehat{s}$-shape statistic under the uniform and Yule-Harding distributions of trees, and answer questions for the related $Q$-shape statistic as well.
2024-08-13 v2
A complete characterization of pairs of binary phylogenetic trees with identical $A_k$-alignments
Phylogenetic trees play a key role in the reconstruction of evolutionary relationships. Typically, they are derived from aligned sequence data (like DNA, RNA, or proteins) by using optimization criteria like, e.g., maximum parsimony (MP). It is believed that the latter is able to reconstruct the \enquote{true} tree, i.e., the tree that generated the data, whenever the number of substitutions required to explain the data with that tree is relatively small compared to the size of the tree (measured in the number $n$ of leaves of the tree, which represent the species under investigation). However, reconstructing the correct tree from any alignment first and foremost requires the given alignment to perform differently on the \enquote{correct} tree than on others. A special type of alignments, namely so-called $A_k$-alignments, has gained considerable interest in recent literature. These alignments consist of all binary characters (\enquote{sites}) which require precisely $k$ substitutions on a given tree. It has been found that whenever $k$ is small enough (in comparison to $n$), $A_k$-alignments uniquely characterize the trees that generated them. However, recent literature has left a significant gap between $n\leq 2k+2$ -- namely the cases in which no such characterization is possible -- and $n\geq 4k$ -- namely the cases in which this characterization works. It is the main aim of the present manuscript to close this gap, i.e., to present a full characterization of all pairs of trees that share the same $A_k$-alignment. In particular, we show that indeed every binary phylogenetic tree with $n$ leaves is uniquely defined by its $A_k$-alignments if $n\geq 2k+3$. By closing said gap, we also ensure that our result is optimal.
2024-03-02 v5
On the correctness of Maximum Parsimony for data with few substitutions in the NNI neighborhood of phylogenetic trees
Estimating phylogenetic trees, which depict the relationships between different species, from aligned sequence data (such as DNA, RNA, or proteins) is one of the main aims of evolutionary biology. However, tree reconstruction criteria like maximum parsimony do not necessarily lead to unique trees and in some cases even fail to recognize the \enquote{correct} tree (i.e., the tree on which the data was generated). On the other hand, a recent study has shown that for an alignment containing precisely those binary characters (sites) which require up to two substitutions on a given tree, this tree will be the unique maximum parsimony tree. It is the aim of the present paper to generalize this recent result in the following sense: We show that for a tree $T$ with $n$ leaves, as long as $k<\frac{n}{8}+\frac{11}{9}-\frac{1}{18}\sqrt{9\cdot \left(\frac{n}{4}\right)^2+16}$ (or, equivalently, $n>9 k-11+\sqrt{9k^2-22 k+17} $, which in particular holds for all $n\geq 12k$), the maximum parsimony tree for the alignment containing all binary characters which require (up to or precisely) $k$ substitutions on $T$ will be unique in the NNI neighborhood of $T$ and it will coincide with $T$, too. In other words, within the NNI neighborhood of $T$, $T$ is the unique most parsimonious tree for the said alignment. This partially answers a recently published conjecture affirmatively. Additionally, we show that for $n\geq 8$ and for $k$ being in the order of $\frac{n}{2}$, there is always a pair of phylogenetic trees $T$ and $T'$ which are NNI neighbors, but for which the alignment of characters requiring precisely $k$ substitutions each on $T$ in total requires fewer substitutions on $T'$.
Measuring 3D tree imbalance of plant models using graph-theoretical approaches
Imbalance in the 3D structure of plants can be an important indicator of insufficient light or nutrient supply, as well as excessive wind, (formerly present) physical barriers, neighbor or storm damage. It can also be a simple means to detect certain illnesses, since some diseases like the apple proliferation disease, an infection with the barley yellow dwarf virus or plant canker can cause abnormal growth, like \enquote{witches' brooms} or burls, resulting in a deviating 3D plant architecture. However, quantifying imbalance of plant growth is not an easy task, and it requires a mathematically sound 3D model of plants to which imbalance indices can be applied. Current models of plants are often based on stacked cylinders or voxel matrices and do not allow for measuring the degree of 3D imbalance in the branching structure of the whole plant. On the other hand, various imbalance indices are readily available for so-called graph-theoretical trees and are frequently used in areas like phylogenetics and computer science. While only some basic ideas of these indices can be transferred to the 3D setting, graph-theoretical trees are a logical foundation for 3D plant models that allow for elegant and natural imbalance measures. In this manuscript, our aim is thus threefold: We first present a new graph-theoretical 3D model of plants and discuss desirable properties of imbalance measures in the 3D setting. We then introduce and analyze eight different 3D imbalance indices and their properties. Thirdly, we illustrate all our findings using a data set of 63 bush beans. Moreover, we implemented all our indices in the publicly available \textsf{R}-software package \textsf{treeDbalance} accompanying this manuscript.
The weighted total cophenetic index: A novel balance index for phylogenetic networks
Phylogenetic networks play an important role in evolutionary biology as, other than phylogenetic trees, they can be used to accommodate reticulate evolutionary events such as horizontal gene transfer and hybridization. Recent research has provided a lot of progress concerning the reconstruction of such networks from data as well as insight into their graph theoretical properties. However, methods and tools to quantify structural properties of networks or differences between them are still very limited. For example, for phylogenetic trees, it is common to use balance indices to draw conclusions concerning the underlying evolutionary model, and more than twenty such indices have been proposed and are used for different purposes. One of the most frequently used balance index for trees is the so-called total cophenetic index, which has several mathematically and biologically desirable properties. For networks, on the other hand, balance indices are to-date still scarce. In this contribution, we introduce the \textit{weighted} total cophenetic index as a generalization of the total cophenetic index for trees to make it applicable to general phylogenetic networks. As we shall see, this index can be determined efficiently and behaves in a mathematical sound way, i.e., it satisfies so-called locality and recursiveness conditions. In addition, we analyze its extremal properties and, in particular, we investigate its maxima and minima as well as the structure of networks that achieve these values within the space of so-called level-$1$ networks. We finally briefly compare this novel index to the two other network balance indices available so-far.
2023-03-06 v2
Defining binary phylogenetic trees using parsimony: new bounds
Phylogenetic trees are frequently used to model evolution. Such trees are typically reconstructed from data like DNA, RNA, or protein alignments using methods based on criteria like maximum parsimony (amongst others). Maximum parsimony has been assumed to work well for data with only few state changes. Recently, some progress has been made to formally prove this assertion. For instance, it has been shown that each binary phylogenetic tree $T$ with $n \geq 20k$ leaves is uniquely defined by the set $A_k(T)$, which consists of all characters with parsimony score $k$ on $T$. In the present manuscript, we show that the statement indeed holds for all $n \geq 4k$, thus drastically lowering the lower bound for $n$ from $20k$ to $4k$. However, it has been known that for $n \leq 2k$ and $k \geq 3$, it is not generally true that $A_k(T)$ defines $T$. We improve this result by showing that the latter statement can be extended from $n \leq 2k$ to $n \leq 2k+2$. So we drastically reduce the gap of values of $n$ for which it is unknown if trees $T$ on $n$ taxa are defined by $A_k(T)$ from the previous interval of $[2k+1,20k-1]$ to the interval $[2k+3,4k-1]$. Moreover, we close this gap completely for the nearest neighbor interchange (NNI) neighborhood of $T$ in the following sense: We show that as long as $n\geq 2k+3$, no tree that is one NNI move away from $T$ (and thus very similar to $T$) shares the same $A_k$-alignment.
2021-11-08 v2
Defining binary phylogenetic trees using parsimony
Published • View PublicationBIB
Phylogenetic (i.e. leaf-labeled) trees play a fundamental role in evolutionary research. A typical problem is to reconstruct such trees from data like DNA alignments (whose columns are often referred to as characters), and a simple optimization criterion for such reconstructions is maximum parsimony. It is generally assumed that this criterion works well for data in which state changes are rare. In the present manuscript, we prove that each phylogenetic tree $T$ with $n\geq 20 k$ leaves is uniquely defined by the set $A_k(T)$, which consists of all characters with parsimony score $k$ on $T$. This can be considered as a promising first step towards showing that maximum parsimony as a tree reconstruction criterion is justified when the number of changes in the data is relatively small.
Tree balance indices: a comprehensive survey
Tree balance plays an important role in phylogenetics and other research areas, which is why several indices to measure tree balance have been introduced over the years. Nevertheless, a formal definition of what a balance index actually is and what makes it a useful measure of balance (or, in other cases, imbalance), has so far not been introduced in the literature. While the established indices all summarize the (im)balance of a tree in a single number, they vary in their definitions and underlying principles. It is the aim of the present manuscript to introduce formal definitions of balance and imbalance indices that classify desirable properties of such indices and to analyze and categorize established indices accordingly. In this regard, we review 19 established (im)balance indices from the literature, summarize their general, statistical and combinatorial properties (where known), prove numerous additional results and indicate directions for future research by making explicit open questions and gaps in the literature. We also prove that a few tree shape statistics that have been used to measure tree balance in the literature do not fulfill our definition of an (im)balance index, which might indicate that their properties are not as useful for practical purposes. Moreover, we show that five additional tree shape statistics from other contexts actually are tree (im)balance indices according to our definition. The manuscript is accompanied by the website \url{treebalance.wordpress.com} containing fact sheets of the discussed indices. Moreover, we introduce the software package \verb|treebalance| implemented in $\mathsf{R}$ that can be used to calculate all indices discussed.
Phylogenetic Diversity Rankings in the Face of Extinctions: the Robustness of the Fair Proportion Index
Published • View PublicationBIB
Planning for the protection of species often involves difficult choices about which species to prioritize, given constrained resources. One way of prioritizing species is to consider their "evolutionary distinctiveness", i.e. their relative evolutionary isolation on a phylogenetic tree. Several evolutionary isolation metrics or phylogenetic diversity indices have been introduced in the literature, among them the so-called Fair Proportion index (also known as the "evolutionary distinctiveness" score). This index apportions the total diversity of a tree among all leaves, thereby providing a simple prioritization criterion for conservation. Here, we focus on the prioritization order obtained from the Fair Proportion index and analyze the effects of species extinction on this ranking. More precisely, we analyze the extent to which the ranking order may change when some species go extinct and the Fair Proportion index is re-computed for the remaining taxa. We show that for each phylogenetic tree, there are edge lengths such that the extinction of one leaf per cherry completely reverses the ranking. Moreover, we show that even if only the lowest ranked species goes extinct, the ranking order may drastically change. We end by analyzing the effects of these two extinction scenarios (extinction of the lowest ranked species and extinction of one leaf per cherry) for a collection of empirical and simulated trees. In both cases, we can observe significant changes in the prioritization orders, highlighting the empirical relevance of our theoretical findings.
2021-05-03 v2
Measuring tree balance using symmetry nodes -- a new balance index and its extremal properties
Published • View PublicationBIB
Effects like selection in evolution as well as fertility inheritance in the development of populations can lead to a higher degree of asymmetry in evolutionary trees than expected under a null hypothesis. To identify and quantify such influences, various balance indices were proposed in the phylogenetic literature and have been in use for decades. However, so far no balance index was based on the number of \emph{symmetry nodes}, even though symmetry nodes play an important role in other areas of mathematical phylogenetics and despite the fact that symmetry nodes are a quite natural way to measure balance or symmetry of a given tree. The aim of this manuscript is thus twofold: First, we will introduce the \emph{symmetry nodes index} as an index for measuring balance of phylogenetic trees and analyze its extremal properties. We also show that this index can be calculated in linear time. This new index turns out to be a generalization of a simple and well-known balance index, namely the \emph{cherry index}, as well as a specialization of another, less established, balance index, namely \emph{Rogers' $J$ index}. Thus, it is the second objective of the present manuscript to compare the new symmetry nodes index to these two indices and to underline its advantages. In order to do so, we will derive some extremal properties of the cherry index and Rogers' $J$ index along the way and thus complement existing studies on these indices. Moreover, we used the programming language \textsf{R} to implement all three indices in the software package \textsf{symmeTree}, which has been made publicly available.
2021-02-08 v4
A survey of the monotonicity and non-contradiction of consensus methods and supertree methods
Published • View PublicationBIB
In a recent study, Bryant, Francis and Steel investigated the concept of \enquote{future-proofing} consensus methods in phylogenetics. That is, they investigated if such methods can be robust against the introduction of additional data like added trees or new species. In the present manuscript, we analyze consensus methods under a different aspect of introducing new data, namely concerning the discovery of new clades. In evolutionary biology, often formerly unresolved clades get resolved by refined reconstruction methods or new genetic data analyses. In our manuscript we investigate which properties of consensus methods can guarantee that such new insights do not disagree with previously found consensus trees, but merely refine them, a property termed \emph{monotonicity}. Along the lines of analyzing monotonicity, we also study two {established} supertree methods, namely Matrix Representation with Parsimony (MRP) and Matrix Representation with Compatibility (MRC), which have also been suggested as consensus methods in the literature. While we (just like Bryant, Francis and Steel in their recent study) unfortunately have to conclude some negative answers concerning general consensus methods, we also state some relevant and positive results concerning the majority rule ($\mathtt{MR}$) and strict consensus methods, which are amongst the most frequently used consensus methods. Moreover, we show that there exist infinitely many consensus methods which are monotonic and have some other desirable properties. \textbf{Keywords:} consensus tree, phylogenetics, majority rule, tree refinement, matrix representation with parsimony \textbf{MSC:} C92B05, 05C05
2020-11-26 v2
Combinatorial perspectives on Dollo-$k$ characters in phylogenetics
Published • View PublicationBIB
Recently, the perfect phylogeny model with persistent characters has attracted great attention in the literature. It is based on the assumption that complex traits or characters can only be gained once and lost once in the course of evolution. Here, we consider a generalization of this model, namely Dollo parsimony, that allows for multiple character losses. More precisely, we take a combinatorial perspective on the notion of Dollo-$k$ characters, i.e. traits that are gained at most once and lost precisely $k$ times throughout evolution. We first introduce an algorithm based on the notion of spanning subtrees for finding a Dollo-$k$ labeling for a given character and a given tree in linear time. We then compare persistent characters (consisting of the union of Dollo-0 and Dollo-1 characters) and general Dollo-$k$ characters. While it is known that there is a strong connection between Fitch parsimony and persistent characters, we show that Dollo parsimony and Fitch parsimony are in general very different. Moreover, while it is known that there is a direct relationship between the number of persistent characters and the Sackin index of a tree, a popular index of tree balance, we show that this relationship does not generalize to Dollo-$k$ characters. In fact, determining the number of Dollo-$k$ characters for a given tree is much more involved than counting persistent characters, and we end this manuscript by introducing a recursive approach for the former. This approach leads to a polynomial time algorithm for counting the number of Dollo-$k$ characters, and both this algorithm as well as the algorithm for computing Dollo-$k$ labelings are publicly available in the Babel package for BEAST 2.
Non-binary universal tree-based networks
Published • View PublicationBIB
A tree-based network $N$ on $X$ is called universal if every phylogenetic tree on $X$ is a base tree for $N$. Recently, binary universal tree-based networks have attracted great attention in the literature and their existence has been analyzed in various studies. In this note, we extend the analysis to non-binary networks and show that there exist both a rooted and an unrooted non-binary universal tree-based network with $n$ leaves for all positive integers $n$.
2019-10-13
The space of tree-based phylogenetic networks
Published • View PublicationBIB
Phylogenetic networks are generalizations of phylogenetic trees that allow the representation of reticulation events such as horizontal gene transfer or hybridization, and can also represent uncertainty in inference. A subclass of these, tree-based phylogenetic networks, have been introduced to capture the extent to which reticulate evolution nevertheless broadly follows tree-like patterns. Several important operations that change a general phylogenetic network have been developed in recent years, and are important for allowing algorithms to move around spaces of networks; a vital ingredient in finding an optimal network given some biological data. A key such operation is the Nearest Neighbor Interchange, or NNI. While it is already known that the space of unrooted phylogenetic networks is connected under NNI, it has been unclear whether this also holds for the subspace of tree-based networks. In this paper we show that the space of unrooted tree-based phylogenetic networks is indeed connected under the NNI operation. We do so by explicitly showing how to get from one such network to another one without losing tree-basedness along the way. Moreover, we introduce some new concepts, for instance ``shoat networks'', and derive some interesting aspects concerning tree-basedness. Last, we use our results to derive an upper bound on the size of the space of tree-based networks.