arXiv++ Combinatorics

Browse math.CO papers from arXiv

q-bio.PE ↗ arXiv

15 papers in this category
2026-09-04
A Short Combinatorial Proof of the Pons-Batle Identity for Counting Tree-Child Networks
Tree-child networks are a useful class of binary phylogenetic networks. The Pons--Batle identity (Pons and Batle, \textit{Scientific Reports}, 2021) states that the number $a_{n,k}$ of tree-child networks with $k$ reticulations on $n$ taxa satisfies \[ a_{n,k}=(n-k+1)a_{n,k-1} +\frac{n(2n+k-3)}{n-k}a_{n-1,k}. \] In this paper, we present a short combinatorial proof of this identity.
2026-08-26
Tree Buckets and the Reconstruction of Pairs of Phylogenetic Trees
Phylogenetic trees are used in evolutionary biology to represent the evolutionary history of a collection of taxa. As we have incomplete information about any evolutionary history, recovering trees from partial information is a focus of phylogenetic combinatorics. However, in some cases the available data does not describe a single phylogenetic tree. We consider recovery of pairs of phylogenetic trees from their combined subtrees with $k$ leaves, which we call a $k$-bucket. We establish the exact cases in which these pairs of trees are recoverable from their subtrees with a single leaf removed, both when just considering the structure of the trees, and when additionally considering the set of taxa on the leaves. We also consider recovery of pairs of trees with labelled leaves from their rooted triples, and establish that they are recoverable up to a sequence of subtree swaps.
2026-08-24
On the maximum size of 2-weakly compatible split systems
We consider a Turán-type problem arising in phylogenetics: determining the maximum size of a 2-weakly compatible split system. This compatibility condition arises in the reconstruction of phylogenetic networks from quartet weights. It was previously shown that a 2-weakly compatible split system has size at most \[ 3\binom{n}{4}+\binom{n}{2}. \] We prove that the maximum size is $O(n^{5/2})$.
Enumerating monophyletic characters in mathematical phylogenetics
Grouping species according to their phylogenetic relationships often results in different groups than grouping them according to their shared traits. Monophyletic groups play an important role in this regard, as they are groups of species sharing the same trait and being uniquely defined by a joint phylogenetic subtree. This immediately leads to the question of how to identify possible monophyletic groups in characters, which assign each present-day species a certain trait and which are typically used for phylogenetic tree reconstruction. In our manuscript, we provide a general formula to quantify how many different characters are monophyletic on any given tree and provide simple formulae for binary characters and for certain tree shapes. We also investigate relations between monophyly and the well-known phylogenetic tree reconstruction criterion maximum parsimony by providing a linear-time algorithm which determines the parsimony score together with the monophyly type of a character on a tree.
Classes of phylogenetic networks that are robust to root placement
Standard phylogenetic reconstruction techniques often yield unrooted phylogenetic networks; these are subsequently rooted to infer evolutionary history. A common problem in this process is to determine the structural classes to which the resulting network will belong. In this paper, we investigate unrooted networks in which the choice of any root results in a valid rooted phylogenetic network, a property we define as {\em robustly orientable}. We then establish a strict structural condition for this class, specifically, that an unrooted network is robustly orientable if and only if it contains no sink components. We also show that if an unrooted network is level-$2$ or less, or if it is tree-based, then it is robustly orientable. Furthermore, we define an unrooted network to be {\em robustly class $\mathcal C$} if the choice of any root results in a network belonging to class $\mathcal C$. We demonstrate that an unrooted network is robustly tree-child or robustly stack-free if and only if it is level-$1$ or less. Finally, we show that a phylogenetic network is robustly normal if and only if it is a phylogenetic tree.
Proximity Measures for Classes of Phylogenetic Networks
Phylogenetic networks are used to represent the evolutionary history of species. Due to biological interpretations and computational advantages, researchers have focused on restricted classes of phylogenetic networks, such as tree-child, orchard, and tree-based. These classes capture different notions of tree-likeness: tree-child networks require every internal vertex to have a taxon reachable by a tree path, orchard networks are trees with horizontal arcs (for modelling histories rife with horizontal gene transfers), and tree-based networks are trees with additional (not-necessarily horizontal) arcs. A natural question to ask is ``how far is a given network from belonging to a particular class?'' This motivates the study of proximity measures, which measure the minimum number of graph modifications required to transform a network into one belonging to a particular class. In this paper, we consider three proximity measures based on leaf addition, valid arc deletion, and arc deletion. We study pairwise comparability of the proximity measures, prove complexity results, and derive extremal bounds for the classes of tree, tree-child, orchard, and tree-based networks.
2026-07-06
Polynomial encoding of rooted trees with branch lengths
Phylogenetic trees are rooted trees with branch lengths that record genetic divergence or elapsed time, and quantifying differences between them is central to a wide range of evolutionary and epidemiological analyses. Graph-polynomial encodings of rooted trees provide an accurate, interpretable, and computationally efficient way to compare tree shapes, but existing polynomial encodings must be paired with auxiliary structures to study rooted trees with branch lengths. We introduce a bivariate polynomial encoding that incorporates branch lengths directly into a recursive computation from the leaf vertices to the root vertex of a tree. We prove that, for rooted trees with branch lengths and no vertices of degree two, which include all standard phylogenetic trees, two trees have the same polynomial if and only if their underlying unlabeled trees are isomorphic and the branch lengths of corresponding edges are equal. We apply the polynomial encoding to three published HIV-1 phylogenies sampled in different epidemiological settings and show that it accurately separates the three datasets based on their tree topologies and branch lengths, outperforming previous polynomial-based approaches for analyzing rooted trees with branch lengths.
2026-06-23 v2
Exact Enumeration of Phylogenetic Networks: The Tree-Child, Reticulation-Visible and Orchard Hierarchy
We develop a unified framework for the exact enumeration and asymptotic analysis of the three most studied classes of phylogenetic networks: tree-child (TC), reticulation-visible (RV) and orchard networks, whose cardinalities satisfy the strict ordering $|\mathrm{TC}_{\ell,k}|<|\mathrm{RV}_{\ell,k}|<|\mathrm{Orch}_{\ell,k}|$ for reticulation number $k\geq2$ (with $\mathrm{TC}\subsetneq\mathrm{RV}$ and $\mathrm{TC}\subsetneq\mathrm{Orch}$, while $\mathrm{RV}$ and $\mathrm{Orch}$ are incomparable as sets). Using the Chang--Fuchs structural theorem, we derive a two-level master functional equation for the RV bivariate generating function and obtain exact closed-form identities for the differences $Δ_k(\ell):=|RV_{\ell,k}|-|TC_{\ell,k}|$ for $k=2,3$, with the asymptotic universality $Δ_k(\ell)/|TC_{\ell,k}|\sim k!/\ell$. For orchard networks, we prove a \emph{universal hypergeometric law} that resolves the exact enumeration problem for all $\ell$: the column generating function $F_\ell(v)$ is rational with denominator $D_\ell(v)=\prod_{j=2}^\ell X_j(v)$, where \[ X_\ell(v) = \sum_{k=0}^{\lfloor\ell/2\rfloor}(-1)^k\, \frac{\ell!}{(\ell-2k)!\,k!}\,v^k \] is the matching polynomial of the complete graph $K_\ell$ and a rescaled Jacobi polynomial. This immediately resolves the intractable $\ell=9$ case: $D_9$ has degree 20, dominant growth rate $\approx40.73$, and all spectral roots are positive real. A complete enumeration table is provided extending the published data of Cardona, Ribas and Pons.
A parameterized family of balance indices for phylogenetic networks
We introduce a new family of balance indices for phylogenetic networks: the $H_α$ indices, where $α$ is a positive real number. This family includes the $B_2$ index as a special case ($α= 1$) and provides a natural extension of the Sackin index to phylogenetic networks. We show that the $H_α$ indices share many structural properties with the $B_2$ index, most notably a "grafting property" that makes it possible to express the $H_α$ index of a network in terms of the $H_α$ indices of its biconnected components. These properties allow us to identify networks that minimize / maximize $H_α$ for various classes of phylogenetic networks, and to study its distribution for several models of random trees and networks (in particular, Galton-Watson trees and binary Markov branching trees, with a focus on the Yule and PDA models). Finally, we show how local limits can be used to analyze the asymptotic behavior of $H_α$ for large trees and networks, and we obtain general results for the moments of $H_α$ for a broad class of random phylogenetic networks known as blowups of Galton-Watson trees.
2026-06-12
Note on the Maximum Number of Trees Displayed by a Tree-Child Network
In this note, we show that, for all $n\ge 2$, the number of distinct rooted binary phylogenetic $X$-trees displayed by a binary tree-child network $\mathcal{N}$ on $X$ with $n$ leaves is at most $2^{n-1}-1$ and that this upper bound is sharp. Furthermore, if $\mathcal{N}$ displays exactly $2^{n-1}-1$ such trees, then exactly one rooted binary phylogenetic $X$-tree is displayed twice, and this tree can be canonically found by iteratively replacing a reticulated cherry with a cherry.
2026-05-20
Regularizing and Normalizing DAGs and Phylogenetic Networks
Phylogenetic networks and, more generally, directed acyclic graphs (DAGs) represent hierarchical structure beyond trees, for instance in the presence of reticulate evolutionary events such as hybridization or horizontal gene transfer. A central question is which parts of such graphs are essential with respect to leaf-observable information, and which parts can be removed without changing this information. Resolving this question can lead to principled simplification methods for phylogenetic networks, such as the recent normalization approach of Francis et al. In this paper, we study this question from three related perspectives: clusters displayed by a DAG $G$, least common ancestors (LCAs) of subsets of its leaf set, and visibility, a path-based property of vertices. We first introduce an LCA-based simplification procedure called $i$-regularization. For a DAG $G$ and $i\geq 1$, the DAG $\reg_i(G)$ retains precisely those vertices that occur as unique LCAs of leaf subsets of size at most $i$, removes the remaining non-leaf vertices by a graph-editing operation $\ominus$, and then deletes shortcuts. We show that $\reg_i(G)$ preserves all such LCAs, is $i$-lca-relevant, and admits a cluster-level description: it is regular, i.e., isomorphic to the Hasse diagram of the corresponding lca-clusters. We then compare LCA-based regularization with normalization. Using the same $\ominus$-operator, we describe the cover construction underlying normalization, identify visible vertices that are nevertheless removed, and characterize when regularization and normalization coincide. Together, these results provide a unified framework for cluster-based, LCA-based, and visibility-based simplifications of DAGs and phylogenetic networks.
2026-05-07
A $μ$-distance for semidirected orchard phylogenetic networks
In evolutionary biology, phylogenetic networks are now widely used to represent the historical relationships between species and population, when this history includes reticulation events such as hybridization, gene flow and admixture between populations. Semidirected phylogenetic networks are appropriate models when the direction of some edges and the root position are not identifiable from data. Comparing semidirected networks is important in many applications. For rooted and directed networks, a $μ$-representation was originally introduced to distinguish tree-child networks, and has since been extended in two different directions: to the larger class of orchard directed networks by adding an extra component that counts paths to reticulations; and to semidirected networks, through an edge-based variant. However, the latter does not provide a distance between semidirected and orchard networks. We introduce here a new edge-based $μ$-representation capable of distinguishing distinct orchard binary semidirected networks. For this class, we provide a reconstruction algorithm and therefore obtain a true distance that is computable in polynomial time.
2026-04-06
Nested tree space: a geometric framework for co-phylogeny
Nested (or reconciled) phylogenetic trees model co-evolutionary systems in which one evolutionary history is embedded within another. We introduce a geometric framework for such systems by defining $σ$-space, a moduli space of fully nested ultrametric phylogenetic trees with a fixed leaf map. Generalizing the $τ$-space of Gavryushkin and Drummond, $σ$-space is constructed as a cubical complex parametrised by nested ranked tree topologies and inter-event time coordinates of the combined host and parasite speciation events. We characterise admissible orderings via binary \textit{nesting sequences} and organise them into a natural poset. We show that $σ$-space is contractible and satisfies Gromov's cube condition, and is therefore CAT(0). In particular, it admits unique geodesics and well-defined Fréchet means. We further describe its geometric structure, including boundary strata corresponding to cospeciation events, and relate it to products of ultrametric tree spaces via natural forgetful maps.
A Class of Unrooted Phylogenetic Networks Inspired by the Properties of Rooted Tree-Child Networks
A directed phylogenetic network is tree-child if every non-leaf vertex has a child that is not a reticulation. As a class of directed phylogenetic networks, tree-child networks are very useful from a computational perspective. For example, several computationally difficult problems in phylogenetics become tractable when restricted to tree-child networks. At the same time, the class itself is rich enough to contain quite complex networks. Furthermore, checking whether a directed network is tree-child can be done in polynomial time. In this paper, we seek a class of undirected phylogenetic networks that is rich and computationally useful in a similar way to the class tree-child directed networks. A natural class to consider for this role is the class of tree-child-orientable networks which contains all those undirected phylogenetic networks whose edges can be oriented to create a tree-child network. However, we show here that recognizing such networks is NP-hard, even for binary networks, and as such this class is inappropriate for this role. Towards finding a class of undirected networks that fills a similar role to directed tree-child networks, we propose new classes called $q$-cuttable networks, for any integer $q\geq 1$. We show that these classes have many of the desirable properties, similar to tree-child networks in the rooted case, including being recognizable in polynomial time, for all $q\geq 1$. Towards showing the computational usefulness of the class, we show that the NP-hard problem Tree Containment is polynomial-time solvable when restricted to $q$-cuttable networks with $q\geq 3$.
2026-02-25
A kernel for the maximum agreement forest problem on multiple binary phylogenetic trees
The maximum agreement forest (MAF) problem in phylogenetics takes as input a set t >=2 of binary phylogenetic trees T on the same set of taxa X. It asks for a partition X into the smallest number of blocks such that the subtrees induced by these blocks are disjoint and have common topology across all the trees in T. We produce a modified version of the well-known chain reduction rule in order to prove the existence of a kernel of size O( t * r * k ) where k is the natural parameter (the number of blocks) and r=min{max{k,3},t+1}}. We prove this bound for both the unrooted and rooted version of the problem, and demonstrate that the bound r, the length to which common chains are truncated, is tight. Our results constitute the first kernels for MAF in the t > 2 regime.