species
265 papers tagged with this keyword
A dichotomy law for certain classes of phylogenetic networks
Many classes of phylogenetic networks have been proposed in the literature. A feature of several of these classes is that if one restricts a network in the class to a subset of its leaves, then the resulting network may no longer lie within this class. This has implications for their biological applicability, since some species -- which are the leaves of an underlying evolutionary network -- may be missing (e.g., they may have become extinct, or there are no data available for them) or we may simply wish to focus attention on a subset of the species. On the other hand, certain classes of networks are `closed' when we restrict to subsets of leaves, such as (i) the classes of all phylogenetic networks or all phylogenetic trees; (ii) the classes of galled networks, simplicial networks, galled trees; and (iii) the classes of networks that have some parameter that is monotone-under-leaf-subsampling (e.g., the number of reticulations, height, etc.) bounded by some fixed value. It is easily shown that a closed subclass of phylogenetic trees is either all trees or a vanishingly small proportion of them (as the number of leaves grows). In this short paper, we explore whether this dichotomy phenomenon holds for other classes of phylogenetic networks, and their subclasses.
Yet Another Species of Forbidden-distances Chromatic Number
Published in Geombinatorics, vol. 10 no. 3 (2001), pp. 89-95
• Search Publication
This 2001 paper introduces a new type of chromatic number for point sets.
Covariance Decomposition for Distance Based Species Tree Estimation
In phylogenomics, species-tree methods must contend with two major sources of noise; stochastic gene-tree variation under the multispecies coalescent model (MSC) and finite-sequence substitutional noise. Fast agglomerative methods such as GLASS, STEAC, and METAL combine multi-locus information via distance-based clustering. We derive the exact covariance matrix of these pairwise distance estimates under a joint MSC-plus-substitution model and leverage it for reliable confidence estimation, and we algebraically decompose it into components attributable to coalescent variation versus sequence-level stochasticity. Our theory identifies parameter regimes where one source of variance greatly exceeds the other. For both very low and very high mutation rates, substitutional noise dominates, while coalescent variance is the primary contributor at intermediate mutation rates. Moreover, the interval over which coalescent variance dominates becomes narrower as the species-tree height increases. These results imply that in some settings one may legitimately ignore the weaker noise source when designing methods or collecting data. In particular, when gene-tree variance is dominant, adding more loci is most beneficial, while when substitution noise dominates, longer sequences or imputation are needed. Finally, leveraging the derived covariance matrix, we implement a Gaussian-sampling procedure to generate split support values for METAL trees and demonstrate empirically that this approach yields more reliable confidence estimates than traditional bootstrapping.
Higher Order Bell Symmetric Functions
We study symmetric function analogues of the higher order Bell numbers. Their construction involves iterated plethystic exponential towers mimicking the single variable exponential generating functions for the higher order Bell numbers. We derive explicit recurrence relations for the expansion coefficients of the Bell functions into the monomial and power sum bases of the ring of symmetric functions. Using the machinery of combinatorial species, the Bell functions are proven to be the Frobenius characteristics of the permutation representations of symmetric groups on hyper-partitions of certain orders and sizes. In the order 1 case, we are able to give more details about the expansion coefficients of the Bell functions in terms of vector partitions and divisor sums as well as give a recurrence relation analogous to the well known recursion for the Bell numbers. Lastly, we use Littlewood's reciprocity theorem and the Hardy-Littlewood Tauberian theorem to prove that the Schur expansion coefficients of the order 1 Bell functions are certain asymptotic averages of restriction coefficients.
Lie-operads and operadic modules from poset cohomology
As observed by Joyal, the cohomology groups of the partition posets are naturally identified with the components of the operad encoding Lie algebras. This connection was explained in terms of operadic Koszul duality by Fresse, and later generalized by Vallette to the setting of decorated partitions. In this article, we set up and study a general formalism which produces a priori operadic structures (operads and operadic modules) on the cohomology of families of posets equipped with some natural recursive structure, that we call "operadic poset species". This framework goes beyond decorated partitions and operadic Koszul duality, and contains the metabelian Lie operad and Kontsevich's operad of trees as two simple instances. In forthcoming work, we will apply our results to the hypertree posets and their connections to post-Lie and pre-Lie algebras.
Unexpectedly, a symmetry on unlabeled graphs
We exhibit the joint symmetric distribution of the following two parameters on the set of unlabeled, simple, connected graphs with $n$ vertices. The first parameter is the maximal number of leaves attached to a vertex. The second parameter is the size of the largest set of vertices sharing the same closed neighborhood minus $1$.
Apparently, this is the first example of a natural, non-trivial equidistribution of graph parameters on unlabeled connected graphs on a fixed set of vertices.
Our proof is enumerative, using the theory of species. Exhibiting an explicit bijection interchanging the two parameters remains an open problem.
Coconvex characters on collections of phylogenetic trees
In phylogenetics, a key problem is to construct evolutionary trees from collections of characters where, for a set X of species, a character is simply a function from X onto a set of states. In this context, a key concept is convexity, where a character is convex on a tree with leaf set X if the collection of subtrees spanned by the leaves of the tree that have the same state are pairwise disjoint. Although collections of convex characters on a single tree have been extensively studied over the past few decades, very little is known about coconvex characters, that is, characters that are simultaneously convex on a collection of trees. As a starting point to better understand coconvexity, in this paper we prove a number of extremal results for the following question: What is the minimal number of coconvex characters on a collection of n-leaved trees taken over all collections of size t >= 2, also if we restrict to coconvex characters which map to k states? As an application of coconvexity, we introduce a new one-parameter family of tree metrics, which range between the coarse Robinson-Foulds distance and the much finer quartet distance. We show that bounds on the quantities in the above question translate into bounds for the diameter of the tree space for the new distances. Our results open up several new interesting directions and questions which have potential applications to, for example, tree spaces and phylogenomics.
A New Representation of Ewens-Pitman's Partition Structure and Its Characterization via Riordan Array Sums
Ewens-Pitman's partition structure arises as a system of sampling consistent probability distributions on set partitions induced by the Pitman-Yor process. It is widely used in statistical applications, particularly in species sampling models in Bayesian nonparametrics. Drawing references from the area of representation theory of the infinite symmetric group, we view Ewens-Pitman's partition structure as an example of a non-extreme harmonic function on a branching graph, specifically, the Kingman graph. Taking this perspective enables us to obtain combinatorial and algebraic constructions of this distribution using the interpolation polynomial approach proposed by Borodin and Olshanski (The Electronic Journal of Combinatorics, 7, 2000). We provide a new explicit representation of Ewens-Pitman's partition structure using modern umbral interpolation based on Sheffer polynomial sequences. In addition, we show that a certain type of marginals of this distribution can be computed using weighted row sums of a Riordan array. In this way, we show that some summary statistics and estimators derived from Ewens-Pitman's partition structure can be obtained using methods of generating functions. This approach simplifies otherwise cumbersome calculations of these quantities often involving various special combinatorial functions. In addition, it has the added benefit of being amenable to symbolic computation.
Multispecies inhomogeneous $t$-PushTASEP from antisymmetric fusion
Published in Electron. J. Probab. 30: 1-28 (2025)
• View Publication
• BIB
We investigate the recently introduced inhomogeneous $n$-species $t$-PushTASEP, a long-range stochastic process on a periodic lattice. A Baxter-type formula is established, expressing the Markov matrix as an alternating sum of commuting transfer matrices over all the fundamental representations of $U_t(\widehat{sl}_{n+1})$. This superposition acts as an inclusion-exclusion principle, selectively extracting the sequential particle transitions characteristic of the PushTASEP, while canceling forbidden channels. The homogeneous specialization connects the PushTASEP to ASEP, showing that the two models share eigenstates and a common integrability structure.
Orthology and Near-Cographs in the Context of Phylogenetic Networks
Orthologous genes, which arise through speciation, play a key role in comparative genomics and functional inference. In particular, graph-based methods allow for the inference of orthology estimates without prior knowledge of the underlying gene or species trees. This results in orthology graphs, where each vertex represents a gene, and an edge exists between two vertices if the corresponding genes are estimated to be orthologs. Orthology graphs inferred under a tree-like evolutionary model must be cographs. However, real-world data often deviate from this property, either due to noise in the data, errors in inference methods or, simply, because evolution follows a network-like rather than a tree-like process. The latter, in particular, raises the question of whether and how orthology graphs can be derived from or, equivalently, are explained by phylogenetic networks. Here, we study the constraints imposed on orthology graphs when the underlying evolutionary history follows a phylogenetic network instead of a tree. We show that any orthology graph can be represented by a sufficiently complex level-k network. However, such networks lack biologically meaningful constraints. In contrast, level-1 networks provide a simpler explanation, and we establish characterizations for level-1 explainable orthology graphs, i.e., those derived from level-1 evolutionary histories. To this end, we employ modular decomposition, a classical technique for studying graph structures. Specifically, an arbitrary graph is level-1 explainable if and only if each primitive subgraph is a near-cograph (a graph in which the removal of a single vertex results in a cograph). Additionally, we present a linear-time algorithm to recognize level-1 explainable orthology graphs and to construct a level-1 network that explains them, if such a network exists.
Predicting the depth of the most recent common ancestor of a random sample of $k$ species: the impact of phylogenetic tree shape
We consider the following question: how close to the ancestral root of a phylogenetic tree is the most recent common ancestor of $k$ species randomly sampled from the tips of the tree? For trees having shapes predicted by the Yule-Harding model, it is known that the most recent common ancestor is likely to be close to (or equal to) the root of the full tree, even as $n$ becomes large (for $k$ fixed). However, this result does not extend to models of tree shape that more closely describe phylogenies encountered in evolutionary biology. We investigate the impact of tree shape (via the Aldous $β-$splitting model) to predict the number of edges that separate the most recent common ancestor of a random sample of $k$ tip species and the root of the parent tree they are sampled from. Both exact and asymptotic results are presented. We also briefly consider a variation of the process in which a random number of tip species are sampled.
Species of Rota-Baxter algebras by rooted trees, twisted bialgebras and Fock functors
As a fundamental and ubiquitous combinatorial notion, species has attracted sustained interest, generalizing from set-theoretical combinatorial to algebraic combinatorial and beyond. The Rota-Baxter algebra is one of the algebraic structures with broad applications from Renormalization of quantum field theory to integrable systems and multiple zeta values. Its interpretation in terms of monoidal categories has also recently appeared. This paper studies species of Rota-Baxter algebras, making use of the combinatorial construction of free Rota-Baxter algebras in terms of angularly decorated trees and forests. The notion of simple angularly decorated forests is introduced for this purpose and the resulting Rota-Baxter species is shown to be free. Furthermore, a twisted bialgebra structure, as the bialgebra for species, is established on this free Rota-Baxter species. Finally, through the Fock functor, another proof of the bialgebra structure on free Rota-Baxter algebras is obtained.
Integro-differential rings on species and derived structures
In the theory of species, differential as well as integral operators are known to arise in a natural way. In this paper, we shall prove that they precisely fit together in the algebraic framework of integro-differential rings, which are themselves an abstraction of classical calculus (incorporating its Fundamental Theorem). The results comprise (set) species as well as linear species. Localization of (set) species leads to the more general structure of modified integro-differential rings, previously employed in the algebraic treatment of Volterra integral equations. Furthermore, the ring homomorphism from species to power series via taking generating series is shown to be a (modified) integro-differential ring homomorphism. As an application, a topology and further algebraic operations are imported to virtual species from the general theory of integro-differential rings.
Enumerative aspects of Caylerian polynomials
Eulerian polynomials record the distribution of descents over permutations. Caylerian polynomials likewise record the distribution of descents over Cayley permutations, where a Cayley permutation is a word of positive integers such that if a number appears in the word then all positive integers less than that number also appear in the word. Using combinatorial species and sign-reversing involutions we derive counting formulas and generating functions for the Caylerian polynomials as well as for related refined polynomials.
Simplifying and Characterizing DAGs and Phylogenetic Networks via Least Common Ancestor Constraints
Rooted phylogenetic networks, or more generally, directed acyclic graphs (DAGs), are widely used to model species or gene relationships that traditional rooted trees cannot fully capture, especially in the presence of reticulate processes or horizontal gene transfers. Such networks or DAGs are typically inferred from observable data (e.g. genomic sequences of extant species), providing only an estimate of the true evolutionary history. However, these inferred DAGs are often complex and difficult to interpret. In particular, many contain vertices that do not serve as least common ancestors (LCAs) for any subset of the underlying genes or species, thus may lack direct support from the observable data. In contrast, LCA vertices are witnessed by historical traces justifying their existence and thus represent ancestral states substantiated by the data. To reduce unnecessary complexity and eliminate unsupported vertices, we aim to simplify a DAG to retain only LCA vertices while preserving essential evolutionary information.
In this paper, we characterize $\mathrm{LCA}$-relevant and $\mathrm{lca}$-relevant DAGs, defined as those in which every vertex serves as an LCA (or unique LCA) for some subset of taxa. We introduce methods to identify LCAs in DAGs and efficiently transform any DAG into an $\mathrm{LCA}$-relevant or $\mathrm{lca}$-relevant one while preserving key structural properties of the original DAG or network. This transformation is achieved using a simple operator ``$\ominus$'' that mimics vertex suppression.
Characterising rooted and unrooted tree-child networks
Rooted phylogenetic networks are used by biologists to infer and represent complex evolutionary relationships between species that cannot be accurately explained by a phylogenetic tree. Tree-child networks are a particular class of rooted phylogenetic networks that has been extensively investigated in recent years. In this paper, we give a novel characterisation of a tree-child network $\mathcal{R}$ in terms of cherry-picking sequences that are sequences on the leaves of $\mathcal{R}$ and reduce it to a single vertex by repeatedly applying one of two reductions to its leaves. We show that our characterisation extends to unrooted tree-child networks which are mostly unexplored in the literature and, in turn, also offers a new approach to settling the computational complexity of deciding if an unrooted phylogenetic network can be oriented as a rooted tree-child network.
A complete characterization of pairs of binary phylogenetic trees with identical $A_k$-alignments
Phylogenetic trees play a key role in the reconstruction of evolutionary relationships. Typically, they are derived from aligned sequence data (like DNA, RNA, or proteins) by using optimization criteria like, e.g., maximum parsimony (MP). It is believed that the latter is able to reconstruct the \enquote{true} tree, i.e., the tree that generated the data, whenever the number of substitutions required to explain the data with that tree is relatively small compared to the size of the tree (measured in the number $n$ of leaves of the tree, which represent the species under investigation). However, reconstructing the correct tree from any alignment first and foremost requires the given alignment to perform differently on the \enquote{correct} tree than on others.
A special type of alignments, namely so-called $A_k$-alignments, has gained considerable interest in recent literature. These alignments consist of all binary characters (\enquote{sites}) which require precisely $k$ substitutions on a given tree. It has been found that whenever $k$ is small enough (in comparison to $n$), $A_k$-alignments uniquely characterize the trees that generated them. However, recent literature has left a significant gap between $n\leq 2k+2$ -- namely the cases in which no such characterization is possible -- and $n\geq 4k$ -- namely the cases in which this characterization works. It is the main aim of the present manuscript to close this gap, i.e., to present a full characterization of all pairs of trees that share the same $A_k$-alignment. In particular, we show that indeed every binary phylogenetic tree with $n$ leaves is uniquely defined by its $A_k$-alignments if $n\geq 2k+3$. By closing said gap, we also ensure that our result is optimal.
Compactifications of phylogenetic systems and species of electrical networks
We describe new spaces and maps. Our graphical map is a visual and numerical correspondence between spaces of circular electrical networks and circular planar split systems. When restricted to the planar circular electrical case, this graphical map finds the split system uniquely associated with the Kalmanson resistance distance of the dual network, matching the induced split system familiar from phylogenetics. This correspondence is extended to compactifications of the respective spaces, taking cactus networks to the cactus split systems defined herein. The graphical map preserves both network components and cactus structure, allowing an elegant enumeration of induced phylogenetic split systems via combinatorial species. We introduce the global spaces of circular planar electrical networks and circular split systems. These new spaces are also CW complexes, but the 0-cells of each are counted by the Bell numbers as opposed to the Catalan numbers. As species, the two sorts of global cacti are seen to be compositions in complementary ways.
Pattern-avoiding Cayley permutations via combinatorial species
A Cayley permutation is a word of positive integers such that if a letter appears in this word, then all positive integers smaller than that letter also appear. We initiate a systematic study of pattern avoidance on Cayley permutations adopting a combinatorial species approach. Our methods lead to species equations, generating series, and counting formulas for Cayley permutations avoiding any pattern of length at most three. We also introduce the species of primitive structures as a generalization of Cayley permutations with no "flat steps". Finally, we explore various notions of Wilf equivalence arising in this context.
Colouring isonemal fabrics with more than two colours and non-twilly redundancy
Perfect colouring of isonemal fabrics by thin and thick striping of warp and weft with more than two colours is examined where the cells with warps and wefts of the same colour do not appear along diagonal lines (not twilly redundancy). In species 33 to 39, all fabrics of specific orders (linear periods) can be perfectly coloured by thin striping with satin redundancy. In fewer species, all fabrics of specific orders can be perfectly coloured by thick striping in two different ways with doubled satin redundancy. The isonemal fabric 6-1-1 can be used as a redundancy configuration. All these colourings allow the coloured weaving of flat tori and -- a few of them -- cubes.