lyndon word
58 papers tagged with this keyword
Necklaces and Lyndon words in colexicographic order
We present the first constant-amortized-time algorithms for generating all length-$n$ necklaces and Lyndon words over a $k$-letter alphabet in colexicographic order, for arbitrary $k\geq 2$. Our approach introduces a novel class of words called \emph{quasinecklaces}, which serve as an easily generated superset of necklaces through which all necklaces can be efficiently identified. We derive a formula for the number $Q_k(n)$ of length-$n$ quasinecklaces and show that $Q_k(n)$ is proportional to the number of length-$n$ necklaces, which is the key property needed to achieve constant amortized time. We also apply our results to efficiently generate a well-known de Bruijn sequence and efficiently generate necklaces and Lyndon words subject to a weight constraint.
An extremal problem for completely unclustered Burrows-Wheeler images
The Burrows--Wheeler transform is usually viewed as a clustering transform: it tends to group equal letters into long runs. We study the opposite extremal regime, where the BWT output is completely unclustered, that is, has as many equal-letter runs as positions. Known results imply, on the one hand, that the number of runs in the BWT of a Lyndon word can increase by at most a factor of two, and, on the other hand, that over every alphabet of size at least three completely unclustered BWT images exist in every length. This leads to the extremal problem lying between these two facts. For \(k\ge3\), let \(U_k(n)\) be the minimum cyclic run number of a primitive necklace of length \(n\) whose BWT has \(n\) runs.
We prove the universal lower bound \(U_k(n)\ge\lceil n/2\rceil\), reduce the sharpness problem for one-cycle BWT images \(L\) to the Hamming identity \[
\cruns(\BWT^{-1}(L))=\dH(L,\sort(L)), \] and develop a natural multiset-of-necklaces relaxation with an explicit constant-cycle correction. We compute the small values, including the exceptional value \(U_k(6)=4\), prove a parity obstruction for the Parikh vectors of sharp examples, and determine the multiset relaxation exactly. Finally, for every prime \(p\equiv5\pmod8\) for which \(2\) is a primitive root modulo \(p\), we prove sharpness in the adjacent lengths \(p-1\) and \(p\). Under the corresponding Artin-type infinitude hypothesis, this gives infinitely many adjacent sharp pairs.
Chains of affine standard Lyndon words
In this note, we establish the periodicity of chains of affine standard Lyndon words in all types and determine tight bounds on that periodicity, greatly generalizing the $A$-type results of arXiv:2305.16299. Our approach crucially utilizes the convexity and monotonicity of arXiv:2505.15432 together with the new idea to consider the polarization of the root system given by increasing and decreasing chains.
Applications of Combinatorics on Words with Symbolic Dynamics
In this paper, we explore applications of combinatorics on words across various domains, including data compression, error detection, cryptographic protocols, and pseudorandom number generation. The examination of the theoretical foundations enabling these applications, emphasizing important concepts of mathematical relationships and algorithms. In data compression, we discuss the Lempel-Ziv family of algorithms and Lyndon factorization, with the number of Lyndon words of length \( n \) over an alphabet of size \( k \) given by \[ L(n,k) = \frac{1}{n} \sum_{d|n} μ(d) k^{n/d}. \] We address cryptographic protocols and pseudorandom number generation, highlighting the role of pseudorandomness theory and complexity measures. Also, by explore de Bruijn sequences, topological entropy, and synchronizing words in their practical contexts, demonstrating their contributions to optimizing information storage, ensuring data integrity, and enhancing cybersecurity.
Affine standard Lyndon words
In this note, we establish the convexity and monotonicity for affine standard Lyndon words in all types, generalizing the $A$-type results of arXiv:2305.16299. We also derive partial results on the structure of imaginary standard Lyndon words and present a conjecture for their general form. Additionally, we provide computer code in Appendix which, in particular, allows to efficiently compute affine standard Lyndon words in exceptional types for all orders.
Number of partitions of modular integers (with an Appendix by P. Deligne)
For integers $n,k,s$, we give a formula for the number $T(n,k,s)$ of order $k$ subsets of the ring $\mathbb{Z}/n\mathbb{Z}$ whose sum of elements is $s$ modulo $n$. To do so, we describe explicitly a sequence of matrices $M(k)$, for positive integers $k$, such that the size of $M(k)$ is the number of divisors of $k$, and for two coprime integers $k_{1},k_{2}$, the matrix $M(k_{1}k_{2})$ is the Kronecker product of $M(k_{1})$ and $M(k_{2})$. For $s=0, 1, 2$, and for $s=k/2$ when $k$ is even, the sequences $T(n,k,s)$ are related to the number of necklaces with $k$ black beads and $n-k$ white beads, and to Lyndon words. This work begins with empirical determinations of $M(k)$ up to $k=10000$, from which we infer a closed formula that encompasses many entries in the Encyclopedia of Integer Sequences. Its proof comes from work on Ramanujan sums, by Ramanathan, with a generalization to wider problems linked to representation theory and recently described by Deligne.
Periodic orbits on 2-regular circulant digraphs
Periodic orbits (equivalence classes of closed paths up to cyclic shifts) play an important role in applications of graph theory. For example, they appear in the definition of the Ihara zeta function and exact trace formulae for the spectra of quantum graphs. Circulant graphs are Cayley graphs of $\mathbb{Z}_n$. Here we consider directed Cayley graphs with two generators (2-regular Cayley digraphs). We determine the number of primitive periodic orbits of a given length (total number of directed edges) in terms of the number of times edges corresponding to each generator appear in the periodic orbit (the step count). Primitive periodic orbits are those periodic orbits that cannot be written as a repetition of a shorter orbit. We describe the lattice structure of lengths and step counts for which periodic orbits exist and characterize the repetition number of a periodic orbit by its winding number (the sum of the step sequence divided by the number of vertices) and the repetition number of its step sequence. To obtain these results, we also evaluate the number of Lyndon words on an alphabet of two letters with a given length and letter count.
Standard Lyndon loop words: weighted orders
Published in International Mathematics Research Notices (2025)
• View Publication
• BIB
We generalize the study of standard Lyndon loop words from [A.Negut, A.Tsymbaliuk, "Quantum loop groups and shuffle algebras via Lyndon words", Adv. Math. 439 (2024), Paper No. 109482] to a more general class of orders on the underlying alphabet, as suggested in Remark 3.15 of loc.cit. The main new ingredient is the exponent-tightness of these words, which also allows to generalize the construction of PBW bases of the untwisted quantum loop algebra via the combinatorics of loop words.
Perfectly Clustering Words and Iterated Palindromes over a Ternary Alphabet
Published in EPTCS 403, 2024, pp. 134-138
• View Publication
• BIB
Recently, a new characterization of Lyndon words that are also perfectly clustering was proposed by Lapointe and Reutenauer (2024). A word over a ternary alphabet {a,b,c} is called perfectly clustering Lyndon if and only if it is the product of two palindromes and it can be written as apbqc where p and q are palindromes. We study the properties of palindromes appearing as factors p and q and their links with iterated palindromes over a ternary alphabet.
Lyndon pairs and the lexicographically greatest perfect necklace
Published in Comb. Number Th. 13 (2024) 361-375
• View Publication
• BIB
Fix a finite alphabet. A necklace is a circular word. For positive integers $n$ and~$k$, a necklace is $(n,k)$-perfect if all words of length $n$ occur $k$ times but at positions with different congruence modulo $k$, for any convention of the starting position. We define the notion of a Lyndon pair and we use it to construct the lexicographically greatest $(n,k)$-perfect necklace, for any $n$ and $k$ such that $n$ divides~$k$ or $k$ divides~$n$. Our construction generalizes Fredricksen and Maiorana's construction of the lexicographically greatest de Bruijn sequence of order $n$, based on the concatenation of the Lyndon words whose length divide $n$.
From the Lyndon factorization to the Canonical Inverse Lyndon factorization: back and forth
The notion of inverse Lyndon word is related to the classical notion of Lyndon word. More precisely, inverse Lyndon words are all and only the nonempty prefixes of the powers of the anti-Lyndon words, where an anti-Lyndon word with respect to a lexicographical order is a classical Lyndon word with respect to the inverse lexicographic order. Each word $w$ admits a factorization in inverse Lyndon words, named the canonical inverse Lyndon factorization $\ICFL(w)$, which maintains the main properties of the Lyndon factorization of $w$. Although there is a huge literature on the Lyndon factorization, the relation between the Lyndon factorization $\CFL_{in}$ with respect to the inverse order and the canonical inverse Lyndon factorization $\ICFL$ has not been thoroughly investigated. In this paper, we address this question and we show how to obtain one factorization from the other via the notion of grouping. This result naturally opens new insights in the investigation of the relationship between $\ICFL$ and other notions, e.g., variants of Burrows Wheeler Transform, as already done for the Lyndon factorization.
The dimension of the region of feasible tournament profiles
Erd\H os, Lovász and Spencer showed in the late 1970s that the dimension of the region of $k$-vertex graph profiles, i.e., the region of feasible densities of $k$-vertex graphs in large graphs, is equal to the number of non-trivial connected graphs with at most $k$ vertices. We determine the dimension of the region of $k$-vertex tournament profiles. Our result, which explores an interesting connection to Lyndon words, yields that the dimension is much larger than just the number of strongly connected tournaments, which would be the answer expected as the analogy to the setting of graphs.
The dimension of the feasible region of pattern densities
Published in Math. Proc. Camb. Phil. Soc. 178 (2025) 1-14
• View Publication
• BIB
A classical result of Erdős, Lovász and Spencer from the late 1970s asserts that the dimension of the feasible region of densities of graphs with at most k vertices in large graphs is equal to the number of non-trivial connected graphs with at most k vertices. Indecomposable permutations play the role of connected graphs in the realm of permutations, and Glebov et al. showed that pattern densities of indecomposable permutations are independent, i.e., the dimension of the feasible region of densities of permutation patterns of size at most k is at least the number of non-trivial indecomposable permutations of size at most k. However, this lower bound is not tight already for k=3. We prove that the dimension of the feasible region of densities of permutation patterns of size at most k is equal to the number of non-trivial Lyndon permutations of size at most k. The proof exploits an interplay between algebra and combinatorics inherent to the study of Lyndon words.
Concatenation trees: A framework for efficient universal cycle and de Bruijn sequence constructions
Classic cycle-joining techniques have found widespread application in creating universal cycles for a diverse range of combinatorial objects, such as shorthand permutations, weak orders, orientable sequences, and various subsets of $k$-ary strings, including de Bruijn sequences. In the most favorable scenarios, these algorithms operate with a space complexity of $O(n)$ and require $O(n)$ time to generate each symbol in the sequences. In contrast, concatenation-based methods have been developed for a limited selection of universal cycles. In each of these instances, the universal cycles can be generated far more efficiently, with an amortized time complexity of $O(1)$ per symbol, while still using $O(n)$ space.
This paper introduces $\mathit{concatenation~trees}$, which serve as the fundamental structures needed to bridge the gap between cycle-joining constructions based on the pure cycle register and corresponding concatenation-based approaches. They immediately demystify the relationship between the classic Lyndon word concatenation construction of de Bruijn sequences and a corresponding cycle-joining based construction. To underscore their significance, concatenation trees are applied to construct universal cycles for shorthand permutations and weak orders in $O(1)$-amortized time per symbol. Moreover, we provide insights as to how similar results can be obtained for other universal cycles including cut-down de Bruijn sequences and orientable sequences.
Decompositions of Nonlinear Input-Output Systems to Zero the Output
Published in Systems & Control Letters, Volume 187, May 2024, 105783
• View Publication
• BIB
Consider an input-output system where the output is the tracking error given some desired reference signal. It is natural to consider under what conditions the problem has an exact solution, that is, the tracking error is exactly the zero function. If the system has a well defined relative degree and the zero function is in the range of the input-output map, then it is well known that the system is locally left invertible, and thus, the problem has a unique exact solution. A system will fail to have relative degree when more than one exact solution exists. The general goal of this paper is to describe a decomposition of an input-output system having a Chen-Fliess series representation into a parallel product of subsystems in order to identify possible solutions to the problem of zeroing the output. For computational purposes, the focus is on systems whose generating series are polynomials. It is shown that the shuffle algebra on the set of generating polynomials is a unique factorization domain so that any polynomial can be uniquely factored modulo a permutation into its irreducible elements for the purpose of identifying the subsystems in a parallel product decomposition. This is achieved using the fact that this shuffle algebra is isomorphic to the symmetric algebra over the vector space spanned by Lyndon words. A specific algorithm for factoring generating polynomials into its irreducible factors is presented based on the Chen-Fox-Lyndon factorization of words.
Affine Standard Lyndon words: A-type
Published in International Mathematics Research Notices (2024), 37pp
• View Publication
• BIB
We generalize an algorithm of Leclerc describing explicitly the bijection of Lalonde-Ram from finite to affine Lie algebras. In type $A_n^{(1)}$, we compute all affine standard Lyndon words for any order of the simple roots, and establish some properties of the induced orders on the positive affine roots.
Necklaces and bracelets in R
This note introduces a code snippet in R aiming to generate necklaces as well as bracelets. Among various uses, necklaces are useful tools to manage traces of products of random matrices. Functionality for necklaces and bracelets is provided with some examples of applications such as Lyndon words and de Bruijn sequences. The routines are collected in the Necklaces package available from the Comprehensive R Archive Network.
Cut-Down de Bruijn Sequences
Published
• View Publication
• BIB
A cut-down de Bruijn sequence is a cyclic string of length $L$, where $1 \leq L \leq k^n$, such that every substring of length $n$ appears at most once. Etzion [Theor. Comp. Sci 44 (1986)] gives an algorithm to construct binary cut-down de Bruijn sequences that requires $o(n)$ simple $n$-bit operations per symbol generated. In this paper, we simplify the algorithm and improve the running time to $\mathcal{O}(n)$ time per symbol generated using $\mathcal{O}(n)$ space. We then provide the first successor-rule approach for constructing a binary cut-down de Bruijn sequence by leveraging recent ranking algorithms for fixed-density Lyndon words. Finally, we develop an algorithm to generate cut-down de Bruijn sequences for $k>2$ that runs in $\mathcal{O}(n)$ time per symbol using $\mathcal{O}(n)$ space after some initialization. While our $k$-ary algorithm is based on our simplified version of Etzion's binary algorithm, a number of non-trivial adaptations are required to generalize to larger alphabets.
Representations of orientifold Khovanov-Lauda-Rouquier algebras and the Enomoto-Kashiwara algebra
Published in Pacific J. Math. 322 (2023) 407-441
• View Publication
• BIB
We consider an "orientifold" generalization of Khovanov-Lauda-Rouquier algebras, depending on a quiver with an involution and a framing. Their representation theory is related, via a Schur-Weyl duality type functor, to Kac-Moody quantum symmetric pairs, and, via a categorification theorem, to highest weight modules over an algebra introduced by Enomoto and Kashiwara. Our first main result is a new shuffle realization of these highest weight modules and a combinatorial construction of their PBW and canonical bases in terms of Lyndon words. Our second main result is a classification of irreducible representations of orientifold KLR algebras and a computation of their global dimension in the case when the framing is trivial.
Counting Lyndon Subsequences
Counting substrings/subsequences that preserve some property (e.g., palindromes, squares) is an important mathematical interest in stringology. Recently, Glen et al. studied the number of Lyndon factors in a string. A string $w = uv$ is called a Lyndon word if it is the lexicographically smallest among all of its conjugates $vu$. In this paper, we consider a more general problem "counting Lyndon subsequences". We show (1) the maximum total number of Lyndon subsequences in a string, (2) the expected total number of Lyndon subsequences in a string, (3) the expected number of distinct Lyndon subsequences in a string.