arXiv++ Combinatorics

Browse math.CO papers from arXiv

random word

36 papers tagged with this keyword
2025-10-05
Results on long twins in random words and permutations
We study long $r$-twins in random words and permutations. Motivated by questions posed in works of Dudek-Grytczuk-Ruciński, we obtain the following. For a uniform word in $[k]^n$ we prove sharp one-sided tail bounds showing that the maximum $r$-power length (the longest contiguous block that can be partitioned into $r$ identical subblocks) is concentrated around $\frac{\log n}{(r-1)\log k}$. For random permutations, we prove that for fixed $k$ and $r\to\infty$, a uniform permutation of $[rk]$ a.a.s. contains $r$ disjoint increasing subsequences of length $k$, generalizing a previous result that proves this for $k=2$. Finally, we use a computer-aided pattern count to improve the best known lower bound on the length of alternating twins in a random permutation to $α_n \ge \left(\tfrac{1}{3}+0.0989-o(1)\right)n$, strengthening the previous constant.
2024-12-30 v2
Random Fibonacci Words via Clone Schur Functions
Published in Forum of Mathematics, Sigma 14 (2026) e15 • View PublicationBIB
We study positivity and probabilistic properties arising from the Young--Fibonacci lattice $\mathbb{YF}$, a 1-differential poset on binary (Fibonacci) words of 1's and 2's, graded by digit sum. Building on Okada's theory of clone Schur functions (Trans. Amer. Math. Soc. 346 (1994), 549--568), we define clone coherent measures on $\mathbb{YF}$ that generate random Fibonacci words of increasing length; unlike for the Young lattice (powered by the classical Schur functions), clone coherent measures are generally not extremal on $\mathbb{YF}$. Our first main result is a complete characterization of Fibonacci positive specializations -- parameter sequences which yield positive clone Schur functions on $\mathbb{YF}$. Second, we connect Fibonacci positivity with: (i) total positivity of tridiagonal matrices; (ii) Stieltjes moment sequences; (iii) the combinatorics of set partitions; and (iv) families of univariate orthogonal polynomials from the (q-)Askey scheme. We further link moment sequences of orthogonal polynomials to combinatorial structures on Fibonacci words, a connection that may be of independent interest. Third, we analyze scaling limits of the induced random words, obtaining stick-breaking-type limits (linked to GEM laws), new dependent stick-breaking limits, and limits supported on the discrete part of the Martin boundary of $\mathbb{YF}$. These results significantly extend the asymptotics of the Plancherel measure on $\mathbb{YF}$ proved by Gnedin--Kerov (Math. Proc. Camb. Philos. Soc. 129 (2000), 433--446). Finally, we prove Cauchy-type identities for clone Schur functions with quadridiagonal-determinant right-hand side (in contrast to the product form for classical Schur functions), and construct models of random permutations and involutions from Fibonacci-positive specializations together with a Robinson--Schensted correspondence adapted to $\mathbb{YF}$.
2022-07-28 v2
Short Synchronizing Words for Random Automata
Published • View PublicationBIB
We prove that a uniformly random automaton with $n$ states on a 2-letter alphabet has a synchronizing word of length $O(n^{1/2}\log n)$ with high probability (w.h.p.). That is to say, w.h.p. there exists a word $ω$ of such length, and a state $v_0$, such that $ω$ sends all states to $v_0$. Prior to this work, the best upper bound was the quasilinear bound $O(n\log^3n)$ due to Nicaud (2016). The correct scaling exponent had been subject to various estimates by other authors between $0.5$ and $0.56$ based on numerical simulations, and our result confirms that the smallest one indeed gives a valid upper bound (with a log factor). Our proof introduces the concept of $w$-trees, for a word $w$, that is, automata in which the $w$-transitions induce a (loop-rooted) tree. We prove a strong structure result that says that, w.h.p., a random automaton on $n$ states is a $w$-tree for some word $w$ of length at most $(1+ε)\log_2(n)$, for any $ε>0$. The existence of the (random) word $w$ is proved by the probabilistic method. This structure result is key to proving that a short synchronizing word exists.
Long twins in random words
Published • View PublicationBIB
Twins in a finite word are formed by a pair of identical subwords placed at disjoint sets of positions. We investigate the maximum length of twins in a random word over a $k$-letter alphabet. The obtained lower bounds for small values of $k$ significantly improve the best estimates known in the deterministic case. Bukh and Zhou in 2016 showed that every ternary word of length $n$ contains twins of length at least $0.34n$. Our main result states that in a random ternary word of length $n$, with high probability, one can find twins of length at least $0.41n$. In the general case of alphabets of size $k\geq 3$ we obtain analogous lower bounds of the form $\frac{1.64}{k+1}n$ which are better than the known deterministic bounds for $k\leq 354$. In addition, we present similar results for multiple twins in random words.
Asymptotic bit frequency in Fibonacci words
Published • View PublicationBIB
It is known that binary words containing no $k$ consecutive 1s are enumerated by $k$-step Fibonacci numbers. In this note we discuss the expected value of a random bit in a random word of length $n$ having this property.
2020-09-12 v2
Longest common subsequences between words of very unequal length
We consider the expected length of the longest common subsequence between two random words of lengths $n$ and $(1-\varepsilon)kn$ over $k$-symbol alphabet. It is well-known that this quantity is asymptotic to $γ_{k,\varepsilon} n$ for some constant $γ_{k,\varepsilon}$. We show that $γ_{k,\varepsilon}$ is of the order $1-c\varepsilon^2$ uniformly in $k$ and $\varepsilon$. In addition, for large $k$, we give evidence that $γ_{k,\varepsilon}$ approaches $1-\tfrac{1}{4}\varepsilon^2$, and prove a matching lower bound.
2020-03-07 v3
Quasi-random words and limits of word sequences
Published in European Journal of Combinatorics, Volume 98, 2021 • View PublicationBIB
Words are sequences of letters over a finite alphabet. We study two intimately related topics for this object: quasi-randomness and limit theory. With respect to the first topic we investigate the notion of uniform distribution of letters over intervals, and in the spirit of the famous Chung--Graham--Wilson theorem for graphs we provide a list of word properties which are equivalent to uniformity. In particular, we show that uniformity is equivalent to counting 3-letter subsequences. Inspired by graph limit theory we then investigate limits of convergent word sequences, those in which all subsequence densities converge. We show that convergent word sequences have a natural limit, namely Lebesgue measurable functions of the form $f:[0,1]\to[0,1]$. Via this theory we show that every hereditary word property is testable, address the problem of finite forcibility for word limits and establish as a byproduct a new model of random word sequences. Along the lines of the proof of the existence of word limits, we can also establish the existence of limits for higher dimensional structures. In particular, we obtain an alternative proof of the result by Hoppen, Kohayakawa, Moreira, Ráth and Sampaio [{\it J. Combin. Theory Ser. B 103(1):93--113, 2013}] establishing the existence of permutons.
2019-12-07 v3
Periodic words, common subsequences and frogs
Published • View PublicationBIB
Let $W^{(n)}$ be the $n$-letter word obtained by repeating a fixed word $W$, and let $R_n$ be a random $n$-letter word over the same alphabet. We show several results about the length of the longest common subsequence (LCS) between $W^{(n)}$ and $R_n$; in particular, we show that its expectation is $γ_W n-O(\sqrt{n})$ for an efficiently-computable constant $γ_W$. This is done by relating the problem to a new interacting particle system, which we dub "frog dynamics". In this system, the particles (`frogs') hop over one another in the order given by their labels. Stripped of the labeling, the frog dynamics reduces to a variant of the PushTASEP. In the special case when all symbols of $W$ are distinct, we obtain an explicit formula for the constant $γ_W$ and a closed-form expression for the stationary distribution of the associated frog dynamics. In addition, we propose new conjectures about the asymptotic of the LCS of a pair of random words. These conjectures are informed by computer experiments using a new heuristic algorithm to compute the LCS. Through our computations, we found periodic words that are more random-like than a random word, as measured by the LCS.
Staircase patterns in words: subsequences, subwords, and separation number
Published • View PublicationBIB
We revisit staircases for words and prove several exact as well as asymptotic results for longest left-most staircase subsequences and subwords and staircase separation number, the latter being defined as the number of consecutive maximal staircase subwords packed in a word. We study asymptotic properties of the sequence $h_{r,k}(n),$ the number of $n$-array words with $r$ separations over alphabet $[k]$ and show that for any $r\geq 0,$ the growth sequence $\big(h_{r,k}(n)\big)^{1/n}$ converges to a characterized limit, independent of $r.$ In addition, we study the asymptotic behavior of the random variable $\mathcal{S}_k(n),$ the number of staircase separations in a random word in $[k]^n$ and obtain several limit theorems for the distribution of $\mathcal{S}_k(n),$ including a law of large numbers, a central limit theorem, and the exact growth rate of the entropy of $\mathcal{S}_k(n).$ Finally, we obtain similar results, including growth limits, for longest $L$-staircase subwords and subsequences.
2019-04-17
Circularly squarefree words and unbordered conjugates: a new approach
Using a new approach based on automatic sequences, logic, and a decision procedure, we reprove some old theorems about circularly squarefree words and unbordered conjugates in a new and simpler way. Furthermore, we prove three new results about unbordered conjugates: we complete the classification, due to Harju and Nowotka, of binary words with the maximum number of unbordered conjugates; we prove that for every possible number, up to the maximum, there exists a word having that number of unbordered conjugates, and finally, we determine the expected number of unbordered conjugates in a random word.
2017-05-15 v2
A $q$-deformation of the symplectic Schur functions and the Berele insertion algorithm
Published • View PublicationBIB
A randomisation of the Berele insertion algorithm is proposed, where the insertion of a letter to a symplectic Young tableau leads to a distribution over the set of symplectic Young tableaux. Berele's algorithm provides a bijection between words from an alphabet and a symplectic Young tableau along with a recording oscillating tableau. The randomised version of the algorithm is achieved by introducing a parameter $0 < q < 1$. The classic Berele algorithm corresponds to letting the parameter $q \to 0$. The new version provides a probabilistic framework that allows to prove Littlewood-type identities for a $q$-deformation of the symplectic Schur functions. These functions correspond to multilevel extensions of the continuous $q$-Hermite polynomials. Finally, we show that when both the original and the $q$-modified insertion algorithms are applied to a random word then the shape of the symplectic Young tableau evolves as a Markov chain on the set of partitions.
2016-05-11 v2
Doob-Martin compactification of a Markov chain for growing random words sequentially
Published • View PublicationBIB
We consider a Markov chain that iteratively generates a sequence of random finite words in such a way that the $n^{\mathrm{th}}$ word is uniformly distributed over the set of words of length $2n$ in which $n$ letters are $a$ and $n$ letters are $b$: at each step an $a$ and a $b$ are shuffled in uniformly at random among the letters of the current word. We obtain a concrete characterization of the Doob-Martin boundary of this Markov chain. Writing $N(u)$ for the number of letters $a$ (equivalently, $b$) in the finite word $u$, we show that a sequence $(u_n)_{n \in \mathbb{N}}$ of finite words converges to a point in the boundary if, for an arbitrary word $v$, there is convergence as $n$ tends to infinity of the probability that the selection of $N(v)$ letters $a$ and $N(v)$ letters $b$ uniformly at random from $u_n$ and maintaining their relative order results in $v$. We exhibit a bijective correspondence between the points in the boundary and ergodic random total orders on the set $\{a_1, b_1, a_2, b_2, \ldots \}$ that have distributions which are separately invariant under finite permutations of the indices of the $a'$s and those of the $b'$s. We establish a further bijective correspondence between the set of such random total orders and the set of pairs $(μ,ν)$ of diffuse probability measures on $[0,1]$ such that $\frac{1}{2}(μ+ν)$ is Lebesgue measure: the restriction of the random total order to $\{a_1, b_1, \ldots, a_n, b_n\}$ is obtained by taking $X_1, \ldots, X_n$ (resp. $Y_1, \ldots, Y_n$) i.i.d. with common distribution $μ$ (resp. $ν$), letting $(Z_1, \ldots, Z_{2n})$ be $\{X_1, Y_1, \ldots, X_n, Y_n\}$ in increasing order, and declaring that the $k^{\mathrm{th}}$ smallest element in the restricted total order is $a_i$ (resp. $b_j$) if $Z_k = X_i$ (resp. $Z_k = Y_j$).
2015-12-17 v2
A Central Limit Theorem for the Optimal Alignments Score in Multiple Random Words
Let $\mathbf{X}^{(1)}_{n},\ldots,\mathbf{X}^{(m)}_{n}$, where $\mathbf{X}^{(i)}_{n}=(X^{(i)}_{1},\ldots,X^{(i)}_{n})$, $i=1,\ldots,m$, be $m$ independent sequences of independent and identically distributed random variables taking their values in a finite alphabet $\mathcal{A}$. Let the score function $S$, defined on $\mathcal{A}^{m}$, be non-negative, bounded, permutation-invariant, and satisfy a bounded differences condition. Under a variance lower-bound assumption, a central limit theorem is proved for the optimal alignments score of the $m$ random words.
2015-09-17
Periods and borders of random words
Published in STACS 2016, LIPIcs 47, 44:1-44:10 • View PublicationBIB
We investigate the behavior of the periods and border lengths of random words over a fixed alphabet. We show that the asymptotic probability that a random word has a given maximal border length $k$ is a constant, depending only on $k$ and the alphabet size $\ell$. We give a recurrence that allows us to determine these constants with any required precision. This also allows us to evaluate the expected period of a random word. For the binary case, the expected period is asymptotically about $n-1.641$. We also give explicit formulas for the probability that a random word is unbordered or has maximum border length one.
2015-09-15
Toward the Combinatorial Limit Theory of Free Words
Free words are elements of a free monoid, generated over an alphabet via the binary operation of concatenation. Casually speaking, a free word is a finite string of letters. Henceforth, we simply refer to them as words. Motivated by recent advances in the combinatorial limit theory of graphs-notably those involving flag algebras, graph homomorphisms, and graphons-we investigate the extremal and asymptotic theory of pattern containment and avoidance in words. Word V is a factor of word W provided V occurs as consecutive letters within W. W is an instance of V provided there exists a nonerasing monoid homomorphsism φ with φ(V) = W. For example, using the homomorphism φ defined by φ(P) = Ror, φ(h) = a, and φ(D) = baugh, we see that Rorabaugh is an instance of PhD. W avoids V if no factor of W is an instance of V. V is unavoidable provided, over any finite alphabet, there are only finitely many words that avoid V. Unavoidable words were classified by Bean, Ehrenfeucht, and McNulty (1979) and Zimin (1982). We briefly address the following Ramsey-theoretic question: For unavoidable word V and a fixed alphabet, what is the longest a word can be that avoids V? The density of V in W is the proportion of nonempty substrings of W that are instances of V. Since there are 45 substrings in Rorabaugh and 28 of them are instances of PhD, the density of PhD in Rorabaugh is 28/45. We establish a number of asymptotic results for word densities, including the expected density of a word in arbitrarily long, random words and the minimum density of an unavoidable word over arbitrarily long words. This is joint work with Joshua Cooper.
2015-05-29
The Number of Distinct Subpalindromes in Random Words
Published • View PublicationBIB
We prove that a random word of length $n$ over a $k$-ary fixed alphabet contains, on expectation, $Θ(\sqrt{n})$ distinct palindromic factors. We study this number of factors, $E(n,k)$, in detail, showing that the limit $\lim_{n\to\infty}E(n,k)/\sqrt{n}$ does not exist for any $k\ge2$, $\liminf_{n\to\infty}E(n,k)/\sqrt{n}=Θ(1)$, and $\limsup_{n\to\infty}E(n,k)/\sqrt{n}=Θ(\sqrt{k})$. Such a complicated behaviour stems from the asymmetry between the palindromes of even and odd length. We show that a similar, but much simpler, result on the expected number of squares in random words holds. We also provide some experimental data on the number of palindromic factors in random words.
2015-04-17 v2
Density dichotomy in random words
Published • View PublicationBIB
Word $W$ is said to encounter word $V$ provided there is a homomorphism $φ$ mapping letters to nonempty words so that $φ(V)$ is a substring of $W$. For example, taking $φ$ such that $φ(h)=c$ and $φ(u)=ien$, we see that "science" encounters "huh" since $cienc=φ(huh)$. The density of $V$ in $W$, $δ(V,W)$, is the proportion of substrings of $W$ that are homomorphic images of $V$. So the density of "huh" in "science" is $2/{8 \choose 2}$. A word is doubled if every letter that appears in the word appears at least twice. The dichotomy: Let $V$ be a word over any alphabet, $Σ$ a finite alphabet with at least 2 letters, and $W_n \in Σ^n$ chosen uniformly at random. Word $V$ is doubled if and only if $\mathbb{E}(δ(V,W_n)) \rightarrow 0$ as $n \rightarrow \infty$. We further explore convergence for nondoubled words and concentration of the limit distribution for doubled words around its mean.
Application of Smirnov Words to Waiting Time Distributions of Runs
Published in Electron. J. Combin., 24 (3), #P3.55, 2017 • View PublicationBIB
Consider infinite random words over a finite alphabet where the letters occur as an i.i.d. sequence according to some arbitrary distribution on the alphabet. The expectation and the variance of the waiting time for the first completed $h$-run of any letter (i.e., first occurrence of $h$ subsequential equal letters) is computed. The expected waiting time for the completion of $h$-runs of $j$ arbitrary distinct letters is also given.
2014-08-07 v6
A Central Limit Theorem for the Length of the Longest Common Subsequences in Random Words
Published in Electronic Journal of Probability 2023, Vol. 28, paper no. 3, 1-24 • View PublicationBIB
Let $(X_i)_{i \geq 1}$ and $(Y_i)_{i\geq1}$ be two independent sequences of independent identically distributed random variables taking their values in a common finite alphabet and having the same law. Let $LC_n$ be the length of the longest common subsequences of the two random words $X_1\cdots X_n$ and $Y_1\cdots Y_n$. Under a lower bound assumption on the order of its variance, $LC_n$ is shown to satisfy a central limit theorem. This is in contrast to the limiting distribution of the length of the longest common subsequences in two independent uniform random permutations of $\{1, \dots, n\}$, which is shown to be the Tracy-Widom distribution.
2013-03-10
Descent-Inversion Statistics in Riffle Shuffles
This paper studies statistics of riffle shuffles by relating them to random word statistics with the use of inverse shuffles. Asymptotic normality of the number of descents and inversions in riffle shuffles with convergence rates of order $1/\sqrt{n}$ in the Kolmogorov distance are proven. Results are also given about the lengths of the longest alternating subsequences of random permutations resulting from riffle shuffles. A sketch of how the theory of multisets can be useful for statistics of a variation of top $m$ to random shuffles is presented.