arXiv++ Combinatorics

Browse math.CO papers from arXiv

Papers by Arseny M. Shur

14 paper(s) by this author · All BibTeX
2024-10-22 v2
Expected Density of Random Minimizers
Minimizer schemes, or just minimizers, are a very important computational primitive in sampling and sketching biological strings. Assuming a fixed alphabet of size $σ$, a minimizer is defined by two integers $k,w\ge2$ and a total order $ρ$ on strings of length $k$ (also called $k$-mers). A string is processed by a sliding window algorithm that chooses, in each window of length $w+k-1$, its minimal $k$-mer with respect to $ρ$. A key characteristic of the minimizer is the expected density of chosen $k$-mers among all $k$-mers in a random infinite $σ$-ary string. Random minimizers, in which the order $ρ$ is chosen uniformly at random, are often used in applications. However, little is known about their expected density $\mathcal{DR}_σ(k,w)$ besides the fact that it is close to $\frac{2}{w+1}$ unless $w\gg k$. We first show that $\mathcal{DR}_σ(k,w)$ can be computed in $O(kσ^{k+w})$ time. Then we attend to the case $w\le k$ and present a formula that allows one to compute $\mathcal{DR}_σ(k,w)$ in just $O(w \log w)$ time. Further, we describe the behaviour of $\mathcal{DR}_σ(k,w)$ in this case, establishing the connection between $\mathcal{DR}_σ(k,w)$, $\mathcal{DR}_σ(k+1,w)$, and $\mathcal{DR}_σ(k,w+1)$. In particular, we show that $\mathcal{DR}_σ(k,w)<\frac{2}{w+1}$ (by a tiny margin) unless $w$ is small. We conclude with some partial results and conjectures for the case $w>k$.
2023-10-23 v3
Power-free Complementary Binary Morphisms
We revisit the topic of power-free morphisms, focusing on the properties of the class of complementary morphisms. Such morphisms are defined over a $2$-letter alphabet, and map the letters 0 and 1 to complementary words. We prove that every prefix of the famous Thue-Morse word $\mathbf{t}$ gives a complementary morphism that is $3^+$-free and hence $α$-free for every real number $α>3$. We also describe, using a 4-state binary finite automaton, the lengths of all prefixes of $\mathbf{t}$ that give cubefree complementary morphisms. Next, we show that $3$-free (cubefree) complementary morphisms of length $k$ exist for all $k\not\in \{3,6\}$. Moreover, if $k$ is not of the form $3\cdot2^n$, then the images of letters can be chosen to be factors of $\mathbf{t}$. Finally, we observe that each cubefree complementary morphism is also $α$-free for some $α<3$; in contrast, no binary morphism that maps each letter to a word of length 3 (resp., a word of length 6) is $α$-free for any $α<3$. In addition to more traditional techniques of combinatorics on words, we also rely on the Walnut theorem-prover. Its use and limitations are discussed.
2023-08-29
Distance Labeling for Families of Cycles
For an arbitrary finite family of graphs, the distance labeling problem asks to assign labels to all nodes of every graph in the family in a way that allows one to recover the distance between any two nodes of any graph from their labels. The main goal is to minimize the number of unique labels used. We study this problem for the families $\mathcal{C}_n$ consisting of cycles of all lengths between 3 and $n$. We observe that the exact solution for directed cycles is straightforward and focus on the undirected case. We design a labeling scheme requiring $\frac{n\sqrt{n}}{\sqrt{6}}+O(n)$ labels, which is almost twice less than is required by the earlier known scheme. Using the computer search, we find an optimal labeling for each $n\le 17$, showing that our scheme gives the results that are very close to the optimum.
On minimal critical exponent of balanced sequences
We study the threshold between avoidable and unavoidable repetitions in infinite balanced sequences over finite alphabets. The conjecture stated by Rampersad, Shallit and Vandomme says that the minimal critical exponent of balanced sequences over the alphabet of size $d \geq 5$ equals $\frac{d-2}{d-3}$. This conjecture is known to hold for $d\in \{5, 6, 7,8,9,10\}$. We refute this conjecture by showing that the picture is different for bigger alphabets. We prove that critical exponents of balanced sequences over an alphabet of size $d\geq 11$ are lower bounded by $\frac{d-1}{d-2}$ and this bound is attained for all even numbers $d\geq 12$. According to this result, we conjecture that the least critical exponent of a balanced sequence over $d$ letters is $\frac{d-1}{d-2}$ for all $d\geq 11$.
2021-05-06
Branching Frequency and Markov Entropy of Repetition-Free Languages
Published • View PublicationBIB
We define a new quantitative measure for an arbitrary factorial language: the entropy of a random walk in the prefix tree associated with the language; we call it Markov entropy. We relate Markov entropy to the growth rate of the language and to the parameters of branching of its prefix tree. We show how to compute Markov entropy for a regular language. Finally, we develop a framework for experimental study of Markov entropy by modelling random walks and present the results of experiments with power-free and Abelian-power-free languages.
2018-12-28
Transition Property For Cube-Free Words
Published • View PublicationBIB
We study cube-free words over arbitrary non-unary finite alphabets and prove the following structural property: for every pair $(u,v)$ of $d$-ary cube-free words, if $u$ can be infinitely extended to the right and $v$ can be infinitely extended to the left respecting the cube-freeness property, then there exists a "transition" word $w$ over the same alphabet such that $uwv$ is cube free. The crucial case is the case of the binary alphabet, analyzed in the central part of the paper. The obtained "transition property", together with the developed technique, allowed us to solve cube-free versions of three old open problems by Restivo and Salemi. Besides, it has some further implications for combinatorics on words; e.g., it implies the existence of infinite cube-free words of very big subword (factor) complexity.
2018-01-16
Subword complexity and power avoidance
Published • View PublicationBIB
We begin a systematic study of the relations between subword complexity of infinite words and their power avoidance. Among other things, we show that -- the Thue-Morse word has the minimum possible subword complexity over all overlap-free binary words and all $(\frac 73)$-power-free binary words, but not over all $(\frac 73)^+$-power-free binary words; -- the twisted Thue-Morse word has the maximum possible subword complexity over all overlap-free binary words, but no word has the maximum subword complexity over all $(\frac 73)$-power-free binary words; -- if some word attains the minimum possible subword complexity over all square-free ternary words, then one such word is the ternary Thue word; -- the recently constructed 1-2-bonacci word has the minimum possible subword complexity over all \textit{symmetric} square-free ternary words.
Lower Bounds on Words Separation: Are There Short Identities in Transformation Semigroups?
Published • View PublicationBIB
The words separation problem, originally formulated by Goralcik and Koubek (1986), is stated as follows. Let $Sep(n)$ be the minimum number such that for any two words of length $\le n$ there is a deterministic finite automaton with $Sep(n)$ states, accepting exactly one of them. The problem is to find the asymptotics of the function $Sep$. This problem is inverse to finding the asymptotics of the length of the shortest identity in full transformation semigroups $T_k$. The known lower bound on $Sep$ stems from the unary identity in $T_k$. We find the first series of identities in $T_k$ which are shorter than the corresponding unary identity for infinitely many values of $k$, and thus slightly improve the lower bound on $Sep(n)$. Then we present some short positive identities in symmetric groups, improving the lower bound on separating words by permutational automata by a multiplicative constant. Finally, we present the results of computer search for short identities for small $k$.
2015-05-29
The Number of Distinct Subpalindromes in Random Words
Published • View PublicationBIB
We prove that a random word of length $n$ over a $k$-ary fixed alphabet contains, on expectation, $Θ(\sqrt{n})$ distinct palindromic factors. We study this number of factors, $E(n,k)$, in detail, showing that the limit $\lim_{n\to\infty}E(n,k)/\sqrt{n}$ does not exist for any $k\ge2$, $\liminf_{n\to\infty}E(n,k)/\sqrt{n}=Θ(1)$, and $\limsup_{n\to\infty}E(n,k)/\sqrt{n}=Θ(\sqrt{k})$. Such a complicated behaviour stems from the asymmetry between the palindromes of even and odd length. We show that a similar, but much simpler, result on the expected number of squares in random words holds. We also provide some experimental data on the number of palindromic factors in random words.
2014-09-29
On Searching Zimin Patterns
Published in Theoretical Computer Science, Vol. 571 (2015), 50-57 • View PublicationBIB
In the area of pattern avoidability the central role is played by special words called Zimin patterns. The symbols of these patterns are treated as variables and the rank of the pattern is its number of variables. Zimin type of a word $x$ is introduced here as the maximum rank of a Zimin pattern matching $x$. We show how to compute Zimin type of a word on-line in linear time. Consequently we get a quadratic time, linear-space algorithm for searching Zimin patterns in words. Then we how the Zimin type of the length $n$ prefix of the infinite Fibonacci word is related to the representation of $n$ in the Fibonacci numeration system. Using this relation, we prove that Zimin types of such prefixes and Zimin patterns inside them can be found in logarithmic time. Finally, we give some bounds on the function $f(n,k)$ such that every $k$-ary word of length at least $f(n,k)$ has a factor that matches the rank $n$ Zimin pattern.
Binary Patterns in Binary Cube-Free Words: Avoidability and Growth
Published in RAIRO-Theor. Inf. Appl. 48 (2014) 369-389 • View PublicationBIB
The avoidability of binary patterns by binary cube-free words is investigated and the exact bound between unavoidable and avoidable patterns is found. All avoidable patterns are shown to be D0L-avoidable. For avoidable patterns, the growth rates of the avoiding languages are studied. All such languages, except for the overlap-free language, are proved to have exponential growth. The exact growth rates of languages avoiding minimal avoidable patterns are approximated through computer-assisted upper bounds. Finally, a new example of a pattern-avoiding language of polynomial growth is given.
2010-10-26
Combinatorial Characterization of Formal Languages
This paper is an extended abstract of the dissertation presented by the author for the doctoral degree in physics and mathematics (in Russia). The main characteristic studied in the dissertation is combinatorial complexity, which is a "counting" function associated with a language and returning the number of words of given length in this language. For several classes of languages, a variety of problems about combinatorial complexity and its connections to other parameters of languages are studied. A brief introduction to the topic and the formulations of results are presented. No proofs are given; instead, the papers containing the proofs are cited.
2010-09-29
On ternary square-free circular words
Published in Electronic Journal of Combinatorics 2010 Vol. 17(1) #R140 • View PublicationBIB
Circular words are cyclically ordered finite sequences of letters. We give a computer-free proof of the following result by Currie: square-free circular words over the ternary alphabet exist for all lengths $l$ except for 5, 7, 9, 10, 14, and 17. Our proof reveals an interesting connection between ternary square-free circular words and closed walks in the $K_{3{,}3}$ graph. In addition, our proof implies an exponential lower bound on the number of such circular words of length $l$ and allows one to list all lengths $l$ for which such a circular word is unique up to isomorphism.
2010-09-22 v2
Numerical values of the growth rates of power-free languages
Published • View PublicationBIB
We present upper and two-sided bounds of the exponential growth rate for a wide range of power-free languages. All bounds are obtained with the use of algorithms previously developed by the author.