arXiv++ Combinatorics

Browse math.CO papers from arXiv

Papers by Jeffrey Shallit

126 paper(s) by this author · All BibTeX
2013-04-12
Sets Represented as the Length-n Factors of a Word
Published • View PublicationBIB
In this paper we consider the following problems: how many different subsets of Sigma^n can occur as set of all length-n factors of a finite word? If a subset is representable, how long a word do we need to represent it? How many such subsets are represented by words of length t? For the first problem, we give upper and lower bounds of the form alpha^(2^n) in the binary case. For the second problem, we give a weak upper bound and some experimental data. For the third problem, we give a closed-form formula in the case where n <= t < 2n. Algorithmic variants of these problems have previously been studied under the name "shortest common superstring".
2013-04-10
Shortest Repetition-Free Words Accepted by Automata
Published • View PublicationBIB
We consider the following problem: given that a finite automaton $M$ of $N$ states accepts at least one $k$-power-free (resp., overlap-free) word, what is the length of the shortest such word accepted? We give upper and lower bounds which, unfortunately, are widely separated.
2012-12-01 v3
Repetition Avoidance in Circular Factors
Published • View PublicationBIB
We consider the following novel variation on a classical avoidance problem from combinatorics on words: instead of avoiding repetitions in all factors of a word, we avoid repetitions in all factors where each individual factor is considered as a "circular word", i.e., the end of the word wraps around to the beginning. We determine the best possible avoidance exponent for alphabet size 2 and 3, and provide a lower bound for larger alphabets.
2012-11-06
On the Number of Unbordered Factors
We illustrate a general technique for enumerating factors of k-automatic sequences by proving a conjecture on the number f(n) of unbordered factors of the Thue-Morse sequence. We show that f(n) <= n for n >= 4 and that f(n) = n infinitely often. We also give examples of automatic sequences having exactly 2 unbordered factors of every length.
2012-07-21 v3
Primitive Words and Lyndon Words in Automatic and Linearly Recurrent Sequences
We investigate questions related to the presence of primitive words and Lyndon words in automatic and linearly recurrent sequences. We show that the Lyndon factorization of a k-automatic sequence is itself k-automatic. We also show that the function counting the number of primitive factors (resp., Lyndon factors) of length n in a k-automatic sequence is k-regular. Finally, we show that the number of Lyndon factors of a linearly recurrent sequence is bounded.
2012-06-23 v4
Subword Complexity and k-Synchronization
We show that the subword complexity function p_x(n), which counts the number of distinct factors of length n of a sequence x, is k-synchronized in the sense of Carpi if x is k-automatic. As an application, we generalize recent results of Goldstein. We give analogous results for the number of distinct factors of length n that are primitive words or powers. In contrast, we show that the function that counts the number of unbordered factors of length n is not necessarily k-synchronized for k-automatic sequences.
2012-04-23
Enumerating regular expressions and their languages
Published • View PublicationBIB
In this chapter we discuss the problem of enumerating distinct regular expressions by size and the regular languages they represent. We discuss various notions of the size of a regular expression that appear in the literature and their advantages and disadvantages. We consider a formal definition of regular expressions using a context-free grammar. We then show how to enumerate strings generated by an unambiguous context-free grammar using the Chomsky-Schützenberger theorem. This theorem allows one to construct an algebraic equation whose power series expansion provides the enumeration. Classical tools from complex analysis, such as singularity analysis, can then be used to determine the asymptotic behavior of the enumeration. We use these algebraic and analytic methods to obtain asymptotic estimates on the number of regular expressions of size n. A single regular language can often be described by several regular expressions, and we estimate the number of distinct languages denoted by regular expressions of size n. We also give asymptotic estimates for these quantities. For the first few values, we provide exact enumeration results.
Avoiding Three Consecutive Blocks of the Same Size and Same Sum
Published • View PublicationBIB
We show that there exists an infinite word over the alphabet {0, 1, 3, 4} containing no three consecutive blocks of the same size and the same sum. This answers an open problem of Pirillo and Varricchio from 1994.
2011-01-18
Avoiding 3/2-powers over the natural numbers
Published in Discrete Mathematics 312 (2012) 1282-1288 • View PublicationBIB
In this paper we answer the following question: what is the lexicographically least sequence over the natural numbers that avoids 3/2-powers?
Thue-Morse at Multiples of an Integer
Published • View PublicationBIB
Let (t_n) be the classical Thue-Morse sequence defined by t_n = s_2(n) (mod 2), where s_2 is the sum of the bits in the binary representation of n. It is well known that for any integer k>=1 the frequency of the letter "1" in the subsequence t_0, t_k, t_{2k}, ... is asymptotically 1/2. Here we prove that for any k there is a n<=k+4 such that t_{kn}=1. Moreover, we show that n can be chosen to have Hamming weight <=3. This is best in a twofold sense. First, there are infinitely many k such that t_{kn}=1 implies that n has Hamming weight >=3. Second, we characterize all k where the minimal n equals k, k+1, k+2, k+3, or k+4. Finally, we present some results and conjectures for the generalized problem, where s_2 is replaced by s_b for an arbitrary base b>=2.
2010-08-14 v2
Inverse Star, Borders, and Palstars
Published • View PublicationBIB
A language L is closed if L = L*. We consider an operation on closed languages, L-*, that is an inverse to Kleene closure. It is known that if L is closed and regular, then L-* is also regular. We show that the analogous result fails to hold for the context-free languages. Along the way we find a new relationship between the unbordered words and the prime palstars of Knuth, Morris, and Pratt. We use this relationship to enumerate the prime palstars, and we prove that neither the language of all unbordered words nor the language of all prime palstars is context-free.
2009-01-12
Avoiding Squares and Overlaps Over the Natural Numbers
Published • View PublicationBIB
We consider avoiding squares and overlaps over the natural numbers, using a greedy algorithm that chooses the least possible integer at each step; the word generated is lexicographically least among all such infinite words. In the case of avoiding squares, the word is 01020103..., the familiar ruler function, and is generated by iterating a uniform morphism. The case of overlaps is more challenging. We give an explicitly-defined morphism phi : N* -> N* that generates the lexicographically least infinite overlap-free word by iteration. Furthermore, we show that for all h,k in N with h <= k, the word phi^{k-h}(h) is the lexicographically least overlap-free word starting with the letter h and ending with the letter k, and give some of its symmetry properties.
2008-12-12 v5
Van der Waerden's Theorem and Avoidability in Words
Published • View PublicationBIB
Pirillo and Varricchio, and independently, Halbeisen and Hungerbuhler considered the following problem, open since 1994: Does there exist an infinite word w over a finite subset of Z such that w contains no two consecutive blocks of the same length and sum? We consider some variations on this problem in the light of van der Waerden's theorem on arithmetic progressions.
2008-08-19 v2
Morphic and Automatic Words: Maximal Blocks and Diophantine Approximation
Published • View PublicationBIB
Let $\mb w$ be a morphic word over a finite alphabet $Σ$, and let $Δ$ be a nonempty subset of $Σ$. We study the behavior of maximal blocks consisting only of letters from $Δ$ in $\mb w$, and prove the following: let $(i_k,j_k)$ denote the starting and ending positions, respectively, of the $k$'th maximal $Δ$-block in $\mb w$. Then $\limsup_{k\to\infty} (j_k/i_k)$ is algebraic if $\mb w$ is morphic, and rational if $\mb w$ is automatic. As a result, we show that the same conclusion holds if $(i_k,j_k)$ are the starting and ending positions of the $k$'th maximal zero block, and, more generally, of the $k$'th maximal $x$-block, where $x$ is an arbitrary word. This enables us to draw conclusions about the irrationality exponent of automatic and morphic numbers. In particular, we show that the irrationality exponent of automatic (resp., morphic) numbers belonging to a certain class that we define is rational (resp., algebraic).
2007-10-05 v2
Hamming Distance for Conjugates
Published • View PublicationBIB
Let x, y be strings of equal length. The Hamming distance h(x,y) between x and y is the number of positions in which x and y differ. If x is a cyclic shift of y, we say x and y are conjugates. We consider f(x,y), the Hamming distance between the conjugates xy and yx. Over a binary alphabet f(x,y) is always even, and must satisfy a further technical condition. By contrast, over an alphabet of size 3 or greater, f(x,y) can take any value between 0 and |x|+|y|, except 1; furthermore, we can always assume that the smaller string has only one type of letter.
2007-08-23
The Frobenius Problem in a Free Monoid
The classical Frobenius problem is to compute the largest number g not representable as a non-negative integer linear combination of non-negative integers x_1, x_2, ..., x_k, where gcd(x_1, x_2, ..., x_k) = 1. In this paper we consider generalizations of the Frobenius problem to the noncommutative setting of a free monoid. Unlike the commutative case, where the bound on g is quadratic, we are able to show exponential or subexponential behavior for an analogue of g, depending on the particular measure chosen.
Words avoiding repetitions in arithmetic progressions
Published • View PublicationBIB
Carpi constructed an infinite word over a 4-letter alphabet that avoids squares in all subsequences indexed by arithmetic progressions of odd difference. We show a connection between Carpi's construction and the paperfolding words. We extend Carpi's result by constructing uncountably many words that avoid squares in arithmetic progressions of odd difference. We also construct infinite words avoiding overlaps and infinite words avoiding arbitrarily large squares in arithmetic progressions of odd difference. We use these words to construct labelings of the 2-dimensional integer lattice such that any line through the lattice encounters a squarefree (resp. overlapfree) sequence of labels.
2005-11-16
Binary words containing infinitely many overlaps
Published • View PublicationBIB
We characterize the squares occurring in infinite overlap-free binary words and construct various alpha power-free binary words containing infinitely many overlaps.
2003-11-07
Words avoiding reversed subwords
We examine words w satisfying the following property: if x is a subword of w and |x| is at least k for some fixed k, then the reversal of x is not a subword of w.
2003-10-10
A Generalization of Repetition Threshold
Published • View PublicationBIB
Brandenburg and (implicitly) Dejean introduced the concept of repetition threshold: the smallest real number alpha such that there exists an infinite word over a k-letter alphabet that avoids beta-powers for all beta>alpha. We generalize this concept to include the lengths of the avoided words. We give some conjectures supported by numerical evidence and prove one of these conjectures.