sequence
6845 papers tagged with this keyword
Existence proofs in combinatorics using independence
Published in Mat. Prosveschenie, 19 (2015)
• Search Publication
This note is purely expository and is in Russian. We show how to prove interesting combinatorial results using the local Lovasz lemma. The note is accessible for students having basic knowledge of combinatorics; the notion of independence is defined and the Lovasz lemma is stated and proved. Our exposition follows `Probabilistic methods' of N. Alon and J. Spencer. The main difference is that we show how the proof could have been invented. The material is presented as a sequence of problems, which is peculiar not only to Zen monasteries but also to advanced mathematical education; most problems are presented with hints or solutions.
Partially user-irrepressible sequence sets and conflict-avoiding codes
Published in Des. Codes Cryptogr. (2016) 78:679-691
• View Publication
• BIB
In this paper we give a partial shift version of user-irrepressible sequence sets and conflict-avoiding codes. By means of disjoint difference sets, we obtain an infinite number of such user-irrepressible sequence sets whose lengths are shorter than known results in general. Subsequently, the newly defined partially conflict-avoiding codes are discussed.
Doob--Martin boundary of Rémy's tree growth chain
Published in Ann. Probab. 45 (2017), 225-277
• View Publication
• BIB
Rémy's algorithm is a Markov chain that iteratively generates a sequence of random trees in such a way that the $n^{\mathrm{th}}$ tree is uniformly distributed over the set of rooted, planar, binary trees with $2n+1$ vertices. We obtain a concrete characterization of the Doob--Martin boundary of this transient Markov chain and thereby delineate all the ways in which, loosely speaking, this process can be conditioned to "go to infinity" at large times. A (deterministic) sequence of finite rooted, planar, binary trees converges to a point in the boundary if for each $m$ the random rooted, planar, binary tree spanned by $m+1$ leaves chosen uniformly at random from the $n^{\mathrm{th}}$ tree in the sequence converges in distribution as $n$ tends to infinity -- a notion of convergence that is analogous to one that appears in the recently developed theory of graph limits.
We show that a point in the Doob--Martin boundary may be identified with the following ensemble of objects: a complete separable $\mathbb{R}$-tree that is rooted and binary in a suitable sense, a diffuse probability measure on the $\mathbb{R}$-tree that allows us to make sense of sampling points from it, and a kernel on the $\mathbb{R}$-tree that describes the probability that the first of a given pair of points is below and to the left of their most recent common ancestor while the second is below and to the right. The Doob--Martin boundary corresponds bijectively to the set of extreme points of the closed convex set of normalized nonnegative harmonic functions, in other words, the minimal and full Doob--Martin boundaries coincide. These results are in the spirit of the identification of graphons as limit objects in the theory of graph limits.
Length distribution of sequencing by synthesis: fixed flow cycle model
Published in Journal of mathematical biology 67 (2), 389-410, 2013
• View Publication
• BIB
Sequencing by synthesis is the underlying technology for many next-generation DNA sequencing platforms. We developed a new model, the fixed flow cycle model, to derive the distributions of sequence length for a given number of flow cycles under the general conditions where the nucleotide incorporation is probabilistic and may be incomplete, as in some single-molecule sequencing technologies. Unlike the previous model, the new model yields the probability distribution for the sequence length. Explicit closed form formulas are derived for the mean and variance of the distribution.
Statistical distributions of pyrosequencing
Published in Journal of Computational Biology 16 (1), 31-42, 2009
• View Publication
• BIB
Pyrosequencing is emerging as one of the important next-generation sequencing technologies. We derive the statistical distributions of this technique in terms of nucleotide probabilities of the target sequences. We give exact distributions both for fixed number of flow cycles and for fixed sequence length. Explicit formulas are derived for the mean and variance of these distributions. In both cases, the distributions can be approximated accurately by normal distributions with the same mean and variance. The statistical distributions will be useful for instrument and software development for pyrosequencing platforms.
Statistical distributions of sequencing by synthesis with probabilistic nucleotide incorporation
Published in Journal of Computational Biology. June 2009, 16(6): 817-827
• View Publication
• BIB
Sequencing by synthesis is used in many next-generation DNA sequencing technologies. Some of the technologies, especially those exploring the principle of single-molecule sequencing, allow incomplete nucleotide incorporation in each cycle. We derive statistical distributions for sequencing by synthesis by taking into account the possibility that nucleotide incorporation may not be complete in each flow cycle. The statistical distributions are expressed in terms of nucleotide probabilities of the target sequences and the nucleotide incorporation probabilities for each nucleotide. We give exact distributions both for fixed number of flow cycles and for fixed sequence length. Explicit formulas are derived for the mean and variance of these distributions. The results are generalizations of our previous work for pyrosequencing. Incomplete nucleotide incorporation leads to significant change in the mean and variance of the distributions, but still they can be approximated by normal distributions with the same mean and variance. The results are also generalized to handle sequence context dependent incorporation. The statistical distributions will be useful for instrument and software development for sequencing by synthesis platforms.
Distributions of positive signals in pyrosequencing
Published in Journal of mathematical biology 69 (1), 39-54, 2014
• View Publication
• BIB
Pyrosequencing is one of the important next-generation sequencing technologies. We derive the distribution of the number of positive signals in pyrograms of this sequencing technology as a function of flow cycle numbers and nucleotide probabilities of the target sequences. As for the distribution of sequence length, we also derive the distribution of positive signals for the fixed flow cycle model. Explicit formulas are derived for the mean and variance of the distributions. A simple result for the mean of the distribution is that the mean number of positive signals in a pyrogram is approximately twice the number of flow cycles, regardless of nucleotide probabilities. The statistical distributions will be useful for instrument and software development for pyrosequencing and other related platforms.
Frustrated Triangles
Published
• View Publication
• BIB
A triple of vertices in a graph is a \emph{frustrated triangle} if it induces an odd number of edges. We study the set $F_n\subset[0,\binom{n}{3}]$ of possible number of frustrated triangles $f(G)$ in a graph $G$ on $n$ vertices. We prove that about two thirds of the numbers in $[0,n^{3/2}]$ cannot appear in $F_n$, and we characterise the graphs $G$ with $f(G)\in[0,n^{3/2}]$. More precisely, our main result is that, for each $n\geq 3$, $F_n$ contains two interlacing sequences $0=a_0\leq b_0\leq a_1\leq b_1\leq \dots \leq a_m\leq b_m\sim n^{3/2}$ such that $F_n\cap(b_t,a_{t+1})=\emptyset$ for all $t$, where the gaps are $|b_t-a_{t+1}|=(n-2)-t(t+1)$ and $|a_t-b_t|=t(t-1)$. Moreover, $f(G)\in[a_t,b_t]$ if and only if $G$ can be obtained from a complete bipartite graph by flipping exactly $t$ edges/nonedges. On the other hand, we show, for all $n$ sufficiently large, that if $m\in[f(n),\binom{n}{3}-f(n)]$, then $m\in F_n$ where $f(n)$ is asymptotically best possible with $f(n)\sim n^{3/2}$ for $n$ even and $f(n)\sim \sqrt{2}n^{3/2}$ for $n$ odd. Furthermore, we determine the graphs with the minimum number of frustrated triangles amongst those with $n$ vertices and $e\leq n^2/4$ edges.
Odd behavior in the coefficients of reciprocals of binary power series
Published
• View Publication
• BIB
Let $\mathcal{A}$ be a finite subset of $\mathbb{N}$ including $0$ and $f_\mathcal{A}(n)$ be the number of ways to write $n=\sum_{i=0}^{\infty}ε_i2^i$, where $ε_i\in\mathcal{A}$. The sequence $\left(f_\mathcal{A}(n)\right) \bmod 2$ is always periodic, and $f_\mathcal{A}(n)$ is typically more often even than odd. We give four families of sets $\left(\mathcal{A}_m\right)$ with $\left|\mathcal{A}_m\right|=4$ such that the proportion of odd $f_{\mathcal{A}_m}(n)$'s goes to $1$ as $m\to\infty$.
A structure theorem for strong immersions
A graph H is strongly immersed in G if H is obtained from G by a sequence of vertex splittings (i.e., lifting some pairs of incident edges and removing the vertex) and edge removals. Equivalently, vertices of H are mapped to distinct vertices of G (branch vertices) and edges of H are mapped to pairwise edge-disjoint paths in G, each of them joining the branch vertices corresponding to the ends of the edge and not containing any other branch vertices. We describe the structure of graphs avoiding a fixed graph as a strong immersion.
Codes for DNA Storage Channels
Published
• View Publication
• BIB
We consider the problem of assembling a sequence based on a collection of its substrings observed through a noisy channel. The mathematical basis of the problem is the construction and design of sequences that may be discriminated based on a collection of their substrings observed through a noisy channel. We explain the connection between the sequence reconstruction problem and the problem of DNA synthesis and sequencing, and introduce the notion of a DNA storage channel. We analyze the number of sequence equivalence classes under the channel mapping and propose new asymmetric coding techniques to combat the effects of synthesis and sequencing noise. In our analysis, we make use of restricted de Bruijn graphs and Ehrhart theory for rational polytopes.
On Weak Hamiltonicity of a Random Hypergraph
A {\it weak (Berge) cycle} is an alternating sequence of vertices and (hyper)edges $C=(v_0, e_1, v_1, ..., v_{\ell-1}, e_\ell, v_{\ell}=v_0)$ such that the vertices $v_0, ..., v_{\ell-1}$ are distinct with $v_k, v_{k+1} \in e_{k}$ for each $k$, but the edges $e_1, ..., e_\ell$ are not necessarily distinct. We prove that the main barrier to the random $d$-uniform hypergraph $H_d(n,p),$ where each of the potential edges of cardinality $d$ is present with probability $p$, developing a weak Hamilton cycle is the presence of isolated vertices. In particular, for $d \geq 3$ fixed and $p=(d-1)! \frac{\ln n + c}{n^{d-1}}$, the probability that $H_d(n, p)$ has a weak Hamilton cycle tends to $e^{-e^{-c}}$, which is also the limiting probability that $H_d(n,p)$ has no isolated vertices. As a consequence, the probability that the random hypergraph $H_d(n, m=\frac{n(\ln n + c)}{d}),$ where $m$ potential edges are chosen uniformly at random to be present, is weak Hamiltonian also tends to $e^{-e^{-c}}$.
A monotonicity property for generalized Fibonacci sequences
Published
• View Publication
• BIB
Given k>1, let a_n be the sequence defined by the recurrence a_n=c_1a_{n-1}+c_2a_{n-2}+...+c_ka_{n-k} for n>=k, with initial values a_0=a_1=...=a_{k-2}=0 and a_{k-1}= 1. We show under a couple of assumptions concerning the constants c_i that the ratio of the n-th root of a_n to the (n-1)-st root of a_{n-1} is strictly decreasing for all n>=N, for some N depending on the sequence, and has limit 1. In particular, this holds in the cases when all of the c_i are unity or when all of the c_i are zero except for the first and last, which are unity. Furthermore, when k=3 or k=4, it is shown that one may take N to be an integer less than 12 in each of these cases.
On the number of 5-cycles in a tournament
Published
• View Publication
• BIB
We find an exact formula for the number of directed 5-cycles in a tournament in terms of its edge score sequence. We use this formula to find both upper and lower bounds on the number of 5-cycles in any $n$-tournament. In particular, we show that the maximum number of 5-cycles is asymptotically equal to $\frac{3}{4}{n \choose 5}$, the expected number 5-cycles in a random tournament ($p=\frac{1}{2}$), with equality (up to order of magnitude) for almost all tournaments. Note that this means that almost all $n$-tournaments contain the maximum number of $5$-cycles.
Properties of stochastic Kronecker graphs
Published
• View Publication
• BIB
The stochastic Kronecker graph model introduced by Leskovec et al. is a random graph with vertex set $\mathbb Z_2^n$, where two vertices $u$ and $v$ are connected with probability $α^{{u}\cdot{v}}γ^{(1-{u})\cdot(1-{v})}β^{n-{u}\cdot{v}-(1-{u})\cdot(1-{v})}$ independently of the presence or absence of any other edge, for fixed parameters $0<α,β,γ<1$. They have shown empirically that the degree sequence resembles a power law degree distribution. In this paper we show that the stochastic Kronecker graph a.a.s. does not feature a power law degree distribution for any parameters $0<α,β,γ<1$. In addition, we analyze the number of subgraphs present in the stochastic Kronecker graph and study the typical neighborhood of any given vertex.
Szemerédi's regularity lemma via martingales
Published in The Electronic Journal of Combinatorics 23 (2016), Research Paper P3.11, 1-24
• View Publication
• BIB
We prove a variant of the abstract probabilistic version of Szemerédi's regularity lemma, due to Tao, which applies to a number of structures (including graphs, hypergraphs, hypercubes, graphons, and many more) and works for random variables in $L_p$ for any $p>1$. Our approach is based on martingale difference sequences.
On homology of finite topological spaces
Published in Topology and its Applications 217 (2017), 1-19
• View Publication
• BIB
We develop a new method to compute the homology groups of finite topological spaces (or equivalently of finite partially ordered sets) by means of spectral sequences giving a complete and simple description of the corresponding differentials. Our method proves to be powerful and involves far fewer computations than the standard one. We derive many applications of our technique which include a generalization of Hurewicz theorem for regular CW-complexes, results in homological Morse theory and formulas to compute the Möbius function of posets.
Generalized Fibonacci and Lucas cubes arising from powers of paths and cycles
Published in Discrete Mathematics, Volume 339, Issue 1, 6 January 2016, Pages 270-282, ISSN 0012-365X
• View Publication
• BIB
The paper deals with some generalizations of Fibonacci and Lucas sequences, arising from powers of paths and cycles, respectively.
In the first part of the work we provide a formula for the number of edges of the Hasse diagram of the independent sets of the h-th power of a path ordered by inclusion. For h=1 such a diagram is called a Fibonacci cube, and for h>1 we obtain a generalization of the Fibonacci cube. Consequently, we derive a generalized notion of Fibonacci sequence, called h-Fibonacci sequence. Then, we show that the number of edges of a generalized Fibonacci cube is obtained by convolution of an h-Fibonacci sequence with itself.
In the second part we consider the case of cycles. We evaluate the number of edges of the Hasse diagram of the independent sets of the hth power of a cycle ordered by inclusion. For h=1 such a diagram is called Lucas cube, and for h>1 we obtain a generalization of the Lucas cube. We derive then a generalized version of the Lucas sequence, called h-Lucas sequence. Finally, we show that the number of edges of a generalized Lucas cube is obtained by an appropriate convolution of an h-Fibonacci sequence with an h-Lucas sequence.
Sturmian words and the Stern sequence
Published
• View Publication
• BIB
Central, standard, and Christoffel words are three strongly interrelated classes of binary finite words which represent a finite counterpart of characteristic Sturmian words. A natural arithmetization of the theory is obtained by representing central and Christoffel words by irreducible fractions labeling respectively two binary trees, the Raney (or Calkin-Wilf) tree and the Stern-Brocot tree. The sequence of denominators of the fractions in Raney's tree is the famous Stern diatomic numerical sequence. An interpretation of the terms $s(n)$ of Stern's sequence as lengths of Christoffel words when $n$ is odd, and as minimal periods of central words when $n$ is even, allows one to interpret several results on Christoffel and central words in terms of Stern's sequence and, conversely, to obtain a new insight in the combinatorics of Christoffel and central words by using properties of Stern's sequence. One of our main results is a non-commutative version of the "alternating bit sets theorem" by Calkin and Wilf. We also study the length distribution of Christoffel words corresponding to nodes of equal height in the tree, obtaining some interesting bounds and inequalities.
Intervals of permutation class growth rates
Published in Combinatorica, 38(2):279-303, 2018
• View Publication
• BIB
We prove that the set of growth rates of permutation classes includes an infinite sequence of intervals whose infimum is $θ_B\approx2.35526$, and that it also contains every value at least $λ_B\approx2.35698$. These results improve on a theorem of Vatter, who determined that there are permutation classes of every growth rate at least $λ_A\approx2.48187$. Thus, we also refute his conjecture that the set of growth rates below $λ_A$ is nowhere dense. Our proof is based upon an analysis of expansions of real numbers in non-integer bases, the study of which was initiated by Rényi in the 1950s. In particular, we prove two generalisations of a result of Pedicini concerning expansions in which the digits are drawn from sets of allowed values.