uniform distribution
173 papers tagged with this keyword
An enumerative formula for the spherical cap discrepancy
Published
• View Publication
• BIB
The spherical cap discrepancy is a widely used measure for how uniformly a sample of points on the sphere is distributed. Being hard to compute, this discrepancy measure is typically replaced by some lower or upper estimates when designing optimal sampling schemes for the uniform distribution on the sphere. In this paper, we provide a fully explicit, easy to implement enumerative formula for the spherical cap discrepancy. Not surprisingly, this formula is of combinatorial nature and, thus, its application is limited to spheres of small dimension and moderate sample sizes. Nonetheless, it may serve as a useful calibrating tool for testing the efficiency of sampling schemes and its explicit character might be useful also to establish necessary optimality conditions when minimizing the discrepancy with respect to a sample of given size.
Towards the sampling Lovász Local Lemma
Published
• View Publication
• BIB
Let $Φ= (V, \mathcal{C})$ be a constraint satisfaction problem on variables $v_1,\dots, v_n$ such that each constraint depends on at most $k$ variables and such that each variable assumes values in an alphabet of size at most $[q]$. Suppose that each constraint shares variables with at most $Δ$ constraints and that each constraint is violated with probability at most $p$ (under the product measure on its variables). We show that for $k, q = O(1)$, there is a deterministic, polynomial time algorithm to approximately count the number of satisfying assignments and a randomized, polynomial time algorithm to sample from approximately the uniform distribution on satisfying assignments, provided that \[C\cdot q^{3}\cdot k \cdot p \cdot Δ^{7} < 1, \quad \text{where }C \text{ is an absolute constant.}\] Previously, a result of this form was known essentially only in the special case when each constraint is violated by exactly one assignment to its variables.
For the special case of $k$-CNF formulas, the term $Δ^{7}$ improves the previously best known $Δ^{60}$ for deterministic algorithms [Moitra, J.ACM, 2019] and $Δ^{13}$ for randomized algorithms [Feng et al., arXiv, 2020]. For the special case of properly $q$-coloring $k$-uniform hypergraphs, the term $Δ^{7}$ improves the previously best known $Δ^{14}$ for deterministic algorithms [Guo et al., SICOMP, 2019] and $Δ^{9}$ for randomized algorithms [Feng et al., arXiv, 2020].
Intransitive dice tournament is not quasirandom
Published
• View Publication
• BIB
We settle a version of the conjecture about intransitive dice posed by Conrey, Gabbard, Grant, Liu and Morrison in 2016 and Polymath in 2017. We consider generalized dice with $n$ faces and we say that a die $A$ beats $B$ if a random face of $A$ is more likely to show a higher number than a random face of $B$. We study random dice with faces drawn iid from the uniform distribution on $[0,1]$ and conditioned on the sum of the faces equal to $n/2$. Considering the "beats" relation for three such random dice, Polymath showed that each of eight possible tournaments between them is asymptotically equally likely. In particular, three dice form an intransitive cycle with probability converging to $1/4$. In this paper we prove that for four random dice not all tournaments are equally likely and the probability of a transitive tournament is strictly higher than $3/8$.
Glauber dynamics for colourings of chordal graphs and graphs of bounded treewidth
The Glauber dynamics on the colourings of a graph is a random process which consists in recolouring at each step a random vertex of a graph with a new colour chosen uniformly at random among the colours not already present in its neighbourhood. It is known that when the total number of colours available is at least $Δ+2$, where $Δ$ is the maximum degree of the graph, this process converges to a uniform distribution on the set of all the colourings. Moreover, a well known conjecture is that the time it takes for the convergence to happen, called the mixing time, is polynomial in the size of the graph. Many weaker variants of this conjecture have been studied in the literature by allowing either more colours, or restricting the graphs to particular classes, or both. This paper follows this line of research by studying the mixing time of the Glauber dynamics on chordal graphs, as well as graphs of bounded treewidth. We show that the mixing time is polynomial in the size of the graph in the two following cases:
- on graphs with bounded treewidth, and at least $Δ+2$ colours,
- on chordal graphs if the number of colours is at least $(1+\varepsilon) (Δ+1)$, for any fixed constant $\varepsilon$.
Combinatorial Methods for Minkowski Tensors of Polytopes
In this paper we use a generating function approach to record and calculate entries of the Minkowski tensors of a polytope. We focus on ''surface tensors'', extending the methods used in arXiv:1807.10258 for moments of the uniform distribution which correspond to volume tensors. In this context we also extend the definition of the adjoint polynomial to the boundary complex of a polytope with simplicial facets. In the case of simplicial polytopes we give an explicit formulation for these surface tensors.
On sampling symmetric Gibbs distributions on sparse random graphs and hypergraphs
We introduce efficient algorithms for approximate sampling from symmetric Gibbs distributions on the sparse random (hyper)graph. The examples we consider include (but are not restricted to) important distributions on spin systems and spin-glasses such as the q state antiferromagnetic Potts model for $q\geq 2$, including the colourings, the uniform distributions over the Not-All-Equal solutions of random k-CNF formulas. Finally, we present an algorithm for sampling from the spin-glass distribution called the k-spin model. To our knowledge this is the first, rigorously analysed, efficient algorithm for spin-glasses which operates in a non trivial range of the parameters.
Our approach builds on the one that was introduced in [Efthymiou: SODA 2012]. For a symmetric Gibbs distribution $μ$ on a random (hyper)graph whose parameters are within an certain range, our algorithm has the following properties: with probability $1-o(1)$ over the input instances, it generates a configuration which is distributed within total variation distance $n^{-Ω(1)}$ from $μ$. The time complexity is $O((n\log n)^2)$.
The algorithm requires a range of the parameters which, for the graph case, coincide with the tree-uniqueness region, parametrised w.r.t. the expected degree d. For the hypergraph case, where uniqueness is less restrictive, we go beyond uniqueness. Our approach utilises in a novel way the notion of contiguity between Gibbs distributions and the so-called teacher-student model.
Enumeration of standard barely set-valued tableaux of shifted shapes
Published
• View Publication
• BIB
A standard barely set-valued tableau of shape $λ$ is a filling of the Young diagram $λ$ with integers $1,2,\dots,|λ|+1$ such that the integers are increasing in each row and column, and every cell contains one integer except one cell that contains two integers. Counting standard barely set-valued tableaux is closely related to the coincidental down-degree expectations (CDE) of lower intervals in Young's lattice. Using $q$-integral techniques we give a formula for the number of standard barely set-valued tableaux of arbitrary shifted shape. We show how it can be used to recover two formulas, originally conjectured by Reiner, Tenner and Yong, and proved by Hopkins, for numbers of standard barely set valued tableaux of particular shifted-balanced shapes. We also prove a conjecture of Reiner, Tenner and Yong on the CDE property of the shifted shape $(n,n-2,n-4,\dots,n-2k+2)$. Finally, in the Appendix we raise a conjecture on an $\mathsf a;q$-analogue of the down-degree expectation with respect to the uniform distribution for a specific class of lower order ideals of Young's lattice.
On the expected number of perfect matchings in cubic planar graphs
Published in Publicacions Matemàtiques, 2022, Vol. 66, Núm. 1, p. 325-353
• View Publication
• BIB
A well-known conjecture by Lovász and Plummer from the 1970s asserted that a bridgeless cubic graph has exponentially many perfect matchings. It was solved in the affirmative by Esperet et al. (Adv. Math. 2011). On the other hand, Chudnovsky and Seymour (Combinatorica 2012) proved the conjecture in the special case of cubic planar graphs. In our work we consider random bridgeless cubic planar graphs with the uniform distribution on graphs with $n$ vertices. Under this model we show that the expected number of perfect matchings in labeled bridgeless cubic planar graphs is asymptotically $cγ^n$, where $c>0$ and $γ\sim 1.14196$ is an explicit algebraic number. We also compute the expected number of perfect matchings in (non necessarily bridgeless) cubic planar graphs and provide lower bounds for unlabeled graphs. Our starting point is a correspondence between counting perfect matchings in rooted cubic planar maps and the partition function of the Ising model in rooted triangulations.
Stolarsky's invariance principle for finite metric spaces
Published in Mathematika, vol. 67, no. 1, 2021, pp. 158-186
• View Publication
• BIB
Stolarsky's invariance principle quantifies the deviation of a subset of a metric space from the uniform distribution. Classically derived for spherical sets, it has been recently studied in a number of other situations, revealing a general structure behind various forms of the main identity. In this work we consider the case of finite metric spaces, relating the quadratic discrepancy of a subset to a certain function of the distribution of distances in it. Our main results are related to a concrete form of the invariance principle for the Hamming space. We derive several equivalent versions of the expression for the discrepancy of a code, including expansions of the discrepancy and associated kernels in the Krawtchouk basis. Codes that have the smallest possible quadratic discrepancy among all subsets of the same cardinality can be naturally viewed as energy minimizing subsets in the space. Using linear programming, we find several bounds on the minimal discrepancy and give examples of minimizing configurations. In particular, we show that all binary perfect codes have the smallest possible discrepancy.
Quasi-random words and limits of word sequences
Published in European Journal of Combinatorics, Volume 98, 2021
• View Publication
• BIB
Words are sequences of letters over a finite alphabet. We study two intimately related topics for this object: quasi-randomness and limit theory. With respect to the first topic we investigate the notion of uniform distribution of letters over intervals, and in the spirit of the famous Chung--Graham--Wilson theorem for graphs we provide a list of word properties which are equivalent to uniformity. In particular, we show that uniformity is equivalent to counting 3-letter subsequences.
Inspired by graph limit theory we then investigate limits of convergent word sequences, those in which all subsequence densities converge. We show that convergent word sequences have a natural limit, namely Lebesgue measurable functions of the form $f:[0,1]\to[0,1]$. Via this theory we show that every hereditary word property is testable, address the problem of finite forcibility for word limits and establish as a byproduct a new model of random word sequences.
Along the lines of the proof of the existence of word limits, we can also establish the existence of limits for higher dimensional structures. In particular, we obtain an alternative proof of the result by Hoppen, Kohayakawa, Moreira, Ráth and Sampaio [{\it J. Combin. Theory Ser. B 103(1):93--113, 2013}] establishing the existence of permutons.
On Littlewood-Offord theory for arbitrary distributions
Let $X_1,\ldots,X_n$ be independent identically distributed random vectors in $\mathbb{R}^d$. We consider upper bounds on $\max_x \mathbb{P}(a_1X_1+\cdots+a_nX_n=x)$ under various restrictions on $X_i$ and the weights $a_i$. When $\mathbb{P}(X_i=\pm 1) = \frac {1} {2}$, this corresponds to the classical Littlewood-Offord problem. We prove that in general for identically distributed random vectors and even values of $n$ the optimal choice for $(a_i)$ is $a_i=1$ for $i\leq \frac{n}{2}$ and $a_i=-1$ for $i > \frac {n} 2$, regardless of the distribution of $X_1$. Applying these results for Bernoulli random variables answers a recent question of Fox, Kwan and Sauermann.
Finally, we provide sharp bounds for concentration probabilities of sums of random vectors under the condition $\sup_{x}\mathbb{P}(X_i=x)\leq α$, where it turns out that the worst case scenario is provided by distributions on an arithmetic progression that are in some sense as close to the uniform distribution as possible. An important feature of this work is that unlike much of the literature on the subject we use neither methods of harmonic analysis nor those from extremal combinatorics.
Successive shortest paths in complete graphs with random edge weights
Published
• View Publication
• BIB
Consider a complete graph $K_n$ with edge weights drawn independently from a uniform distribution $U(0,1)$. The weight of the shortest (minimum-weight) path $P_1$ between two given vertices is known to be $\ln n / n$, asymptotically. Define a second-shortest path $P_2$ to be the shortest path edge-disjoint from $P_1$, and consider more generally the shortest path $P_k$ edge-disjoint from all earlier paths. We show that the cost $X_k$ of $P_k$ converges in probability to $2k/n+\ln n/n$ uniformly for all $k \leq n-1$. We show analogous results when the edge weights are drawn from an exponential distribution. The same results characterise the collectively cheapest $k$ edge-disjoint paths, i.e., a minimum-cost $k$-flow. We also obtain the expectation of $X_k$ conditioned on the existence of $P_k$.
Half-graphs, other non-stable degree sequences, and the switch Markov chain
Published in The Electronic Journal of Combinatorics, Volume 28, Issue 3 (2021) P3.7
• View Publication
• BIB
One of the simplest methods of generating a random graph with a given degree sequence is provided by the Monte Carlo Markov Chain method using switches. The switch Markov chain converges to the uniform distribution, but generally the rate of convergence is not known. After a number of results concerning various degree sequences, rapid mixing was established for so-called $P$-stable degree sequences (including that of directed graphs), which covers every previously known rapidly mixing region of degree sequences.
In this paper we give a non-trivial family of degree sequences that are not $P$-stable and the switch Markov chain is still rapidly mixing on them. This family has an intimate connection to Tyshkevich-decompositions and strong stability as well.
Some new results in random matrices over finite fields
Published
• View Publication
• BIB
In this note we give various characterizations of random walks with possibly different steps that have relatively large discrepancy from the uniform distribution modulo a prime p, and use these results to study the distribution of the rank of random matrices over F_p and the equi-distribution behavior of normal vectors of random hyperplanes. We also study the probability that a random square matrix is eigenvalue-free, or when its characteristic polynomial is divisible by a given irreducible polynomial in the limit n to infinity in F_p. We show that these statistics are universal, extending results of Stong and Neumann-Praeger beyond the uniform model.
Successive minimum spanning trees
In a complete graph $K_n$ with edge weights drawn independently from a uniform distribution $U(0,1)$ (or alternatively an exponential distribution $\operatorname{Exp}(1)$), let $T_1$ be the MST (the spanning tree of minimum weight) and let $T_k$ be the MST after deletion of the edges of all previous trees $T_i$, $i<k$. We show that each tree's weight $w(T_k)$ converges in probability to a constant $γ_k$ with $2k-2\sqrt k <γ_k<2k+2\sqrt k$, and we conjecture that $γ_k = 2k-1+o(1)$. The problem is distinct from that of Frieze and Johansson (2018), finding $k$ MSTs of combined minimum weight, and for $k=2$ ours has strictly larger cost.
Our results also hold (and mostly are derived) in a multigraph model where edge weights for each vertex pair follow a Poisson process; here we additionally have $\mathbb E(w(T_k)) \to γ_k$. Thinking of an edge of weight $w$ as arriving at time $t=n w$, Kruskal's algorithm defines forests $F_k(t)$, each initially empty and eventually equal to $T_k$, with each arriving edge added to the first $F_k(t)$ where it does not create a cycle. Using tools of inhomogeneous random graphs we obtain structural results including that $C_1(F_k(t))/n$, the fraction of vertices in the largest component of $F_k(t)$, converges in probability to a function $ρ_k(t)$, uniformly for all $t$, and that a giant component appears in $F_k(t)$ at a time $t=σ_k$. We conjecture that the functions $ρ_k$ tend to time translations of a single function, $ρ_k(2k+x)\toρ_\infty(x)$ as $k \to \infty$, uniformly in $x\in \mathbb R$.
Simulations and numerical computations give estimated values of $γ_k$ for small $k$, and support the conjectures just stated.
Testing Graphs against an Unknown Distribution
The area of graph property testing seeks to understand the relation between the global properties of a graph and its local statistics. In the classical model, the local statistics of a graph is defined relative to a uniform distribution over the graph's vertex set. A graph property $\mathcal{P}$ is said to be testable if the local statistics of a graph can allow one to distinguish between graphs satisfying $\mathcal{P}$ and those that are far from satisfying it.
Goldreich recently introduced a generalization of this model in which one endows the vertex set of the input graph with an arbitrary and unknown distribution, and asked which of the properties that can be tested in the classical model can also be tested in this more general setting. We completely resolve this problem by giving a (surprisingly "clean") characterization of these properties. To this end, we prove a removal lemma for vertex weighted graphs which is of independent interest.
Counting and sampling gene family evolutionary histories in the duplication-loss and duplication-loss-transfer models
Given a set of species whose evolution is represented by a species tree, a gene family is a group of genes having evolved from a single ancestral gene. A gene family evolves along the branches of a species tree through various mechanisms, including - but not limited to - speciation, gene duplication, gene loss, horizontal gene transfer. The reconstruction of a gene tree representing the evolution of a gene family constrained by a species tree is an important problem in phylogenomics. However, unlike in the multispecies coalescent evolutionary model, very little is known about the search space for gene family histories accounting for gene duplication, gene loss and horizontal gene transfer (the DLT-model). We introduce the notion of evolutionary histories defined as a binary ordered rooted tree describing the evolution of a gene family, constrained by a species tree in the DLT-model. We provide formal grammars describing the set of all evolutionary histories that are compatible with a given species tree, whether it is ranked or unranked. These grammars allow us, using either analytic combinatorics or dynamic programming, to efficiently compute the number of histories of a given size, and also to generate random histories of a given size under the uniform distribution. We apply these tools to obtain exact asymptotics for the number of gene family histories for two species trees, the rooted caterpillar and the complete binary tree, as well as estimates of the range of the exponential growth factor of the number of histories for random species trees of size up to 25. Our results show that including horizontal gene transfer induce a dramatic increase of the number of evolutionary histories. We also show that, within ranked species trees, the number of evolutionary histories in the DLT-model is almost independent of the species tree topology.
An Optimal Algorithm for Stopping on the Element Closest to the Center of an Interval
Real numbers from the interval [0, 1] are randomly selected with uniform distribution. There are $n$ of them and they are revealed one by one. However, we do not know their values but only their relative ranks. We want to stop on recently revealed number maximizing the probability that that number is closest to $\frac{1}{2}$. We design an optimal stopping algorithm achieving our goal and prove that its probability of success is asymptotically equivalent to $\frac{1}{\sqrt{n}}\sqrt{\frac{2}π}$.
The strong circular law: a combinatorial view
Let $N_n$ be an $n\times n$ complex random matrix, each of whose entries is an independent copy of a centered complex random variable $z$ with finite non-zero variance $σ^{2}$. The strong circular law, proved by Tao and Vu, states that almost surely, as $n\to \infty$, the empirical spectral distribution of $N_n/(σ\sqrt{n})$ converges to the uniform distribution on the unit disc in $\mathbb{C}$.
A crucial ingredient in the proof of Tao and Vu, which uses deep ideas from additive combinatorics, is controlling the lower tail of the least singular value of the random matrix $xI - N_{n}/(σ\sqrt{n})$ (where $x\in \mathbb{C}$ is fixed) with failure probability that is inverse polynomial. In this paper, using a simple and novel approach (in particular, not using tools from additive combinatorics or any net arguments), we show that for any fixed matrix $M$ with operator norm at most $n^{0.51}$ and for all $η\geq 0$, $$\Pr\left(s_n(M+N_n) \leq η\right) \lesssim n^{C}η+ \exp(-n^{c}),$$ where $s_n(M+N_n)$ is the least singular value of $M+N_n$ and $C,c$ are absolute constants. Our result is optimal up to the constants $C,c$ and the inverse exponential-type error rate improves upon the inverse polynomial error rate due to Tao and Vu.
During the course of our proof, we extend the solution of the counting problem in inverse Littlewood-Offord theory, recently isolated by the author along with Ferber, Luh, and Samotij, from Rademacher variables to general complex random variables. This significantly improves on estimates for this problem obtained using the optimal inverse Littlewood-Offord theorem of Nguyen and Vu, and may be of independent interest.
Extractors for small zero-fixing sources
A random variable $X$ is an $(n,k)$-zero-fixing source if for some subset $V\subseteq[n]$, $X$ is the uniform distribution on the strings $\{0,1\}^n$ that are zero on every coordinate outside of $V$. An $ε$-extractor for $(n,k)$-zero-fixing sources is a mapping $F:\{0,1\}^n\to\{0,1\}^m$, for some $m$, such that $F(X)$ is $ε$-close in statistical distance to the uniform distribution on $\{0,1\}^m$ for every $(n,k)$-zero-fixing source $X$. Zero-fixing sources were introduced by Cohen and Shinkar in [10] in connection with the previously studied extractors for bit-fixing sources. They constructed, for every $μ>0$, an efficiently computable extractor that extracts a positive fraction of entropy, i.e., $Ω(k)$ bits, from $(n,k)$-zero-fixing sources where $k\geq(\log\log n)^{2+μ}$.
In this paper we present two different constructions of extractors for zero-fixing sources that are able to extract a positive fraction of entropy for $k$ essentially smaller than $\log\log n$. The first extractor works for $k\geq C\log\log\log n$, for some constant $C$. The second extractor extracts a positive fraction of entropy for $k\geq \log^{(i)}n$ for any fixed $i\in \mathbb{N}$, where $\log^{(i)}$ denotes $i$-times iterated logarithm. The fraction of extracted entropy decreases with $i$. The first extractor is a function computable in polynomial time in~$n$ (for $ε=o(1)$, but not too small); the second one is computable in polynomial time when $k\leqα\log\log n/\log\log\log n$, where $α$ is a positive constant.
The subject studied in this paper is closely related to Ramsey theory. We use methods developed in Ramsey theory and our results can also be interpreted as a contribution to this field.