random
6952 papers tagged with this keyword
Concentration estimates for functions of finite high-dimensional random arrays
Published in Random Structures & Algorithms 63 (2023), 997-1053
• View Publication
• BIB
Let $\boldsymbol{X}$ be a $d$-dimensional random array on $[n]$ whose entries take values in a finite set $\mathcal{X}$, that is, $\boldsymbol{X}=\langle X_s:s\in \binom{[n]}{d}\rangle$ is an $\mathcal{X}$-valued stochastic process indexed by the set $\binom{[n]}{d}$ of all $d$-element subsets of $[n]:=\{1,\dots,n\}$. We give easily checked conditions on $\boldsymbol{X}$ that ensure, for instance, that for every function $f\colon \mathcal{X}^{\binom{[n]}{d}}\to\mathbb{R}$ that satisfies $\mathbb{E}[f(\boldsymbol{X})]=0$ and $\|f(\boldsymbol{X})\|_{L_p}=1$ for some $p>1$, the random variable $f(\boldsymbol{X})$ becomes concentrated after conditioning it on a large subarray of $\boldsymbol{X}$. These conditions cover several classes of random arrays with not necessarily independent entries. Applications are given in combinatorics, and examples are also presented that show the optimality of various aspects of the results.
A branching process approach to level-$k$ phylogenetic networks
Published
• View Publication
• BIB
The mathematical analysis of random phylogenetic networks via analytic and algorithmic methods has received increasing attention in the past years. In the present work we introduce branching process methods to their study. This approach appears to be new in this context. Our main results focus on random level-$k$ networks with $n$ labelled leaves. Although the number of reticulation vertices in such networks is typically linear in $n$, we prove that their asymptotic global and local shape is tree-like in a well-defined sense. We show that the depth process of vertices in a large network converges towards a Brownian excursion after rescaling by $n^{-1/2}$. We also establish Benjamini--Schramm convergence of large random level-$k$ networks towards a novel random infinite network.
Localization Game for Random Geometric Graphs
Published
• View Publication
• BIB
The localization game is a two player combinatorial game played on a graph $G=(V,E)$. The cops choose a set of vertices $S_1 \subseteq V$ with $|S_1|=k$. The robber then chooses a vertex $v \in V$ whose location is hidden from the cops, but the cops learn the graph distance between the current position of the robber and the vertices in $S_1$. If this information is sufficient to locate the robber, the cops win immediately; otherwise the cops choose another set of vertices $S_2 \subseteq V$ with $|S_2|=k$, and the robber may move to a neighbouring vertex. The new distances are presented to the robber, and if the cops can deduce the new location of the robber based on all information they accumulated thus far, then they win; otherwise, a new round begins. If the robber has a strategy to avoid being captured, then she wins. The localization number is defined to be the smallest integer $k$ so that the cops win the game. In this paper we determine the localization number (up to poly-logarithmic factors) of the random geometric graph $G \in \mathcal G(n,r)$ slightly above the connectivity threshold.
Regularity method and large deviation principles for the Erdős--Rényi hypergraph
Published
• View Publication
• BIB
We develop a quantitative large deviations theory for random hypergraphs, which rests on tensor decomposition and counting lemmas under a novel family of cut-type norms. As our main application, we obtain sharp asymptotics for joint upper and lower tails of homomorphism counts in the $r$-uniform Erdős--Rényi hypergraph for any fixed $r\ge 2$, generalizing and improving on previous results for the Erdős--Rényi graph ($r=2$). The theory is sufficiently quantitative to allow the density of the hypergraph to vanish at a polynomial rate, and additionally yields tail asymptotics for other nonlinear functionals, such as induced homomorphism counts.
Note on induced paths in sparse random graphs
We show that for $d\ge d_0(ε)$, with high probability, the random graph $G(n,d/n)$ contains an induced path of length $(3/2-ε)\frac{n}{d}\log d$. This improves a result obtained independently by Luczak and Suen in the early 90s, and answers a question of Fernandez de la Vega. Along the way, we generalize a recent result of Cooley, Draganić, Kang and Sudakov who studied the analogous problem for induced matchings.
Distinguishing power-law uniform random graphs from inhomogeneous random graphs through small subgraphs
Published
• View Publication
• BIB
We investigate the asymptotic number of induced subgraphs in power-law uniform random graphs. We show that these induced subgraphs appear typically on vertices with specific degrees, which are found by solving an optimization problem. Furthermore, we show that this optimization problem allows to design a linear-time, randomized algorithm that distinguishes between uniform random graphs and random graph models that create graphs with approximately a desired degree sequence: the power-law rank-1 inhomogeneous random graph. This algorithm uses the fact that some specific induced subgraphs appear significantly more often in uniform random graphs than in rank-1 inhomogeneous random graphs.
On the sampling Lovász Local Lemma for atomic constraint satisfaction problems
We study the problem of sampling an approximately uniformly random satisfying assignment for atomic constraint satisfaction problems i.e. where each constraint is violated by only one assignment to its variables. Let $p$ denote the maximum probability of violation of any constraint and let $Δ$ denote the maximum degree of the line graph of the constraints.
Our main result is a nearly-linear (in the number of variables) time algorithm for this problem, which is valid in a Lovász local lemma type regime that is considerably less restrictive compared to previous works. In particular, we provide sampling algorithms for the uniform distribution on:
(1) $q$-colorings of $k$-uniform hypergraphs with $Δ\lesssim q^{(k-4)/3 + o_{q}(1)}.$
The exponent $1/3$ improves the previously best-known $1/7$ in the case $q, Δ= O(1)$ [Jain, Pham, Vuong; arXiv, 2020] and $1/9$ in the general case [Feng, He, Yin; STOC 2021].
(2) Satisfying assignments of Boolean $k$-CNF formulas with $Δ\lesssim 2^{k/5.741}.$
The constant $5.741$ in the exponent improves the previously best-known $7$ in the case $k = O(1)$ [Jain, Pham, Vuong; arXiv, 2020] and $13$ in the general case [Feng, He, Yin; STOC 2021].
(3) Satisfying assignments of general atomic constraint satisfaction problems with $p\cdot Δ^{7.043} \lesssim 1.$
The constant $7.043$ improves upon the previously best-known constant of $350$ [Feng, He, Yin; STOC 2021].
At the heart of our analysis is a novel information-percolation type argument for showing the rapid mixing of the Glauber dynamics for a carefully constructed projection of the uniform distribution on satisfying assignments. Notably, there is no natural partial order on the space, and we believe that the techniques developed for the analysis may be of independent interest.
Large deviations for the largest eigenvalue of Gaussian networks with constant average degree
Published
• View Publication
• BIB
Large deviation behavior of the largest eigenvalue $λ_1$ of Gaussian networks (Erdős-Rényi random graphs $\mathcal{G}_{n,p}$ with i.i.d. Gaussian weights on the edges) has been the topic of considerable interest. Recently in [6,30], a powerful approach was introduced based on tilting measures by suitable spherical integrals, particularly establishing a non-universal large deviation behavior for fixed $p<1$ compared to the standard Gaussian ($p=1$) case. The case when $p\to 0$ was however completely left open with one expecting the dense behavior to hold only until the average degree is logarithmic in $n$. In this article we focus on the case of constant average degree i.e., $p=\frac{d}{n}$. We prove the following results towards a precise understanding of the large deviation behavior in this setting.
1. (Upper tail probabilities): For $δ>0,$ we pin down the exact exponent $ψ(δ)$ such that $$\mathbb{P}(λ_1\ge \sqrt{2(1+δ)\log n})=n^{-ψ(δ)+o(1)}.$$ Further, we show that conditioned on the upper tail event, with high probability, a unique maximal clique emerges with a very precise $δ$ dependent size (takes either one or two possible values) and the Gaussian weights are uniformly high in absolute value on the edges in the clique. Finally, we also prove an optimal localization result for the leading eigenvector, showing that it allocates most of its mass on the aforementioned clique which is spread uniformly across its vertices.
2. (Lower tail probabilities): The exact stretched exponential behavior of $\mathbb{P}(λ_1\le \sqrt{2(1-δ)\log n})$ is also established.
As an immediate corollary, we get $λ_1 \approx \sqrt{2 \log n}$ typically, a result that surprisingly appears to be new. A key ingredient is an extremal spectral theory for weighted graphs obtained via the classical Motzkin-Straus theorem.
Involutive random walks on total orders and the anti-diagonal eigenvalue property
Published
• View Publication
• BIB
This paper studies a family of random walks defined on the finite ordinals using their order reversing involutions. Starting at $x \in \{0,1,\ldots,n-1\}$, an element $y \le x$ is chosen according to a prescribed probability distribution, and the walk then steps to $n-1-y$. We show that under very mild assumptions these walks are irreducible, recurrent and ergodic. We then find the invariant distributions, eigenvalues and eigenvectors of a distinguished subfamily of walks whose transition matrices have the global anti-diagonal eigenvalue property studied in earlier work by Ochiai, Sasada, Shirai and Tsuboi. We prove that this subfamily of walks is characterised by their reversibility. As a corollary, we obtain the invariant distributions and rate of convergence of the random walk on the set of subsets of $\{1,\ldots, m\}$ in which steps are taken alternately to subsets and supersets, each chosen equiprobably. We then consider analogously defined random walks on the real interval $[0,1]$ and use techniques from the theory of self adjoint compact operators on Hilbert spaces to prove analogues of the main results in the discrete case.
The Phase Transition of Discrepancy in Random Hypergraphs
Published in SIAM Journal on Discrete Mathematics 37(3), 1818-1841, 2023
• View Publication
• BIB
Motivated by the Beck-Fiala conjecture, we study the discrepancy problem in two related models of random hypergraphs on $n$ vertices and $m$ edges. In the first (edge-independent) model, a random hypergraph $H_1$ is constructed by fixing a parameter $p$ and allowing each of the $n$ vertices to join each of the $m$ edges independently with probability $p$. In the parameter range in which $pn \rightarrow \infty$ and $pm \rightarrow \infty$, we show that with high probability (w.h.p.) $H_1$ has discrepancy at least $Ω(2^{-n/m} \sqrt{pn})$ when $m = O(n)$, and at least $Ω(\sqrt{pn \logγ})$ when $m \gg n$, where $γ= \min\{ m/n, pn\}$. In the second (edge-dependent) model, $d$ is fixed and each vertex of $H_2$ independently joins exactly $d$ edges uniformly at random. We obtain analogous results for this model by generalizing the techniques used for the edge-independent model with $p=d/m$. Namely, for $d \rightarrow \infty$ and $dn/m \rightarrow \infty$, we prove that w.h.p. $H_{2}$ has discrepancy at least $Ω(2^{-n/m} \sqrt{dn/m})$ when $m = O(n)$, and at least $Ω(\sqrt{(dn/m) \logγ})$ when $m \gg n$, where $γ=\min\{m/n, dn/m\}$. Furthermore, we obtain nearly matching asymptotic upper bounds on the discrepancy in both models (when $p=d/m$), in the dense regime of $m \gg n$. Specifically, we apply the partial colouring lemma of Lovett and Meka to show that w.h.p. $H_{1}$ and $H_{2}$ each have discrepancy $O( \sqrt{dn/m} \log(m/n))$, provided $d \rightarrow \infty$, $d n/m \rightarrow \infty$ and $m \gg n$. This result is algorithmic, and together with the work of Bansal and Meka characterizes how the discrepancy of each random hypergraph model transitions from $Θ(\sqrt{d})$ to $o(\sqrt{d})$ as $m$ varies from $m=Θ(n)$ to $m \gg n$.
Anti-concentration of random variables from zero-free regions
Published in Discrete Analysis 2022:13
• Search Publication
This paper provides a connection between the concentration of a random variable and the distribution of the roots of its probability generating function. Let $X$ be a random variable taking values in $\{0,\ldots,n\}$ with $\mathbb{P}(X = 0)\mathbb{P}(X = n) > 0$ and with probability generating function $f_X$. We show that if all of the zeros $ζ$ of $f_X$ satisfy $|\arg(ζ)| \geq δ$ and $R^{-1} \leq |ζ| \leq R$ then \[ \operatorname{Var}(X) \geq c R^{-2π/δ}n, \] where $c > 0$ is a absolute constant. We show that this result is sharp, up to the factor $2$ in the exponent of $R$. As a consequence, we are able to deduce a Littlewood--Offord type theorem for random variables that are not necessarily sums of i.i.d.\ random variables.
Expected number of induced subtrees shared by two independent copies of a random tree
Published
• View Publication
• BIB
Consider a rooted tree $T$ with leaf-set $[n]$, and with all non-leaf vertices having out-degree $2$, at least. A rooted tree $\mathcal T$ with leaf-set $S\subset [n]$ is induced by $S$ in $T$ if $\mathcal T$ is the lowest common ancestor subtree for $S$, with all its degree-2 vertices suppressed. A "maximum agreement subtree" (MAST) for a pair of two trees $T'$ and $T"$ is a tree $\mathcal T$ with a largest leaf-set $S\subset [n]$ such that $\mathcal T$ is induced by $S$ both in $T'$ and $T"$. Bryant et al. \cite{BryMcKSte} and Bernstein et al. \cite{Ber} proved, among other results, that for $T'$ and $T"$ being two independent copies of a random binary (uniform or Yule-Harding distributed) tree $T$, the likely magnitude order of $\text{MAST}(T',T")$ is $O(n^{1/2})$. We prove this bound for a wide class of random rooted trees : $T$ is a terminal tree of a branching, Galton--Watson, process with an ordered-offspring distribution of mean $1$, conditioned on "total number of leaves is $n$".
Quantitative twisted patterns in positive density subsets
We make quantitative improvements to recently obtained results on the structure of the image of a large difference set under certain quadratic forms and other homogeneous polynomials. Previous proofs used deep results of Benoist-Quint on random walks in certain subgroups of $\operatorname{SL}_r(\mathbb{Z})$ (the symmetry groups of these quadratic forms) that were not of a quantitative nature. Our new observation relies on noticing that rather than studying random walks, one can obtain more quantitative results by considering polynomial orbits of these group actions that are not contained in cosets of submodules of $\mathbb{Z}^r$ of small index. Our main new technical tool is a uniform Furstenberg-Sárközy theorem that holds for a large class of polynomials not necessarily vanishing at zero, which may be of independent interest and is derived from a density increment argument and Hua's bound on polynomial exponential sums.
Trace Reconstruction with Bounded Edit Distance
Published
• View Publication
• BIB
The trace reconstruction problem studies the number of noisy samples needed to recover an unknown string $\boldsymbol{x}\in\{0,1\}^n$ with high probability, where the samples are independently obtained by passing $\boldsymbol{x}$ through a random deletion channel with deletion probability $q$. The problem is receiving significant attention recently due to its applications in DNA sequencing and DNA storage. Yet, there is still an exponential gap between upper and lower bounds for the trace reconstruction problem. In this paper we study the trace reconstruction problem when $\boldsymbol{x}$ is confined to an edit distance ball of radius $k$, which is essentially equivalent to distinguishing two strings with edit distance at most $k$. It is shown that $n^{O(k)}$ samples suffice to achieve this task with high probability.
Cutoff for non-negatively curved Markov chains
Published
• View Publication
• BIB
Discovered in the context of card shuffling by Aldous, Diaconis and Shahshahani, the cutoff phenomenon has since then been established in a variety of Markov chains. However, proving cutoff remains a delicate affair, which requires a detailed knowledge of the chain. Identifying the general mechanisms underlying this phase transition -- without having to pinpoint its precise location -- remains one of the most fundamental open problems in the area of mixing times. In the present paper, we make a step in this direction by establishing cutoff for Markov chains with non-negative curvature, under a suitably refined product condition. The result applies, in particular, to random walks on abelian Cayley expanders satisfying a mild degree condition, hence in particular to \emph{almost all} abelian Cayley graphs. Our proof relies on a quantitative \emph{entropic concentration principle}, which we believe to lie behind all cutoff phenomena.
A note on invariable generation of nonsolvable permutation groups
We prove a result on the asymptotic proportion of randomly chosen pairs of permutations in the symmetric group $S_n$ which "invariably" generate a nonsolvable subgroup, i.e., whose cycle structures cannot possibly both occur in the same solvable subgroup of $S_n$. As an application, we obtain that for a large degree "random" integer polynomial $f$, reduction modulo two different primes can be expected to suffice to prove the nonsolvability of $Gal(f/\mathbb{Q})$.
Prophet Matching Meets Probing with Commitment
We consider the online stochastic matching problem for bipartite graphs where edges adjacent to an online node must be probed to determine if they exist, based on known edge probabilities. Our algorithms respect commitment, in that if a probed edge exists, it must be used in the matching. We study this matching problem subject to a downward-closed constraint on each online node's allowable edge probes. Our setting generalizes the commonly studied patience (or time-out) constraint which limits the number of probes that can be made to an online node's adjacent edges. We introduce a new LP that we prove is a relaxation of an optimal offline probing algorithm (the adaptive benchmark) and which overcomes the limitations of previous LP relaxations.
(1) A tight $\frac{1}{2}$ ratio when the stochastic graph is generated from a known stochastic type graph where the $t^{th}$ online node is drawn independently from a known distribution $\scr{D}_{π(t)}$ and $π$ is chosen adversarially. We refer to this setting as the known i.d. stochastic matching problem with adversarial arrivals.
(2) A $1-1/e$ ratio when the stochastic graph is generated from a known stochastic type graph where the $t^{th}$ online node is drawn independently from a known distribution $\scr{D}_{π(t)}$ and $π$ is a random permutation. We refer to this setting as the known i.d. stochastic matching problem with random order arrivals.
Our results improve upon the previous best competitive ratio of $0.46$ in the known i.i.d. setting against the standard adaptive benchmark. Moreover, we are the first to study the prophet secretary matching problem in the context of probing, where we match the best known classical result.
A shape theorem for exploding sandpiles
Published
• View Publication
• BIB
We study scaling limits of exploding Abelian sandpiles using ideas from percolation and front propagation in random media. We establish sufficient conditions under which a limit shape exists and show via a family of counterexamples that convergence may not occur in general. A corollary of our proof is a simple criteria for determining if a sandpile is explosive; this strengthens a result of Fey, Levine, and Peres (2010).
Statistical Enumeration of Groups by Double Cosets
Published
• View Publication
• BIB
Let $H$ and $K$ be subgroups of a finite group $G$. Pick $g \in G$ uniformly at random. We study the distribution induced on double cosets. Three examples are treated in detail: 1) $H = K = $ the Borel subgroup in $GL_n(\mathbb{F}_q)$. This leads to new theorems for Mallows measure on permutations and new insights into the LU matrix factorization. 2) The double cosets of the hyperoctahedral group inside $S_{2n}$, which leads to new applications of the Ewens's sampling formula of mathematical genetics. 3) Finally, if $H$ and $K$ are parabolic subgroups of $S_n$, the double cosets are `contingency tables', studied by statisticians for the past 100 years.
Discrete Max-Linear Bayesian Networks
Published in Alg. Stat. 12 (2021) 213-225
• View Publication
• BIB
Discrete max-linear Bayesian networks are directed graphical models specified by the same recursive structural equations as max-linear models but with discrete innovations. When all of the random variables in the model are binary, these models are isomorphic to the conjunctive Bayesian network (CBN) models of Beerenwinkel, Eriksson, and Sturmfels. Many of the techniques used to study CBN models can be extended to discrete max-linear models and similar results can be obtained. In particular, we extend the fact that CBN models are toric varieties after linear change of coordinates to all discrete max-linear models.