arXiv++ Combinatorics

Browse math.CO papers from arXiv

random

6952 papers tagged with this keyword
2020-10-29 v2
A probabilistic way to discover the rainbow
Published • View PublicationBIB
"No two rainbows are the same. Neither are two packs of Skittles. Enjoy an odd mix!". Using an interpretation via spatial random walks, we quantify the probability that two randomly selected packs of Skittles candy are identical and determine the expected number of packs one has to purchase until the first match. We believe this problem to be appealing for middle and high school students as well as undergraduate students at University.
2020-10-28 v2
Random walks on stochastic uniform growth trees: Analytical formula for mean first-passage time
Published • View PublicationBIB
As known, the commonly-utilized ways to determine mean first-passage time $\overline{\mathcal{F}}$ for random walk on networks are mainly based on Laplacian spectra. However, methods of this type can become prohibitively complicated and even fail to work when the Laplacian matrix of network under consideration is difficult to describe in the first place. In this paper, we propose an effective approach to determining quantity $\overline{\mathcal{F}}$ on some widely-studied tree networks. To this end, we first build up a general formula between Wiener index $\mathcal{W}$ and $\overline{\mathcal{F}}$ on a tree. This enables us to convert issues to answer into calculation of $\mathcal{W}$ on networks in question. As opposed to most of previous work focusing on deterministic growth trees, our goal is to consider stochastic case. Towards this end, we establish a principled framework where randomness is introduced into the process of growing trees. As an immediate consequence, the previously published results upon deterministic cases are thoroughly covered by formulas established in this paper. Additionally, it is also straightforward to obtain Kirchhoff index on our tree networks using the proposed approach. Most importantly, our approach is more manageable than many other methods including spectral technique in situations considered herein.
2020-10-28
Two point concentration of maximum degree in sparse random planar graphs
Let $P(n,m)$ be a graph chosen uniformly at random from the class of all planar graphs on vertex set $\left\{1, \ldots, n\right\}$ with $m=m(n)$ edges. We show that in the sparse regime, when $\limsup_{n \to \infty} m/n<1$, with high probability the maximum degree of $P(n,m)$ takes at most two different values.
The p-Airy distribution
In this manuscript we consider the set of Dyck paths equipped with the uniform measure, and we study the statistical properties of a deformation of the observable "area below the Dyck path" as the size $N$ of the path goes to infinity. The deformation under analysis is apparently new: while usually the area is constructed as the sum of the heights of the steps of the Dyck path, here we regard it as the sum of the lengths of the connected horizontal slices under the path, and we deform it by applying to the lengths of the slices a positive regular function $ω(\ell)$ such that $ω(\ell) \sim \ell^p$ for large argument. This shift of paradigm is motivated by applications to the Euclidean Random Assignment Problem in Random Combinatorial Optimization, and to Tree Hook Formulas in Algebraic Combinatorics. For $p \in \mathbb{R}^+ \smallsetminus \left\{ \frac{1}{2}\right\}$, we characterize the statistical properties of the deformed area as a function of the deformation function $ω(\ell)$ by computing its integer moments, finding a generalization of a well-known recursion for the moments of the area-Airy distribution, due to Takács. Most of the properties of the distribution of the deformed area are \emph{universal}, meaning that they depend on the deformation parameter $p$, but not on the microscopic details of the function $ω(\ell)$. We call \emph{$p$-Airy distribution} this family of universal distributions.
2020-10-26 v2
Gaussian Asymptotics of Jack Measures on Partitions from Weighted Enumeration of Ribbon Paths
Published • View PublicationBIB
In this paper we determine two asymptotic results for Jack measures on partitions, a model defined by two specializations of Jack polynomials proposed by Borodin-Olshanski in [European J. Combin. 26.6 (2005): 795-834]. Assuming these two specializations are the same, we derive limit shapes and Gaussian fluctuations for the anisotropic profiles of these random partitions in three asymptotic regimes associated to diverging, fixed, and vanishing values of the Jack parameter. To do so, we introduce a generalization of Motzkin paths we call "ribbon paths", show for general Jack measures that certain joint cumulants are weighted sums of connected ribbon paths on $n$ sites with $n-1+g$ pairings, and derive our two results from the contributions of $(n,g)=(1,0)$ and $(2,0)$, respectively. Our analysis makes use of Nazarov-Sklyanin's spectral theory for Jack polynomials. As a consequence, we give new proofs of several results for Schur measures, Plancherel measures, and Jack-Plancherel measures. In addition, we relate our weighted sums of ribbon paths to the weighted sums of ribbon graphs of maps on non-oriented real surfaces recently introduced by Chapuy-Dolęga.
2020-10-26 v2
Asymptotic Enumeration and Distributional Properties of Galled Networks
Published • View PublicationBIB
We show a first-order asymptotics result for the number of galled networks with $n$ leaves. This is the first class of phylogenetic networks of {\it large} size for which an asymptotic counting result of such strength can be obtained. In addition, we also find the limiting distribution of the number of reticulation nodes of a galled networks with $n$ leaves chosen uniformly at random. These results are obtained by performing an asymptotic analysis of a recent approach of Gunawan, Rathin, and Zhang (2020) which was devised for the purpose of (exactly) counting galled networks. Moreover, an old result of Bender and Richmond (1984) plays a crucial role in our proofs, too.
2020-10-26 v2
The training accuracy of two-layer neural networks: its estimation and understanding using random datasets
Published • View PublicationBIB
Although the neural network (NN) technique plays an important role in machine learning, understanding the mechanism of NN models and the transparency of deep learning still require more basic research. In this study, we propose a novel theory based on space partitioning to estimate the approximate training accuracy for two-layer neural networks on random datasets without training. There appear to be no other studies that have proposed a method to estimate training accuracy without using input data and/or trained models. Our method estimates the training accuracy for two-layer fully-connected neural networks on two-class random datasets using only three arguments: the dimensionality of inputs (d), the number of inputs (N), and the number of neurons in the hidden layer (L). We have verified our method using real training accuracies in our experiments. The results indicate that the method will work for any dimension, and the proposed theory could extend also to estimate deeper NN models. The main purpose of this paper is to understand the mechanism of NN models by the approach of estimating training accuracy but not to analyze their generalization nor their performance in real-world applications. This study may provide a starting point for a new way for researchers to make progress on the difficult problem of understanding deep learning.
2020-10-22 v2
A Generalized Faulhaber Inequality, Improved Bracketing Covers, and Applications to Discrepancy
Published in Mathematics of Computation 90 (2021), 2873-2898 • View PublicationBIB
We prove a generalized Faulhaber inequality to bound the sums of the $j$-th powers of the first $n$ (possibly shifted) natural numbers. With the help of this inequality we are able to improve the known bounds for bracketing numbers of $d$-dimensional axis-parallel boxes anchored in $0$ (or, put differently, of lower left orthants intersected with the $d$-dimensional unit cube $[0,1]^d$). We use these improved bracketing numbers to establish new bounds for the star-discrepancy of negatively dependent random point sets and its expectation. We apply our findings also to the weighted star-discrepancy.
Asymmetric Ramsey Properties of Random Graphs for Cliques and Cycles
Published • View PublicationBIB
We say that $G \to (F,H)$ if, in every edge colouring $c: E(G) \to \{1,2\}$, we can find either a $1$-coloured copy of $F$ or a $2$-coloured copy of $H$. The well-known Kohayakawa--Kreuter conjecture states that the threshold for the property $G(n,p) \to (F,H)$ is equal to $n^{-1/m_{2}(F,H)}$, where $m_{2}(F,H)$ is given by \[ m_{2}(F,H):= \max \left\{\dfrac{e(J)}{v(J)-2+1/m_2(H)} : J \subseteq F, e(J)\ge 1 \right\}. \] In this paper, we show the $0$-statement of the Kohayakawa--Kreuter conjecture for every pair of cycles and cliques.
2020-10-21
Random polynomials: the closest roots to the unit circle
Let $f = \sum_{k=0}^n \varepsilon_k z^k$ be a random polynomial, where $\varepsilon_0,\ldots ,\varepsilon_n$ are iid standard Gaussian random variables, and let $ζ_1,\ldots,ζ_n$ denote the roots of $f$. We show that the point process determined by the magnitude of the roots $\{ 1-|ζ_1|,\ldots, 1-|ζ_n| \}$ tends to a Poisson point process at the scale $n^{-2}$ as $n\rightarrow \infty$. One consequence of this result is that it determines the magnitude of the closest root to the unit circle. In particular, we show that \[ \min_{k} ||ζ_k| - 1|n^2 \rightarrow \mathrm{Exp}(1/6),\] in distribution, where $\mathrm{Exp}(λ)$ denotes an exponential random variable of mean $λ^{-1}$. This resolves a conjecture of Shepp and Vanderbei from 1995 that was later studied by Konyagin and Schlag.
2020-10-21
Real roots near the unit circle of random polynomials
Published • View PublicationBIB
Let $f_n(z) = \sum_{k = 0}^n \varepsilon_k z^k$ be a random polynomial where $\varepsilon_0,\ldots,\varepsilon_n$ are i.i.d. random variables with $\mathbb{E} \varepsilon_1 = 0$ and $\mathbb{E} \varepsilon_1^2 = 1$. Letting $r_1, r_2,\ldots, r_k$ denote the real roots of $f_n$, we show that the point process defined by $\{|r_1| - 1,\ldots, |r_k| - 1 \}$ converges to a non-Poissonian limit on the scale of $n^{-1}$ as $n \to \infty$. Further, we show that for each $δ> 0$, $f_n$ has a real root within $Θ_δ(1/n)$ of the unit circle with probability at least $1 - δ$. This resolves a conjecture of Shepp and Vanderbei from 1995 by confirming its weakest form and refuting its strongest form.
2020-10-21 v2
On the robustness of the metric dimension of grid graphs to adding a single edge
Published • View PublicationBIB
The metric dimension (MD) of a graph is a combinatorial notion capturing the minimum number of landmark nodes needed to distinguish every pair of nodes in the graph based on graph distance. We study how much the MD can increase if we add a single edge to the graph. The extra edge can either be selected adversarially, in which case we are interested in the largest possible value that the MD can take, or uniformly at random, in which case we are interested in the distribution of the MD. The adversarial setting has already been studied by [Eroh et. al., 2015] for general graphs, who found an example where the MD doubles on adding a single edge. By constructing a different example, we show that this increase can be as large as exponential. However, we believe that such a large increase can occur only in specially constructed graphs, and that in most interesting graph families, the MD at most doubles on adding a single edge. We prove this for $d$-dimensional grid graphs, by showing that $2d$ appropriately chosen corners and the endpoints of the extra edge can distinguish every pair of nodes, no matter where the edge is added. For the special case of $d=2$, we show that it suffices to choose the four corners as landmarks. Finally, when the extra edge is sampled uniformly at random, we conjecture that the MD of 2-dimensional grids converges in probability to $3+\mathrm{Ber}(8/27)$, and we give an almost complete proof.
2020-10-20
Area Statistics for Large Oscillating Tableaux
In this note we show that the area of the partitions making up an oscillating tableaux is described by a random walk on the first quadrant of $\mathbb{Z}^2$ with certain position dependent weights. We are able to recursively calculate the moments of the walk. As the length of the oscillating tableaux becomes large we show that this random walk converges to a Gaussian stochastic process.
2020-10-20 v2
Sparse reconstruction in spin systems I: iid spins
Published • View PublicationBIB
For a sequence of Boolean functions $f_n : \{-1,1\}^{V_n} \longrightarrow \{-1,1\}$, defined on increasing configuration spaces of random inputs, we say that there is sparse reconstruction if there is a sequence of subsets $U_n \subseteq V_n$ of the coordinates satisfying $|U_n| = o(|V_n|)$ such that knowing the coordinates in $U_n$ gives us a non-vanishing amount of information about the value of $f_n$. We first show that, if the underlying measure is a product measure, then no sparse reconstruction is possible for any sequence of transitive functions. We discuss the question in different frameworks, measuring information content in $L^2$ and with entropy. We also highlight some interesting connections with cooperative game theory. Beyond transitive functions, we show that the left-right crossing event for critical planar percolation on the square lattice does not admit sparse reconstruction either. Some of these results answer questions posed by Itai Benjamini.
2020-10-19 v2
independence: Fast Rank Tests
In 1948 Hoeffding devised a nonparametric test that detects dependence between two continuous random variables X and Y, based on the ranking of n paired samples (Xi,Yi). The computation of this commonly-used test statistic takes O(n log n) time. Hoeffding's test is consistent against any dependent probability density f(x,y), but can be fooled by other bivariate distributions with continuous margins. Variants of this test with full consistency have been considered by Blum, Kiefer, and Rosenblatt (1961), Yanagimoto (1970), Bergsma and Dassios (2010). The so far best known algorithms to compute these stronger independence tests have required quadratic time. Here we improve their run time to O(n log n), by elaborating on new methods for counting ranking patterns, from a recent paper by the author and Leng (SODA'21). Therefore, in all circumstances under which the classical Hoeffding independence test is applicable, we provide novel competitive algorithms for consistent testing against all alternatives. Our R package, independence, offers a highly optimized implementation of these rank-based tests. We demonstrate its capabilities on large-scale datasets.
2020-10-18 v2
On the permanent of a random symmetric matrix
Published • View PublicationBIB
Let $M_{n}$ denote a random symmetric $n\times n$ matrix, whose entries on and above the diagonal are i.i.d. Rademacher random variables (taking values $\pm 1$ with probability $1/2$ each). Resolving a conjecture of Vu, we prove that the permanent of $M_{n}$ has magnitude $n^{n/2+o(n)}$ with probability $1-o(1)$. Our result can also be extended to more general models of random matrices.
Revisiting Shao and Sokal's $B_2$ index of phylogenetic balance
Published in Journal of Mathematical Biology 83:52 (2021) • View PublicationBIB
Measures of phylogenetic balance, such as the Colless and Sackin indices, play an important role in phylogenetics. Unfortunately, these indices are specifically designed for phylogenetic trees, and do not extend naturally to phylogenetic networks (which are increasingly used to describe reticulate evolution). This led us to consider a lesser-known balance index, whose definition is based on a probabilistic interpretation that is equally applicable to trees and to networks. This index, known as the $B_2$ index, was first proposed by Shao and Sokal in 1990. Surprisingly, it does not seem to have been studied mathematically since. Likewise, it is used only sporadically in the biological literature, where it tends to be viewed as arcane. In this paper, we study mathematical properties of $B_2$ such as its expectation and variance under the most common models of random trees and its extremal values over various classes of phylogenetic networks. We also assess its relevance in biological applications, and find it to be comparable to that of the Colless and Sackin indices. Altogether, our results call for a reevaluation of the status of this somewhat forgotten measure of phylogenetic balance.
2020-10-16
Central Limit Theorem for Majority Dynamics: Bribing Three Voters Suffices
Published • View PublicationBIB
Given a graph $G$ and some initial labelling $σ: V(G) \to \{Red, Blue\}$ of its vertices, the \textit{majority dynamics model} is the deterministic process where at each stage, every vertex simultaneously replaces its label with the majority label among its neighbors (remaining unchanged in the case of a tie). We prove---for a wide range of parameters---that if an initial assignment is fixed and we independently sample an Erdős--Rényi random graph, $G_{n,p}$, then after one step of majority dynamics, the number of vertices of each label follows a central limit law. As a corollary, we provide a strengthening of a theorem of Benjamini, Chan, O'Donnell, Tamuz, and Tan about the number of steps required for the process to reach unanimity when the initial assignment is also chosen randomly. Moreover, suppose there are initially three more red vertices than blue. In this setting, we prove that if we independently sample the graph $G_{n,1/2}$, then with probability at least $51\%$, the majority dynamics process will converge to every vertex being red. This improves a result of Tran and Vu who addressed the case that the initial lead is at least 10.
2020-10-16
The threshold for the square of a Hamilton cycle
Published • View PublicationBIB
Resolving a conjecture of Kühn and Osthus from 2012, we show that $p= 1/\sqrt{n}$ is the threshold for the random graph $G_{n,p}$ to contain the square of a Hamilton cycle.
2020-10-15
On sparse random combinatorial matrices
Published • View PublicationBIB
Let $Q_{n,d}$ denote the random combinatorial matrix whose rows are independent of one another and such that each row is sampled uniformly at random from the subset of vectors in $\{0,1\}^n$ having precisely $d$ entries equal to $1$. We present a short proof of the fact that $\Pr[\det(Q_{n,d})=0] = O\left(\frac{n^{1/2}\log^{3/2} n}{d}\right)=o(1)$, whenever $d=ω(n^{1/2}\log^{3/2} n)$. In particular, our proof accommodates sparse random combinatorial matrices in the sense that $d = o(n)$ is allowed. We also consider the singularity of deterministic integer matrices $A$ randomly perturbed by a sparse combinatorial matrix. In particular, we prove that $\Pr[\det(A+Q_{n,d})=0]=O\left(\frac{n^{1/2}\log^{3/2} n}{d}\right)$, again, whenever $d=ω(n^{1/2}\log^{3/2} n)$ and $A$ has the property that $(1,-d)$ is not an eigenpair of $A$.