Papers by Pu Gao
46 paper(s) by this author
· All BibTeX
Hamiltonicity of random graphs in the stochastic block model
Published
• View Publication
• BIB
We study the Hamiltonicity of the following model of a random graph. Suppose that we partition [n] into V_1,V_2,...,V_k and add edge {x,y} to our graph with probability p if there exists i such that x,y\in V_i. Otherwise, we add the edge with probbability q. We denote this model by G(n, p,q) and give tight results for Hamiltonicity, including a critical window analysis, under various conditions.
The rank of sparse random matrices
Published
• View Publication
• BIB
We determine the rank of a random matrix over an arbitrary field with prescribed numbers of non-zero entries in each row and column. As an application we obtain a formula for the rate of low-density parity check codes. This formula vindicates a conjecture of Lelarge (2013). The proofs are based on coupling arguments and a novel random perturbation, applicable to any matrix, that diminishes the number of short linear relations.
Sandwiching random regular graphs between binomial random graphs
Published
• View Publication
• BIB
Kim and Vu made the following conjecture (\textit{Advances in Mathematics}, 2004): if $d\gg \log n$, then the random $d$-regular graph $\mathcal G(n,d)$ can asymptotically almost surely be "sandwiched" between $\mathcal G(n,p_1)$ and $\mathcal G(n,p_2)$ where $p_1$ and $p_2$ are both $(1+o(1))d/n$. They proved this conjecture for $\log n\ll d\le n^{1/3-o(1)}$, with a defect in the sandwiching: $\mathcal G(n,d)$ contains $\mathcal G(n,p_1)$ perfectly, but is not completely contained in $\mathcal G(n,p_2)$. Recently, the embedding $\mathcal G(n,p_1) \subseteq \mathcal G(n,d)$ was improved by Dudek, Frieze, Ruciński and Šileikis to $d=o(n)$. In this paper, we prove Kim--Vu's sandwich conjecture, with perfect containment on both sides, for all $d\gg n/\sqrt{\log n}$. For $d=O(n/\sqrt{\log n})$, we prove a weaker version of the sandwich conjecture with $p_2$ approximately equal to $(d/n)\log n$, without any defect. In addition to sandwiching regular graphs, our results cover graphs whose degrees are asymptotically equal. The proofs rely on estimates for the probability that a random factor of a pseudorandom graph contains a given edge, which is of independent interest.
As applications, we obtain new results on the properties of random graphs with given near-regular degree sequences, including Hamiltonicity and universality in subgraph containment. We also determine several graph parameters in these random graphs, such as the chromatic number, small subgraph counts, the diameter, and the independence number. We are also able to characterise many phase transitions in edge percolation on these random graphs, such as the threshold for the appearance of a giant component.
Fast uniform generation of random graphs with given degree sequences
In this paper we provide an algorithm that generates a graph with given degree sequence uniformly at random. Provided that $Δ^4=O(m)$, where $Δ$ is the maximal degree and $m$ is the number of edges,the algorithm runs in expected time $O(m)$. Our algorithm significantly improves the previously most efficient uniform sampler, which runs in expected time $O(m^2Δ^2)$ for the same family of degree sequences. Our method uses a novel ingredient which progressively relaxes restrictions on an object being generated uniformly at random, and we use this to give fast algorithms for uniform sampling of graphs with other degree sequences as well. Using the same method, we also obtain algorithms with expected run time which is (i) linear for power-law degree sequences in cases where the previous best was $O(n^{4.081})$, and (ii) $O(nd+d^4)$ for $d$-regular graphs when $d=o(\sqrt n)$, where the previous best was $O(nd^3)$.
The rank of random matrices over finite fields
We determine the rank of a random matrix A over a finite field with prescribed numbers of non-zero entries in each row and column. As an application we obtain a formula for the rate of low-density parity check codes. This formula verifies a conjecture of Lelarge [Proc. IEEE Information Theory Workshop 2013]. The proofs are based on coupling arguments and the interpolation method from mathematical physics.
Uniform generation of spanning regular subgraphs of a dense graph
Published
• View Publication
• BIB
Let $H_n$ be a graph on $n$ vertices and let $\ber{H_n}$ denote the complement of $H_n$. Suppose that $Δ= Δ(n)$ is the maximum degree of $\ber{H_n}$. We analyse three algorithms for sampling $d$-regular subgraphs ($d$-factors) of $H_n$. This is equivalent to uniformly sampling $d$-regular graphs which avoid a set $E(\ber{H_n})$ of forbidden edges. Here $d=d(n)$ is a positive integer which may depend on $n$.
Two of these algorithms produce a uniformly random $d$-factor of $H_n$ in expected runtime which is linear in $n$ and low-degree polynomial in $d$ and $Δ$. The first algorithm applies when $(d+Δ)dΔ= o(n)$. This improves on an earlier algorithm by the first author, which required constant $d$ and at most a linear number of edges in $\ber{H_n}$. The second algorithm applies when $H_n$ is regular and $d^2+Δ^2 = o(n)$, adapting an approach developed by the first author together with Wormald. The third algorithm is a simplification of the second, and produces an approximately uniform $d$-factor of $H_n$ in time $O(dn)$. Here the output distribution differs from uniform by $o(1)$ in total variation distance, provided that $d^2+Δ^2 = o(n)$.
The satisfiability threshold for random linear equations
Published
• View Publication
• BIB
Let $A$ be a random $m\times n$ matrix over the finite field $F_q$ with precisely $k$ non-zero entries per row and let $y\in F_q^m$ be a random vector chosen independently of $A$. We identify the threshold $m/n$ up to which the linear system $A x=y$ has a solution with high probability and analyse the geometry of the set of solutions. In the special case $q=2$, known as the random $k$-XORSAT problem, the threshold was determined by [Dubois and Mandler 2002, Dietzfelbinger et al. 2010, Pittel and Sorkin 2016], and the proof technique was subsequently extended to the cases $q=3,4$ [Falke and Goerdt 2012]. But the argument depends on technically demanding second moment calculations that do not generalise to $q>3$. Here we approach the problem from the viewpoint of a decoding task, which leads to a transparent combinatorial proof.
Counterexamples on matchings in hypergraphs and full rainbow matchings in graphs
A graph $G$ whose edges are coloured (not necessarily properly) contains a full rainbow matching if there is a matching $M$ that contains exactly one edge of each colour. We refute several conjectures on matchings in hypergraphs and full rainbow matchings in graphs, made by Aharoni and Berger and others.
Full rainbow matchings in graphs and hypergraphs
Published in Combinatorics, Probability and Computing 30, (2021) 762-780
• View Publication
• BIB
Let $G$ be a simple graph that is properly edge coloured with $m$ colours and let $\M=\{M_1,\ldots, M_m\}$ be the set of $m$ matchings induced by the colours in $G$. Suppose that $m\le n-n^{c}$, where $c>9/10$, and every matching in $\M$ has size $n$. Then $G$ contains a full rainbow matching, i.e.\ a matching that contains exactly one edge from $M_i$ for each $1\le i\le m$. This answers an open problem of Pokrovskiy and gives an affirmative answer to a generalisation of a special case of a conjecture of Aharoni and Berger.
Related results are also found for multigraphs with edges of bounded multiplicity, and for hypergraphs.
Finally, we provide counterexamples to several conjectures on full rainbow matchings made by Aharoni and Berger.
Uniform generation of random graphs with power-law degree sequences
Published
• View Publication
• BIB
We give a linear-time algorithm that approximately uniformly generates a random simple graph with a power-law degree sequence whose exponent is at least 2.8811. While sampling graphs with power-law degree sequence of exponent at least 3 is fairly easy, and many samplers work efficiently in this case, the problem becomes dramatically more difficult when the exponent drops below 3; ours is the first provably practicable sampler for this case. We also show that with an appropriate rejection scheme, our algorithm can be tuned into an exact uniform sampler. The running time of the exact sampler is O(n^{2.107}) with high probability, and O(n^{4.081}) in expectation.
Inside the clustering window for random linear equations
Published
• View Publication
• BIB
We study a random system of cn linear equations over n variables in GF(2), where each equation contains exactly r variables; this is equivalent to r-XORSAT. Previous work has established a clustering threshold, c^*_r for this model: if c=c_r^*-εfor any constant ε>0 then with high probability all solutions form a well-connected cluster; whereas if c=c^*_r+ε, then with high probability the solutions partition into well-connected, well-separated clusters (with probability tending to 1 as n goes to infinity). This is part of a general clustering phenomenon which is hypothesized to arise in most of the commonly studied models of random constraint satisfaction problems, via sophisticated but mostly non-rigorous techniques from statistical physics. We extend that study to the range c=c^*_r+o(1), and prove that the connectivity parameters of the r-XORSAT clusters undergo a smooth transition around the clustering threshold.
Uniform generation of random regular graphs
Published
• View Publication
• BIB
We develop a new approach for uniform generation of combinatorial objects, and apply it to derive a uniform sampler REG for d-regular graphs. REG can be implemented such that each graph is generated in expected time O(nd^3), provided that d=o(n^{1/2}). Our result significantly improves the previously best uniform sampler, which works efficiently only when d=O(n^{1/3}), with essentially the same running time for the same d. We also give a linear-time approximate sampler REG*, which generates a random d-regular graph whose distribution differs from the uniform by o(1) in total variation distance, when d=o(n^{1/2}).
The stripping process can be slow: part II
Published
• View Publication
• BIB
This paper is a continuation of the previous results on the stripping number of a random uniform hypergraph, and the maximum depth over all non-k-core vertices. The previous results focus on the supercritical case, whereas this work analyses these parameters in the subcritical regime and inside the critical window.
The stripping process can be slow: part I
Published
• View Publication
• BIB
Given an integer k, we consider the parallel k-stripping process applied to a hypergraph H: removing all vertices with degree less than k in each iteration until reaching the k-core of H. Take H as H_r(n,m): a random r-uniform hypergraph on n vertices and m hyperedges with the uniform distribution. Fixing k,r\ge 2 with (k,r)\neq (2,2), it has previously been proved that there is a constant c_{r,k} such that for all m=cn with constant c\neq c_{r,k}, with high probability, the parallel k-stripping process takes O(\log n) iterations. In this paper we investigate the critical case when c=c_{r,k}+o(1). We show that the number of iterations that the process takes can go up to some power of n, as long as c approaches c_{r,k} sufficiently fast. A second result we show involves the depth of a non-k-core vertex v: the minimum number of steps required to delete v from H_r(n,m) where in each step one vertex with degree less than k is removed. We will prove lower and upper bounds on the maximum depth over all non-k-core vertices.
Enumeration of graphs with a heavy-tailed degree sequence
Published
• View Publication
• BIB
In this paper, we asymptotically enumerate graphs with a given degree sequence d=(d_1,...,d_n) satisfying restrictions designed to permit heavy-tailed sequences in the sparse case (i.e. where the average degree is rather small). Our general result requires upper bounds on functions of M_k= \sum_{i=1}^n [d_i]_k for a few small integers k\ge 1. Note that M_1 is simply the total degree of the graphs. As special cases, we asymptotically enumerate graphs with (i) degree sequences satisfying M_2=o(M_1^{ 9/8}); (ii) degree sequences following a power law with parameter gamma>5/2; (iii) power-law degree sequences that mimic independent power-law "degrees" with parameter gamma>1+\sqrt{3}\approx 2.732; (iv) degree sequences following a certain "long-tailed" power law; (v) certain bi-valued sequences. A previous result on sparse graphs by McKay and the second author applies to a wide range of degree sequences but requires Delta =o(M_1^{1/3}), where Delta is the maximum degree. Our new result applies in some cases when Delta is only barely o(M_1^ {3/5}). Case (i) above generalises a result of Janson which requires M_2=O(M_1) (and hence M_1=O(n) and Delta=O(n^{1/2})). Cases (ii) and (iii) provide the first asymptotic enumeration results applicable to degree sequences of real-world networks following a power law, for which it has been empirically observed that 2<gamma<3.
Analysis of the parallel peeling algorithm: a short proof
A recent paper by Jiang, Mitzenmacher and Thaler upper bounded the number of rounds needed in a parallel peeling algorithm applied to a random hypergraph whose edge density is below the k-core emergence threshold. I gave a very short proof of their result in this note.
On the Geometric Ramsey Number of Outerplanar Graphs
Published in Discrete and Computational Geometry 53 (1): 64-79 (2015)
• View Publication
• BIB
We prove polynomial upper bounds of geometric Ramsey numbers of pathwidth-2 outerplanar triangulations in both convex and general cases. We also prove that the geometric Ramsey numbers of the ladder graph on $2n$ vertices are bounded by $O(n^{3})$ and $O(n^{10})$, in the convex and general case, respectively. We then apply similar methods to prove an $n^{O(\log(n))}$ upper bound on the Ramsey number of a path with $n$ ordered vertices.
Inside the clustering threshold for random linear equations
We study a random system of $cn$ linear equations over $n$ variables in GF(2), where each equation contains exactly $r$ variables; this is equivalent to $r$-XORSAT. \cite{ikkm,amxor} determined the clustering threshold, $c^*_r$: if $c=c^*_r+\e$ for any constant $\e>0$, then \aas the solutions partition into well-connected, well-separated {\em clusters} (with probability tending to 1 as $n\rightarrow\infty$). This is part of a general clustering phenomenon which is hypothesized to arise in most of the commonly studied models of random constraint satisfaction problems, via sophisticated but mostly non-rigorous techniques from statistical physics. We extend that study to the range $c=c^*_r+o(1)$, showing that if $c=c^*_r+n^{-\d}, \d>0$, then the connectivity parameter of each $r$-XORSAT cluster is $n^{Θ(\d)}$, as compared to $O(\log n)$ when $c=c^*_r+\e$. This means that one can move between any two solutions in the same cluster via a sequence of solutions where consecutive solutions differ on at most $n^{Θ(\d)}$ variables; this is tight up to the implicit constant. In contrast, moving to a solution in another cluster requires that some pair of consecutive solutions differ in at least $n^{1-O(\d)}$ variables.
Along the way, we prove that in a random $r$-uniform hypergraph with edge-density $n^{-\d}$ above the $k$-core threshold, \aas every vertex not in the $k$-core can be removed by a sequence of $n^{Θ(\d)}$ vertex-deletions in which the deleted vertex has degree less than $k$; again, this is tight up to the implicit constant.
On the geometric Ramsey numbers of trees
Published
• View Publication
• BIB
In this paper, we obtain upper bounds for the geometric Ramsey numbers of trees. We prove that $R_c(T_n,H_m)=(n-1)(m-1)+1$ if $T_n$ is a caterpillar and $H_m$ is a Hamiltonian outerplanar graph on $m$ vertices. Moreover, if $T_n$ has at most two non-leaf vertices, then $R_g(T_n,H_m)=(n-1)(m-1)+1$. We also prove that $R_c(T_n,H_m)=O(n^2m)$ and $R_g(T_n,H_m)=O(n^3m^2)$ if $T_n$ is an arbitrary tree on $n$ vertices and $H_m$ is an outerplanar triangulation with pathwidth 2. %Further, we prove a uniform polynomial upper bound for the geometric Ramsey numbers of caterpillars and we also give an upper bound for $R_g(T_n)$ where $T_n$ is an arbitrary tree.
A transition of limiting distributions of large matchings in random graphs
Published
• View Publication
• BIB
We study the asymptotic distribution of the number of matchings of size $\ell=\ell(n)$ in $G(n,p)$ for a wide range of $p=p(n)\in (0,1)$ and for every $1\le \ell\le \lfloor n/2\rfloor$. We prove that this distribution changes from normal to log-normal as $\ell$ increases, and we determine the critical value of $\ell$, as a function of $n$ and $p$, at which the transition of the limiting distribution occurs.