Papers by David Gamarnik
21 paper(s) by this author
· All BibTeX
The stochastic block model has the overlap graph property for modularity
The overlap gap property (OGP) is a statement about the geometry of near-optimal solutions. Exhibiting OGP implies failure of a class of local algorithms; and has been observed to coincide with conjectured algorithmic limits in problems with statistical computational gap.
We consider the Stochastic Block Model (SBM), where the graph has a planted partition with $k$ equal-size blocks which form the `communities', and where, for parameters $p>q$, vertices within the same community connect with probability $p$, while vertices in different communities connect with probability $q$, independently across pairs of vertices. Modularity--based clustering algorithms have become ubiquitous in applications. This article studies theoretical limits of local algorithms based on the modularity score on the SBM.
We establish that modularity exhibits OGP on the SBM. This rules out a class of local algorithms based on modularity for recovery in the SBM, and shows slow mixing time for a related Markov Chain. Theoretically this is one of the few instances where OGP has been established for a `planted' model, as most such analyses to date consider the `null' model.
As part of our analysis, we extend a result by Bickel and Chen 2009, who established that with high probability, the modularity optimal partition of SBM is $o(n)$ local moves away from the planted partition, where $n$ is the graph size. We show that, with high probability, any partition with modularity score sufficiently near the optimal value is close to the planted partition.
Optimal Hardness of Online Algorithms for Large Common Induced Subgraphs
We study the problem of efficiently finding large common induced subgraphs of two independent Erdős--Rényi random graphs $G_1, G_2 \sim \mathbb{G}(n,1/2)$. Recently, Chatterjee and Diaconis showed that the largest common induced subgraph of $G_1$ and $G_2$ has size $(4-o(1))\log_2 n$ with high probability. We first show that a simple greedy online algorithm finds a common induced subgraph of $G_1$ and $G_2$ of size $(2-o(1)) \log_2 n$ with high probability. Our main result shows that no online algorithm can find a common induced subgraph of $G_1$ and $G_2$ of size at least $(2+\varepsilon) \log_2 n$ with probability bounded away from $0$ as $n \to \infty$. Together, these results provide evidence that this problem exhibits a computation-to-optimization gap. To prove the impossibility result, we show that the solution space of the problem exhibits a version of the (multi) overlap gap property (OGP), and utilize an interpolation argument recently developed by Gamarnik, Kizildağ, and Warnke that connects OGP and online algorithms.
Minimum Number of Monochromatic Subgraphs of a Random Graph
We consider the problem of minimizing the number of monochromatic subgraphs of a random graph, when each node of the host graph is assigned one of the two colors. Using a recently discovered contiguity between appearance of strictly balanced subgraphs $F$ in a random graph, and random hypergraphs where copies of $F$ are generated independently, we show that the minimum value converges to a limit, when the expected number of copies of $F$ is linear in the number of nodes $|V|$. Furthermore, using the connections with mean field spin glass models, we obtain an asymptotic expression for this limit as the normalized expected number of copies of $F$ and the size of $F$ diverge to infinity.
Optimal Hardness of Online Algorithms for Large Independent Sets
We study the algorithmic problem of finding a large independent set in the Erd{ö}s-Rényi random graph $G(n,p)$. For constant $p$ and $b=1/(1-p)$, the largest independent set has size $2\log_b n$, while a simple greedy algorithm - revealing vertices sequentially and making decisions based only on previously seen vertices - finds an independent set of size $\log_b n$. In his seminal 1976 paper, Karp challenged to either improve this guarantee or establish its hardness. Decades later, this problem remains open - one of the most prominent algorithmic problems in the theory of random graphs.
In this paper, we establish that a broad class of online algorithms fails to find an independent set of size $(1+ε)\log_b n$ whp. This class includes Karp's algorithm as a special case, and extends it by allowing the algorithm to query exceptional edges, not yet "seen" by the algorithm. Our lower bound holds for $p\in [d/n,1-n^{-1/d}]$. In the dense regime (constant $p$), we also prove that our result is asymptotically tight with respect to the number of exceptional edges queried, by designing an online algorithm which beats the half-optimality threshold when the number of exceptional edges slightly exceeds our bound.
Our result provides evidence for the algorithmic hardness of Karp's problem, by supporting the conjectured optimality of the greedy algorithm and establishing it within the class of online algorithms. Our proof relies on a refined analysis of the geometric structure of large independent sets, establishing a variant of the Overlap Gap Property (OGP). While OGP has predominantly served as a barrier to stable algorithms, online algorithms are inherently unstable, necessitating new ideas. Our proof refines the OGP framework by incorporating several new ideas (including temporal interpolation paths and stopping-times) that we expect to be useful for other online models.
Integrating High-Dimensional Functions Deterministically
We design a Quasi-Polynomial time deterministic approximation algorithm for computing the integral of a multi-dimensional separable function, supported by some underlying hyper-graph structure, appropriately defined. Equivalently, our integral is the partition function of a graphical model with continuous potentials. While randomized algorithms for high-dimensional integration are widely known, deterministic counterparts generally do not exist. We use the correlation decay method applied to the Riemann sum of the function to produce our algorithm. For our method to work, we require that the domain is bounded and the hyper-edge potentials are positive and bounded on the domain. We further assume that upper and lower bounds on the potentials separated by a multiplicative factor of $1 + O(1/Δ^2)$, where $Δ$ is the maximum degree of the graph. When $Δ= 3$, our method works provided the upper and lower bounds are separated by a factor of at most $1.0479$. To the best of our knowledge, our algorithm is the first deterministic algorithm for high-dimensional integration of a continuous function, apart from the case of trivial product form distributions.
Computing the Volume of a Restricted Independent Set Polytope Deterministically
We construct a quasi-polynomial time deterministic approximation algorithm for computing the volume of an independent set polytope with restrictions. Randomized polynomial time approximation algorithms for computing the volume of a convex body have been known now for several decades, but the corresponding deterministic counterparts are not available, and our algorithm is the first of this kind. The class of polytopes for which our algorithm applies arises as linear programming relaxation of the independent set problem with the additional restriction that each variable takes value in the interval $[0,1-α]$ for some $α<1/2$. (We note that the $α\ge 1/2$ case is trivial).
We use the correlation decay method for this problem applied to its appropriate and natural discretization. The method works provided $α> 1/2-O(1/Δ^2)$, where $Δ$ is the maximum degree of the graph. When $Δ=3$ (the sparsest non-trivial case), our method works provided $0.488<α<0.5$. Interestingly, the interpolation method, which is based on analyzing complex roots of the associated partition functions, fails even in the trivial case when the underlying graph is a singleton.
Maximally-stable Local Optima in Random Graphs and Spin Glasses: Phase Transitions and Universality
We consider $h$-stable local optima of Ising spin glass models, defined as spin configurations such that for nearly all of the spins, flipping their values results in increasing energy by at least a given amount $h$. Spins satisfying this condition are referred to as $h$-stable spins for that configuration. Similarly, we consider a very related notion of $h$-friendly partitions of a graph. These are defined as bi-partitionings such that for most nodes, the normalized number of neighbors within the node's partition exceed the normalized number of neighbors outside the partition by a certain amount $h$. For spin glasses as well as sparse and dense random graphs, while restricting to bisections, we prove the existence of a phase transition for the normalized energy level $h$ around a universal value $h^*$. For $h$ below the phase transition value $h^*$, bisections exist where the number of spins (nodes) which are not $h$-stable (not $h$-friendly) is sublinear. Above the phase transition level $h^*$ the smallest number of spins that are not $h$-stable (not $h$-friendly) is linear. This confirms a conjecture from Behrens et al. (2022). Our results also allow the characterization of possible energy values of stable local optima for varying $h$. In particular, for $h=0$, this rigorously proves seminal results in statistical physics regarding the so-called metastable states, such as in the work of Bray and Moore (1981). Our results extend a recent proof of the so-called Friendly Partition Conjecture in Ferber et al. (2022) from the case $h=0$ to the case when $h$ takes general values. Our proofs are obtained by analyzing the model on sparse random graphs and adopting Lindeberg's type universality method to lift the results from sparse to dense graphs and spin systems.
Cliques, Chromatic Number, and Independent Sets in the Semi-random Process
The semi-random graph process is a single player game in which the player is initially presented an empty graph on $n$ vertices. In each round, a vertex $u$ is presented to the player independently and uniformly at random. The player then adaptively selects a vertex $v$, and adds the edge $uv$ to the graph. For a fixed monotone graph property, the objective of the player is to force the graph to satisfy this property with high probability in as few rounds as possible. In this paper, we investigate the following three properties: containing a complete graph of order $k$, having the chromatic number at least $k$, and not having an independent set of size at least $k$.
Correlation Decay and the Absence of Zeros Property of Partition Functions
Published
• View Publication
• BIB
Absence of (complex) zeros property is at the heart of the interpolation method developed by Barvinok \cite{barvinok2017combinatorics} for designing deterministic approximation algorithms for various graph counting and computing partition functions problems. Earlier methods for solving the same problem include the one based on the correlation decay property. Remarkably, the classes of graphs for which the two methods apply sometimes coincide or nearly coincide. In this paper we show that this is more than just a coincidence. We establish that if the interpolation method is valid for a family of graphs satisfying the self-reducibility property, then this family exhibits a form of correlation decay property which is asymptotic Strong Spatial Mixing (SSM) at distances $ω(\log n)$, where $n$ is the number of nodes of the graph. This applies in particular to amenable graphs, such as graphs which are finite subsets of lattices.
Our proof is based on a certain graph polynomial representation of the associated partition function. This representation is at the heart of the design of the polynomial time algorithms underlying the interpolation method itself. We conjecture that our result holds for all, and not just amenable graphs.
Finding cliques using few probes
Published
• View Publication
• BIB
Consider algorithms with unbounded computation time that probe the entries of the adjacency matrix of an $n$ vertex graph, and need to output a clique. We show that if the input graph is drawn at random from $G_{n,\frac{1}{2}}$ (and hence is likely to have a clique of size roughly $2\log n$), then for every $δ< 2$ and constant $\ell$, there is an $α< 2$ (that may depend on $δ$ and $\ell$) such that no algorithm that makes $n^δ$ probes in $\ell$ rounds is likely (over the choice of the random graph) to output a clique of size larger than $α\log n$.
Suboptimality of local algorithms for a class of max-cut problems
Published in Annals of Probability 2019, Vol. 47, No. 3, 1587-1618
• View Publication
• BIB
We show that in random $K$-uniform hypergraphs of constant average degree, for even $K \geq 4$, local algorithms defined as factors of i.i.d. can not find nearly maximal cuts, when the average degree is sufficiently large. These algorithms have been used frequently to obtain lower bounds for the max-cut problem on random graphs, but it was not known whether they could be successful in finding nearly maximal cuts. This result follows from the fact that the overlap of any two nearly maximal cuts in such hypergraphs does not take values in a certain non-trivial interval - a phenomenon referred to as the overlap gap property - which is proved by comparing diluted models with large average degree with appropriate fully connected spin glass models and showing the overlap gap property in the latter setting.
Performance of the Survey Propagation-guided decimation algorithm for the random NAE-K-SAT problem
Published
• View Publication
• BIB
We show that the Survey Propagation-guided decimation algorithm fails to find satisfying assignments on random instances of the "Not-All-Equal-$K$-SAT" problem if the number of message passing iterations is bounded by a constant independent of the size of the instance and the clause-to-variable ratio is above $(1+o_K(1)){2^{K-1}\over K}\log^2 K$ for sufficiently large $K$. Our analysis in fact applies to a broad class of algorithms described as "sequential local algorithms". Such algorithms iteratively set variables based on some local information and then recurse on the reduced instance. Survey Propagation-guided as well as Belief Propagation-guided decimation algorithms - two widely studied message passing based algorithms, fall under this category of algorithms provided the number of message passing iterations is bounded by a constant. Another well-known algorithm falling into this category is the Unit Clause algorithm. Our work constitutes the first rigorous analysis of the performance of the SP-guided decimation algorithm.
The approach underlying our paper is based on an intricate geometry of the solution space of random NAE-$K$-SAT problem. We show that above the $(1+o_K(1)){2^{K-1}\over K}\log^2 K$ threshold, the overlap structure of $m$-tuples of satisfying assignments exhibit a certain clustering behavior expressed in the form of constraints on distances between the $m$ assignments, for appropriately chosen $m$. We further show that if a sequential local algorithm succeeds in finding a satisfying assignment with probability bounded away from zero, then one can construct an $m$-tuple of solutions violating these constraints, thus leading to a contradiction. Along with (citation), this result is the first work which directly links the clustering property of random constraint satisfaction problems to the computational hardness of finding satisfying assignments.
Limits of local algorithms over sparse random graphs
Published
• View Publication
• BIB
Local algorithms on graphs are algorithms that run in parallel on the nodes of a graph to compute some global structural feature of the graph. Such algorithms use only local information available at nodes to determine local aspects of the global structure, while also potentially using some randomness. Recent research has shown that such algorithms show significant promise in computing structures like large independent sets in graphs locally. Indeed the promise led to a conjecture by Hatami, \Lovasz and Szegedy \cite{HatamiLovaszSzegedy} that local algorithms may be able to compute maximum independent sets in (sparse) random $d$-regular graphs. In this paper we refute this conjecture and show that every independent set produced by local algorithms is multiplicative factor $1/2+1/(2\sqrt{2})$ smaller than the largest, asymptotically as $d\rightarrow\infty$.
Our result is based on an important clustering phenomena predicted first in the literature on spin glasses, and recently proved rigorously for a variety of constraint satisfaction problems on random graphs. Such properties suggest that the geometry of the solution space can be quite intricate. The specific clustering property, that we prove and apply in this paper shows that typically every two large independent sets in a random graph either have a significant intersection, or have a nearly empty intersection. As a result, large independent sets are clustered according to the proximity to each other. While the clustering property was postulated earlier as an obstruction for the success of local algorithms, such as for example, the Belief Propagation algorithm, our result is the first one where the clustering property is used to formally prove limits on local algorithms.
Convergent sequences of sparse graphs: A large deviations approach
Published
• View Publication
• BIB
In this paper we introduce a new notion of convergence of sparse graphs which we call Large Deviations or LD-convergence and which is based on the theory of large deviations. The notion is introduced by "decorating" the nodes of the graph with random uniform i.i.d. weights and constructing random measures on $[0,1]$ and $[0,1]^2$ based on the decoration of nodes and edges. A graph sequence is defined to be converging if the corresponding sequence of random measures satisfies the Large Deviations Principle with respect to the topology of weak convergence on bounded measures on $[0,1]^d, d=1,2$. We then establish that LD-convergence implies several previous notions of convergence, namely so-called right-convergence, left-convergence, and partition-convergence. The corresponding large deviation rate function can be interpreted as the limit object of the sparse graph sequence. In particular, we can express the limiting free energies in terms of this limit object.
Strong spatial mixing for list coloring of graphs
Published
• View Publication
• BIB
The property of spatial mixing and strong spatial mixing in spin systems has been of interest because of its implications on uniqueness of Gibbs measures on infinite graphs and efficient approximation of counting problems that are otherwise known to be #P hard. In the context of coloring, strong spatial mixing has been established for regular trees when $q \geq α^{*} Δ+ 1$ where $q$ the number of colors, $Δ$ is the degree and $α^* = 1.763..$ is the unique solution to $xe^{-1/x} = 1$. It has also been established for bounded degree lattice graphs whenever $q \geq α^* Δ- β$ for some constant $β$, where $Δ$ is the maximum vertex degree of the graph. The latter uses a technique based on recursively constructed coupling of Markov chains whereas the former is based on establishing decay of correlations on the tree. We establish strong spatial mixing of list colorings on arbitrary bounded degree triangle-free graphs whenever the size of the list of each vertex $v$ is at least $αΔ(v) + β$ where $Δ(v)$ is the degree of vertex $v$ and $α> α^*$ and $β$ is a constant that only depends on $α$. We do this by proving the decay of correlations via recursive contraction of the distance between the marginals measured with respect to a suitably chosen error function.
Sequential cavity method for computing free energy and surface pressure
Published
• View Publication
• BIB
We propose a new method for the problems of computing free energy and surface pressure for various statistical mechanics models on a lattice $\Z^d$. Our method is based on representing the free energy and surface pressure in terms of certain marginal probabilities in a suitably modified sublattice of $\Z^d$. Then recent deterministic algorithms for computing marginal probabilities are used to obtain numerical estimates of the quantities of interest. The method works under the assumption of Strong Spatial Mixing (SSP), which is a form of a correlation decay.
We illustrate our method for the hard-core and monomer-dimer models, and improve several earlier estimates. For example we show that the exponent of the monomer-dimer coverings of $\Z^3$ belongs to the interval $[0.78595,0.78599]$, improving best previously known estimate of (approximately) $[0.7850,0.7862]$ obtained in \cite{FriedlandPeled},\cite{FriedlandKropLundowMarkstrom}. Moreover, we show that given a target additive error $ε>0$, the computational effort of our method for these two models is $(1/ε)^{O(1)}$ \emph{both} for free energy and surface pressure. In contrast, prior methods, such as transfer matrix method, require $\exp\big((1/ε)^{O(1)}\big)$ computation effort.
A Deterministic Approximation Algorithm for Computing a Permanent of a 0,1 matrix
We construct a deterministic approximation algorithm for computing a permanent of a $0,1$ $n$ by $n$ matrix to within a multiplicative factor $(1+ε)^n$, for arbitrary $ε>0$. When the graph underlying the matrix is a constant degree expander our algorithm runs in polynomial time (PTAS). In the general case the running time of the algorithm is $\exp(O(n^{2\over 3}\log^3n))$. For the class of graphs which are constant degree expanders the first result is an improvement over the best known approximation factor $e^n$ obtained in \cite{LinialSamorodnitskyWigderson}.
Our results use a recently developed deterministic approximation algorithm for counting partial matchings of a graph Bayati et al., and Jerrum-Vazirani decomposition method.
Correlation decay and deterministic FPTAS for counting list-colorings of a graph
Published
• View Publication
• BIB
We propose a deterministic algorithm for approximately counting the number of list colorings of a graph. Under the assumption that the graph is triangle free, the size of every list is at least $αΔ$, where $α$ is an arbitrary constant bigger than $α^{**}=2.8432...$, and $Δ$ is the maximum degree of the graph, we obtain the following results. For the case when the size of the each list is a large constant, we show the existence of a \emph{deterministic} FPTAS for computing the total number of list colorings. The same deterministic algorithm has complexity $2^{O(\log^2 n)}$, without any assumptions on the sizes of the lists, where $n$ is the instance size. We further extend our method to a discrete Markov random field (MRF) model. Under certain assumptions relating the size of the alphabet, the degree of the graph and the interacting potentials we again construct a deterministic FPTAS for computing the partition function of a MRF.
Our results are not based on the most powerful existing counting technique -- rapidly mixing Markov chain method. Rather we build upon concepts from statistical physics, in particular, the decay of correlation phenomena and its implication for the uniqueness of Gibbs measures in infinite graphs. This approach was proposed in two recent papers \cite{BandyopadhyayGamarnikCounting} and \cite{weitzCounting}. The principle insight of this approach is that the correlation decay property can be established with respect to certain \emph{computation tree}, as opposed to the conventional correlation decay property with respect to graph theoretic neighborhoods of a given node. This allows truncation of computation at a logarithmic depth in order to obtain polynomial accuracy in polynomial time.
Maximum Weight Independent Sets and Matchings in Sparse Random Graphs. Exact Results using the Local Weak Convergence Method
Published
• View Publication
• BIB
Let $G(n,c/n)$ and $G_r(n)$ be an $n$-node sparse random graph and a sparse random $r$-regular graph, respectively, and let ${\cal I}(n,r)$ and ${\cal I}(n,c)$ be the sizes of the largest independent set in $G(n,c/n)$ and $G_r(n)$. The asymptotic value of ${\cal I}(n,c)/n$ as $n\to\infty$, can be computed using the Karp-Sipser algorithm when $c\leq e$. For random cubic graphs, $r=3$, it is only known that $.432\leq\liminf_n {\cal I}(n,3)/n \leq \limsup_n {\cal I}(n,3)\leq .4591$ with high probability (w.h.p.) as $n\to\infty$, as shown by Frieze and Suen and by Bollobas, respectively.
In this paper we assume in addition that the nodes of the graph are equipped with non-negative weights, independently generated according to some common distribution, and we consider instead the maximum weight of an independent set. Surprisingly, we discover that for certain weight distributions, the limit $\lim_n {\cal I}(n,c)/n$ can be computed exactly even when $c>e$, and $\lim_n {\cal I}(n,r)/n$ can be computed exactly for some $r\geq 2$. For example, when the weights are exponentially distributed with parameter 1, $\lim_n {\cal I}(n,2e)/n\approx .5517$, and $\lim_n {\cal I}(n,3)/n\approx .6077$. Our results are established using the recently developed local weak convergence method further reduced to a certain local optimality property exhibited by the models we consider.
Random MAX SAT, Random MAX CUT, and Their Phase Transitions
Published
• View Publication
• BIB
Given a 2-SAT formula $F$ consisting of $n$ variables and $\cn$ random clauses, what is the largest number of clauses $\max F$ satisfiable by a single assignment of the variables? We bound the answer away from the trivial bounds of $(3/4)cn$ and $cn$. We prove that for $c<1$, the expected number of clauses satisfiable is $\cn-Θ(1/n)$; for large $c$, it is $((3/4)c + Θ(\sqrt{c}))n$; for $c = 1+\eps$, it is at least $(1+\eps-O(\eps^3))n$ and at most $(1+\eps-Ω(\eps^3/\ln \eps))n$; and in the ``scaling window'' $c= 1+Θ(n^{-1/3})$, it is $cn-Θ(1)$. In particular, just as the decision problem undergoes a phase transition, our optimization problem also undergoes a phase transition at the same critical value $c=1$.
Nearly all of our results are established without reference to the analogous propositions for decision 2-SAT, and as a byproduct we reproduce many of those results, including much of what is known about the 2-SAT scaling window.
We consider ``online'' versions of MAX-2-SAT, and show that for one version, the obvious greedy algorithm is optimal.
We can extend only our simplest MAX-2-SAT results to MAX-k-SAT, but we conjecture a ``MAX-k-SAT limiting function conjecture'' analogous to the folklore satisfiability threshold conjecture, but open even for $k=2$. Neither conjecture immediately implies the other, but it is natural to further conjecture a connection between them.
Finally, for random MAXCUT (the size of a maximum cut in a sparse random graph) we prove analogous results.