random
6952 papers tagged with this keyword
Upper bounds on the average edit distance between two random strings
We study the average edit distance between two random strings. More precisely, we adapt a technique introduced by Lueker in the context of the average longest common subsequence of two random strings to improve the known upper bound on the average edit distance. We improve all the known upper bounds for small alphabets. We also provide a new implementation of Lueker technique to improve the lower bound on the average length of the longest common subsequence of two random strings for all small alphabets of size other than $2$ and $4$.
Distance Reconstruction of Sparse Random Graphs
In the distance query model, we are given access to the vertex set of a $n$-vertex graph $G$, and an oracle that takes as input two vertices and returns the distance between these two vertices in $G$. We study how many queries are needed to reconstruct the edge set of $G$ when $G$ is sampled according to the $G(n,p)$ Erdős-Renyi-Gilbert distribution. Our approach applies to a large spectrum of values for $p$ starting slightly above the connectivity threshold: $p \geq \frac{2000 \log n}{n}$. We show that there exists an algorithm that reconstructs $G \sim G(n,p)$ using $O( Δ^2 n \log n )$ queries in expectation, where $Δ$ is the expected average degree of $G$. In particular, for $p \in [\frac{2000 \log n}{n}, \frac{\log^2 n}{n}]$ the algorithm uses $O(n \log^5 n)$ queries.
Large matchings and nearly spanning, nearly regular subgraphs of random subgraphs
Given a graph $G$ and $p\in [0,1]$, the random subgraph $G_p$ is obtained by retaining each edge of $G$ independently with probability $p$. We show that for every $ε>0$, there exists a constant $C>0$ such that the following holds. Let $d\ge C$ be an integer, let $G$ be a $d$-regular graph and let $p\ge \frac{C}{d}$. Then, with probability tending to one as $|V(G)|$ tends to infinity, there exists a matching in $G_p$ covering at least $(1-ε)|V(G)|$ vertices.
We further show that for a wide family of $d$-regular graphs $G$, which includes the $d$-dimensional hypercube, for any $p\ge \frac{\log^5d}{d}$ with probability tending to one as $d$ tends to infinity, $G_p$ contains an induced subgraph on at least $(1-o(1))|V(G)|$ vertices, whose degrees are tightly concentrated around the expected average degree $dp$.
Inference of rankings planted in random tournaments
We consider the problem of inferring an unknown ranking of $n$ items from a random tournament on $n$ vertices whose edge directions are correlated with the ranking. We establish, in terms of the strength of these correlations, the computational and statistical thresholds for detection (deciding whether an observed tournament is purely random or drawn correlated with a hidden ranking) and recovery (estimating the hidden ranking with small error in Spearman's footrule or Kendall's tau metric on permutations). Notably, we find that this problem provides a new instance of a detection-recovery gap: solving the detection problem requires much weaker correlations than solving the recovery problem. In establishing these thresholds, we also identify simple algorithms for detection (thresholding a degree 2 polynomial) and recovery (outputting a ranking by the number of "wins" of a tournament vertex, i.e., the out-degree) that achieve optimal performance up to constants in the correlation strength. For detection, we find that the above low-degree polynomial algorithm is superior to a natural spectral algorithm. We also find that, whenever it is possible to achieve strong recovery (i.e., to estimate with vanishing error in the above metrics) of the hidden ranking, then the above "Ranking By Wins" algorithm not only does so, but also outputs a close approximation of the maximum likelihood estimator, a task that is NP-hard in the worst case.
Repeated Block Averages: entropic time and mixing profiles
We consider randomized dynamics over the $n$-simplex, where at each step a random set, or block, of coordinates is evenly averaged. When all blocks have size 2, this reduces to the repeated averages studied in [CDSZ22], a version of the averaging process on a graph [AL12]. We study the convergence to equilibrium of this process as a function of the distribution of the block size, and provide sharp conditions for the emergence of the cutoff phenomenon. Moreover, we characterize the size of the cutoff window and provide an explicit Gaussian cutoff profile. To complete the analysis, we study in detail the simplified case where the block size is not random. We show that the absence of a cutoff is equivalent to having blocks of size $n^{Ω(1)}$, in which case we provide a convergence in distribution for the total variation distance at any given time, showing that, on the proper time scale, it remains constantly 1 up to an exponentially distributed random time, after which it decays following a Poissonian profile.
A note on the probability of a groupoid having deficient sets
A subset $X$ of a groupoid is said to be deficient if $|X \cdot X|\leq |X|$. It is well-known that the probability that a random groupoid has a deficient $t$-element set with $t\geq 3$ is zero. However, as conjectured in [4], we show that the probability is not zero in the case of sets of two elements and calculate the exact value. We explore some generalisations on deficient sets and their likelihoods.
Phase transition for tree-rooted maps
Published in 35th International Conference on Probabilistic, Combinatorial and Asymptotic Methods for the Analysis of Algorithms (AofA 2024). Leibniz International Proceedings in Informatics (LIPIcs), Volume 302, pp. 6:1-6:14
• View Publication
• BIB
We introduce a model of tree-rooted planar maps weighted by their number of $2$-connected blocks. We study its enumerative properties and prove that it undergoes a phase transition. We give the distribution of the size of the largest $2$-connected blocks in the three regimes (subcritical, critical and supercritical) and further establish that the scaling limit is the Brownian Continuum Random Tree in the critical and supercritical regimes, with respective rescalings $\sqrt{n/\log(n)}$ and $\sqrt{n}$.
Distributional limits of graph cuts on discretized grids
Published in Electron. J. Statist. 19(2): 5925-5978 (2025)
• View Publication
• BIB
Graph cuts are among the most prominent tools for clustering and classification analysis. While intensively studied from geometric and algorithmic perspectives, graph cut-based statistical inference still remains elusive to a certain extent. Distributional limits are fundamental in understanding and designing such statistical procedures on randomly sampled data. We provide explicit limiting distributions for balanced graph cuts in general on a fixed but arbitrary discretization. In particular, we show that Minimum Cut, Ratio Cut and Normalized Cut behave asymptotically as the minimum of Gaussians as sample size increases. Interestingly, our results reveal a dichotomy for Cheeger Cut: The limiting distribution of the optimal objective value is the minimum of Gaussians only when the optimal partition yields two sets of unequal volumes, while otherwise the limiting distribution is the minimum of a random mixture of Gaussians. Further, we show the bootstrap consistency for all types of graph cuts by utilizing the directional differentiability of cut functionals. We validate these theoretical findings by Monte Carlo experiments, and examine differences between the cuts and the dependency on the underlying distribution. Additionally, we expand our theoretical findings to the Xist algorithm, a computational surrogate of graph cuts recently proposed in Suchan, Li and Munk (arXiv, 2023), thus demonstrating the practical applicability of our findings e.g. in statistical tests.
Large random matrices with given margins
We study large random matrices with i.i.d. entries conditioned to have prescribed row and column sums (margins), a problem connected to relative entropy minimization, Schrödinger bridges, contingency tables, and random graphs with given degree sequences. Our central result is a `transference principle': the complex margin-conditioned matrix can be closely approximated by a simpler matrix whose entries are independent and drawn from an exponential tilting of the original model. The tilt parameters are determined by the sum of two potentials. We establish phase diagrams for `tame margins', where these potentials are uniformly bounded. This framework resolves a 2011 conjecture by Chatterjee, Diaconis, and Sly on $δ$-tame degree sequences and generalizes a sharp phase transition in contingency tables obtained by Dittmer, Lyu, and Pak in 2020. For tame margins, we show that a generalized Sinkhorn algorithm can compute the potentials at a dimension-free exponential rate. Our limit theory further establishes that for a convergent sequence of tame margins, the potentials converge as fast as the margins converge.
We apply this framework and obtain several key results for the conditioned matrix: The marginal distribution of any single entry is asymptotically an exponential tilting of the base measure, resolving a 2010 conjecture by Barvinok on contingency tables. The conditioned matrix concentrates in cut norm around a `typical table' (the expectation of the tilted model), which acts as a static Schrödinger bridge between the margins. The empirical singular value distribution of the rescaled matrix converges to an explicit law determined by the variance profile of the tilted model. In particular, we confirm the universality of the Marchenko-Pastur law for constant linear margins.
Coprime networks of the composite numbers: pseudo-randomness and synchronizability
Published in Discrete Applied Mathematics, 355(2024)96
• View Publication
• BIB
In this paper, we propose a network whose nodes are labeled by the composite numbers and two nodes are connected by an undirected link if they are relatively prime to each other. As the size of the network increases, the network will be connected whenever the largest possible node index $n\geq 49$. To investigate how the nodes are connected, we analytically describe that the link density saturates to $6/π^2$, whereas the average degree increases linearly with slope $6/π^2$ with the size of the network. To investigate how the neighbors of the nodes are connected to each other, we find the shortest path length will be at most 3 for $49\leq n\leq 288$ and it is at most 2 for $n\geq 289$. We also derive an analytic expression for the local clustering coefficients of the nodes, which quantifies how close the neighbors of a node to form a triangle. We also provide an expression for the number of $r$-length labeled cycles, which indicates the existence of a cycle of length at most $O(\log n)$. Finally, we show that this graph sequence is actually a sequence of weakly pseudo-random graphs. We numerically verify our observed analytical results. As a possible application, we have observed less synchronizability (the ratio of the largest and smallest positive eigenvalue of the Laplacian matrix is high) as compared to Erdős-Rényi random network and Barabási-Albert network. This unusual observation is consistent with the prolonged transient behaviors of ecological and predator-prey networks which can easily avoid the global synchronization.
Rainbow connectivity of multilayered random geometric graphs
An edge-colored multigraph $G$ is rainbow connected if every pair of vertices is joined by at least one rainbow path, i.e., a path where no two edges are of the same color.
In the context of multilayered networks we introduce the notion of multilayered random geometric graphs, from $h\ge 2$ independent random geometric graphs $G(n,r)$ on the unit square. We define an edge-coloring by coloring the edges according to the copy of $G(n,r)$ they belong to and study the rainbow connectivity of the resulting edge-colored multigraph. We show that $r(n)=\left(\frac{\log n}{n}\right)^{\frac{h-1}{2h}}$ is a threshold of the radius for the property of being rainbow connected. This complements the known analogous results for the multilayerd graphs defined on the Erdős-R\' enyi random model.
A concentration inequality for random combinatorial optimisation problems
Given a finite set $S$, i.i.d. random weights $\{X_i\}_{i\in S}$, and a family of subsets $\mathcal{F}\subseteq 2^S$, we consider the minimum weight of an $F\in \mathcal{F}$: \[ M(\mathcal{F}):= \min_{F\in \mathcal{F}} \sum_{i\in F}X_i. \] In particular, we investigate under what conditions this random variable is sharply concentrated around its mean.
We define the patchability of a family $\mathcal{F}$: essentially, how expensive is it to finish an almost-complete $F$ (that is, $F$ is close to $\mathcal{F}$ in Hamming distance) if the edge weights are re-randomized? Combining the patchability of $\mathcal{F}$, applying the Talagrand inequality to a dual problem, and a sprinkling-type argument, we prove a concentration inequality for the random variable $M(\mathcal{F})$.
A threshold for relative hyperbolicity in random right-angled Coxeter groups
We consider the random right-angled Coxeter group $W_Γ$ whose presentation graph $Γ\sim \mathcal{G}_{n,p}$ is an Erd{\H o}s--Rényi random graph on $n$ vertices with edge probability $p=p(n)$. We establish that $p=1/\sqrt{n}$ is a threshold for relative hyperbolicity of the random group $W_Γ$. As a key step in the proof, we determine the minimal number of pairs of generators that must commute in a right-angled Coxeter group which is not relatively hyperbolic, a result which is of independent interest.
We also show that there is an interval of edge probabilities of width $Ω(1/\sqrt{n})$ in which the random right-angled Coxeter group has precisely cubic divergence. This interval is between the thresholds for relative hyperbolicity (whence exponential divergence) and quadratic divergence. Moreover, a simple random walk on any Cayley graph of the random right-angled Coxeter group for $p$ in this interval satisfies a central limit theorem.
Graph-theoretical estimates of the diameters of the Rubik's Cube groups
A strict lower bound for the diameter of a symmetric graph is proposed, which is calculable with the order $n$ and other local parameters of the graph such as the degree $k\,(\geq 3)$, even girth $g\,(\geq 4)$, and number of $g$-cycles traversing a vertex, which are easily determined by inspecting a small portion of the graph (unless the girth is large). It is applied to the symmetric Cayley graphs of some Rubik's Cube groups of various sizes and metrics, yielding slightly tighter lower bounds of the diameters than those for random $k$-regular graphs proposed by Bollobás and de la Vega. They range from 60% to 77% of the correct diameters of large-$n$ graphs.
Long cycles in percolated expanders
Given a graph $G$ and probability $p$, we form the random subgraph $G_p$ by retaining each edge of $G$ independently with probability $p$. Given $d\in\mathbb{N}$ and constants $0<c<1, \varepsilon>0$, we show that if every subset $S\subseteq V(G)$ of size exactly $\frac{c|V(G)|}{d}$ satisfies $|N(S)|\ge d|S|$ and $p=\frac{1+\varepsilon}{d}$, then the probability that $G_p$ does not contain a cycle of length $Ω(\varepsilon^2c^2|V(G)|)$ is exponentially small in $|V(G)|$. As an intermediate step, we also show that given $k,d\in \mathbb{N}$ and a constant $\varepsilon>0$, if every subset $S\subseteq V(G)$ of size exactly $k$ satisfies $|N(S)|\ge kd$ and $p=\frac{1+\varepsilon}{d}$, then the probability that $G_p$ does not contain a path of length $Ω(\varepsilon^2 kd)$ is exponentially small. We further discuss applications of these results to $K_{s,t}$-free graphs of maximal density.
Limit theorems for walks and triangles on Erdös-Rényi random graphs with large interaction radius
We study cumulants of numbers of $q$-step walks on Erdös-Rényi-type random graphs of long-range percolation radius model in the limit when the number of vertices $N$, concentration $c$, and the interaction radius $R$ tend to infinity. These cumulants can be associated with a formal cumulant expansion of the free energy of matrix models of exponential random graphs widely known in mathematical and theoretical physics.
We show that in three different asymptotic regimes, the limiting values of $k$-th cumulants ${\cal F}_k^{(q)}$ exist and can be associated with one or another family of tree-type diagrams, in dependence of the asymptotic behavior of parameters $cR/N$ for $q$-step non-closed walks and $c^2R/N^2$ for 3-step closed walks, respectively. In certain cases, we obtain ${\cal F}_k^{(q)}$ in explicit form.
These results allow us to prove Limit Theorems for the number of non-closed walks and for the number of triangles in corresponding ensembles of large random graphs. As a consequence, we indicate an asymptotic regime when in random graphs that we consider, the average vertex degree remains bounded while the total number of triangles infinitely increases, thus rigorously solving a graph collapse problem known in applications.
Topological Minors in Typical Lifts
An $\ell$-lift of a graph $G$ is any graph obtained by replacing every vertex of $G$ with an independent set of size $\ell$, and connecting every pair of two such independent sets that correspond to an edge in $G$ by a matching of size $\ell$. Graph lifts have found numerous interesting applications and connections to a variety of areas over the years. Of particular importance is the random graph model obtained by considering an $\ell$-lift of a graph sampled uniformly at random. This model was first introduced by Amit and Linial in 1999, and has been extensively investigated since. In this paper, we study the size of the largest topological clique in random lifts of complete graphs.
In 2006, Drier and Linial raised the conjecture that almost all $\ell$-lifts of the complete graph on $n$ vertices contain a subdivision of a clique of order $Ω(n)$ as a subgraph provided $\ell$ is at least linear in $n$. We confirm their conjecture in a strong form by showing that for $\ell \ge (1+o(1))n$, one can almost surely find a subdivision of a clique of order $n$. We prove that this is tight by showing that for $\ell \le (1-o(1))n$, almost all $\ell$-lifts do not contain subdivisions of cliques of order $n$.
Finally, for $2 \le \ell \ll n$, we show that almost all $\ell$-lifts of $K_n$ contain a subdivision of a clique on $(1-o(1))\sqrt{\frac{2n \ell}{1-1/\ell}}$ vertices and that this is tight up to the lower order term.
Kohayakawa-Nagle-R{ö}dl-Schacht conjecture for subdivisions
In this paper, we study the well-known Kohayakawa-Nagle-R{ö}dl-Schacht (KNRS) conjecture, with a specific focus on graph subdivisions. The KNRS conjecture asserts that for any graph $H$, locally dense graphs contain asymptotically at least the number of copies of $H$ found in a random graph with the same edge density. We prove the following results about $k$-subdivisions of graphs (obtained by replacing edges with paths of length $k+1$): (1). If $H$ satisfies the KNRS conjecture, then its $(2k-1)$-subdivision satisfies Sidorenko's conjecture, extending a prior result of Conlon, Kim, Lee and Lee; (2). If $H$ satisfies the KNRS conjecture, then its $2k$-subdivision satisfies a constant-fraction version of the KNRS conjecture; (3). If $H$ is regular and satisfies the KNRS conjecture, then its $2k$-subdivision also satisfies the KNRS conjecture. These findings imply that all balanced subdivisions of cliques satisfy the KNRS conjecture, improving upon a recent result of Bradač, Sudakov and Wigerson. Our work provides new insights into this pivotal conjecture in extremal graph theory.
Central Limit Theorem on the Conjugacy Measure of Symmetric Groups
Regarding the conjugacy representation on symmetric groups, we initiate a normalized measure emerging from this representation, namely the conjugacy measure. A central limit theorem for character ratios of random representations of the symmetric group on the conjugacy measure is obtained.
Explicit estimates for the Stirling numbers of the second kind
We give explicit estimates for the Stirling numbers of the second kind $S(n,m)$. With a few exceptions, such estimates are asymptotically sharp. The form of these estimates varies according to $m$ lying in the central or non-central regions of $\{1,\ldots ,n\}$. In each case, we use a different probabilistic representation of $S(n,m)$ in terms of well known random variables to show the corresponding results.