arXiv++ Combinatorics

Browse math.CO papers from arXiv

probability distribution

284 papers tagged with this keyword
2023-12-26
Cycle structure of Mallows permutation model with the $L^1$ distance
Introduced by Mallows as a ranking model in statistics, Mallows permutation model is a class of non-uniform probability distributions on the symmetric group $S_n$. The model depends on a distance metric on $S_n$ and a scale parameter $β$. In this paper, we take the distance metric to be the $L^1$ distance (also known as Spearman's footrule in the statistics literature), and investigate the cycle structure of random permutations drawn from Mallows permutation model with the $L^1$ distance. We focus on the parameter regime where $β>0$. We show that the expected length of the cycle containing a given point is of order $\min\{\max\{β^{-2},1\},n\}$, and the expected diameter of the cycle containing a given point is of order $\min\{e^{-2β}\max\{β^{-2},1\}, n-1\}$. Moreover, when $β\ll n^{-1\slash 2}$, the sorted cycle lengths (in descending order) normalized by $n$ converge in distribution to the Poisson-Dirichlet law with parameter $1$. The proofs of the results rely on the hit and run algorithm, a Markov chain for sampling from the model.
Leading All The Way
Xavier and Yushi run a ``random race'' as follows. A continuous probability distribution $μ$ on the real line is chosen. The runners begin at zero. At time $i$ Xavier draws $\mathbf{X}_i$ from $μ$ and advances that distance, while Yushi advances by an independent drawing $\mathbf{Y}_i$. After $n$ such moves, Xavier wins a valuable prize provided he not only wins the race but leads after every step; that is, $\sum_{i=1}^k \mathbf{X}_i > \sum_{i=1}^k \mathbf{Y}_i$ for all $k = 1,2, \dots, n$. What distribution is best for Xavier, and what then is his probability of getting the prize?
2023-11-30 v3
$ε$-Uniform Mixing in Discrete Quantum Walks
We study whether the probability distribution of a discrete quantum walk can get arbitrarily close to uniform, given that the walk starts with a uniform superposition of the outgoing arcs of some vertex. We establish a characterization of this phenomenon on regular non-bipartite graphs in terms of their adjacency eigenvalues and eigenprojections. Using theory from association schemes, we show this phenomenon happens on a strongly regular graph $X$ if and only if $X$ or $\overline{X}$ has parameters $(4m^2, 2m^2\pm m, m^2\pm m, m^2\pm m)$ where $m\ge 2$.
2023-10-12
Algebraic Connectivity Characterization of Ensemble Random Hypergraphs
Random hypergraph is a broad concept used to describe probability distributions over hypergraphs, which are mathematical structures with applications in various fields, e.g., complex systems in physics, computer science, social sciences, and network science. Ensemble methods, on the other hand, are crucial both in physics and machine learning. In physics, ensemble theory helps bridge the gap between the microscopic and macroscopic worlds, providing a statistical framework for understanding systems with a vast number of particles. In machine learning, ensemble methods are valuable because they improve predictive accuracy, reduce overfitting, lower prediction variance, mitigate bias, and capture complex relationships in data. However, there is limited research on applying ensemble methods to a set of random hypergraphs. This work aims to study the connectivity behavior of an ensemble of random hypergraphs. Specifically, it focuses on quantifying the random behavior of the algebraic connectivity of these ensembles through tail bounds. We utilize Laplacian tensors to represent these ensemble random hypergraphs and establish mathematical theorems, such as Courant-Fischer and Lieb-Seiringer theorems for tensors, to derive tail bounds for the algebraic connectivity. We derive three different tail bounds, i.e., Chernoff, Bennett, and Bernstein bounds, for the algebraic connectivity of ensemble hypergraphs with respect to different random hypergraphs assumptions.
Estimating systemic importance with missing data in input-output graphs
In the context of the Cobb-Douglas productivity model we consider the $N \times N$ input-output linkage matrix $W$ for a network of $N$ firms $f_1, f_2, \cdots, f_N$. The associated influence vector $v_w$ of $W$ is defined in terms of the Leontief inverse $L_W$ of $W$ as $v_W = \fracα{N} L_W \vec{\mathbf{1}}$ where $L_W = (I - (1-α) W')^{-1}$, $W'$ denotes the transpose of $W$ and $I$ is the identity matrix. Here $\vec{\mathbf{1}}$ is the $N \times 1$ vector whose entries are all one. The influence vector is a metric of the importance for the firms in the production network. Under the realistic assumption that the data to compute the influence vector is incomplete, we prove bounds on the worst-case error for the influence vector that are sharp up to a constant factor. We also consider the situation where the missing data is binomially distributed and contextualize the bound on the influence vector accordingly. We also investigate how far off the influence vector can be when we only have data on nodes and connections that are within distance $k$ of some source node. A comparison of our results is juxtaposed against PageRank analogues. We close with a discussion on a possible extension beyond Cobb-Douglas to the Constant Elasticity of Substitution model, as well as the possibility of considering other probability distributions for missing data.
2023-09-21 v3
Six-Vertex Model and Random Matrix Distributions
We survey the connections between the six-vertex (square ice) model of 2d statistical mechanics and random matrix theory. We highlight the same universal probability distributions appearing on both sides, and also indicate related open questions and conjectures. We present full proofs of two asymptotic theorems for the six-vertex model: in the first one the Gaussian Unitary Ensemble and GUE-corners process appear; the second one leads to the Tracy-Widom distribution $F_2$. While both results are not new, we found shorter transparent proofs for this text. On our way we introduce the key tools in the study of the six-vertex model, including the Yang-Baxter equation and the Izergin-Korepin formula.
Unexpected Averages of Mixing Matrices
The (standard) average mixing matrix of a continuous-time quantum walk is computed by taking the expected value of the mixing matrices of the walk under the uniform sampling distribution on the real line. In this paper we consider alternative probability distributions, either discrete or continuous, and first we show that several algebraic properties that hold for the average mixing matrix still stand for this more general setting. Then, we provide examples of graphs and choices of distributions where the average mixing matrix behaves in an unexpected way: for instance, we show that there are probability distributions for which the average mixing matrices of the paths on three or four vertices have constant entries, opening a significant line of investigation about how to use classical probability distributions to sample quantum walks and obtain desired quantum effects. We present results connecting the trace of the average mixing matrix and quantum walk properties, and we show that the Gram matrix of average states is the average mixing matrix of a certain related distribution. Throughout the text, we employ concepts of classical probability theory not usually seen in texts about quantum walks.
2023-08-28
Distribution of the number of zeros of polynomials over a finite field
Published in Involve 18 (2025) 707-718 • View PublicationBIB
We study the probability distribution of the number of zeros of multivariable polynomials with bounded degree over a finite field. We find the probability generating function for each set of bounded degree polynomials. In particular, in the single variable case, we show that as the degree of the polynomials and the order of the field simultaneously approach infinity, the distribution converges to a Poisson distribution.
2023-08-08
A Littlewood-Offord kind of problem in $\mathbb{Z}_p$ and $Γ$-sequenceability
The Littlewood-Offord problem is a classical question in probability theory and discrete mathematics, proposed, firstly by Littlewood and Offord in the 1940s. Given a set $A$ of integer, this problem asks for an upper bound on the probability that a randomly chosen subset $X$ of $A$ sums to an integer $x$. This article proposes a variation of the problem, considering a subset $A$ of a cyclic group of prime order and examining subsets $X\subseteq A$ of a given cardinality $\ell$. The main focus of this paper is then on bounding the probability distribution of the sum $Y$ of $\ell$ i.i.d. $Y_1,\dots, Y_{\ell}$ whose support is contained in $\mathbb{Z}_p$. The main result here presented is that, if the probability distributions of the variables $Y_i$ are bounded by $λ\leq 9/10$, then, assuming that $p> \frac{2}λ\left(\frac{\ell_0}{3}\right)^ν$ (for some $\ell_0\leq\ell$), the distribution of $Y$ is bounded by $λ\left(\frac{3}{\ell_0}\right)^ν$ for some positive absolute constant $ν$. Then an analogous result is implied for the Littlewood-Offord problem over $\mathbb{Z}_p$ on subsets $X$ of a given cardinality $\ell$ in the regime where $n$ is large enough. Finally, as an application of our results, we propose a variation of the set-sequenceability problem: that of $Γ$-sequenceability. Given a graph $Γ$ on the vertex set $\{1,2,\dots,n\}$ and given a subset $A\subseteq \mathbb{Z}_p$ of size $n$, here we want to find an ordering of $A$ such that the partial sums $s_i$ and $s_j$ are different whenever $\{i,j\}\in E(Γ)$. As a consequence of our results on the Littlewood-Offord problem, we have been able to prove that, if the maximum degree of $Γ$ is at most $d$, $n$ is large enough, and $p>n^2$, any subset $A\subseteq \mathbb{Z}_p$ of size $n$ is $Γ$-sequenceable.
Probability Metrics for Tropical Spaces of Different Dimensions
The problem of comparing probability distributions is at the heart of many tasks in statistics and machine learning. Established comparison methods treat the standard setting that the distributions are supported in the same space. Recently, a new geometric solution has been proposed to address the more challenging problem of comparing measures in Euclidean spaces of differing dimensions. Here, we study the same problem of comparing probability distributions of different dimensions in the tropical setting, which is becoming increasingly relevant in applications involving complex data structures such as phylogenetic trees. Specifically, we construct a Wasserstein distance between measures on different tropical projective tori -- the focal metric spaces in both theory and applications of tropical geometry -- via tropical mappings between probability measures. We prove equivalence of the directionality of the maps, whether mapping from a low dimensional space to a high dimensional space or vice versa. As an important practical implication, our work provides a framework for comparing probability distributions on the spaces of phylogenetic trees with different leaf sets. We demonstrate the computational feasibility of our approach using existing optimisation techniques on both simulated and real data.
2023-06-13 v2
Connectivity threshold for superpositions of Bernoulli random graphs
Let $G_1,\dots, G_m$ be independent Bernoulli random subgraphs of the complete graph ${\cal K}_n$ having variable sizes $x_1,\dots, x_m\in [n]$ and densities $q_1,\dots, q_m\in [0,1]$. Letting $n,m\to+\infty$, we study the connectivity threshold for the union $\cup_{i=1}^mG_i$ defined on the vertex set of ${\cal K}_n$. Assuming that the empirical distribution $P_{n,m}$ of the pairs $(x_1,q_1),\dots, (x_m,q_m)$ converges to a probability distribution $P$ we show that the threshold is defined by the mixed moments $κ_n=\iint x(1-(1-q)^{|x-1|})P_{n,m}(dx,dq)$. For $\ln n-\frac{m}{n}κ_n\to-\infty$ we have $P\{\cup_{i=1}^mG_i$ is connected$\}\to 1$ and for $\ln n-\frac{m}{n}κ_n\to+\infty$ we have $P\{\cup_{i=1}^mG_i$ is connected$\}\to 0$. Interestingly, this dichotomy only holds if the mixed moment $\iint x(1-(1-q)^{|x-1|})\ln(1+x)P(dx,dq)<\infty$.
2023-05-26 v3
On the maximum of the weighted binomial sum $(1+a)^{-r}\sum_{i=0}^{r}\binom{m}{i}a^{i}$
Recently, Glasby and Paseman considered the following sequence of binomial sums $\{2^{-r}\sum_{i=0}^{r}\binom{m}{i}\}_{r=0}^{m}$ and showed that this sequence is unimodal and attains its maximum value at $r=\lfloor\frac{m}{3}\rfloor+1$ for $m\in\mathbb{Z}_{\geq0}\setminus\{0,3,6,9,12\}$. They also analyzed the asymptotic behavior of the maximum value of the sequence as $m$ approaches infinity. In the present work, we generalize their results by considering the sequence $\{(1+a)^{-r}\sum_{i=0}^{r}\binom{m}{i}a^{i}\}_{r=0}^{m}$ for integers $a \geq 1$. We also consider a family of discrete probability distributions that naturally arises from this sequence.
2023-05-04
Stirling numbers with higher level and records
In this present paper, we show that the Stirling numbers of the first kind with higher level connected with the probability distribution of the number of records and record times in the so-called F^α-scheme. In addition, we determine the location of the maximum of the Stirling numbers of the first kind with higher level.
Evolutionary quantum feature selection
Effective feature selection is essential for enhancing the performance of artificial intelligence models. It involves identifying feature combinations that optimize a given metric, but this is a challenging task due to the problem's exponential time complexity. In this study, we present an innovative heuristic called Evolutionary Quantum Feature Selection (EQFS) that employs the Quantum Circuit Evolution (QCE) algorithm. Our approach harnesses the unique capabilities of QCE, which utilizes shallow depth circuits to generate sparse probability distributions. Our computational experiments demonstrate that EQFS can identify good feature combinations with quadratic scaling in the number of features. To evaluate EQFS's performance, we counted the number of times a given classical model assesses the cost function for a specific metric, as a function of the number of generations.
2022-12-12 v4
The one-sided cycle shuffles in the symmetric group algebra
Published in Shortened version in: Algebraic Combinatorics, Volume 7 (2024) no. 2, pp. 275-326 • View PublicationBIB
We study a family of shuffling operators on the symmetric group $S_n$, which includes the top-to-random shuffle. The general shuffling scheme consists of removing one card at a time from the deck (according to some probability distribution) and re-inserting it at a (uniformly) random position further below. Rewritten in terms of the group algebra $\mathbb{R}[S_n]$, our shuffle corresponds to right multiplication by a linear combination of the elements \[t_i:=\text{cyc}_{i}+\text{cyc}_{i,i+1}+\text{cyc}_{i,i+1,i+2}+\cdots+\text{cyc}_{i,i+1,\ldots,n}\in \mathbb{R}[S_n]\] for all $i\in\{1,2,\ldots,n\}$ (where $\text{cyc}_{j_1,j_2,\ldots,j_p}$ stands for a $p$-cycle). We compute the eigenvalues of these shuffling operators and of all their linear combinations. In particular, we show that the eigenvalues of right multiplication by a linear combination $λ_1t_1+λ_2t_2+\cdots+λ_nt_n$ are the numbers $λ_1m_{I,1}+λ_2m_{I,2}+\cdots+λ_nm_{I,n}$, where $I$ ranges over the subsets of $\{1,2,\ldots,n-1\}$ that contain no two consecutive integers; here $m_{I,i}$ are certain integers. We compute the multiplicities of these eigenvalues and show that if they are all distinct, the shuffling operator is diagonalizable. To this purpose, we show that the operators of right multiplication by $t_1,t_2,\ldots,t_n$ on $\mathbb{R}[S_n]$ are simultaneously triangularizable (via a combinatorially defined basis). The results stated here over $\mathbb{R}$ for convenience are actually stated and proved over an arbitrary commutative ring $\mathbf{k}$. We finish by describing a strong stationary time for the random-to-below shuffle, which is the shuffle in which the card that moves below is selected uniformly at random, and we give the waiting time for this event to happen.
Ergodicity of a generalized probabilistic cellular automaton with parity-based neighbourhoods
We study a one-dimensional generalized probabilistic cellular automaton $E_{p, q}$ with universe $\mathbb Z$, alphabet $\mathcal A = \{0, 1\}$, parameters $p$ and $q$ such that $0 < p+q \leq 1$ and two neighbourhoods $\mathcal N_0 = \{0, 1\}$ and $\mathcal N = \{1, 2\}$. The state $E_{p, q} η(x)$ of any $x \in \mathbb Z$ under the application of $E_{p, q}$ is a random variable whose probability distribution depends on the states $η(x + y)$ for $y \in \mathcal N_i$ where $i$ has the same parity as $x$. We establish ergodicity of this GPCA for various ranges of values of $p$ and $q$ via its connection with a suitable percolation game on a two-dimensional lattice. For these same ranges of values of $p$ and $q$, we show that the above-mentioned game has probability $0$ of resulting in a draw.
2022-11-09 v2
Decomposition of Probability Marginals for Security Games in Max-Flow/Min-Cut Systems
Published • View PublicationBIB
Given a set system $(E, \mathcal{P})$ with $ρ\in [0, 1]^E$ and $π\in [0,1]^{ \mathcal{P}}$, our goal is to find a probability distribution for a random set $S \subseteq E$ such that $\operatorname{Pr}[e \in S] = ρ_e$ for all $e \in E$ and $\operatorname{Pr}[P \cap S \neq \emptyset] \geq π_P$ for all $P \in \mathcal{P}$. We extend the results of Dahan, Amin, and Jaillet (MOR 2022) who studied this problem motivated by a security game in a directed acyclic graph (DAG). We focus on the setting where $π$ is of the affine form $π_P = 1 - \sum_{e \in P} μ_e$ for $μ\in [0, 1]^E$. A necessary condition for the existence of the desired distribution is that $\sum_{e \in P} ρ_e \geq π_P$ for all $P \in \mathcal{P}$. We show that this condition is sufficient if and only if $\mathcal{P}$ has the weak max-flow/min-cut property. We further provide an efficient combinatorial algorithm for computing the corresponding distribution in the special case where $(E, \mathcal{P})$ is an abstract network. As a consequence, equilibria for the security game by Dahan et al. can be efficiently computed in a wide variety of settings (including arbitrary digraphs). As a subroutine of our algorithm, we provide a combinatorial algorithm for computing shortest paths in abstract networks, partially answering an open question by McCormick (SODA 1996). We further show that a conservation law proposed by Dahan et al. for the requirement vector $π$ in DAGs can be reduced to the setting of affine requirements described above.
2022-11-03 v2
Average Mixing in Quantum Walks of Reversible Markov Chains
Published in Discrete Mathematics, Volume 348, Issue 1, 2025, 114196 • View PublicationBIB
The Szegedy quantum walk is a discrete time quantum walk model which defines a quantum analogue of any Markov chain. The long-term behavior of the quantum walk can be encoded in a matrix called the average mixing matrix, whose columns give the limiting probability distribution of the walk given an initial state. We define a version of the average mixing matrix of the Szegedy quantum walk which allows us to more readily compare the limiting behavior to that of the chain it quantizes. We prove a formula for our mixing matrix in terms of the spectral decomposition of the Markov chain and show a relationship with the mixing matrix of a continuous quantum walk on the chain. In particular, we prove that average uniform mixing in the continuous walk implies average uniform mixing in the Szegedy walk. We conclude by giving examples of Markov chains of arbitrarily large size which admit average uniform mixing in both the continuous and Szegedy quantum walk.
2022-11-01 v2
On distribution of runs and patterns in four state trials
Published • View PublicationBIB
From a mathematical and statistical point of view, a segment of a DNA strand can be viewed as a sequence of four-state (A, C, G, T) trials. We consider distributions of runs and patterns related to run lengths of multi-state sequences, especially for four states (A, B, C, D). Let $X_{1}, X_{2}, \ldots$ be a sequence of four state i.i.d.\ trials taking values in the set $\mathscr{S}=\{A,\ B,\ C,\ D\}$ of four symbols with probability $P(A)=P_{a}$, $P(B)=P_{b}$, $P(C)=P_{c}$ and $P(D)=P_{d},$ respectively. In this paper, we obtain exact formulae for the probability distribution function for runs of B's the discrete distribution of order $k$, longest run statistics, shortest run statistics, waiting time distribution and the distribution of run lengths.
2022-10-17
Recurrence algorithms of waiting time for the success run of length $k$ in relation to generalized Fibonacci sequences
Let $V(k)$ denote the waiting time, the number of trials needed to get a consecutive $k$ ones. We propose recurrence algorithms for the probability distribution function (pdf) and the probability generating function (pgf) of $V(k)$ in sequences of independent and Markov dependent Bernoulli trials using generalized Fibonacci sequences of order $k$. Maximum likelihood estimation (MLE) methods for the probability distributions are presented in both cases with simulation examples.