cs.IT ↗ arXiv
154 papers in this category
Large values in time series and additive combinatorics
It is well-known in industrial data science that large values of real-life time series tend to be structured and often follow concrete and visible patterns. In this paper, we use ideas from additive combinatorics and discrete Fourier analysis to give this heuristic a mathematical foundation. Our main tool is the Fourier ratio, a complexity measure previously used in compressed sensing, combined with a generalized version of Chang's lemma from additive combinatorics. Together, these yield a precise prediction: when the Fourier ratio of a time series is small, the set of its largest values can be additively generated by a very small set using only $\{-1,0,1\}$ coefficients. We test this prediction on US inflation data and Delhi climate data, both in their original form and after mean-centering. The numerical results confirm the predicted structure: a generating set of size $4$--$7$ suffices to span large spectra containing dozens of points, even when the Fourier ratio is large enough that our theoretical bounds become loose. These findings provide a rigorous explanation for why extreme values in real-world data are information-rich and structurally significant.
Entropy lower bounds and sum-product phenomena
Various lower bounds are established for the entropy of sums, products and their combinations. First, we derive a prime-field analogue of a version of the entropy power inequality established by Tao over torsion-free groups. Next, we prove an entropy sum-product statement: For independent and identically distributed random variables $X,X'$, the maximum of ${\bf H}(X+X')$ and ${\bf H}(XX')$ is bounded below by a linear combination of the entropy and the min-entropy (Rényi entropy of order~$\infty$) of $X$. This result, obtained by bounding entropies of the form ${\bf H}\bigl( X(Y+Z)\bigr)$ from above and below, is valid over arbitrary fields $F$. Over $F={\bf R}$, a slightly stronger inequality is derived. Finally, a weak version of a purely Shannon-entropic sum-product result is developed: If the entropic additive doubling of a random variable $X$ over an arbitrary field is $O(1)$, then its multiplicative doubling is at least proportional to ${\bf H}(X)$.
On the independence number of de Bruijn graphs
We derive the asymptotic formula $α(k,q)=λ_{k-1}q^k+o(q^k)$, where $α(k,q)$ is the independence number of the de Bruijn graph $B(k,q)$, and $λ_{k-1}$ is a constant arising from a variational problem on the unit $(k-1)$-dimensional cube. When $k=4$, we show the bounds $91/240\le λ_3\le 11/28$. For odd prime $k$, we analyse the binary case $q=2$ via a phase reduction on rotation orbits. For $k=11$ and $k=13$ this yields certified optimal constructions, which combined with a lifting theorem by Lichiardopol give exact formulas for $α(11,q)$ and $α(13,q)$ for all $q\ge2$, extending the known cases $k=3,5,7$.
Transfer Operators and Independence Polynomials for Strong Powers of Circulant Graphs
We study independent sets in strong powers of circulant graphs using a transfer matrix formulation. The compatibility constraints separate into intra-layer and inter-layer components, yielding a transfer operator that is equivariant under the dihedral group action. The characteristic polynomial of the transfer operator factors into an \emph{anomalous} component (arising from the trivial isotypic component, with rational coefficients) and a \emph{cyclotomic} component (arising from nontrivial Fourier modes, splitting over the maximal real cyclotomic subfield). We show that the spectral radius is attained in the trivial isotypic component, so the dominant exponential growth is governed by a low-dimensional orbit-compressed operator. The independence polynomial is computed exactly for strong cylinders and tori, with the cyclotomic sector contributing a sparse correction confined to high-weight coefficients. All results are verified for $C_7$.
Explicit Rank Extractors and Subspace Designs via Function Fields, with Applications to Strong Blocking Sets
We give new explicit constructions of several fundamental objects in linear-algebraic pseudorandomness and combinatorics, including lossless rank extractors, weak subspace designs, and strong $s$-blocking sets over finite fields.
Our focus is on the small-field regime, where the field size depends only on a secondary parameter (such as the rank or codimension) and is independent of the ambient dimension. This regime is central to several applications, yet remains poorly understood from the perspective of explicit constructions.
In this setting, we obtain the first explicit constructions of lossless rank extractors and weak subspace designs for $r\ll k$, where $r$ denotes the rank (or codimension), over finite fields $\mathbb{F}_q$ with $q \ge \mathrm{poly}(r)$ and $q$ non-prime, with near-optimal parameters. For other finite fields, including prime fields and small fields, we obtain weaker but still improved bounds.
As a consequence, we construct explicit strong $s$-blocking sets in $\mathrm{PG}(k-1,q)$ of size $O(s(k-s)q^s)$ for all sufficiently large non-prime fields $q \ge \mathrm{poly}(s)$, matching the best known non-explicit bounds up to constant factors. This significantly improves the previous best bound $2^{O(s^2 \log s)} q^s k$ of Bishnoi and Tomon (Combinatorica, 2026), which requires $q \ge 2^{Ω(s)}$.
Our approach is primarily algebraic, combining techniques from function fields and polynomial identity testing. In addition, we develop a complementary Fourier-analytic framework based on $\varepsilon$-biased sets, which yields improved explicit constructions of strong $s$-blocking sets over small fields.
Turán-Theoretic Bounds on Several Elementary Trapping Sets in LDPC Codes
LDPC codes have attracted significant attention because of their superior performance close to the Shannon limit. Elementary trapping sets are the main cause of the error floor phenomenon in LDPC codes. We consider typical graphs related to trapping sets, including theta graphs, dumbbell graphs, and short cycles with chords. Based on the Turán numbers of $θ(2,2,2)$, $θ(1,3,3)$ and $D(4,4;0)$, we prove that any $(a,b)$-ETS with $g=8$ variable-regular $γ$ satisfies the inequality $b\geq aγ-\frac{a(\sqrt{24a-23}-1)}{4}$, provided that any two 8-cycles in the Tanner graph do not share common variable node. In addition, we can also eliminate ETSs by removing certain short-cycle structures with chords. The minimum sizes of ETSs obtained through these methods are significantly increased. To assess practical impact , we analyze spectral radii of the ETSs and construct QC-LDPC codes to show frame error rates in the error floor region.
On additive averaging kernels for finite Markov chains
We study additive mixtures of Markov kernels of the form $A_α= αP + (1-α)G$, where $α\in [0,1]$, $P$ is a baseline sampler and $G$ is a Gibbs kernel induced by a partition of the state space. We first motivate the study of $A_α$, which can be interpreted as the projection of a lifted Markov chain. We then consider the minimisation of distance to stationarity under two objectives: the squared Frobenius norm and the Kullback-Leibler (KL) divergence. For the Frobenius objective, we derive explicit trace formulas and identify a Cheeger-type functional that characterises optimal two-block partitions. This yields a structured combinatorial optimisation problem admitting a difference-of-submodular decomposition, enabling efficient approximation via majorisation-minimisation. We also obtain geometric decay rates governed by the absolute spectral gap of $P$. For the KL divergence, we establish convexity-based bounds showing that the divergence of $A_α$ is controlled by those of both $P$ and $G$, thereby reducing partition selection to the Gibbs component. Numerical experiments on the Curie-Weiss model demonstrate that suitable choice of both the partition and the parameter $α$ can significantly accelerate convergence in total variation distance. We observe a consistent trade-off between local exploration and global averaging, with intermediate values of $α$ achieving the best performance across regimes.
Asymptotic enumeration of admixed arrays and a different independence heuristic
We introduce a class of paired binary matrices called admixed arrays, which arise in analyses of large-scale genetic data and can be viewed as weighted edge colorings of complete bipartite graphs. This combinatorial structure gives rise to two natural families of marginal constraints: a row-sum constraint and a paired column-sum constraint, the latter inducing an inequality among entries of the matrix pair. We study the enumeration of admixed arrays under these constraints in dense regimes. First, we obtain exact formulas for the sizes of the families defined by each constraint in isolation and derive a finite-size criterion characterizing when one constraint is more restrictive than the other. In the large-dimension limit, this comparison simplifies to an entropy inequality, yielding an information-theoretic interpretation and a quantifiable error bound in the semi-regular case. We then analyze the asymptotic enumeration of the doubly constrained family in a semi-regular setting. Using saddle-point approximation and probabilistic techniques, we derive a detailed asymptotic expansion for the logarithm of the count, isolating an explicit fourth-moment contribution and establishing quantitative control of the higher-order remainder. A consequence of this analysis is a phenomenon absent from classical binary and integer matrix models: in the regime $N=Θ(P)$ with uniform margins and density bounded away from zero, the two constraint families obey the independence heuristic with a correction factor $1/\sqrt[4]{e}$ rather than the familiar $e^{\pm1/2}$. Numerical experiments corroborate the analytical approximations, and we implement and extend an algorithm of Miller and Harrison (2013) as open-source software to enumerate constrained admixed arrays.
The Random Subsequence Model and Uniform Codes for the Deletion Channel
We introduce the Random Subsequence Model, a spin glass model on pairs of random strings $(X,Y) \in \{0,1\}^N \times \{0,1\}^M$ whose partition function counts subsequence embeddings of $Y$ into $X$. We study two variants: the null model, where $X$ and $Y$ are independent and uniform, and the planted model, where $X$ is uniform and $Y$ is a uniformly-random length-$M$ subsequence of $X$. We connect the Random Subsequence Model to longstanding problems in various fields, including the best rate achievable by uniformly-random codes in the deletion channel, the longest common subsequence problem between two random strings, and models of directed polymers in statistical physics.
In the regime where $N,M\to\infty$ at a fixed ratio $α= M/N \in (0,1)$, we exhibit strict asymptotic separations between the null annealed free energy and the quenched free energies of the null and planted models at all values of the density parameter $α$. This suggests that these models are in a spin glass phase at zero temperature throughout the entire dense regime. As a consequence, we show that uniformly-random codes achieve a positive rate in the deletion channel for all deletion probabilities $p\in [0,1),$ settling multiple conjectures of the second author, Isik and Weissman (2024) and proving the first such positive rate result for the regime $p \geq 1/2$.
We also give an exact analytic formula for the annealed free energy of the planted model for all values of the density parameter. This implies a corresponding analytic upper bound on the best rate achievable by uniformly-random codes in the deletion channel, complementing the lower bound from our first result. Our upper and lower bounds for the capacity of the deletion channel under uniform codes are far closer to each other than the best known upper and lower bounds for the capacity of the deletion channel.
Binary Caps and LCD Codes with Large Dimensions
We establish a connection between linear complementary dual (LCD) codes and caps in projective space. Using this framework and the structure theory of maximal caps, we derive nonexistence theorems for LCD codes with minimum distance at least $4$, providing computation-free proofs that were previously obtained only through exhaustive search. As an application, we completely determine the optimal minimum distances for codimensions $7$ and $8$ for the first time.
Overconstrained character sums over finite abelian groups and decompositions of generalized bent, plateaued and landscape functions
Generalized bent (gbent) functions from an $n$-variable Boolean space to $\mathbb{Z}_{2^k}$ are central in cryptography and sequence design. Instead of the usual binary decomposition, we introduce a $2^\ell$-adic representation, for $k=\ell r$, writing such functions as linear combinations of $r$ component functions valued in $\mathbb{Z}_{2^\ell}$. We prove a general result on overconstrained character sums over finite abelian groups: under a common-argument hypothesis, sequences with two-level Fourier magnitude spectra must be extremely sparse, with a conditional extension to multi-level spectra. As an application, we derive consequences for generalized plateaued functions under suitable assumptions. We then show that if $f:\mathbb{F}_2^n\to\mathbb{Z}_{2^k}$ is landscape, then under the $2^\ell$-adic decomposition every function in a certain affine space over $\mathbb{Z}_{2^\ell}$ is again landscape with the same Walsh magnitudes. This gives an unconditional necessity result, with no structural assumptions on $f$, together with a complete characterization using only a small subset of these maps. For generalized bent and generalized plateaued functions, sufficiency is also obtained from linear combinations of lower components under natural assumptions; a counterexample shows these assumptions are essential. Our method reduces verification for landscape functions from $2^{2^{k-1}}$ checks to fewer than $2^{k-\ell+1}+1$ conditions; for gbent functions this drops to a single basis function under the common-argument hypothesis, and for generalized plateaued functions, under additional assumptions, to $2^{k-\ell}$ checks. The $2^\ell$-adic framework also preserves key properties, including duality and differential uniformity.
On the existence of linear rank-metric intersecting codes
Intersecting codes are a classical object in coding theory whose rank-metric analogue has recently been introduced. Although the definition formally parallels the Hamming-metric case, the structure and parameter constraints of rank-metric intersecting codes exhibit substantially different behavior. It was previously shown that a nondegenerate $[n,k,d]_{q^m/q}$ rank-metric intersecting code must satisfy $2k-1 \le n \le 2m-3$, and the tightness of the upper bound was left open. Using the geometric interpretation of rank-metric codes via $q$-systems, we prove that the dual subspace associated with a rank-metric intersecting code must satisfy strong evasiveness properties. This connection allows us to derive new restrictions on the parameters of such codes and to show that the bound $n=2m-3$ can be attained only when $k=3$ and $m\ge 6$. More generally, we show that $n \leq 2m-\lfloor(k+4)/2\rfloor$. Moreover, we obtain a geometric characterization of these extremal codes in terms of scattered $\mathbb{F}_q$-subspaces of $\mathbb{F}_{q^m}^3$. As a consequence, the existence problem for $[2m-3,3,d]_{q^m/q}$ rank-metric intersecting codes is reduced to the existence of scattered subspaces of dimension $m+3$. Using known constructions of maximum scattered subspaces, we derive existence results when $m$ is even. Finally, we prove that $[6,3,3]_{q^5/q}$ rank-metric intersecting codes do not exist for any prime power $q$, thus resolving an open problem posed by Bartoli et al. in 2025.
On the Construction of Recursively Differentiable Quasigroups and an Example of a Recursive $[4,2,3]_{26}$-Code
In 1998, E. Couselo, S. González, V. T. Markov, and A. A. Nechaev introduced the notions of recursive codes and recursively differentiable quasigroups. They conjectured that recursive MDS codes of dimension $2$ and length $4$ exist over every finite alphabet of size $q \not\in \{2, 6\}$, and verified this conjecture in all cases except $q \in \{14, 18, 26, 42\}$. In 2008, V. T. Markov, A. A. Nechaev, S. S. Skazhenik, and E. O. Tveritinov resolved the case $q=42$ by providing an explicit construction. The present paper settles the outstanding case $q=26$. The construction rests upon methods for producing recursively differentiable quasigroups and recursive MDS codes via perfect cyclic Mendelsohn designs. Moreover, we sharpen several known bounds concerning the existence of recursively $n$-differentiable quasigroups of small orders.
Randomstrasse101: Open Problems of 2025
Randomstrasse101 is a blog dedicated to Open Problems in Mathematics, with a focus on Probability Theory, Computation, Combinatorics, Statistics, and related topics. This manuscript serves as a stable record of the Open Problems posted in 2025, with the goal of easing academic referencing. The blog can currently be accessed at randomstrasse101.math.ethz.ch
Nonvanishing $k$-flats of Boolean and vectorial functions
$k$th-order sum-free functions are a natural generalization of APN functions using the concept of (non)vanishing flats. In this paper, we introduce a new combinatorial technique to study the nonvanishing flats of Boolean functions. This approach allows us to determine the number of nonvanishing flats for an infinite family of Boolean functions. We moreover use it to show that any $k$th-order sum-free $(n,n)$-function of algebraic degree $k$ gives rise to an $(n-k)$th-order sum-free $(n,n)$-function of algebraic degree $n-k$. This implies the existence of millions of $(n-2)$th-order sum-free functions.
An infinite family of non-extendable MRD codes
In the realm of rank-metric codes, Maximum Rank Distance (MRD) codes are optimal algebraic structures attaining the Singleton-like bound. A major open problem in this field is determining whether an MRD code can be extended to a longer one while preserving its optimality. This work investigates $\mathbb{F}_{q^m}$-linear MRD codes that are non-extendable but do not attain the maximum possible length. Geometrically, these correspond to scattered subspaces with respect to hyperplanes that are maximal with respect to inclusion but not of maximum dimension. By exploiting this geometric connection, we introduce the first infinite family of non-extendable $[4,2,3]_{q^5/q}$ MRD codes. Furthermore, we prove that these codes are self-dual up to equivalence.
Double Toeplitz codes and their average weight enumerators
Recently, double Toeplitz codes have been introduced as a generalization of double circulant codes. In this paper, we study the average weight enumerators of double Toeplitz codes. As an application, we consider the existence of double Toeplitz codes over $\mathbb{F}_q$ with some specified minimum weights for $q \in \{2,3,4\}$. We also give a classification of double Toeplitz codes over $\mathbb{F}_q$ with the largest minimum weights for modest lengths and $q \in \{2,3,4\}$.
On the number of inequivalent linearized Reed-Solomon codes
Linearized Reed-Solomon (LRS) codes form an important family of maximum sum-rank distance (MSRD) codes that generalize both Reed--Solomon codes and Gabidulin codes. In this paper we study the equivalence problem for LRS codes and determine the number of inequivalent codes within this family. Using the correspondence between sum-rank metric codes and systems of $\mathbb{F}_q$-subspaces, we analyze the stabilizer of the Gabidulin system and derive a characterization of equivalence between LRS codes. In particular, we prove that two LRS codes are equivalent if and only if the sets of norms that define the codes coincide up to multiplication by an element of $\mathbb{F}_q^\ast$. This description allows us to reduce the classification problem to the action of $\mathbb{F}_q^\ast$ on subsets of $\mathbb{F}_q^\ast$. As a consequence, we derive formulas for the number of inequivalent linearized Reed-Solomon codes and illustrate the results with explicit examples.
Coded Information Retrieval for Block-Structured DNA-Based Data Storage
We study the problem of coded information retrieval for block-structured data, motivated by DNA-based storage systems where a database is partitioned into multiple files that must each be recoverable as an atomic unit. We initiate and formalize the block-structured retrieval problem, wherein $k$ information symbols are partitioned into two files $F_1$ and $F_2$ of sizes $s_1$ and $s_2 = k - s_1$. The objective is to characterize the set of achievable expected retrieval time pairs $\bigl(E_1(G), E_2(G)\bigr)$ over all $[n,k]$ linear codes with generator matrix $G$. We derive a family of linear lower bounds via mutual exclusivity of recovery sets, and develop a nonlinear geometric bound via column projection. For codes with no mixed columns, this yields the hyperbolic constraint $s_1/E_1 + s_2/E_2 \le 1$, which we conjecture to hold universally whenever $\max\{s_1,s_2\} \ge 2$. We analyze explicit codes, such as the identity code, file-dedicated MDS codes, and the systematic global MDS code, and compute their exact expected retrieval times. For file-dedicated codes we prove MDS optimality within the family and verify the hyperbolic constraint. For global MDS codes, we establish dominance by the proportional local MDS allocation via a combinatorial subset-counting argument, providing a significantly simpler proof compared to recent literature and formally extending the result to the asymmetric case. Finally, we characterize the limiting achievability region as $n \to \infty$: the hyperbolic boundary is asymptotically achieved by file-dedicated MDS codes, and is conjectured to be the exact boundary of the limiting achievability region.
Aperiodic Structures Never Collapse: Fibonacci Hierarchies for Lossless Compression
We study whether an aperiodic hierarchy can provide a structural advantage for lossless compression over periodic alternatives. We show that Fibonacci quasicrystal tilings avoid the finite-depth collapse that affects periodic hierarchies: usable $n$-gram lookup positions remain non-zero at every level, while periodic tilings collapse after $O(\log p)$ levels for period $p$. This yields an aperiodic hierarchy advantage: dictionary reuse remains available across all scales instead of vanishing beyond a finite depth.
Our analysis gives four main consequences. First, the Golden Compensation property shows that the exponential decay in the number of positions is exactly balanced by the exponential growth in phrase length, so potential coverage remains scale-invariant with asymptotic value $W\varphi/\sqrt{5}$. Second, using the Sturmian complexity law $p(n)=n+1$, we show that Fibonacci/Sturmian hierarchies maximize codebook coverage efficiency among binary aperiodic tilings. Third, under long-range dependence, the resulting hierarchy achieves lower coding entropy than comparable periodic hierarchies. Fourth, redundancy decays super-exponentially with depth, whereas periodic systems remain locked at the depth where collapse occurs.
We validate these results with Quasicryth, a lossless text compressor built on a ten-level Fibonacci hierarchy with phrase lengths ${2,3,5,8,13,21,34,55,89,144}$. In controlled A/B experiments with identical codebooks, the aperiodic advantage over a Period-5 baseline grows from $1{,}372$ B at 3 MB to $1{,}349{,}371$ B at 1 GB, explained by the activation of deeper hierarchy levels. On enwik9, Quasicryth achieves $359{,}883{,}431$ B $(35.99%)$, with $45{,}608{,}715$ B attributable to the quasicrystal tiling itself.