Papers by Henry Kvinge
10 paper(s) by this author
· All BibTeX
Even with AI, Bijection Discovery is Still Hard: The Opportunities and Challenges of OpenEvolve for Novel Bijection Construction
Evolutionary program synthesis systems such as AlphaEvolve, OpenEvolve, and ShinkaEvolve offer a new approach to AI-assisted mathematical discovery. These systems utilize teams of large language models (LLMs) to generate candidate solutions to a problem as human readable code. These candidate solutions are then 'evolved' with the goal of improving them beyond what an LLM can produce in a single shot. While existing mathematical applications have mostly focused on problems of establishing bounds (e.g., sphere packing), the program synthesis approach is well suited to any problem where the solution takes the form of an explicit construction. With this in mind, in this paper we explore the use of OpenEvolve for combinatorial bijection discovery. We describe the results of applying OpenEvolve to three bijection construction problems involving Dyck paths, two of which are known and one of which is open. We find that while systems like OpenEvolve show promise as a valuable tool for combinatorialists, the problem of finding novel, research-level bijections remains a challenging task for current frontier systems, reinforcing the need for human mathematicians in the loop. We describe some lessons learned for others in the field interested in exploring the use of these systems.
Machine Learning meets Algebraic Combinatorics: A Suite of Datasets Capturing Research-level Conjecturing Ability in Pure Mathematics
With recent dramatic increases in AI system capabilities, there has been growing interest in utilizing machine learning for reasoning-heavy, quantitative tasks, particularly mathematics. While there are many resources capturing mathematics at the high-school, undergraduate, and graduate level, there are far fewer resources available that align with the level of difficulty and open endedness encountered by professional mathematicians working on open problems. To address this, we introduce a new collection of datasets, the Algebraic Combinatorics Dataset Repository (ACD Repo), representing either foundational results or open problems in algebraic combinatorics, a subfield of mathematics that studies discrete structures arising from abstract algebra. Further differentiating our dataset collection is the fact that it aims at the conjecturing process. Each dataset includes an open-ended research-level question and a large collection of examples (up to 10M in some cases) from which conjectures should be generated. We describe all nine datasets, the different ways machine learning models can be applied to them (e.g., training with narrow models followed by interpretability analysis or program synthesis with LLMs), and discuss some of the challenges involved in designing datasets like these.
Machines and Mathematical Mutations: Using GNNs to Characterize Quiver Mutation Classes
Machine learning is becoming an increasingly valuable tool in mathematics, enabling one to identify subtle patterns across collections of examples so vast that they would be impossible for a single researcher to feasibly review and analyze. In this work, we use graph neural networks to investigate \emph{quiver mutation} -- an operation that transforms one quiver (or directed multigraph) into another -- which is central to the theory of cluster algebras with deep connections to geometry, topology, and physics. In the study of cluster algebras, the question of \emph{mutation equivalence} is of fundamental concern: given two quivers, can one efficiently determine if one quiver can be transformed into the other through a sequence of mutations? In this paper, we use graph neural networks and AI explainability techniques to independently discover mutation equivalence criteria for quivers of type $\tilde{D}$. Along the way, we also show that even without explicit training to do so, our model captures structure within its hidden representation that allows us to reconstruct known criteria from type $D$, adding to the growing evidence that modern machine learning models are capable of learning abstract and parsimonious rules from mathematical data.
Fast computation of permutation equivariant layers with the partition algebra
Linear neural network layers that are either equivariant or invariant to permutations of their inputs form core building blocks of modern deep learning architectures. Examples include the layers of DeepSets, as well as linear layers occurring in attention blocks of transformers and some graph neural networks. The space of permutation equivariant linear layers can be identified as the invariant subspace of a certain symmetric group representation, and recent work parameterized this space by exhibiting a basis whose vectors are sums over orbits of standard basis elements with respect to the symmetric group action. A parameterization opens up the possibility of learning the weights of permutation equivariant linear layers via gradient descent. The space of permutation equivariant linear layers is a generalization of the partition algebra, an object first discovered in statistical physics with deep connections to the representation theory of the symmetric group, and the basis described above generalizes the so-called orbit basis of the partition algebra. We exhibit an alternative basis, generalizing the diagram basis of the partition algebra, with computational benefits stemming from the fact that the tensors making up the basis are low rank in the sense that they naturally factorize into Kronecker products. Just as multiplication by a rank one matrix is far less expensive than multiplication by an arbitrary matrix, multiplication with these low rank tensors is far less expensive than multiplication with elements of the orbit basis. Finally, we describe an algorithm implementing multiplication with these basis elements.
Hypergraph Models of Biological Networks to Identify Genes Critical to Pathogenic Viral Response
Published
• View Publication
• BIB
Background: Representing biological networks as graphs is a powerful approach to reveal underlying patterns, signatures, and critical components from high-throughput biomolecular data. However, graphs do not natively capture the multi-way relationships present among genes and proteins in biological systems. Hypergraphs are generalizations of graphs that naturally model multi-way relationships and have shown promise in modeling systems such as protein complexes and metabolic reactions. In this paper we seek to understand how hypergraphs can more faithfully identify, and potentially predict, important genes based on complex relationships inferred from genomic expression data sets.
Results: We compiled a novel data set of transcriptional host response to pathogenic viral infections and formulated relationships between genes as a hypergraph where hyperedges represent significantly perturbed genes, and vertices represent individual biological samples with specific experimental conditions. We find that hypergraph betweenness centrality is a superior method for identification of genes important to viral response when compared with graph centrality.
Conclusions: Our results demonstrate the utility of using hypergraphs to represent complex biological systems and highlight central important responses in common to a variety of highly pathogenic viruses.
Multi-Dimensional Scaling on Groups
Leveraging the intrinsic symmetries in data for clear and efficient analysis is an important theme in signal processing and other data-driven sciences. A basic example of this is the ubiquity of the discrete Fourier transform which arises from translational symmetry (i.e. time-delay/phase-shift). Particularly important in this area is understanding how symmetries inform the algorithms that we apply to our data. In this paper we explore the behavior of the dimensionality reduction algorithm multi-dimensional scaling (MDS) in the presence of symmetry. We show that understanding the properties of the underlying symmetry group allows us to make strong statements about the output of MDS even before applying the algorithm itself. In analogy to Fourier theory, we show that in some cases only a handful of fundamental "frequencies" (irreducible representations derived from the corresponding group) contribute information for the MDS Euclidean embedding.
Coherent systems of probability measures on graphs for representations of free Frobenius towers
First formally defined by Borodin and Olshanski, a coherent system on a graded graph is a sequence of probability measures which respect the action of certain down/up transition functions between graded components. In one common example of such a construction, each measure is the Plancherel measure for the symmetric group $S_{n}$ and the down transition function is induced from the inclusions $S_{n} \hookrightarrow S_{n+1}$.
In this paper we generalize the above framework to the case where $\{A_n\}_{n \geq 0}$ is any free Frobenius tower and $A_n$ is no longer assumed to be semisimple. In particular, we describe two coherent systems on graded graphs defined by the representation theory of $\{A_n\}_{n \geq 0}$ and connect one of these systems to a family of central elements of $\{A_n\}_{n \geq 0}$. When the algebras $\{A_n\}_{n \geq 0}$ are not semisimple, the resulting coherent systems reflect the duality between simple $A_n$-modules and indecomposable projective $A_n$-modules.
The center of the twisted Heisenberg category, factorial Schur $Q$-functions, and transition functions on the Schur graph
We establish an isomorphism between the center of the twisted Heisenberg category and the subalgebra of the symmetric functions $Γ$ generated by odd power sums. We give a graphical description of the factorial Schur $Q$-functions as closed diagrams in the twisted Heisenberg category and show that the bubble generators of the center correspond to two sets of generators of $Γ$ which encode data related to up/down transition functions on the Schur graph. Finally, we describe an action of the trace of the twisted Heisenberg category, the $W$-algebra $W^-\subset W_{1+\infty}$, on $Γ$.
Khovanov's Heisenberg category, moments in free probability, and shifted symmetric functions
Published
• View Publication
• BIB
We establish an isomorphism between the center of the Heisenberg category defined by Khovanov and the algebra $Λ^*$ of shifted symmetric functions defined by Okounkov-Olshanski. We give a graphical description of the shifted power and Schur bases of $Λ^*$ as elements of the center, and describe the curl generators of the center in the language of shifted symmetric functions. This latter description makes use of the transition and co-transition measures of Kerov and the noncommutative probability spaces of Biane.
Categorifying the tensor product of the Kirillov-Reshetikhin crystal $B^{1,1}$ and a fundamental crystal
Published
• View Publication
• BIB
We use Khovanov-Lauda-Rouquier (KLR) algebras to categorify a crystal isomorphism between a fundamental crystal and the tensor product of a Kirillov-Reshetikhin crystal and another fundamental crystal, all in affine type. The nodes of the Kirillov-Reshetikhin crystal correspond to a family of "trivial" modules. The nodes of the fundamental crystal correspond to simple modules of the corresponding cyclotomic KLR algebra. The crystal operators correspond to socle of restriction and behave compatibly with the rule for tensor product of crystal graphs.