Papers
Topics
Authors
Recent
Search
2000 character limit reached

Contravariance Theory: Strong Alignment for Minimal Solutions to Hard Tasks

Published 9 Jul 2026 in cs.LG and q-bio.NC | (2607.08561v1)

Abstract: A series of results from the NeuroAI over the past fifteen years have raised core questions both about how to compare Deep Neural Network (DNN) models to the brain, and about how much convergent evolution to expect between artificial networks and real brain networks. Here, we show that for any two minimal DNN solutions to a sufficiently hard task: (i) "weak" alignment of network representations based on affine mappings guarantees "strong" alignment of privileged axes, and (ii) alignment "zippers" up the network hierarchy, causing the emergence of privileged axes from end-to-end task optimization. These results formalize the notion of contravariance from Cao and Yamins [2024], and illustrate important consequences for the theory of NeuroAI: with sufficiently strong tasks, choice of metric for inter-network comparison is not all that sensitive, and that convergent evolution is probably inevitable.

Authors (2)

Summary

  • The paper demonstrates that increasing task hardness contracts the solution space in minimal DNNs, resulting in strong axes-level alignment.
  • It introduces weak–strong equivalence theorems, proving that diminishing weak-alignment errors guarantee near-universal strong alignment across network layers.
  • The work offers empirical guidelines for neuroscience by showing that task-induced minimality leads to robust model–brain convergence.

Contravariance Theory and the Emergence of Strong Alignment in Minimal DNN Solutions

Introduction and Motivation

A series of results in NeuroAI have demonstrated that task-optimized DNNs, across sensory, cognitive, and motor tasks, develop internal representations that are tightly aligned with those observed in biological neural systems—often up to a linear transformation. Intriguingly, the degree of model-neural similarity is monotonically correlated with the DNN's performance on the task at hand, and such alignment is detectable across different modalities and brain regions (Figure 1). Figure 1

Figure 1: Cross-domain correlations between DNN task performance and fit to neural data, using linear similarity metrics across multiple sensory and language areas.

However, this empirical alignment—previously considered surprising given the nonlinearity and high-dimensionality of both model and brain representations—challenges prevailing notions regarding comparability metrics across models and the sufficiency of mere performance matching as a criterion for theoretical or mechanistic parity.

The Contravariance Principle

Previous conceptual work (see Cao & Yamins, 2024) introduced the contravariance principle: as task complexity increases, the solution space for a fixed-capacity DNN contracts, such that multiple independently optimized models must converge toward similar representational structures. This work rigorously formalizes the principle and demonstrates nontrivial consequences: with sufficiently "hard" tasks and appropriately "minimal" solutions, task performance enforces representational alignment not only under weak (e.g., affine) metrics, but also under much stricter axes-level criteria. Figure 2

Figure 2: The contravariance principle: solution set dispersion is inversely related to task constraint strength, tightening as task difficulty increases.

Weak–Strong Equivalence: Theoretical Contribution

This framework distinguishes between weak alignment (linear/affine similarity between representations) and strong alignment (alignment at privileged axes or single-neuron level). Contrary to naive expectations, the paper establishes—under minimality and task hardness assumptions—the Weak–Strong Equivalence Theorems: weak alignment across adjacent network layers mathematically enforces strong axes alignment, provided all nonlinear "gates" (ReLU/softplus) are used by the task.

Fine-grained analysis of activation signatures (e.g., ReLU kinks, softplus curvature ridges) demonstrates that affine transformations cannot destroy task-derived nonlinear signatures at the axes level unless those axes are unused, degenerate, or redundant. Formally, the lower bound on axes-level alignment depends on a task-induced nonlinear axis budget mℓ(ϵ)m_\ell(\epsilon) relative to the layer width dℓAd_\ell^A, quantifying task “hardness.”

Asymptotic Equivalence and Learning Dynamics

Since models rarely achieve exact weak alignment, the authors prove robust, quantitative versions of these theorems: as weak-alignment errors decrease across training, the fraction of axes achieving strong alignment approaches unity. The error in axes-level alignment is bounded above by terms dependent on the layerwise weak-alignment errors and constants determined by architecture class, formalized via a family-dependent constant κK(θ)\kappa_\mathcal{K}(\theta).

This provides a rigorous explanation for observed layerwise progression of representational similarity across independently initialized, identically-trained networks, and allows for precise empirical testing via soft matching metrics.

Zippering and Hierarchical Propagation

Extending further, the Zippering Theorem demonstrates that minimality and task hardness do not merely ensure terminal alignment; rather, weak alignment at the final layer “zippers up” through the full hierarchy (Figure 3), aligning axes recursively at all previous layers. Here, the propagation of alignment is both an empirical observation and a theoretical necessity given the network’s minimality. Figure 3

Figure 3: The layerwise "zippering" pattern: nonlinearity increases divergence, post-nonlinearity similarity is regained upstream, establishing axes-level alignment recursively.

This theorem formalizes the process by which hierarchical and area-specific matches between DNN and brain representations—and between independently trained DNNs—are both expected and inevitable for hard tasks and minimal solutions.

Empirical and Theoretical Implications

The theorems render the oft-observed DNN–brain representational consistency and robust correspondence between model layers and cortical areas expected rather than exceptional (contingent on task hardness and minimality). This conclusion is robust to choice of metric: for hard tasks, even “loose” metrics (e.g., linear regression fit or task performance) enforce “strict” alignment (e.g., axes or unit-level).

Minimality as a Crucial Property

Model solutions must be minimal—free of functionally superfluous axes or affine redundancies—to guarantee contravariance-induced alignment. The paper formalizes the related gate budget concept, which offers a principled proxy for minimality and can potentially be estimated empirically via bottlenecking or ablation.

Guidance for Neuroscience Experimental Design

The results offer a normative prescription: neuroscientists and NeuroAI practitioners should employ ethologically valid, computationally hard tasks to maximize the probability of meaningful model-brain convergence and avoid alignment driven by trivial or degenerate solutions.

Consequences for Comparisons Across Metrics

These findings moot the search for a “correct” metric among affine, Procrustes, or even representational similarity analysis (RSA) frameworks when modeling hard tasks with minimal networks: all enforce approximately the same alignment. Nevertheless, multiplicity and redundancy may confound RSA in ways that linear metrics avoid.

Relevance for Diverse Architectures

While the main results are proven for affine–ReLU and affine–softplus nets, the underlying mechanism is expected to generalize to contemporary architectures (CNNs, ViTs, Transformers, RNNs), pending activation-specific identifiability and task-layer axis minimality.

Contradictions and Strong Claims

  • Privileged axes (e.g., Gabor-like units, V1 tuning) emerge inevitably from hard-task constraints: axes-level structure is not an artifact of initialization or optimization idiosyncrasies.
  • Convergent evolution of DNN and brain representations is expected for hard tasks under minimality: the number of minimal solutions is so restricted that independent optimizations converge on similar axes.
  • Choice of metric is irrelevant for hard tasks/minimal models: weak similarity forces strong similarity, rendering more nuanced metrics redundant.

Future Directions

  • Systematic empirical estimation of the minimal gate budgets and testing the practical tightness of theoretical bounds in large models (e.g., LLMs, ViTs) and real neural systems.
  • Assessing whether SGD-based optimization actually avoids rare degenerate solutions (exceptional sets) in practice.
  • Extending formalism to architectures with tied parameters, multi-branch modules, or complex non-affine nonlinearities.
  • Investigating implications for the Platonic Representation Hypothesis and cross-modal representational alignment across tasks.

Conclusion

This work provides a rigorous mathematical foundation for the principle that “hard” cognitive or sensorimotor tasks strongly constrain the space of admissible DNN solutions to such an extent that even weak performance or linear-affine similarity metrics almost uniquely determine deep representational axes. Theoretical results justify, and in fact necessitate, the robust alignment phenomena widely observed in both ML and neuroscience, shifting the paradigm from surprise at convergent evolution to an expectation based on the geometry of high-dimensional constrained optimization and the structure of minimal solutions.

The contravariance framework thus consolidates a key principle for model–brain comparison, metrics selection, task design, and the future theoretical landscape of NeuroAI.


Figure 2

Figure 2: The contravariance principle’s implications for the solution space geometry as task hardness increases.

Figure 3

Figure 3: Empirical and theoretical “zippering”: recurring pattern of high similarity pre-nonlinearity, with nonlinearity-induced divergence, and re-alignment in the next affine stage, layerwise up the hierarchy.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 2 tweets with 16 likes about this paper.