- The paper presents NEvo, a two-stage neural-guided evolutionary framework that synthesizes ROI-specific dynamic videos to robustly activate target cortical regions.
- It couples static and dynamic optimization through semantic gene encoding and model-based prompt search, outperforming traditional gradient-based synthesis methods.
- Results reveal distinct dynamic selectivity patterns across visual cortical regions, with significant activation improvements in areas like MT and FFA.
Neural-Guided Evolutionary Video Synthesis for Probing Dynamic Visual Selectivity
The paper "NEvo: Neural-Guided Evolutionary Video Synthesis for Dynamic Visual Selectivity" (2607.02317) presents NEvo, a neural-guided evolutionary framework for synthesizing video stimuli tailored to maximally activate target cortical regions. NEvo interrogates dynamic selectivity along the visual hierarchy via model-based video prompt optimization, providing an in silico approach to reveal both established and previously uncharacterized functional gradients in the human visual cortex under temporally structured input.
Figure 1: A. NEvo’s synthesis pipeline iteratively refines video prompts to maximize predicted ROI activation. B. Results highlight synthesis of highly activating stimuli, dynamic selectivity, and trajectory analysis along the lateral stream.
Methodological Framework
NEvo couples a dynamic fMRI encoding model, trained on large-scale naturalistic video datasets, with an evolutionary search over a structured, interpretable prompt space. Each prompt is encoded as a set of semantic, appearance, and motion "genes" that are translated into text for video diffusion models. The framework employs a two-stage optimization: (1) image prompt search for optimal static anchor, followed by (2) video prompt search to tailor motion/event descriptors. Each candidate’s video, rendered via current state-of-the-art video diffusion architectures, is scored by the encoding model with ROI-level mean predicted activation as the synthesis objective.
Figure 2: A. Evolutionary prompting encodes visual/motion attributes as genes guiding generative search. B. Two-stage search first selects optimal static content, then dynamically varies motion. C. Multi-layer encoding selects best-predictive model feature per voxel.
NEvo’s model-based search outperforms direct gradient-based video synthesis, which fails to generate coherent, meaningful dynamics under high prompt dimensionality and generator nonlinearity.
Results: Dynamic Selectivity Across the Visual Cortex
Applying NEvo to classical and novel ROIs—including FFA, PPA, EBA, MT, V3A, and pSTS—demonstrates robust recovery of expected region-specific selectivity. The synthesized stimuli for each ROI display content and dynamics characteristic of the region’s encoding: face-optimized videos for FFA, body-centric motion for EBA, coherent global motion for MT/V3A, and coordinated social interaction for pSTS.
NEvo reliably exceeds both dataset-retrieval and functional localizer benchmarks in predicted activation, achieving >99.5th percentile of Moment-in-Time dataset responses and surpassing dynamic localizer videos:
Figure 3: A. NEvo generates ROI-specific dynamic videos. B. Search trajectories show mean activation gains. C. Synthesized videos consistently outperform dataset-retrieval and hand-designed localizers.
Crucially, dynamic stimuli synthesized by NEvo drive predicted activation significantly higher than matched static-frame controls, with the effect size maximized in motion-selective regions (e.g., MT: mean improvement of 0.61±0.05 a.u.) but clearly present even in classically static-selective ventral areas (FFA: 0.14±0.02 a.u.).
The two-stage optimization is particularly efficient, outperforming single-stage search variants both in computation and outcome, with major gains observed for regions with high selectivity for static content anchored by subsequent dynamic optimization.
Figure 4: A. Search trajectories over time show dynamic evaluation boosts. B. Dynamic videos consistently evoke higher activation than static anchors. C–F. NEvo ablations: Two-stage and genetic search outperform alternatives; multi-layer V-JEPA 2 encoding surpasses last-layer/CLIP baselines and gradient-based synthesis.
Revealing Gradients of Dynamic Social Feature Selectivity
A searchlight mapping along the lateral pathway (V1→MT→EBA→pSTS→aSTS) exposes a continuous increase in complexity for synthesized video-evoked features. Early patches are selectively driven by texture and simple global motion, mid-level regions by body and coordinated movement, and anterior regions by rich interpersonal and social interaction cues. Automated annotation and correlation analyses show region-specific visual property–activation relationships, with, e.g., biological motion most predictive for MT and joint action/social communication peaking in pSTS.
Figure 5: A. Searchlight-synthesized lateral stream stimuli and semantic content summary. B. Correlation between annotated visual properties and predicted activations for key regions and properties.
Controlled Synthesis Experiments
By anchoring the synthesis to abstract, non-naturalistic initial frames, NEvo isolates dynamic selectivity independent of static scene attributes. For instance, pSTS-optimized dynamics add face-like, socially coordinated motion, whereas MT-optimized sequences favor non-biological, high-motion energy content, underscoring the framework’s potential for hypothesis-driven stimulus generation.
Figure 6: Fixing static anchor content, NEvo can drive region-specific dynamic selectivity (e.g., pSTS: social/face features; MT: high-energy motion).
Implications and Future Perspectives
NEvo offers a generalized, neuro-aligned platform to map and probe dynamic selectivity in the human visual system, extending model-based cortical stimulus design beyond statics to full spatiotemporal complexity. It enables direct testing of hypotheses regarding hierarchical organization, specialization, and region-specific temporal integration—particularly along the lateral visual pathway, but extensible to parietal/dorsal and ventral processing.
The robust superiority of semantic evolutionary search over gradient-driven optimization for dynamic synthesis highlights the importance of discrete, interpretable structure in guiding high-dimensional generative models for neurobiological discovery.
Key claim: Dynamic video selectivity is widespread, with temporal structure contributing non-trivially to activation in regions previously considered static-sensitive, and regional specialization for higher-order social features can be systematically revealed through in silico search.
NEvo’s predictions are positioned for prospective validation via closed-loop neuroimaging or behavioral experiments, with potential utility in designing paradigm-shifting localizers and clinical diagnostics.
Methodological Strengths and Limitations
The approach depends critically on the dynamic encoding model’s fidelity and generalization. Any bias or error in the model will propagate into synthesized stimuli and hypotheses. Furthermore, the current pipeline is strictly visual: regions with multimodal selectivity (e.g., aSTS, multimodal temporal-parietal junction) require encoding models with integrated audio/language/video inputs for full exploration.
Prospects
Advances in spatiotemporal generative models and integration with multimodal brain-aligned encoding architectures—especially in the vein of V-JEPA 2 and large-scale video-text foundation models—are expected to further enhance the specificity, diversity, and scalability of neural-guided dynamic synthesis. Model-in-the-loop experiments ("inception loops" for videos) can close the loop from computational prediction to empirical validation, potentially revealing new principles of hierarchical and functional organization across temporal, spatial, and semantic domains.
Conclusion
NEvo provides a scalable, interpretable, and region-agnostic framework for synthesizing dynamic spatiotemporal stimuli optimized for neural selectivity in the human visual cortex. Integrating evolutionary prompt search with multi-layer brain modeling, NEvo demonstrates that temporal structure expands our understanding of feature representation and regional specialization. This work opens avenues for high-throughput discovery and hypothesis testing in systems neuroscience, offering a template for future AI research in dynamic stimulus design, brain-model alignment, and the study of vision’s hierarchical organization.