Papers
Topics
Authors
Recent
Search
2000 character limit reached

Short-Term Synaptic Plasticity Stabilizes Goal-Conditioned Dynamics in a PFC-Inspired Reservoir Model for Multistep Goal-Directed Action Planning

Published 2 Jun 2026 in q-bio.NC and cs.NE | (2606.03481v1)

Abstract: The prefrontal cortex (PFC) maintains goal information for action planning, but how recurrent circuits preserve it in an action-usable form over behavioral timescales remains unclear. Here we ask whether short-term synaptic plasticity (STP) can stabilize goal information as action-usable, goal-conditioned dynamics. We incorporated STP into a PFC-inspired reservoir computing model with basal-ganglia-inspired temporal-difference readout learning, and evaluated paired models with and without STP across 100 independently generated networks in a multistep goal-directed action-selection task with delayed execution. Goal identity was highly decodable during the delay even without STP, so STP was not required to form a linearly readable goal representation. Under state noise, however, success without STP fell from 75.8% to 49.5%, whereas the model with STP remained essentially unchanged (91.8% without noise versus 89.2% under noise; paired Cohen's dz=1.31). Time-resolved decoding, state-space separability, and action-value-difference analyses showed that STP preserved goal information as action-relevant goal-conditioned dynamics available at later action opportunities. Gain-matched and STP-state perturbation controls argued against a simple fixed recurrent-scaling explanation and supported online, history-dependent synaptic modulation. Effective-connectivity analyses showed delay-period goal-specific patterning that increased toward the later part of the trial with STP, where it should be read as goal- and task-state-conditioned patterning; effective connectivity without STP was time-invariant. A grid search identified a facilitation-dominant range of STP time constants associated with high success rates. These results suggest that STP supports robust goal-conditioned dynamics through dynamic modulation of goal-dependent effective recurrent connectivity.

Authors (2)

Summary

  • The paper demonstrates that incorporating short-term synaptic plasticity into a PFC-inspired reservoir model improves retention and robustness of goal-conditioned dynamics.
  • The model uses a reservoir framework with TD learning and the Tsodyks–Markram STP rule to sustain goal information under noise and perturbations.
  • Results reveal that a facilitation-dominant regime enhances multistep action planning by dynamically reorganizing recurrent connectivity.

Short-Term Synaptic Plasticity Stabilizes Goal-Conditioned Dynamics in a PFC-Inspired Reservoir Model


Introduction

The paper "Short-Term Synaptic Plasticity Stabilizes Goal-Conditioned Dynamics in a PFC-Inspired Reservoir Model for Multistep Goal-Directed Action Planning" (2606.03481) presents a systematic investigation of the computational role of short-term synaptic plasticity (STP) in goal representation and action planning within prefrontal cortex (PFC)-inspired recurrent circuits. The study employs a reservoir computing (RC) framework augmented with STP, coupled with reward-driven temporal-difference (TD) learning, in a multistep goal-directed action-selection task. The core inquiry is whether STP can support the retention of goal information in a dynamically usable form across behavioral delays, specifically under state noise, and how STP modulates effective recurrent connectivity to produce robust goal-conditioned dynamics. Through extensive paired statistical evaluations, representational analyses, and mechanistic controls, this work delineates the role of STP as a substrate for working memory and action preparation.


Model Architecture and Biological Motivation

The model architecture closely follows the paradigm of reservoir computing, where only output weights are learned and all internal recurrence remains fixed, aligning with the limited synaptic plasticity observed in cortical circuits outside learning periods. The architecture includes an input layer encoding environmental state and goal cues, a central reservoir with fixed sparse recurrence, and an output layer generating Q-values for action selection. STP is incorporated into the recurrent layer via the Tsodyks–Markram model, providing history-dependent modulation of synaptic efficacy. The network receives task inputs comprising a goal identity, current agent position, and GO signal, processed through the reservoir to produce Q-values for downstream action selection. Reward feedback is injected to facilitate reinforcement learning via TD error updates. Figure 1

Figure 1: Network architecture of the proposed model, including STP-modulated recurrence and reward-based TD learning for output weights.

The modeling rationale blends biological plausibility with computational tractability: the RC constraint that only output weights adapt matches cortico-basal ganglia systems, where recurrent structure is thought to be largely developmentally constrained, and output pathways are shaped via reward signals. STP provides a candidate mechanism for activity-silent memory, a key challenge in PFC-based models of working memory.


Experimental Paradigm and Methodology

The task is derived from primate goal-directed action planning experiments, requiring the agent to maintain a cued goal across a delay and execute a sequence of actions to reach it. A 5×55{\times}5 grid world with eight non-corner perimeter goals and four sequential GO windows is used. Key features: goal input is removed after cue onset (placing demand on internal retention), successful completion requires at most three actions, and state noise is injected to simulate biological variability. Evaluations are conducted with and without STP, across 100 independent random seeds (distinct network architectures), with paired controls ensuring all other factors remain fixed.

TD learning is applied to output weights only; actions are selected greedily based on Q-values, and reward shaping penalizes non-progressive moves. STP variables (xj,uj)(x_j, u_j) are neuron-wise and updated via presynaptic activity, introducing facilitation and depression with experimentally-motivated time constants.


Results: Learning, Robustness, and Noise Sensitivity

Learning dynamics show rapid convergence in the With-STP condition, with superior success rates and reduced variance across seeds compared to the Without-STP baseline. Under state noise (σ=0.001\sigma = 0.001), the success rate plummets in the Without-STP condition (from 75.8%75.8\% to 49.5%49.5\%), but is essentially preserved with STP (91.8%91.8\% vs 89.2%89.2\%). The paired effect size under noise is substantial (dz=1.31d_z = 1.31). Figure 2

Figure 2: STP preserves post-training task success under state noise, markedly outperforming the baseline.

Control experiments reject the hypothesis that STP's benefit is attributable solely to increased recurrent gain or fixed scaling. A gain-matched Without-STP condition does not replicate the performance, and acute perturbation of STP state (reset, freezing) post-training significantly degrades robustness.


Internal Representation and Goal-Conditioned Dynamics

PCA projections and linear decoding demonstrate that goal identity is highly decodable during the delay in both conditions; however, only the With-STP model maintains decodability up to later GO windows and under noise. Scatter ratio analyses reveal that STP preserves geometric separation of goal-conditioned trajectories in state space, confirmed to be due to directional separation rather than amplitude scaling. Figure 3

Figure 3: Behavioral and PCA trajectories illustrate robust separation of goal representations with STP versus diffusion under baseline.

Figure 4

Figure 4: Time-resolved goal decoding accuracy remains near ceiling only in the With-STP condition, especially under noise.

Action-value difference analyses at GO opportunities indicate that STP enhances the readout-level expression of goal-conditioned structure, favoring target-consistent actions even when the model is evaluated under perturbations.


Mechanistic Analysis: Effective Connectivity and Spectral Structure

Beyond behavioral and representational evidence, STP induces dynamic, goal-specific modulation of effective recurrent connectivity—quantified via the cosine similarity index $\Delta_{\mathrm{sim}(t)$. This index rises after cue presentation, persists through delay, and peaks near action execution, supporting the hypothesis that STP reorganizes the recurrent substrate in a history-dependent, goal-centric manner. Figure 5

Figure 5: Representative time-series of scatter ratio demonstrates maintained goal separability with STP.

Figure 6

Figure 6: Temporal profile of the relative effective spectral radius ρrel\rho_\mathrm{rel} reflects persistent STP-mediated scaling modulation.

Figure 7

Figure 7: Eigenvalue-magnitude distribution of effective connectivity demonstrates time-varying spectral structure with STP.

Figure 8

Figure 8: Temporal evolution of (xj,uj)(x_j, u_j)0 quantifies the goal specificity of effective connectivity, confirming dynamic reorganization.

High success rates are localized in a facilitation-dominant regime, where STP’s facilitation time constant (xj,uj)(x_j, u_j)1 exceeds resource recovery (xj,uj)(x_j, u_j)2. This parameter band aligns with short-term facilitation observed experimentally in PFC, lending physiological plausibility to the computational model. Figure 9

Figure 9: Exploratory parameter grid reveals a facilitation-dominant band (low (xj,uj)(x_j, u_j)3, high (xj,uj)(x_j, u_j)4) associated with robust goal retention.


Theoretical and Practical Implications

The findings support a view of working memory where delay-period goal information is not merely statically encoded but dynamically stabilized by activity-silent modulation of effective recurrent connectivity via STP. This mechanism enhances robustness to noise and supports multistep action selection beyond classical persistent-firing models. The results refine the dynamical-systems interpretation of recurrent neural circuits, highlighting transient reconfiguration as a computational motif in the PFC.

Reservoir-based architectures with STP provide a biologically plausible substrate for planning and memory in AI systems, especially for tasks demanding temporally extended retention and multi-action sequencing. The discovered facilitation-dominant operational regime offers guidance for the design of recurrent models with improved robustness and memory characteristics.

Future developments include testing generalization to variable delays, distractor inputs, more complex environments, and comparative studies with alternative memory mechanisms (gated units, attractor networks, slow time constants). Integrated local Jacobian analyses can reveal how STP-modulated connectivity shapes the stability and expressivity of recurrent flows.


Conclusion

The study rigorously establishes that short-term synaptic plasticity, as instantiated in PFC-inspired reservoir models, offers a powerful mechanism for stabilizing goal-conditioned dynamics critical to multistep action planning. STP achieves this not by merely enhancing decodability during the delay, but by dynamically maintaining goal information in an action-relevant form resilient to noise and perturbation. The facilitation-dominant regime identified provides a biologically consistent operational window. The implications extend to both theoretical neuroscience and the engineering of memory-robust recurrent architectures for AI.


Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 1 tweet with 2 likes about this paper.