- The paper introduces DIReCT++, a VLM-modulated rectified flow framework that synthesizes multi-tracer PET images from MRI for precise MCI stratification.
- Experimental evaluations show significant PSNR gains and high classification accuracy compared to baselines such as CycleGAN and diffusion models.
- The method accurately reproduces disease-specific biomarker patterns, supporting non-invasive diagnosis and risk stratification in neurodegenerative disorders.
Precision Synthesis of Multi-Tracer PET via VLM-Modulated Rectified Flow for Stratifying Mild Cognitive Impairment
Introduction and Motivation
The multi-modal characterization of Alzheimer's disease (AD) pathology using PET tracers for amyloid-ฮฒ (18F-AV-45) and neuronal metabolism (18F-FDG) is well-established, yet routine clinical acquisition is limited by cost, radiation risk, and logistical barriers. Accurate, subject-specific PET proxy generation from structural MRI stands as a critical bottleneck for scalable, non-invasive biomarker-driven diagnosis and the prognostic stratification of mild cognitive impairment (MCI). Existing generative models, notably GANs and diffusion-based architectures, have demonstrated visual realism but have failed to resolve the ill-posed inverse mapping from MRI to PET, especially given the heterogeneous pathological burden in early AD stages. The current work presents DIReCT++, a domain-informed rectified flow (RF) framework modulated by a vision-LLM (VLM), designed to synthesize multi-tracer PET from MRI with subject-wise clinical text guidance. This approach directly targets the ambiguity in cross-modal translation, incorporating both global tracer knowledge and granular patient metadata.
Methodological Framework
DIReCT++ integrates a conditional 3D rectified flow model with a domain-adapted VLM, leveraging BiomedCLIP for text-driven guidance. The VLM encodes both general imaging domain knowledge and patient-specific clinical profiles as text embeddings, which are aligned to tracer-specific PET image features via lightweight affine adaptations. These context vectors are incorporated into the velocity field prediction of a multi-task U-Net through cross-attention layers at all spatial scales. The resulting trajectory connecting MRI to PET is governed by an ODE sampled in a single forward pass, ensuring computational efficiency and high-fidelity output. Distillation is applied to collapse the continuous flow into an exact one-step generator. The multi-task formulation allows robust supervision with incomplete dual-tracer availability across real-world clinical cohorts.
Figure 1: Overview of the DIReCT++ framework, illustrating VLM-modulated rectified flow for dual-tracer PET synthesis and its downstream clinical applications.
DIReCT++ was benchmarked against leading architectures: CycleGAN, Swin-UNet, SegGuidedDiff, 3D DDIM, RF (without VLM adaptation), and single-tracer DIReCT. Evaluation employed SSIM, PSNR, MSE, and MAE across ADNI and OASIS datasets. DIReCT++ consistently outperformed all baselines, with substantial quantitative gains such as PSNR improvements of +2.21 to +11.46 dB for 18F-FDG and 180 to 181 dB for 182F-AV-45. Generalization to OASIS yielded PSNR of 183 dB, maintaining superiority outside training distributions.
Figure 2: Radar plots and representative synthetic PET examples demonstrating superior reconstruction quality and dataset generalization for DIReCT++ across diagnostic categories.
Synthesized PET images realistically reproduced metabolic and amyloid patterns observed in disease spectrum subjects, and cross-dataset generalization was robust, affirming the regularization capacity of the VLM-adapted flow.
Regional Biomarker Precision and Disease-Specific Pattern Recapitulation
Anatomical parcellation of both real and synthetic PET using SynthSeg enabled precise ROI-based analyses. Regional SUV comparisons demonstrated nonsignificant paired differences (184) across principal regions, corroborating the quantitative fidelity of DIReCT185. Critically, group-level discriminability between CN, MCI, and AD subjects was preserved: statistical differences between diagnostic groups (e.g., precuneus for amyloid, hippocampus for FDG) identified in real PET were mirrored in synthetic PET, with preserved 186-values and effect sizes.
Figure 3: Magnified regional views show strong anatomical and tracer-wise consistency between synthetic and real PET, and distinct diagnostic group differences.
Figure 4: Violin plots highlight the regional correspondence and disease-specific pattern preservation in synthetic PET compared to real PET for both tracers.
These findings establish the clinical validity of DIReCT187 for robust biomarker reproduction and diagnostic signal maintenance.
Clinical Utility: Stratification and Downstream Classification
DIReCT188 synthetic PET was evaluated in diagnostic and prognostic classification tasks using DenseNet models under 3-fold cross-validation (AD vs. CN, MCI vs. CN, EMCI vs. LMCI). Input modalities included MRI, real PET, synthetic PET, and their combinations. Multi-modal fusion of MRI with synthetic PET yielded classification accuracy of 189 (AD vs. CN), outperforming MRI-only (180) and matching MRI + real PET (181). Sensitivity and specificity were similarly high for synthetic modalities. For MCI stratification, MRI + synthetic PET raised sensitivity to 182 (from 183 MRI-only). Prognostic stratification (EMCI vs. LMCI) reached 184 accuracy for synthetic PET, compared to 185 using MRI, achieving practical utility in risk differentiation.
Figure 5: Bar plots of diagnostic classification metrics demonstrate parity between synthetic and real PET, with pronounced gains in early disease and stratification tasks.
The synthetic PET generated by DIReCT186 thus demonstrated comparable utility to real PET, supporting subject-specific diagnosis and progression risk stratification.
Implications, Limitations, and Future Directions
DIReCT187 marks a methodological advance in domain-informed, VLM-conditioned medical image synthesis. Its subject-level guidance and multi-tracer synergy yield clinically actionable synthetic biomarkers that are both quantitatively precise and biologically discriminatory. Practically, this framework has the potential to democratize PET biomarker profiling and enable scalable screening/intervention in AD, with implications for longitudinal disease monitoring and real-time treatment evaluation without radiation exposure.
Theoretically, DIReCT188 exemplifies the translational value of vision-language conditioning in mitigating cross-modal synthesis ambiguity, establishing a paradigm adaptable to further clinical domains (e.g., tau-PET, other multimodal biomarkers). Future directions include the integration of additional data sources (blood biomarkers, downstream task labels), expansion to broader tracer profiles, and validation in diverse cohorts (including low-field MRI and real-world settings). Incorporation of controllable generation for task-driven personalization and prospective evaluation on clinical populations will be essential to cement its robustness and impact.
Conclusion
DIReCT189 provides a computationally efficient, clinically agile framework for multi-tracer PET synthesis from MRI, leveraging VLM-modulated rectified flow for robust, subject-specific biomarker generation. Its outputs demonstrate superior fidelity, regional precision, disease pattern recapitulation, and practical diagnostic utility, effectively bridging the gap between research prototyping and scalable clinical translation in neurodegenerative disease management.