Papers
Topics
Authors
Recent
Search
2000 character limit reached

Precision Synthesis of Multi-Tracer PET via VLM-Modulated Rectified Flow for Stratifying Mild Cognitive Impairment

Published 13 Apr 2026 in cs.CV | (2604.11176v1)

Abstract: The biological definition of Alzheimer's disease (AD) relies on multi-modal neuroimaging, yet the clinical utility of positron emission tomography (PET) is limited by cost and radiation exposure, hindering early screening at preclinical or prodromal stages. While generative models offer a promising alternative by synthesizing PET from magnetic resonance imaging (MRI), achieving subject-specific precision remains a primary challenge. Here, we introduce DIReCT$++$, a Domain-Informed ReCTified flow model for synthesizing multi-tracer PET from MRI combined with fundamental clinical information. Our approach integrates a 3D rectified flow architecture to capture complex cross-modal and cross-tracer relationships with a domain-adapted vision-LLM (BiomedCLIP) that provides text-guided, personalized generation using clinical scores and imaging knowledge. Extensive evaluations on multi-center datasets demonstrate that DIReCT$++$ not only produces synthetic PET images (${18}$F-AV-45 and ${18}$F-FDG) of superior fidelity and generalizability but also accurately recapitulates disease-specific patterns. Crucially, combining these synthesized PET images with MRI enables precise personalized stratification of mild cognitive impairment (MCI), advancing a scalable, data-efficient tool for the early diagnosis and prognostic prediction of AD. The source code will be released on https://github.com/ladderlab-xjtu/DIReCT-PLUS.

Summary

  • The paper introduces DIReCT++, a VLM-modulated rectified flow framework that synthesizes multi-tracer PET images from MRI for precise MCI stratification.
  • Experimental evaluations show significant PSNR gains and high classification accuracy compared to baselines such as CycleGAN and diffusion models.
  • The method accurately reproduces disease-specific biomarker patterns, supporting non-invasive diagnosis and risk stratification in neurodegenerative disorders.

Precision Synthesis of Multi-Tracer PET via VLM-Modulated Rectified Flow for Stratifying Mild Cognitive Impairment

Introduction and Motivation

The multi-modal characterization of Alzheimer's disease (AD) pathology using PET tracers for amyloid-ฮฒ\beta (18^{18}F-AV-45) and neuronal metabolism (18^{18}F-FDG) is well-established, yet routine clinical acquisition is limited by cost, radiation risk, and logistical barriers. Accurate, subject-specific PET proxy generation from structural MRI stands as a critical bottleneck for scalable, non-invasive biomarker-driven diagnosis and the prognostic stratification of mild cognitive impairment (MCI). Existing generative models, notably GANs and diffusion-based architectures, have demonstrated visual realism but have failed to resolve the ill-posed inverse mapping from MRI to PET, especially given the heterogeneous pathological burden in early AD stages. The current work presents DIReCT++++, a domain-informed rectified flow (RF) framework modulated by a vision-LLM (VLM), designed to synthesize multi-tracer PET from MRI with subject-wise clinical text guidance. This approach directly targets the ambiguity in cross-modal translation, incorporating both global tracer knowledge and granular patient metadata.

Methodological Framework

DIReCT++++ integrates a conditional 3D rectified flow model with a domain-adapted VLM, leveraging BiomedCLIP for text-driven guidance. The VLM encodes both general imaging domain knowledge and patient-specific clinical profiles as text embeddings, which are aligned to tracer-specific PET image features via lightweight affine adaptations. These context vectors are incorporated into the velocity field prediction of a multi-task U-Net through cross-attention layers at all spatial scales. The resulting trajectory connecting MRI to PET is governed by an ODE sampled in a single forward pass, ensuring computational efficiency and high-fidelity output. Distillation is applied to collapse the continuous flow into an exact one-step generator. The multi-task formulation allows robust supervision with incomplete dual-tracer availability across real-world clinical cohorts. Figure 1

Figure 1: Overview of the DIReCT++ framework, illustrating VLM-modulated rectified flow for dual-tracer PET synthesis and its downstream clinical applications.

Quantitative and Qualitative Performance: Fidelity and Generalization

DIReCT++++ was benchmarked against leading architectures: CycleGAN, Swin-UNet, SegGuidedDiff, 3D DDIM, RF (without VLM adaptation), and single-tracer DIReCT. Evaluation employed SSIM, PSNR, MSE, and MAE across ADNI and OASIS datasets. DIReCT++++ consistently outperformed all baselines, with substantial quantitative gains such as PSNR improvements of +2.21+2.21 to +11.46+11.46 dB for 18^{18}F-FDG and 18^{18}0 to 18^{18}1 dB for 18^{18}2F-AV-45. Generalization to OASIS yielded PSNR of 18^{18}3 dB, maintaining superiority outside training distributions. Figure 2

Figure 2: Radar plots and representative synthetic PET examples demonstrating superior reconstruction quality and dataset generalization for DIReCT++ across diagnostic categories.

Synthesized PET images realistically reproduced metabolic and amyloid patterns observed in disease spectrum subjects, and cross-dataset generalization was robust, affirming the regularization capacity of the VLM-adapted flow.

Regional Biomarker Precision and Disease-Specific Pattern Recapitulation

Anatomical parcellation of both real and synthetic PET using SynthSeg enabled precise ROI-based analyses. Regional SUV comparisons demonstrated nonsignificant paired differences (18^{18}4) across principal regions, corroborating the quantitative fidelity of DIReCT18^{18}5. Critically, group-level discriminability between CN, MCI, and AD subjects was preserved: statistical differences between diagnostic groups (e.g., precuneus for amyloid, hippocampus for FDG) identified in real PET were mirrored in synthetic PET, with preserved 18^{18}6-values and effect sizes. Figure 3

Figure 3: Magnified regional views show strong anatomical and tracer-wise consistency between synthetic and real PET, and distinct diagnostic group differences.

Figure 4

Figure 4: Violin plots highlight the regional correspondence and disease-specific pattern preservation in synthetic PET compared to real PET for both tracers.

These findings establish the clinical validity of DIReCT18^{18}7 for robust biomarker reproduction and diagnostic signal maintenance.

Clinical Utility: Stratification and Downstream Classification

DIReCT18^{18}8 synthetic PET was evaluated in diagnostic and prognostic classification tasks using DenseNet models under 3-fold cross-validation (AD vs. CN, MCI vs. CN, EMCI vs. LMCI). Input modalities included MRI, real PET, synthetic PET, and their combinations. Multi-modal fusion of MRI with synthetic PET yielded classification accuracy of 18^{18}9 (AD vs. CN), outperforming MRI-only (18^{18}0) and matching MRI + real PET (18^{18}1). Sensitivity and specificity were similarly high for synthetic modalities. For MCI stratification, MRI + synthetic PET raised sensitivity to 18^{18}2 (from 18^{18}3 MRI-only). Prognostic stratification (EMCI vs. LMCI) reached 18^{18}4 accuracy for synthetic PET, compared to 18^{18}5 using MRI, achieving practical utility in risk differentiation. Figure 5

Figure 5: Bar plots of diagnostic classification metrics demonstrate parity between synthetic and real PET, with pronounced gains in early disease and stratification tasks.

The synthetic PET generated by DIReCT18^{18}6 thus demonstrated comparable utility to real PET, supporting subject-specific diagnosis and progression risk stratification.

Implications, Limitations, and Future Directions

DIReCT18^{18}7 marks a methodological advance in domain-informed, VLM-conditioned medical image synthesis. Its subject-level guidance and multi-tracer synergy yield clinically actionable synthetic biomarkers that are both quantitatively precise and biologically discriminatory. Practically, this framework has the potential to democratize PET biomarker profiling and enable scalable screening/intervention in AD, with implications for longitudinal disease monitoring and real-time treatment evaluation without radiation exposure.

Theoretically, DIReCT18^{18}8 exemplifies the translational value of vision-language conditioning in mitigating cross-modal synthesis ambiguity, establishing a paradigm adaptable to further clinical domains (e.g., tau-PET, other multimodal biomarkers). Future directions include the integration of additional data sources (blood biomarkers, downstream task labels), expansion to broader tracer profiles, and validation in diverse cohorts (including low-field MRI and real-world settings). Incorporation of controllable generation for task-driven personalization and prospective evaluation on clinical populations will be essential to cement its robustness and impact.

Conclusion

DIReCT18^{18}9 provides a computationally efficient, clinically agile framework for multi-tracer PET synthesis from MRI, leveraging VLM-modulated rectified flow for robust, subject-specific biomarker generation. Its outputs demonstrate superior fidelity, regional precision, disease pattern recapitulation, and practical diagnostic utility, effectively bridging the gap between research prototyping and scalable clinical translation in neurodegenerative disease management.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 1 tweet with 1 like about this paper.