- The paper presents a probabilistic framework that integrates multi-omics data, curated regulatory priors, and latent transcription factor activity for network reconstruction.
- It employs modular clustering, consensus filtering, and pseudo-likelihood approximations to enhance biological interpretability and performance metrics such as AUPR and F-score.
- The approach is validated on mouse cellular reprogramming data, successfully capturing both canonical and latent regulatory interactions.
MERLIN-SUITE: Probabilistic, Modular GRN Inference Integrating Multi-Omics, Prior Knowledge, and Latent Transcription Factor Activity
Introduction and Motivation
Accurate reconstruction of gene regulatory networks (GRNs) is a central challenge for understanding the complexity and plasticity of transcriptional control in development, differentiation, and disease. Traditional GRN inference from transcriptomics data alone is limited by the inability of mRNA abundance to capture post-transcriptional and post-translational regulatory dynamics, inherent data noise, and the lack of mechanistic alignment with experimentally validated interactions. MERLIN-SUITE constitutes a comprehensive suite of inference algorithms developed to overcome these limitations by integrating gene expression with curated regulatory priors (motifs, ChIP, perturbations) and, crucially, leveraging latent transcription factor activity (TFA) inferred directly from data and priors.
Theoretical Framework and Algorithmic Components
Probabilistic Modular GRN Inference with MERLIN
The foundation of MERLIN-SUITE is the MERLIN framework, a probabilistic graphical model defining GRN inference as a joint likelihood maximization problem over module assignments, network structure, and model parameters. Here, genes are clustered into co-regulated modules, and both per-gene (fine-grained) and per-module (coarse-grained) regulatory programs are inferred. The inference is regularized by an edge prior incorporating module-specific and external (experimental or computational) prior probabilities and employs a pseudo-likelihood approximation to scale with large omics datasets.
Integration of Regulatory Priors (MERLIN-P)
MERLIN-P augments the inference by directly integrating external prior networksโe.g., those derived from ATAC-seq, ChIP-seq, or TF perturbation screensโinto the graph prior. Edge priors are modeled via a logistic function parameterized by the strength/confidence of prior evidence, enabling systematic up-weighting of biologically plausible interactions. This prior-based constraint substantially increases the biological interpretability and predictive performance of inferred networks with respect to gold-standard references.
Activity-Aware Modeling with MERLIN-P-TFA
MERLIN-P-TFA addresses the central problem that mRNA abundance is frequently a poor proxy for true protein activity, especially for transcription factors subject to extensive post-translational regulation. The key innovation is the regularized estimation of latent TFAs via network component analysis (NCA) or its sparsity-promoting extension, NCA-LASSO. The inferred TFA matrix (regulator activities ร samples/cells) is then concatenated with the original expression matrix and incorporated as an augmented feature set, with updated priors, for the subsequent network inference. This workflow enables robust capture of "hidden" regulatory signals that are undetectable at the transcript level.
Empirical Workflow and Evaluation
Application to Mouse Cellular Reprogramming
The manuscript details a benchmark case study involving single-cell, multi-modal data from a mouse fibroblast-to-embryonic-stem-cell (MEF-to-mESC) reprogramming system, spanning multiple reprogramming conditions and time points. Following estimation of TFA from a regulatory prior based on bulk ATAC-seq and motif data, the combined TFA/expression matrix is subjected to probabilistic network inference over multiple data subsamplings.
Network Consensus and Functional Validation
Edges are filtered by consensus (present in โฅ80% of runs), yielding high-confidence GRNs. Global accuracy is quantitatively assessed via AUPR and F-score against curated mESC gold-standard networks (including KD, ChIP, and Chip+KD sources). Notably, MERLIN-P-TFA achieves improved or comparable AUPR and F-score metrics across diverse prior sources and regularization settings, with the highest performance observed at moderate regularization (ฮป=0.1).
Functional module assignments are derived through co-clustering analysis and show significant enrichment in GO categories such as stem cell development, DNA replication, and macromolecular complex assembly, confirming the biological relevance of the clusters.
Visual Analytics and Multi-Scale Network Interpretation
MERLIN-SUITE includes visualization utilities (MERLIN-VIZ, Cytoscape exports) supporting interrogation of context-specific (e.g., cell-type, condition) network rewiring, hub regulator identification, and temporal dynamics of regulatory architecture. In-depth analyses of module-specific networks (especially Module 921, implicated in mESC identity) demonstrate the progressive rewiring and activation of core pluripotency regulators (Nanog, Esrrb, Tfcp2l1) and reveal staged shifts from early antagonism to late cooperation between key nodes. Critically, the inclusion of TFA enables identification of noncanonical, latent regulators such as Snai2_TFA and Nr6a1_TFA, which are predicted to prime reprogramming by targeting chromatin modifiers and early pluripotency genes.
- MERLIN-P-TFA achieves high consensus network confidence (edges retained in โฅ80% of runs) and shows strong AUPR and F-score against gold-standard regulatory networks, with ฮป=0.1 providing optimal trade-offs (AUPR up to 0.177 against Chip+KD; F-score up to 0.081 against Chip+KD).
- Empirical evidence supports the major claim that joint integration of priors and TFA estimation outperforms or matches expression-only or prior-only models in both edge-level and module-level validation metrics.
- Inferred network structure demonstrates that latent TF activities cannot be captured by mRNA alone, highlighting the importance of activity-aware inference especially in cell-fate transitions.
- Extensive module analyses revealed that a substantial proportion (>51%) of inferred modules are significantly GO-enriched at optimal co-clustering thresholds, supporting the algorithm's ability to reconstruct functionally meaningful regulatory clusters.
Implications and Future Directions
Practically, MERLIN-SUITE establishes a reproducible, scalable pipeline for GRN inference applicable to both bulk and high-dimensional single-cell multi-omics datasets. The integration of priors and regularized TFA estimation increases both the accuracy and biological interpretability of outputs, making it directly applicable to systems biology, developmental biology, disease modeling, and precision medicine contexts.
Theoretically, the approach demonstrates the power of modular, probabilistically regularized frameworks for integrating heterogeneous data streams and resolves ambiguities caused by transcript-protein uncoupling. As high-resolution, multimodal profiling becomes standard, the MERLIN-SUITE paradigmโespecially activity-aware inferenceโwill likely become essential in the mapping of causal regulatory circuits.
Future extensions may include:
- Further automation of cell-type or condition-specific GRN inference to facilitate large-scale atlas construction.
- Integration of phosphoproteomic priors or direct post-translational regulatory evidence for inference of signaling-transcriptional crosstalk.
- Bayesian model averaging to better quantify network uncertainty.
- Downstream causal inference and perturbation prediction leveraging the inferred probabilistic dependency structures.
Conclusion
MERLIN-SUITE provides an advanced, modular computational framework for probabilistic reconstruction of gene regulatory networks leveraging expression data, regulatory priors, and latent TF activity estimation. It demonstrates robust performance for both bulk and single-cell settings and offers biologically validated, interpretable GRNs that capture both canonical and latent regulatory influences. The methodology sets a new standard for integrative, activity-aware network inference, with significant implications for both fundamental and translational genomics research.
Reference:
See "MERLIN-SUITE: Probabilistic modular GRN inference from multi-omics data integrating regulatory priors and transcription factor activity" (2607.01791).