- The paper demonstrates a novel VR platform integrating GenAI to shift heritage narratives from passive remembrance to active, embodied co-creation.
- It employs a modular, spatial negotiation method that resolves design conflicts while enhancing cultural specificity using a Shanghai-tuned LoRA model.
- Findings underline the importance of participatory design in preserving unofficial memories and addressing cultural homogenization in GenAI outputs.
Immersive Co-Creation of Cultural Heritage Narratives in Collaborative VR with Generative AI
Introduction and Motivation
The study "From remembering to shaping: Narrating Shared Experiences by Co-Designing Cultural Heritage Artifacts in Collaborative VR" (2604.15058) systematically investigates the intersection of generative AI (GenAI), collaborative narrative construction, and immersive virtual reality (VR) for expressing and negotiating collective cultural heritage (CH). Recognizing heritage as an inherently communal and narrative process, this work responds to gaps in conventional heritage representation approaches—particularly the marginalization of unofficial memories, affective ties, and local stories by the dominant Authorized Heritage Discourse (AHD) [Smith2006usesofheritage]. Instead of privileging static physical artifacts or expert-curated narratives (cf. [giaccardi2012heritage, Silbernman2014changingvisionsofheritagevalue]), the study seeks to empower communities with participatory mechanisms for collaborative authorship and spatialized negotiation of heritage memory.
The research leverages generative and immersive technologies to instantiate intangible, spatially-grounded collective memories as manipulable 3D objects in shared VR. This enables a shift from passive consumption of heritage to dynamic, co-creative shaping of cultural meaning—addressing the need for socio-technical systems where GenAI operates not merely as a content producer but as an active mediator and enhancer of interpersonal and communal negotiation within digital heritage contexts. The selection of Shanghai’s Haipai culture as the target domain introduces substantial heterogeneity in participants’ place memories, supporting the empirical investigation of negotiation, conflict resolution, and co-constructive practice around CH expression.
System and Study Design
The workflow integrates state-of-the-art text-to-image (FLUX.1-dev) and image-to-3D (Hunyuan3D) models, deployed through the GenSH platform, with real-time immersive spatial interaction on Meta Quest 3 headsets (Figure 1). This design affords two-person, synchronous co-creation in a simplified 3D base map context (central Shanghai) stripped of stylistic guidance, optimizing for unbiased participant expression.
Figure 1: Study setup—two participants with VR headsets interactively co-create and manipulate GenAI-generated 3D models in a shared virtual space, mediated by a facilitating researcher.
Participants can prompt GenAI via naturalistic verbal descriptions (input by the researcher in a Wizard of Oz protocol), iteratively refining both prompt specificity and spatial arrangements in VR. To address semantic gaps and domain bias in foundational GenAI models, a custom Shanghai-style LoRA 3D architecture model can be optionally engaged for more fine-grained and culturally resonant artifact generation.
Data collection triangulates session recordings, system logs, audio transcripts, and post-task semi-structured interviews, with qualitative thematic analysis isolating mechanisms and patterns underlying collaborative memory construction, spatial negotiation, and adaptive responses to model limitations.
Patterns of Co-Constructive Practice
Visualizing Collective Memory in 3D
Participants jointly invoke and elaborate spatially structured memories through iterative generation and manipulation of 3D models. One participant’s contribution—such as a breakfast shop—routinely catalyzes associative recall in the other, leading to increasingly elaborate scenographic assemblies. Direct manipulation (e.g., physically dragging models to simulate “congestion”) is central, enabling fine-grained, embodied adjustment of atmospheres that are typically inarticulable verbally.
Figure 2: Initial object generation prompts associative recall; embodied spatial assembly and manipulation yield richer, collaboratively constructed scenes.
LoRA fine-tuning reduces the prompt burden and increases alignment to local architectural expression, demonstrably streamlining cognitive and interactional demands (Figure 3).
Figure 3: LoRA model enables generation of locally-authentic architecture from minimal prompts, fostering more efficient creative cycles.
Nonverbal Spatial Negotiation
A salient innovation is the frequent use of spatial operations (translation, scaling, modular assembly) as nonverbal negotiation acts—participants resolve conflicts, align perspectives, and suggest aesthetic trade-offs without explicit dialogue. For example, repositioning a contested signage resolves a stylistic disagreement silently, and physically widening a virtual street becomes both a site of reconciliation and a catalyst for subsequent creative divergence (Figure 4).
Figure 4: Action as dialogue—direct spatial manipulation instantly resolves partner’s identified design tensions.
Multi-Perspective Embodiment and Evaluation
Participants strategically switch between global (god’s-eye) and situated (first-person) views for division of labor, macro-level planning, and embodied evaluation of qualitative/aesthetic atmosphere (Figure 5). This supports both top-down and bottom-up negotiation and enables the VR environment to serve as both a planning canvas and an experiential testbed for affective resonance.
Figure 5: Strategic perspective switching—first-person immersion used for collective validation of atmosphere and authenticity.
Modular Construction and Creative Appropriation
A modular, “Lego-like” strategy emerges organically—participants generate architectural components (walls, roofs, details) and assemble/refine them into composite, memory-accurate complexes (Figure 6).
Figure 6: Emergent modular strategy—component-based architectural assembly enables adaptive memory reconstruction and fine-grained co-creation.
This component-level manipulation is crucial for reconciling idiosyncratic, fragmentary memory and for adapting or repurposing unsatisfactory GenAI outputs.
Negotiation of Perspective and Representational Challenges
Establishing Shared Referents and Knowledge Transfer
AI-generated 3D models act as high-bandwidth boundary objects, supporting the transfer of complex, situated knowledge—e.g., using an archetypal “French Concession” building to bridge knowledge gaps between participants with divergent historical or spatial familiarity (Figure 7).
Figure 7: 3D models operate as cognitive bridges, providing tangible, experiential referents during abstract knowledge transfer/negotiation.
Collaborative Conflict Resolution
Friction over aesthetic direction or cultural style (e.g., TV drama vs. film reference) is negotiated through embodied, comparative walkthroughs and spatialized co-design, with outcomes often guided by consensus validation through VR-based affective trial and adjustment.
When the AI-generated output is “out-of-spec” or unsatisfactory, participants consistently respond with adaptive repurposing or iterative refinement—transforming generation “failures” into newly valued design constraints or hybridizations. For instance, an unrequested interior space or stylized neon signage becomes a catalyst for new narratives and aesthetic direction (Figures 11, 14).
Figure 8: Unexpected AI outputs not treated as errors—participants incorporate and elaborate emergent spatial affordances as design opportunities.
Figure 9: Failed building generation yields a stylized neon sign, which is then adopted as a central motif for subsequent creative work.
AI also acts as an impartial mediator in situations of deep style conflict, producing hybridized or compromise outputs that are accepted as formally satisfying (Figure 10).
Figure 10: AI-generated hybrid architecture resolves participant impasse, blending modern and vintage styles into an accepted compromise.
Implications and Theoretical Synthesis
Embodied Construction and Situated Memory
The study demonstrates that immersive, manipulable 3D spaces enable not merely visualization but the materialization of otherwise ineffable, embodied aspects of memory—“atmosphere”, density, intimacy—beyond the reach of 2D visual artifacts [Brady2017memoryconstructive, Paul2001WheretheActionis]. Spatial operations supplement or supersede verbal accounts, allowing for tacit negotiation and enactment of place attachment.
GenAI-Driven Homogenization and the Limits of Universal Models
Empirical evidence reveals persistent domain bias and cultural stereotyping when using generic GenAI models—the default outputs overfit to generic “Chinatown” or Westernized architectural motifs (Figure 11). The LoRA-finetuned model mitigates these gaps but highlights the insufficiency of top-down model adaptation without participatory, community-grounded dataset curation [Yuan2025huayao, Qadri2025nonwesternartworlds]. This supports mounting concerns over cultural homogenization and the erasure of local nuance in LLMs and generative media [Daryani2026homogenizingengine, Agarwal2025homogenize, Zhang2024partialitymisconception].
Figure 11: Direct comparison confirms substantial gains in authenticity and specificity using domain-tuned LoRA models versus generic baselines.
The workflows show that GenAI outputs not only externalize recollections but actively participate in reconstructing memory—sometimes shifting the trajectory of narrative co-construction away from initial participant intent. The mutability of memory and the porous entanglement of human and machinic agency are rendered explicit, foregrounding the need for mechanisms of participant reflection, transparency, and control in heritage-oriented systems [Zhou2026TellMeWhatIMissed, Jin2022fluidheritage].
Participatory and Component-Based Design for Digital Heritage
System-level implications include the clear value of modular generation and assembly tools, multi-perspective navigation, and visible manipulation cues for collaborative negotiation. The necessity for a library of fine-tuned, locally-curated models is echoed, with calls for participatory processes that grant community stakeholders agency in model training and artifact curation.
Limitations and Future Prospects
Primary constraints include the non-ecological nature of lab-based VR versus situated AR experiences; the limited cultural/intellectual scope of a predominantly young, digitally native participant pool; and the need to scale workflows beyond dyads to more complex group configurations where the sociology of negotiation changes (e.g., coalition, leadership, polyvocality) [tang2010communicationcollaboration]. Additionally, even LoRA models narrowly target specific styles, pointing to the ongoing scalability challenge of representing non-canonical and hybridized local narratives within GenAI.
Possible future research directions outlined include:
- In-situ AR heritage workflows: Overlaying GenAI-generated models in real heritage settings for multi-sensory, ecologically valid co-creation.
- Broader, intergenerational co-design studies: Capturing rich transmission and contestation across age, class, and demographic divides [freeman2020use, Wang2024intergenrationaliICH].
- Participatory dataset/model development: Collaboratively curating and fine-tuning local GenAI models in direct partnership with community stakeholders.
- Group dynamics and negotiation at scale: Understanding mediation, coalition-building, and social hierarchy in larger collaborative contexts.
Conclusion
This work advances the theoretical and practical discourse on digital cultural heritage by demonstrating that GenAI-mediated, immersive co-design in VR enables not just the retrieval, but the active shaping and negotiation of collective memory. Spatial interaction becomes as central as linguistic narration for making heritage visible, tangible, and actionable; GenAI operates as both creative surrogate and impromptu mediator, expanding and sometimes constraining the idioms of place-making. The results underscore the need for sociotechnical systems that privilege local specificity, participatory curation, and transparent mediation—and for critical, ongoing examination of how algorithmic actors participate in the shaping of cultural memory and the politics of representation.
Figure 12: Participant first-person view and results of collaborative, agentic construction of heritage-inspired VR streets—demonstrating the compositional diversity and memory-informed design emergent in the workflow.