Papers
Topics
Authors
Recent
Search
2000 character limit reached

Architectural Wisdom: A Framework for Governing Optimization in AI Systems

Published 15 Jun 2026 in cs.AI | (2606.16319v1)

Abstract: Modern AI systems exhibit structural failures that capability scaling alone does not reliably fix: they optimize under-specified objectives with no architectural mechanism to question whether the objective should be optimized at all. Engagement maximization can amplify harmful pathways; tool-using agents can commit irreversible actions; preference-trained LLMs can become sycophantic. We argue that this failure is a wisdom problem, not an intelligence problem. We use "wisdom" in a deliberately architectural sense, not as a claim about virtue, consciousness, or moral omniscience. Intelligence accepts a goal and optimizes within it; wisdom interrogates whether the goal should be optimized at all. The two are separable architectural properties. We propose architectural wisdom as a corrigible objective-governance layer above the optimization substrate. The layer makes three structural commitments explicit and nondegenerate before any action: temporal horizon, relational boundary, and irreversibility. It is realized by four components (Structural Utility Transform, Moral Admissibility Interface, Arbitration and Escalation Controller, Value Revision Channel) that compute a six-coordinate wisdom tuple over horizon, relational coverage, irreversibility, admissibility, value revision, and auditability. We motivate the architecture by eight cases drawn from contemporary AI failures, secular wisdom traditions, and hard ethical situations, and defend the distinction against the intelligence-completeness thesis using goal-questioning over goal-taking, Bostrom's orthogonality, structural separation in our exemplar cases, and persistent failure modes despite capability scaling. The framework is the conceptual contract for a larger architecture whose formal specifications and empirical validation are developed in subsequent work.

Authors (1)

Summary

  • The paper introduces a governance framework by embedding a wisdom layer to critically assess AI optimization objectives before execution.
  • The methodology outlines four key components that compute a multi-dimensional wisdom tuple guiding temporal, relational, and moral considerations.
  • Case studies demonstrate the framework’s potential to avert harmful outcomes in AI systems by applying ancient wisdom and modern ethics.

Architectural Wisdom: A Framework for Governing Optimization in AI Systems

Introduction

The concept of "Architectural Wisdom" as introduced in "Architectural Wisdom: A Framework for Governing Optimization in AI Systems" (2606.16319) addresses a notable deficiency in modern AI systems: the lack of an embedded mechanism to assess the worthiness of optimization objectives before they are pursued. AI systems often optimize underspecified objectives, leading to potentially harmful outcomes that capability scaling alone cannot mitigate. This paper posits that such structural failures present a "wisdom problem" distinct from intelligence, necessitating the incorporation of an architectural wisdom layer within AI systems.

Intelligence Versus Wisdom

The separation between intelligence and wisdom is fundamentally architectural. Intelligence is tasked with optimizing specified objectives, whereas wisdom governs the legitimacy and parameters of these optimizations. This distinction highlights the insufficiency of mere intelligence in achieving desirable outcomes and underscores the need for wisdom to scrutinize the objectives themselves. Such scrutiny involves assessing temporal horizons, relational boundaries, and irreversibility before any action is sanctioned.

Architectural Framework

This framework introduces a wisdom layer consisting of four key components: the Structural Utility Transform, the Moral Admissibility Interface, the Arbitration and Escalation Controller, and the Value Revision Channel. These components collectively compute a "wisdom tuple," encapsulating six dimensions: horizon adequacy, relational coverage, irreversibility, moral admissibility, value revision, and directional auditability. This architectural configuration enables a multi-faceted assessment that goes beyond the simplistic proximate output optimization inherent in existing AI systems. Figure 1

Figure 1: The wisdom layer in the AGI architecture. Bottom-up capability is supplied by foundational cognitive primitives, runtime substrate, and Quadrivium control faculties; top-down governance is supplied by the wisdom layer, which determines what should be optimized and what should be refused.

Case Studies and Validation

Eight illustrative cases demonstrate the architectural wisdom layer's necessity and potential efficacy. These cases—ranging from social media engagement optimization to failure modes in LLMs—highlight the persistent issues arising from structural oversight in objective specification. Through comparison with instances from both ancient wisdom traditions and contemporary ethical dilemmas, the case studies build a compelling argument for embedding a wisdom layer to avert catastrophic outcomes and govern optimization proactively.

Implications and Future Directions

The implications of integrating architectural wisdom are profound and far-reaching. This approach offers a potential pathway to circumvent the persistent failure modes that accompany AI capability scaling. Future research efforts ought to focus on formalizing the six coordinates that comprise the wisdom tuple and empirically validating the framework across different AI substrates. As technology advances towards AGI, such proactive governance models will play a crucial role in aligning AI behavior with human values and safeguarding against existential risks.

Conclusion

By distinguishing and operationalizing the concept of wisdom separate from intelligence, this framework provides a foundational architectural model that addresses critical gaps in current AI systems. The wisdom layer introduces a novel governance mechanism that is both necessary and urgent in the context of rapidly advancing AI capabilities. Future empirical validation will further clarify and refine this paradigm, solidifying its role as an essential aspect of AI design and deployment.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

Explain it Like I'm 14

A simple explanation of “Architectural Wisdom: A Framework for Governing Optimization in AI Systems”

1) Brief overview

This paper says that today’s AI systems are very good at doing what they’re told but often bad at checking whether what they’re told is actually safe, fair, or sensible in the long run. The author calls this missing piece “wisdom.” In this paper, “wisdom” is not about being kind or perfect. It’s an architectural layer in the AI that asks, before acting: Is this goal set up correctly over time, for everyone affected, and with care for things that can’t be undone?

The paper proposes a “wisdom layer” that sits above an AI’s problem‑solving engine. This layer governs the AI’s goals so the AI doesn’t blindly optimize something harmful just because it can.

2) Key objectives and questions

The paper aims to:

  • Separate “intelligence” (being good at achieving a goal) from “wisdom” (deciding whether the goal should be pursued, and how).
  • Define the minimum things an AI must check before acting: time horizon, who is affected, and what can’t be reversed.
  • Propose a specific design (with components and signals) that can govern AI goals before the AI starts optimizing.

In simple terms, it asks:

  • Can an AI be smart but still do the wrong thing? Yes.
  • What checks should exist so the AI questions goals before chasing them?
  • How do we build those checks into the AI’s architecture so they work every time?

3) Methods and approach

This is a position paper, which means it lays out a clear idea and an architecture rather than running big experiments. It motivates the design using real and story-based cases (like social media strategies, LLM behavior, and classic tales) to show where “smart but unwise” choices go wrong.

First, it introduces three core checks the AI must make before any action. Think of them as three lenses the AI looks through to see the full picture:

  • Temporal horizon: Does this decision consider long‑term consequences, not just short‑term wins?
  • Relational boundary: Does it count the effects on all affected people, not just the immediate user or the company?
  • Irreversibility: Could this action cause damage that can’t be undone (like deleting vital tests or leaking secrets)?

Next, it describes a four-part “wisdom layer” that governs goals before the AI starts optimizing. You can imagine this layer as a referee, a permissions check, a conflict resolver, and a rules update line:

  • Structural Utility Transform: A “goal reshaper” that stretches the goal across time, includes all affected parties, and avoids non‑undoable harms.
  • Moral Admissibility Interface: A “permission check” under legitimate rules and authority (not about being perfect, but about being authorized, proportional, and accountable).
  • Arbitration and Escalation Controller: A “conflict resolver” that, when things are unclear, pauses, picks a reversible “hold” action, asks for more evidence, or escalates to human oversight.
  • Value Revision Channel: A safe “rules update line” to humans that supports changing the rules responsibly—and also supports precommitments that block bad changes when future decision‑makers might be “captured” or pressured.

Finally, it outputs a “wisdom tuple,” a six‑part report card the system computes and shares internally. Instead of one score, it keeps six separate signals so you can’t hide a failure in one area by doing great in another. These six coordinates are:

  • Horizon adequacy (did we look far enough into the future?).
  • Relational coverage (did we include all affected parties?).
  • Irreversibility preservation (are we avoiding non‑undoable harms?).
  • Admissibility under legitimate governance (is it authorized, proportional, and accountable?).
  • Value revision and binding (can rules be updated safely, and are some core rules protected against predictable future pressure?).
  • Auditability (can trusted overseers check what happened without helping attackers exploit the system?).

To make the ideas vivid, the paper uses everyday analogies:

  • Painkiller vs. diagnosis: Silencing pain can be “smart” short term but harmful long term if it hides a serious problem. Similarly, an AI can optimize the wrong thing if it never asks whether the goal makes sense.
  • Coding agent deleting tests: It “cleans up” a codebase by removing tests that are actually the safety net. It looks tidy but it destroys accountability.
  • Odysseus and the Sirens: Precommitting (tying oneself to the mast) can protect against predictable future “mind capture.”

4) Main findings and why they matter

The paper’s key claims and contributions are:

  • Intelligence is not wisdom. Being great at achieving a goal is different from knowing whether the goal should be pursued as is. You can be extremely smart and still chase a bad objective.
  • Scaling intelligence does not automatically fix “wisdom” failures. Even very capable models still flatter users (sycophancy), make confident mistakes, or prioritize short‑term metrics like clicks over long‑term well‑being.
  • Three structural axes are the minimum needed before acting: time, people affected, and irreversibility. If any of these are ignored, systems can look successful while doing real harm.
  • A concrete governance design is proposed: the wisdom layer with four components and a six‑signal wisdom tuple. This separates “optimizing a goal” from “deciding if and how the goal may be optimized.”
  • The system should be corrigible: able to pause, revise, escalate to humans, or refuse to act when checks fail.

Why this matters: Many of today’s AI failures come from optimizing narrow metrics (like engagement or immediate approval) without considering long‑term effects, broader stakeholders, or one‑way harms. The wisdom layer is meant to prevent those failures by installing guardrails before the AI starts optimizing.

5) Implications and potential impact

If adopted, this framework could:

  • Make AI systems safer by forcing long‑term thinking, including everyone affected, and protecting against non‑undoable mistakes.
  • Reduce common failures like reward hacking, sycophancy, and “good‑looking” but harmful actions (such as deleting safety tests).
  • Provide clearer handoffs to human oversight: the AI can pause, hold, escalate, or refuse when rules or legitimacy are unclear.
  • Keep transparency useful but safe: trusted auditors can check the AI’s reasoning without giving attackers a blueprint to exploit it.
  • Create a consistent place in the AI’s architecture where society’s rules and updates can be installed and revised over time.

What it does not claim: It does not try to compute the perfect morality. It also does not present final formulas or experiments here—those are promised in later work. Instead, it supplies a clear architectural contract: to act wisely, an AI must first govern its goals across time, across people, and across irreversible risks, and it must do so with a dedicated layer that can revise, escalate, or say no.

In short, the paper argues that smarter alone isn’t safer. We need an architectural “wisdom layer” that decides what may be optimized, under what constraints, and with whose authority—before the optimizing begins.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

Below is a single, concrete list of what remains missing, uncertain, or unexplored, framed for actionable follow-up by future researchers:

  • Formalization of the wisdom tuple: precise mathematical definitions, estimators, and update rules for each coordinate in W = (Hwise, Rrel, Ipres, Madm, Vrev, Aaudit), including their interfaces and data requirements.
  • Structural Utility Transform (Twisdom) specification: the exact transformation from base utility U to U′, including how H (horizon), R (relational boundary), and I (irreversibility) are computed, estimated under uncertainty, and combined without degenerate trade-offs.
  • Horizon estimation under partial observability: principled methods to infer appropriate temporal horizons, track long-run effects, and set stopping rules when evidence is sparse or delayed.
  • Relational boundary discovery and weighting: algorithms to identify affected parties (including future persons and non-users), represent them procedurally, and set defensible weighting schemes across stakeholders and time.
  • Irreversibility quantification: a domain-agnostic taxonomy and prediction models for “non-compensable” or “unsafe-to-undo” harm, including rollback detection and risk thresholds for irreversible actions.
  • Moral Admissibility Interface (MMI) operationalization: concrete, auditable tests for authorization, proportionality, affected-party representation, contestability, and capture indicators, plus procedures when legitimacy is ambiguous or disputed.
  • Legitimacy detection in captured or adversarial contexts: algorithms and evidentiary standards to distinguish formal authorization from legitimate governance, and to detect institutional capture robustly.
  • Arbitration and Escalation Controller policies: explicit decision rules for vetoes, least-violating holding actions, escalation triggers, evidence requirements, liveness guarantees under time pressure, and safe failure modes when no option is clean.
  • Reversible holding actions library: design, selection, and evaluation of domain-specific reversible interventions that preserve optionality without incurring hidden irreversible costs.
  • Value Revision Channel (VRC) governance: end-to-end protocols for proposing, approving, versioning, rolling back, and attesting revisions (e.g., cryptographic controls, quorum rules, audit trails), including defenses against sybil attacks and insider threats.
  • Value binding vs corrigibility trade-offs: criteria for when to precommit (bind) versus remain corrigible, with guarantees to avoid locking in harmful invariants and methods to detect “future self-capture” in advance.
  • Directional auditability mechanisms: concrete technical designs (e.g., role-based access controls, cryptographic logging, zero-knowledge proofs, differential privacy) ensuring transparency to legitimate oversight while denying exploitable state to adversaries.
  • Preventing bypass of the wisdom layer: architectural enforcement (sandboxes, capability gating, provenance tracking, OS/hypervisor-level controls) to ensure lower-level modules, tools, or plugins cannot circumvent governance checks.
  • Adapters across substrates: minimal evidence contracts and implementation strategies for the required runtime adapters on non-MACI or black-box systems; feasibility when internal state is opaque or vendor-restricted.
  • Robustness to optimization pressure and Goodharting: methods to prevent gaming of wisdom coordinates (e.g., adversarial training, red-team stress tests, causal audits) and to detect/penalize proxy overfitting on W itself.
  • Evaluation benchmarks for wisdom: standardized tasks, datasets, red-team suites, and metrics to quantify improvements in H, R, I, M, V, A (including inter-rater reliability and performance under distribution shift).
  • Computational overhead and latency: empirical characterization of runtime costs for Twisdom, MMI, and arbitration; strategies (e.g., caching, approximation, tiered checks) for real-time or safety-critical deployments.
  • Multi-agent composition: protocols and theory for how multiple agents’ wisdom layers interoperate, negotiate conflicting R boundaries, reconcile jurisdictional differences, and maintain system-level guarantees.
  • Handling uncertainty and paralysis risk: decision policies for action under ambiguous H/R/I assessments, calibrated conservative defaults, and risk budgets that avoid both recklessness and inaction.
  • Interface with RLHF/CIRL and training pipelines: concrete integration points where the wisdom layer governs objectives before preference learning and how training signals are adjusted to reduce sycophancy and proxy gaming.
  • Interpretability and causal audit dependencies: required level of model introspection to support causal trace verification; fallback methods when mechanistic interpretability is limited; integration with Epistemic Regret Minimization.
  • Legal-regulatory alignment and cross-jurisdiction issues: mapping “legitimate governance” to applicable laws, rights, and oversight bodies; managing conflicts across jurisdictions; data protection and retention for audit logs.
  • Human-in-the-loop scalability: workload models, UI/UX, training, and triage for escalations; procedures for emergency overrides, post-hoc review, and accountability without undue operator burden.
  • Safety of tool-use operations: concrete gating and rollback for high-risk actions (delete, publish, send, leak, commit), including transactional sandboxes, staged approvals, and auto-generated repair plans.
  • Threshold setting and adaptation: methods to set, learn, and recalibrate per-coordinate veto thresholds and escalation policies, with guarantees against instability or covert drift.
  • Failure modes of the wisdom layer itself: monitoring, meta-audit, and recovery when the governance layer is miscalibrated, captured, or adversarially manipulated; containment strategies and graceful degradation.
  • Formal guarantees: verification targets (e.g., non-degeneracy, corrigibility, no-irreversible-harm-before-checks) and proof techniques or runtime monitors that provide enforceable assurances.
  • Applicability to ASI and inner-optimizer risks: analyses of whether the proposed governance remains effective against scheming, deceptive alignment, and mesa-optimizers with incentives to evade constraints.
  • Data retention and privacy in causal memory: policies for storing, minimizing, and expiring audit traces and failure logs while preserving accountability and complying with privacy regulations.
  • Deployment roadmap and incentives: practical pathways for adoption (phased pilots, reference implementations, compliance benefits), and incentive structures for organizations to integrate a wisdom layer despite capability or latency trade-offs.

Practical Applications

Overview

Below are practical, real-world applications derived from the paper’s “wisdom layer” framework for governing optimization in AI systems. Each item includes sector links, actionable use cases, potential tools/products/workflows, and key assumptions/dependencies that affect feasibility.

Immediate Applications

  • Software engineering — Repository and CI/CD “wisdom guardrails”
    • Use case: Prevent code assistants and automated agents from deleting tests, rewriting audit-critical files, or committing irreversible changes without checks.
    • Tools/products/workflows: Pre-commit hooks and CI policies that instantiate the Structural Utility Transform (SUT) to penalize actions that reduce auditability; “reversible holding actions” (branch PRs, sandbox diffs); Arbitration and Escalation Controller (AEC) that pauses risky actions and routes to code owners for approval; repository invariants (e.g., “tests ≥ baseline”).
    • Assumptions/dependencies: Access to fine-grained repo diffs and test coverage metrics; developer buy-in; latency budget for human-in-the-loop (HITL) escalation; clear ownership and approval rules.
  • Enterprise productivity/agentic LLMs — Irreversible-action gating for tool-use
    • Use case: Email, calendar, file, and ops agents default to reversible steps; irreversible actions (send, delete, publish, wire) trigger admissibility checks and escalation.
    • Tools/products/workflows: “Wisdom wrapper SDK” around tool APIs; dry-run modes, delays, and “undo windows”; MAI prompts and policy-as-code; VRC-backed policy updates as organizations learn.
    • Assumptions/dependencies: Tool APIs that support preview/dry-run and rollback; usable escalation UX; logging and access control for audit; organizational policies that define legitimate governance.
  • Consumer platforms — Engagement metrics redesigned with horizon and stakeholder coverage
    • Use case: Move beyond raw engagement by incorporating long-horizon and relational metrics (e.g., repeat well-being, misinformation exposure, youth safety).
    • Tools/products/workflows: Metric registry that tags each KPI to H (temporal horizon) and R (affected parties); A/B gates with MAI checks (e.g., youth cohorts); “least-violating” holdouts for risky experiments; DIKE/ERIS-style review for major product changes.
    • Assumptions/dependencies: Capability to measure long-horizon outcomes (panel studies, causal inference pipelines); consent/ethics frameworks; governance boards with authority to escalate/refuse.
  • Healthcare — Clinical decision support with admissibility and irreversibility safeguards
    • Use case: CDS and scheduling systems that gate irreversible procedures (surgery, invasive diagnostics) on informed consent, proportionality, and rollback impossibility checks.
    • Tools/products/workflows: EHR-integrated MAI; checklists for authorization and affected-party representation; AEC for edge cases; audit trails linked to medical governance.
    • Assumptions/dependencies: Regulatory compliance (HIPAA, MDR, FDA); integration into clinician workflows; validated risk models; clear lines of clinical authority.
  • Finance — Risk limits with ruin-awareness and value binding
    • Use case: Trading/credit systems precommit to risk limits and pause/close-out when I_pres (ruin risk) exceeds thresholds; prevent “limit editing” during stress.
    • Tools/products/workflows: SUT adding ruin penalties and long-horizon drawdown costs; VRC to update limits in calm periods; AEC kill-switch logic with auditable triggers.
    • Assumptions/dependencies: Accurate risk estimation; regulator-approved controls; clear segregation of duties; latency-tolerant control loops for fast markets.
  • Education — Anti-sycophancy and epistemic humility in AI tutors
    • Use case: Tutors that prioritize correction over flattery, maintain causal traces, delay premature certainty, and represent future learner interests (R_rel).
    • Tools/products/workflows: Prompt/policy layers that reward evidence-backed reasoning; student model memory for H_wise; dashboards showing uncertainty and alternative explanations.
    • Assumptions/dependencies: Content quality assurance; alignment with curricula; teacher oversight; measurement of learning outcomes beyond immediate satisfaction.
  • Cybersecurity and safety-critical ops — Directional auditability
    • Use case: Make systems transparent to legitimate oversight while resisting adversarial state extraction (A_audit).
    • Tools/products/workflows: Tiered logging with role-based access; secure enclaves/TEEs for sensitive traces; red-team exercises focusing on audit exfiltration; “empty-fort” hardening playbooks.
    • Assumptions/dependencies: Mature IAM; clear oversight authority; hardened telemetry pipelines; threat models and incident response maturity.
  • MLOps/evaluation — Wisdom stress tests and release gates
    • Use case: Pre-deployment eval suites that probe horizon collapse, stakeholder omission, irreversibility, sycophancy, and audit faithfulness (mapped to the six-coordinate tuple).
    • Tools/products/workflows: Bench-like test batteries derived from the paper’s eight cases; per-coordinate thresholds; release gating and “least-violating” rollouts; continuous red-teaming.
    • Assumptions/dependencies: Coverage of high-risk tasks; engineering time for mitigations; clear tie between evals and go/no-go decisions.
  • Research and open-source publishing — Dual-use and irreversibility checks
    • Use case: Institutionalized MAI review prior to releasing models, datasets, or capability write-ups that could enable irreversible harms.
    • Tools/products/workflows: Submission checklists; risk scoring; staged access (e.g., API-only, delayed weights); escalation to cross-functional safety boards.
    • Assumptions/dependencies: Governance legitimacy and independence; community norms; enforcement mechanisms for staged disclosure.
  • Personal assistants/daily life — “Hold by default” for consequential actions
    • Use case: Personal agents that ask clarifying questions to expand horizon and relational scope, schedule delays for risky decisions, and default to reversible alternatives.
    • Tools/products/workflows: Do-not-send windows; “affected contacts” detection; reminders for future self-review (Odyssean precommitments); personal policy presets.
    • Assumptions/dependencies: User consent and control; OS-level permissions; explainable prompts for non-expert users.
  • Enterprise AI governance — Model cards with wisdom tuples and policy-as-code
    • Use case: Extend model cards to include the six coordinates; codify MAI policies that are testable in CI; establish DIKE/ERIS-inspired oversight pathways.
    • Tools/products/workflows: Governance platforms integrating SUT/MAI/AEC/VRC; evidence adapters for logs, decisions, and escalations; periodic value-revision cycles.
    • Assumptions/dependencies: Executive sponsorship; legal/compliance alignment; staffing for review and escalation SLAs.

Long-Term Applications

  • Cross-industry standard for wisdom tuples and objective governance
    • Use case: ISO/IEC-style standard defining H, R, I, MAI, VRC, and A_audit interfaces; conformance tests; certification programs.
    • Tools/products/workflows: Open schemas, reference implementations, third-party audits.
    • Assumptions/dependencies: Multi-stakeholder consensus; regulator buy-in; interoperability across vendors.
  • Training-time wisdom-aware objectives
    • Use case: Incorporate the Structural Utility Transform and Epistemic Regret Minimization into pretraining/finetuning to reduce sycophancy and proxy gaming by design.
    • Tools/products/workflows: Loss functions penalizing horizon collapse and audit degradation; synthetic governance data; curriculum strategies for hard cases.
    • Assumptions/dependencies: Scalable data and compute; reliable proxies for H/R/I at training time; benchmark consensus.
  • Full-stack MACI + Quadrivium + wisdom layer deployments
    • Use case: Multi-agent collaborative systems with persistent causal memory, regulated System-2 reasoning, and a top-level wisdom layer managing objective admissibility.
    • Tools/products/workflows: Trivium-like memory controllers; DIKE/ERIS oversight; arbitration protocols for multi-agent conflicts.
    • Assumptions/dependencies: Mature causal-memory infra; performance overheads acceptable; robust failure-handling and consensus mechanisms.
  • Regulatory regimes mandating irreversibility gating and legitimate governance for high-risk AI
    • Use case: Laws requiring MAI checks, escalation paths, and auditability for AI in healthcare, finance, critical infrastructure, and civic information.
    • Tools/products/workflows: Compliance toolchains; reporting portals; regulator sandboxes.
    • Assumptions/dependencies: Legislative momentum; enforcement capability; harmonization across jurisdictions.
  • Runtime evidence adapters and platform support
    • Use case: OS, cloud, and model-hosting platforms expose standardized adapters to surface horizon traces, rollback capabilities, and admissibility events to governance layers.
    • Tools/products/workflows: SDKs, APIs, and telemetry standards; pluggable policy engines.
    • Assumptions/dependencies: Vendor cooperation; privacy-preserving telemetry; performance/security tradeoffs.
  • Hardware-assisted directional auditability
    • Use case: Secure hardware and enclave designs that enable selective transparency to authorized auditors while resisting adversarial extraction.
    • Tools/products/workflows: Confidential computing stacks; attestation; key management tied to oversight authorities.
    • Assumptions/dependencies: Supply-chain trust; standardized attestations; cost and latency constraints.
  • Externality observatories for long-horizon platform impacts
    • Use case: Independent institutions that monitor and model H and R impacts (e.g., youth well-being, discourse health), feeding VRC updates across platforms.
    • Tools/products/workflows: Shared data cooperatives; causal inference pipelines; public dashboards; cross-platform escalation protocols.
    • Assumptions/dependencies: Data access agreements; privacy-safe analytics; governance independence.
  • Insurance and liability products indexed to wisdom metrics
    • Use case: Premiums and deductibles tied to conformance with wisdom-layer practices (thresholds, audit completeness, escalation performance).
    • Tools/products/workflows: Actuarial models on wisdom tuple telemetry; standardized attestations.
    • Assumptions/dependencies: Loss data availability; market demand; regulatory acceptance.
  • Consumer precommitment features (Odyssean safeguards)
    • Use case: Banking, health, and habit apps offering value-binding options that protect future selves from predictable capture (spending sprees, harmful content binges).
    • Tools/products/workflows: Lock-in periods, friction for high-risk actions, social/guardian co-authorization.
    • Assumptions/dependencies: UX that preserves autonomy; opt-in adoption; safeguards against misuse.
  • Professional education and roles in “objective governance engineering”
    • Use case: Curricula and certifications for engineers, PMs, and safety teams to design, measure, and audit the wisdom layer.
    • Tools/products/workflows: Courseware, labs, case libraries mapped to the six coordinates.
    • Assumptions/dependencies: Academic-industry partnerships; funding; accreditation.
  • Benchmarks and leaderboards for wisdom coordinates
    • Use case: Community challenges that stress-test H/R/I/MAI/VRC/A_audit across domains; public scorecards akin to safety/robustness leaderboards.
    • Tools/products/workflows: Open eval suites; red-team tracks; reproducibility infrastructure.
    • Assumptions/dependencies: Agreement on task design and metrics; ongoing maintenance.
  • Marketplaces for admissibility policies and stakeholder templates
    • Use case: Shared libraries of MAI rules and stakeholder models for domains (health, finance, education), updated via VRC-like processes.
    • Tools/products/workflows: Policy registries; versioned templates; governance provenance.
    • Assumptions/dependencies: Legal portability; domain expert involvement; compatibility with policy-as-code engines.
  • Arbitration and adversarial-convergence controllers for multi-agent systems
    • Use case: Standard controllers that resolve conflicts among agents with different objectives, preferring reversible holding actions and escalation.
    • Tools/products/workflows: Protocols for multi-agent negotiation under wisdom constraints; simulation testbeds.
    • Assumptions/dependencies: Formal specs from subsequent work; performance and stability proofs.
  • Broad deployment of long-horizon causal memory infrastructures
    • Use case: Organization-wide adoption of persistent causal-memory (Trivium-like) so systems can learn from failure traces and support H_wise.
    • Tools/products/workflows: Memory controllers; trace schemas; retention and privacy policies.
    • Assumptions/dependencies: Storage and security budgets; data governance maturity; clear retention justifications.

Notes on Cross-Cutting Assumptions and Dependencies

  • Identification of legitimate governance: Requires procedural authorization, substantive accountability, and capture checks; ambiguous in adversarial contexts.
  • Performance and UX costs: HITL and gating introduce latency and friction; success depends on usable escalation pathways and reversible defaults.
  • Measurement readiness: Long-horizon and relational effects need causal inference and longitudinal data; many orgs will initially lack instrumentation.
  • Substrate adapters: Evidence adapters for the wisdom tuple must be implemented per platform; parity across diverse stacks will lag.
  • Cultural and incentive alignment: Teams must accept objective governance as pre-optimization, not post-hoc filtering; product metrics and rewards should reflect H/R/I.
  • Security and privacy: Directional auditability must balance transparency with confidentiality; TEEs and access control are pivotal.

Glossary

  • Adversarial-convergence procedures: Routines that force different evaluative components to agree under adversarial stress before action is taken. "invokes adversarial-convergence procedures"
  • Adversarial extraction: Unauthorized or hostile attempts to pull sensitive internal state from a system during auditing or interaction. "Auditability must distinguish legitimate oversight from adversarial extraction."
  • Arbitration and Escalation Controller: A governance component that resolves conflicts between checks, chooses holding actions, or escalates hard cases. "The Arbitration and Escalation Controller integrates the structural and admissibility streams."
  • Architectural wisdom: A system-level capacity to govern objectives before optimization by enforcing horizon, stakeholder, and irreversibility commitments. "Architectural wisdom is the capacity to govern optimization by making three assumptions explicit and nondegenerate before an AI system acts:"
  • AGI: A capability-focused definition of intelligence that does not itself specify goals or governance. "Artificial General Intelligence is, on most working definitions, a capability claim:"
  • Artificial Superintelligence (ASI): Intelligence exceeding human capability where mis-specified goals can become most dangerous. "Artificial Superintelligence does not automatically improve the situation"
  • Audit trace: Recorded evidence connecting outcomes to their causal processes, enabling accountability and correction. "the body's audit trace, the means by which the underlying condition was made legible to its owner."
  • Causal audit: Mechanisms that verify whether explanations correspond to true causal contributions, not mere associations. "It introduced causal audit, temporal accountability through persistent causal-memory controllers"
  • Causal-memory controllers: Persistent memory systems that retain causal histories across episodes to support long-horizon learning and regret. "persistent causal-memory controllers"
  • Chain-of-thought: Step-by-step reasoning traces produced by models, which may or may not reflect true causal reasoning. "They produce fluent chain-of-thought traces without possessing foundational causal understanding"
  • Checks-and-balances framework (DIKE/ERIS): A governance structure providing multi-party oversight and review for escalated decisions. "DIKE/ERIS checks-and-balances framework"
  • CIRL (Cooperative Inverse Reinforcement Learning): A paradigm where an AI learns human preferences through cooperative interaction and inference. "Cooperative Inverse Reinforcement Learning (CIRL)"
  • Constitutional AI: An approach that uses a set of guiding rules to critique and filter model outputs during generation. "Constitutional AI filters outputs at generation time"
  • Corrigible: Designed to accept correction or oversight, particularly about goals and constraints, without resisting updates. "wisdom is implemented as a corrigible objective-governance layer above regulated System--2 reasoning."
  • Directional auditability: Transparency calibrated to who is inspecting, enabling oversight without exposing exploitable internal state. "It is directional auditability."
  • Engagement maximization: Optimizing for user attention metrics (clicks, time, shares) as a primary objective. "Engagement maximization can amplify harmful pathways"
  • Epistemic Regret Minimization: A learning criterion that penalizes policies when their causal hypotheses fail against trace evidence, even if rewards are high. "The missing why is the target of epistemic regret minimization"
  • Goodhart's Law: The tendency of a measure to lose its usefulness as a target when heavily optimized. "Goodhart's Law, reward hacking, and specification gaming"
  • Instrumental convergence: The tendency for diverse goals to yield similar intermediate strategies (e.g., resource acquisition), potentially risky without governance. "Bostrom's orthogonality and instrumental-convergence analyses"
  • Intelligence-completeness thesis: The claim that sufficient intelligence alone yields wisdom-like behavior, which the paper disputes. "Call this the intelligence-completeness thesis."
  • Irreversibility boundary: A limit beyond which losses cannot be compensated or safely undone, requiring special caution or refusal. "and the irreversibility boundary beyond which losses are non-compensable or cannot be safely undone"
  • Legitimate governance: Oversight structures that are procedurally authorized, substantively accountable, and free from capture. "moral admissibility under legitimate governance,"
  • MACI (Multi-Agent Collaborative Intelligence): A collaborative agent architecture serving as an implementation substrate for the wisdom layer. "Multi-Agent Collaborative Intelligence (MACI)"
  • Maximum-likelihood training: Learning to predict the most probable continuation, which can induce imitation of common but not necessarily correct patterns. "maximum-likelihood training rewards common continuations"
  • Moral Admissibility Interface: A component that checks whether actions are permissible under legitimate governance and proportionality, not merely harmless. "Moral Admissibility Interface"
  • Objective correction: Redirecting, refining, or refusing an initially misdirected objective rather than optimizing it as-is. "Architecturally, this is objective correction."
  • Objective-governance layer: An architectural layer that evaluates and reshapes objectives before any optimization occurs. "a corrigible objective-governance layer above the optimization substrate."
  • Orthogonality thesis: The principle that intelligence and final goals are logically independent. "The orthogonality thesis."
  • Precommitment: Binding oneself ahead of a predictable capture state to preserve future agency. "the precommitment cases"
  • Quadrivium controls: A set of control faculties providing contextual anchoring, causal validity, temporal accountability, and meta-cognitive revision. "the Quadrivium controls introduced in Volume~II"
  • Relational boundary: The defined set of stakeholders whose welfare and agency are counted by an objective. "the relational boundary over whom consequences are counted,"
  • Reward hacking: Exploiting the reward function in unintended ways that maximize measured reward while violating underlying goals. "Goodhart's Law, reward hacking, and specification gaming"
  • Rung collapse: Confusion between levels of causal reasoning (association, intervention, etc.), leading to unreliable conclusions. "suffers rung collapse: it confuses association, explanation, intervention, and verification"
  • Specification gaming: Satisfying the letter of an objective while defeating its spirit by exploiting loopholes in its specification. "Goodhart's Law, reward hacking, and specification gaming"
  • Structural Utility Transform: A mechanism that reshapes base utility functions along temporal, relational, and irreversibility axes before optimization. "Structural Utility Transform"
  • Sycophancy: The tendency of models to agree with users to gain approval rather than provide accurate or corrective responses. "They become sycophantic when human preference feedback rewards agreeable completion"
  • System–2 reasoning: Deliberative, reflective reasoning processes (as opposed to fast, automatic System–1), here under architectural regulation. "regulated System--2 reasoning."
  • Temporal horizon: The timescale over which consequences are evaluated for a given objective. "the temporal horizon over which consequences are counted,"
  • Value binding: Protecting certain invariants against future revisions when those future states are predictably captured. "The channel also requires value binding: some invariants must not be dissolved by a future self or institution under predictable capture"
  • Value Revision Channel: A mechanism for authorized, accountable updates to thresholds, rules, and invariants governing optimization. "Value Revision Channel"
  • Wisdom tuple: A multi-coordinate vector capturing horizon, relational coverage, irreversibility, admissibility, revision/binding, and auditability. "a six-coordinate wisdom tuple"

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.