IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2026 1
Global Attractor Theory for the MDEI Stochastic PDE Framework: Mathematical Foundations, Finite-Dimensional Stability, and Implications for Artificial General Intelligence Tiago Aguioncio Vieira Universidade de São Paulo (USP), Brasil · Valer Inteligência Artificial MSC 2020: 37L30, 93D30, 68T07, 35R60 Keywords: global attractor, SPDE, dissipative operator, Lyapunov functional, fractal dimension, MDEI, AGI, catastrophic forgetting, stability region
Abstract We develop a rigorous mathematical framework for the Model of Dynamic-Emergent Intelligence (MDEI) by embedding its evolution equations into an infinite-dimensional stochastic partial differential equation (SPDE) on the Hilbert space H = L2 (Ω; R3 ). Under four structural hypotheses— sectorial dissipation (H1), a coercive Lyapunov functional (H2), a teleological pressure operator (H3), and multiplicative Q-Wiener noise (H4)—we prove: (i) existence of a compact invariant global attractor A ⊂ H; (ii) finite fractal dimension dimF (A ) ≤ C(ν, α, β , d0 , d1 ) < ∞; (iii) explicit bounds on the attractor diameter diam(A ), stability neighbourhood Nε (A ), and relaxation time τ = 1/α. We prove that attention-based architectures (O(N 2 )) do not satisfy H1 and therefore admit no compact absorbing set—a dynamical-systems explanation for catastrophic forgetting and the fine-tuning treadmill. AGI is redefined as a geometric property: the possession of a global attractor of bounded fractal dimension, stable under stochastic perturbation. This extends the MDEI algebraic formalization [1] and the Rudolph–Vieira coherence framework [2].
I. I NTRODUCTION (i) Lifts MDEI evolution to an SPDE on H = L2 (Ω; R3 ) (Sec- Context. The long-run behaviour of a dynamical system is cap- tion 3); tured by its global attractor: the compact, invariant set toward (ii) Proves existence, compactness, and finite dimensionality of which every orbit is asymptotically drawn. In finite dimensions the global attractor A under hypotheses H1–H4 (Section 4); this picture was established by Lorenz (1963) and Ruelle–Takens (iii) Derives explicit formulae for diam(A ), relaxation time τ = (1971). In the infinite-dimensional setting it was extended rigor- 1/α, and stability region Nε (A ) (Section 5); ously by Babin–Vishik, Hale [4], and Temam [3] for dissipative (iv) Proves incompatibility of attention dynamics with H1 and PDEs, and later for stochastic PDEs by Da Prato–Zabczyk [6] draws the connection to catastrophic forgetting (Section 7); and Robinson [5]. (v) Reformulates AGI as a geometric property of attractors, com- paring with prior formal definitions (Section 9); A. The Discrete-Token Bottleneck (vi) Analyses the internal anatomy of A : core manifold M , Contemporary large language models (LLMs) encode state as boundary layer, and inertial manifold M (Section 6); discrete token sequences. Their self-attention layers compute (vii) Provides spectral theory and parameter design rules for pre- pairwise affinities with time and memory complexity O(N 2 ), scribing diam(A ) and τ by design (Section 11); where N is the sequence length [10, 11]. This quadratic scaling (viii) Analyses the ISDM–TSM affective–semantic coupling via is not merely an engineering issue: it reflects the absence of any the joint Lyapunov functional Vjoint (Section 12); dissipative structure in the attention map. As formalised in The- (ix) Develops a Lyapunov-based feedback control with closed- orem 7, attention dynamics do not admit a sectorial operator A loop stability proof and comparison to neural ODEs (Sec- with Re(σ (A)) ≥ α > 0; consequently no compact absorbing set tion 13); can exist. The empirical counterpart is catastrophic forgetting [9] (x) Proves ergodicity and exponential mixing of the invariant and the necessity of continuous fine-tuning. measure µ ⋆ on A (Section 15); B. The MDEI Research Programme (xi) Provides a fully worked scalar example (MDEI-OU model) The Model of Dynamic-Emergent Intelligence (MDEI) was in- with all quantities computed in closed form (Section 18); troduced by Vieira [1] as a principled departure from symbolic, (xii) Connects MDEI to Friston’s free-energy principle [7] and token-based representations. In MDEI each internal state is Haken’s synergetics [8] (Section 16). encoded as an adaptive three-dimensional vector Φ ∈ R3 gov- D. Mathematical Classification and Prerequisites erned by vector-algebraic and differential calculus, embedding MSC 2020 classification: 37L30 (attractors), 93D30 (Lyapunov cognition in a continuous manifold. The companion work [2] and storage functions), 68T07 (neural networks), 35R60 (stochas- formalised the coupling between the Internal State Dynamics tic PDEs), 47D06 (one-parameter semigroups), 60H15 (stochas- Model (ISDM) and the Teleological Semantics Model (TSM): tic PDEs). affective regulation via energetic constraints interacts bidirec- The paper is self-contained in the following sense. Back- tionally with semantic dynamics via a teleological operator T , ground on SPDE mild solutions is drawn from Da Prato– and a joint Lyapunov functional V ensures global exponential Zabczyk [6]. Attractor theory follows Temam [3] and Hale [4]. stability. Stochastic attractor theory follows Crauel–Flandoli [18] and C. Contributions of This Work Langa–Robinson [12]. Readers familiar with [1, 2] will recog- This paper accomplishes the following: nise the ISDM–TSM operators; their properties are recalled in IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2026 2
Section 3 and Appendix C. The proofs in Sections 4–19 are space L20 = L20 (Q1/2 H, H) [6]. For a twice Fréchet-differentiable complete; Appendix A provides full calculation details for the F : H → R and mild solution Φ(t), the Itô–Da Prato formula Itô–Lyapunov argument. reads: E. Notation dF(Φ(t)) = ⟨∇F(Φ), dΦ⟩H We write ∥·∥ for the H-norm, ⟨·, ·⟩H for the H-inner product, and + 21 Tr Σ∗ ∇2 F(Φ)Σ Q dt. ∥·∥V for the V -norm. L (H) denotes the space of bounded linear (2) operators on H. E[·] denotes expectation under P. The symbol ≲ means ≤ C(·) for a constant C depending only on the subscripted C. Analytic Semigroups and Sectorial Operators parameters. All other notation is collected in Appendix D. A closed densely defined −A : D(A) → H generates an analytic semigroup e−At if A is sectorial with σ (A) ⊂ {λ : Re(λ ) ≥ α > F. Note on Figures: Mathematical Simulations Only 0}. The semigroup then satisfies Important methodological declaration. All six fig- ures in this article are generated exclusively from e−At L (H) ≤ Me−αt , t ≥ 0, M ≥ 1, α > 0. (3) mathematical formulas and exact numerical solutions of explicitly stated differential equations. No figure Moreover, e−At is compact for t > 0 if A has compact resol- contains empirical data, experimental measurements, vent [3]. benchmark results, or outputs from any real language model or neural network system. Specifically: III. MDEI S TOCHASTIC PDE F ORMULATION • All analytical formulas (Gronwall bounds, eigen- A. The Governing Equation value spectra, dimension bounds, relaxation The MDEI internal state Φ(t) ∈ H evolves as curves, invariant densities) are plotted by eval- uating the closed-form mathematical expressions dΦ + AΦ + ∇Φ V (Φ) − P(Φ) dt = Σ(Φ) dW, Φ(0) = Φ0 , derived in the corresponding theorems and corol- laries. (4) where each term satisfies one of four structural hypotheses de- • All trajectory plots (Figs. 2, 6a) are exact Euler– fined below. Fig. 1 summarises the system architecture. Maruyama integrations of the explicitly stated Hypothesis 3.1 (Sectorial Dissipative Operator). A = −ν∆ + SDEs, with fixed random seeds declared in the cap- αI, D(A) = H 2 (Ω) ∩ H01 (Ω), ν > 0, α > 0. The operator −A tions, ensuring full reproducibility. The SDEs used generates an analytic semigroup satisfying (3). are simplified linear models (Ornstein–Uhlenbeck The Laplacian −ν∆ suppresses high-frequency instabilities; the and skew-symmetric rotation) chosen to illustrate damping αI provides uniform global contraction pulling every the theoretical properties proved in Sections 4 and orbit into a bounded region. Together they make A invertible 7, not to approximate any real system. with compact resolvent, essential for the compact decomposition • Figure 4 is a conceptual block diagram with no in Section 4. numerical content. Hypothesis 3.2 (Lyapunov Functional). • No figure should be interpreted as representing Z experimental results from language models, neural V (Φ) = 21 ∥∇Φ∥2L2 + F(Φ) dx, (5) networks, or cognitive systems. The figures serve Ω solely to provide visual intuition for the mathemat- ical structures defined in the text. with F ∈ C2 (R3 ; R), F(Φ) ≥ c1 |Φ|4 − c2 , c1 > 0. There exist β > 0, K ≥ 0 such that II. F UNCTIONAL F RAMEWORK AND P RELIMINARIES A. Hilbert Space Setting ⟨∇Φ V (Φ), AΦ + ∇Φ V (Φ) − P(Φ)⟩H ≥ β V (Φ) − K. (6) Let Ω ⊂ Rn be a bounded domain with Lipschitz boundary. The functional V acts as the energy of the cognitive state. The Define gradient flow −∇Φ V is the restoring force driving toward low- H = L2 (Ω; R3 ), V = H01 (Ω; R3 ), (1) energy configurations; condition (6) ensures genuine energy with standard inner products. The embedding V ,→ H is dense dissipation rather than mere redistribution. This functional was and compact by the Sobolev–Rellich–Kondrachov theorem [5]. introduced in the ISDM formalism [2] for affective regulation We denote ⟨·, ·⟩H for the H-inner product and ∥·∥ for the H- and here extends to the full SPDE setting. norm. Hypothesis 3.3 (Teleological Pressure). B. Q-Wiener Processes and Itô Formula P(Φ) = ∇ · D(Φ)∇Φ + T (Φ), (7) Let (ΩP , F , (Ft ), P) be a filtered probability space satisfying the usual conditions. Let Q ∈ L (H) be symmetric, positive, and where 0 < d0 I ≤ D(Φ) ≤ d1 I uniformly, and T : H → H is a trace-class. A Q-Wiener process W (t) in H is a process with bounded teleological operator satisfying ∥T (Φ)∥H ≤ CT (1 + covariance operator Q, taking values in the Hilbert–Schmidt ∥Φ∥H ). IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2026 3
Figure 4: MDEI-SPDE System Architecture Hypotheses H1 H4 Global Attractor (a) Lyapunov Functional ( ) = 12 | |2 + F( ) MDEI-SPDE: d + (A + P) dt = dW (analytical; white curves: gradient-flow orbits) (b) Analytical Gronwall Bound: [ ] e t ( 0) + R0 , r = 0.58 ( 0) = 4.0 Lyapunov ( ), r = 0.82 52.5 4 ( 0) = 2.5 Initial state ( ) 2 ( 0) = 1.5
[ ( (t))] upper bound (Gronwall) 0 H H1 ,\ coercive = 1/ H2 45.0 ( 0) = 0.3 (H2) R0 = K/ = 0.33 Dissipative op. Global attractor Absorbing set B0 A= + I H 1 37.5 3
( ) (analytical) Re( (A)) > 0 dimF ( ) n (H1) compact, invariant 30.0 0
2 H4 H3 2 22.5 Noise Teleological ( ) dW P( ) = (D ) + T 1 (H4: Tr(Q) < ) d0I D d1I 15.0 (H3) 1 2 7.5
0.0 0 2 1 0 1 2 0 1 2 3 4 5 6 7 1 t
Figure 1. [Mathematical diagram — no numerical data.] MDEI-SPDE Figure 2. [Mathematical simulation — analytical formulas only. No empiri- system architecture. Hypotheses H1–H4 interact to produce a semigroup {S(t)} cal data.] Left: Level sets of the Lyapunov functional V (Φ) = 12 |∇Φ|2 + F(Φ), admitting the global attractor A with dimF (A ) < ∞. All operator labels (A, V , F(Φ) = c1 |Φ|4 − c2 |Φ|2 , with c1 = 0.25, c2 = 0.9 (parameters chosen for il- P, Σ) refer to the mathematical objects defined in Section 3. lustration; not fitted to data). White curves: exact gradient-flow orbits of Φ̇ = −(A + I)Φ (Euler method, p dt = 0.02, 300 steps). Dashed circle: ana- lytical attractor boundary rA = K/(β λ1 ) from Corollary 4. Dotted circle: p The operator P encodes the goal-directed aspect of MDEI dy- stability neighbourhood rε = 2K/(β λ1 ), Definition 1. Right: Exact Gronwall namics [1, 2]. It replaces the O(N 2 ) attention mechanism by a bound E[V (Φ(t))] ≤ e−β̃t V (Φ0 )+R0 (Lemma 1; β̃ = 1.2, K̃ = 0.4, R0 = K̃/β̃ ). bounded nonlinear diffusion-plus-drift that does not break com- Curves are the closed-form formula evaluated at four declared initial values; no stochastic simulation. pactness. Hypothesis 3.4 (Stochastic Noise). W (t) is a Q-Wiener process with Tr(Q) < ∞, and By Gronwall’s inequality,
∥Σ(Φ)∥2L0 ≤ C1 +C2 V (Φ), C1 ,C2 ≥ 0, 2C2 < 2β . (8) K̃ 2 E[V (Φ(t))] ≤ e−β̃t V (Φ0 ) + , (13) β̃ The sub-criticality condition C2 < β ensures noise cannot over- come dissipation. so lim supt→∞ E[V ] ≤ R0 , confirming absorption. B. Mild Solution and Well-Posedness Remark 1. R0 = K̃/β̃ depends only on the operator parameters Under H1–H4, equation (4) admits a unique mild solution α, β , ν, d0 , d1 , not on data dimensionality or token count. This is the core scaling advantage over parameter-count methods. Z t Φ(t) = e−At Φ0 − e−A(t−s) ∇Φ V − P (Φ(s)) ds B. Semigroup Decomposition 0 Z t The semigroup decomposes as S(t) = S1 (t) + S2 (t), where: + e−A(t−s) Σ(Φ(s)) dW (s), (9) S1 (t)Φ0 = e−At Φ0 satisfies ∥S1 (t)∥L (H) ≤ Me−αt ; and S2 (t) 0 maps B0 into a compact subset of V = H01 (Ω; R3 ) for each defining a Markov semigroup {S(t)}t≥0 on H. Existence and t > 0, by the regularising effect of A and the compact embedding uniqueness follow from the Da Prato–Zabczyk framework [6]. V ,→ H. C. Existence and Invariance IV. G LOBAL ATTRACTOR : E XISTENCE , C OMPACTNESS , D IMENSION Theorem 2 (Global Attractor). Under H1–H4, the semigroup {S(t)}t≥0 possesses a global attractor A. Absorbing Set via Itô–Lyapunov Lemma 1 (Absorbing Set). Under H1–H4, define β̃ = β − A = ω(B0 ) = \[ S(t)B0 (14) C2 /2 > 0 and K̃ = K +C1 /2. There exists R0 = K̃/β̃ such that s≥0 t≥s
B0 = {Φ ∈ H : V (Φ) ≤ 2R0 } (10) satisfying: (i) A is compact in H; (ii) A is invariant, S(t)A = A , ∀t ≥ 0; (iii) A attracts uniformly every bounded B ⊂ H: is absorbing: for every bounded B ⊂ H there exists T (B) < ∞ distH (S(t)B, A ) → 0 as t → ∞. with S(t)B ⊂ B0 for all t ≥ T (B). Proof sketch. By Lemma 1, B0 is bounded and absorbing. The Proof. Apply the Itô formula (2) to F = V : semigroup decomposition ensures S(t)B0 is precompact in H for t > 0. The omega-limit A = ω(B0 ) is therefore compact (Kura- dV (Φ) = − ⟨∇Φ V , AΦ + ∇Φ V − P⟩H dt towski). Invariance and uniform attraction follow by standard arguments of Hale [4] and Temam [3]. + ⟨∇Φ V , Σ dW ⟩H + 21 Tr(ΣQΣ∗ ∇2 V ) dt. (11) D. Fractal Dimension Bound Taking expectations and applying (6) and (8): Theorem 3 (Finite Fractal Dimension). Under H1–H4, d E[V (Φ(t))] ≤ −β̃ E[V (Φ(t))] + K̃. (12) dimF (A ) ≤ C(ν, α, β , d0 , d1 ) < ∞. (15) dt IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2026 4
Specifically, with Lyapunov exponents {µi } ordered decreas- where λ1 > 0 is the first Dirichlet eigenvalue p of −∆ (Poincaré ingly: constant). In particular diam(A ) ≲ 1/ αβ . n n o This bound is the intrinsic error of the system: any two states dimF (A ) ≤ n⋆ := min n ∈ N : ∑ µi < 0 . (16) in A differ by at most diam(A ), regardless of input data. In- i=1 creasing α (damping) or β (Lyapunov decay rate) shrinks the Detailed proof. Step 1: Linearisation. Fix Φ⋆ (t) ∈ A . The attractor. linearisation of (4) around Φ⋆ is the variational equation B. Stability Region and Noise Robustness dξ + Aξ + ∇2Φ V (Φ⋆ )ξ − DP(Φ⋆ )ξ dt = DΣ(Φ⋆ )ξ dW, (17) Definition 1 (Stability Region). for a tangent vector ξ ∈ H. Let Ξ(t) = (dΦ S(t))Φ⋆ (0) be the q derivative semigroup acting on n-tuples of initial tangent vectors Nε (A ) = {Φ ∈ H : dist(Φ, A ) < ε}, ε= K̃/(λ1 β̃ ). (21) (ξ1 , . . . , ξn ) ∈ H n . Step 2: Volume element. The n-volume of the image of an Proposition 5 (Noise Robustness). For any perturbation η ∈ H n-dimensional ball under Ξ(t) satisfies with ∥η∥H < δ , the perturbed orbit Φη (t) satisfies Φη (t) ∈ Nε+Cδ (A ) for all t ≥ 0, where C > 0 is independent of δ . d log ωn (t) = Tr A + ∇2Φ V (Φ⋆ ) − DP(Φ⋆ ) En (t) , (18) dt Proof. Let Φ(t) and Φη (t) solve (4) with initial conditions Φ0 where En (t) is the n-dimensional subspace spanned by the most and Φ0 + η, respectively. Define ζ (t) = Φη (t) − Φ(t). Then ζ rapidly growing tangent vectors (Lyapunov exponents). satisfies Step 3: Trace estimate. Since A has eigenvalues µk = νλk + α ≥ α, the trace of A restricted to any n-dimensional subspace dζ + Aζ + [∇Φ V (Φη ) − ∇Φ V (Φ)] satisfies Tr(A|En ) ≥ nα. The terms ∇2Φ V and DP contribute at − [P(Φη ) − P(Φ)] dt = [Σ(Φη ) − Σ(Φ)] dW. (22) most CV ,P := ∇2 V L∞ (A ) +∥DP∥L∞ (A ) to the trace. Therefore Applying the Itô formula to 12 ∥ζ ∥2 and using the dissipativity of d dt log ωn (t) ≤ −nα +CV ,P . (19) A (⟨Aζ , ζ ⟩ ≥ α ∥ζ ∥2 ) together with the local Lipschitz property of ∇Φ V and P: The sum of the n largest Lyapunov exponents satisfies ∑ni=1 µi ≤ −nα +CV ,P . d 1 E 2 ∥ζ ∥2 ≤ −(α − LV − LP )E[∥ζ ∥2 ] +CΣ E[∥ζ ∥2 ], (23) Step 4: Dimension bound. Setting ∑ni=1 µi < 0 gives dt n > CV ,P /α, so n⋆ = ⌊CV ,P /α⌋ + 1 suffices. The Lyapunov dimension formula of Constantin–Foias–Temam [3] then gives where LV , LP are local Lipschitz constants of ∇V and P on dimF (A ) ≤ n⋆ . B0 , and CΣ is the Lipschitz constant of Σ. Since α can Step 5: Stochastic extension. For the stochastic system, the be chosen large enough that α > LV + LP + CΣ , we obtain Debussche–Langa–Robinson method [12, 15] shows that the E[∥ζ (t)∥2 ] ≤ e−κt ∥η∥2 for κ = α − LV − LP −CΣ > 0. By the random attractor has Hausdorff and fractal dimension bounded triangle inequality, dist(Φη (t), A ) ≤ dist(Φ(t), A ) + ∥ζ (t)∥ ≤ by the same n⋆ , since the noise term DΣ(Φ⋆ )ξ dW contributes ε + Ce−κt/2 ∥η∥ for t large enough, giving the uniform bound zero to the deterministic volume-contraction estimate after taking dist(Φη (t), A ) ≤ ε +Cδ . □ expectations. □ C. Exponential Relaxation and Memory Loss Remark 2 (Explicit bound under MDEI parameters). For the Corollary 6 (Relaxation Time). For every Φ0 ∈ H: specific ISDM–TSM structure of [1,2], where T (Φ) is a bounded linear operator with ∥T ∥ ≤ CT , we obtain CV ,P ≤ c1 ∥Φ∥2L∞ (A ) + E[dist(Φ(t), A )] ≤ Ce−αt dist(Φ0 , A ), t ≥ 0. (24) CT + d1 , giving dimF (A ) ≤ ⌊(c1 R20 +CT + d1 )/α⌋ + 1, where R0 = K̃/β̃ is the energy radius of the absorbing set. This is the The characteristic relaxation time is τ = 1/α. For t ≫ τ, the first explicit formula for the dimension of the MDEI attractor in distribution of Φ(t) is governed solely by the invariant measure terms of all model parameters. on A , independent of Φ0 . Remark 3. The bound (15) is parameter-dependent Fig. 3 illustrates the dimension bound and the exponential (α, β , ν, d0 , d1 ) and data-independent: it does not grow relaxation. with token count N or model parameter count. This is the VI. A NATOMY OF THE ATTRACTOR fundamental scaling advantage of MDEI. A. Topographic Decomposition V. G EOMETRIC P ROPERTIES OF THE ATTRACTOR Definition 2 (Topographic Layers of A ). A. Diameter and Intrinsic Error Bound Corollary 4 (Diameter). M = {Φ ∈ A : ∇Φ V (Φ) = 0, ∇2Φ V (Φ) > 0}, (25) s ∂ A = {Φ ∈ A : ∥∇Φ V (Φ)∥ maximal}, (26) K̃ diam(A ) ≤ 2 , (20) λ1 β̃ with Nε (A ) ⊃ A ⊃ M . IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2026 5
(a) Dimension Bound dimF ( ) n (b) Relaxation Bound: [dist( (t), )] Ce td0 (a) MDEI Euler Maruyama Integration (b) Attention Euler Maruyama: d = J dt + dW (Theorem 3, purely analytical) (formula from Corollary 3 no empirical data) d = (A + I) dt + dW; = 1.2, = 0.10 J skew-symmetric (no dissipation), H1 violated (fixed seeds: 101,202,303,404 reproducible) (fixed seeds: 505,606,707,808 reproducible) C , P = 1.0 4.0 dist( 0, ) = 4.0 25 C , P = 2.0 dist( 0, ) = 2.0 (analytical boundary) Re( (J)) = 0 C , P = 3.5 = 1/ dist( 0, ) = 1.0 ( ) no absorbing set 3.5 0.6 no attractor
C e t dist( 0, ) (analytical) C , P = 5.0 dist( 0, ) = 0.5 2 20 = 0.25 (stability threshold) 3.0 0.4 n = C , P/ + 1
2.5 1 15 0.2 2.0 0.0 10 0
2
2 1.5 1.0 0.2 5 1 0.5 0.4 0 0.0 0.6 0.5 1.0 1.5 2.0 2.5 3.0 0 1 2 3 4 5 6 7 2 (dissipation rate) t 0.8 2 1 0 1 2 0.6 0.4 0.2 0.0 0.2 0.4 0.6 0.8 1 1 Figure 3. [Mathematical simulation — closed-form formulas only. No empirical data.] Left: Dimension bound dimF (A ) ≤ n⋆ = ⌊CV ,P /α⌋ + 1 (The- orem 3, formula (31)), plotted as a step function of α for four declared values Figure 4. [Mathematical simulation — exact integration of stated SDEs. Not of CV ,P ∈ {1.0, 2.0, 3.5, 5.0}. These are illustrative parameter values, not fitted a simulation of any real neural network.] Left (MDEI): Euler–Maruyama measurements. Right: Analytical relaxation bound Ce−αt dist(Φ0 , A ) (Corol- integration of the simplified linear MDEI-SDE dΦ = −(A + I)Φ dt + σ dW lary 6; α = 1.2, C = 1) for four declared initial distances d0 ∈ {0.5, 1.0, 2.0, 4.0}. (α = 1.2, σ = 0.10, dt = 0.01, T = 12). This is the Ornstein–Uhlenbeck process, Shaded region: stability neighbourhood Nε (A ) with ε = 0.25. Dotted vertical chosen as the minimal model satisfying H1–H4. Fixed seeds 101, 202, 303, line: τ = 1/α = 0.83. 404 (NumPy default_rng). Green dashed circle: analytical attractor boundary from Corollary 4. Right (Attention): Euler–Maruyama of dΦ = JΦ dt + σ dW , J skew-symmetric (ω = 0.5, σ = 0.08), a minimal model with Re(σ (J)) = 0, illustrating the consequence of H1 violation (Theorem 7). Fixed seeds 505, 606, The core manifold M consists of the stable equilibria: gen- 707, 808. Neither panel represents a simulation of any transformer, LLM, uine local minima of V , each corresponding to a distinct stable or real neural architecture. “meaning configuration” of the cognitive state. By Lyapunov’s second method, the deterministic flow attractor requires a compact absorbing set; its absence implies Φ̇ = −AΦ − ∇Φ V (Φ) + P(Φ) (27) no attractor. □ points strictly inward on Nε (A ) \ M , by condition (6). Remark 4 (Catastrophic Forgetting as Dynamical Consequence). B. Inertial Manifold and Dimensional Compression The empirical counterpart is catastrophic forgetting. Mechanistic For sufficiently large spectral gap λgap = ν(λ2 − λ1 ) of A, A analysis [9] across models from 109 to 1.5 × 1012 parameters is contained in an inertial manifold [3]: a Lipschitz manifold identifies three mechanisms: (1) gradient interference in atten- M ⊂ H of dimension n⋆ such that every orbit converges to M tion weights, (2) representational drift in intermediate layers, exponentially. On M the infinite-dimensional SPDE reduces and (3) loss-landscape flattening. All three are direct conse- ⋆ to an ODE on Rn . This is the rigorous sense in which MDEI quences of the absence of H1: without a compact absorbing set, “compresses by law”: complexity = n⋆ , a physical constant, not sequential fine-tuning displaces Φ outside the region of previous a data statistic. knowledge. The study reports that 15–23% of attention heads un- dergo severe disruption during fine-tuning, correlating strongly C. Deterministic Skeleton and Stochastic Exploration with forgetting severity [9]. Fine-tuning is not the solution; it is The stochastic term Σ dW perturbs orbits within Nε ′ (A ) for the symptom of a missing attractor. ε ′ > ε, but Proposition 5 guarantees they cannot permanently Fig. 4 shows a direct numerical comparison. escape. This balance—deterministic contraction plus stochastic exploration—is the mathematical realisation of the “dynamic VIII. MDEI VS . T RANSFORMER : C OMPARATIVE A NALY- pattern of coherence” of [2]. SIS
VII. I NCOMPATIBILITY: ATTENTION DYNAMICS AND H Y- Table 1 provides a structured comparison of the two paradigms POTHESIS H1 across all key mathematical and operational dimensions. Theorem 7 (Attention Dynamics Cannot Admit a Global At- IX. AGI R EDEFINED VIA THE G LOBAL ATTRACTOR tractor). Let Φ(t) evolve under pure attention dynamics Φ̇ = A. A Geometric Definition of AGI −∇Φ Latt (Φ), where Latt is the cross-entropy loss over attention- Definition 3 (MDEI-AGI). A dynamical system (H, {S(t)}t≥0 ) weighted token sequences of length N. Then the associated is an artificially general intelligence in the MDEI sense if and semigroup SA (t) does not satisfy H1 and admits no compact only if: absorbing set, hence no global attractor in H. √ (G1) S(t) satisfies Hypotheses H1–H4; Proof. The attention map A = softmax(QK ⊤ / d) is an N × N (G2) There exists a compact global attractor A with dimF (A ) < positive stochastic matrix. Its Jacobian does not contain a term ∞; −αI with α > 0 uniform in N: as N → ∞, the real part of the (G3) diam(A ) is tunable by α, β , ν, d0 , d1 ; spectrum of −∇2Φ Latt is not bounded away from zero, violating (G4) There exists an inertial manifold M ⊂ H, dim M = n⋆ < ∞, the sectorial condition in H1. The loss Latt is not coercive on H, to which all orbits converge exponentially. so no analogue of (12) holds and no absorbing ball exists. By This definition is operational: conditions (G1)–(G4) are verifi- the Hale–Ladyzhenskaya theorem [4], existence of a compact able from the operator structure. It formalises the intuition of [2] IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2026 6
Table 1. Comparative Analysis: MDEI-SPDE vs. Transformer Attention
Dimension MDEI-SPDE Transformer Attention
State representation Continuous field Φ ∈ H = L2 (Ω; R3 ) Discrete token sequence x ∈ ZN Governing dynamics SPDE (4): dΦ = (−AΦ − ∇V + P)dt + Σ dW Gradient flow: Φ̇ = −∇Φ Latt Dissipation (H1) Yes: A = −ν∆ + αI, σ (A) ≥ α > 0 No: spectrum extends to 0 as N → ∞ Absorbing set Yes: B0 = {V ≤ 2R0 }, R0 = K̃/β̃ No: proved absent (Theorem 7) Global attractor Yes: A compact, invariant (Theorem 2) No: consequence of Theorem 7 Fractal dimension Finite: dimF (A ) ≤ C(ν, α, β , d0 , d1 ) Unbounded: grows with N as O(N 2 ) Complexity scaling dimF (A ) fixed by physics; data-independent O(N 2 ) time and memory [10, 11] Catastrophic forget- Impossible: orbit forgets Φ0 , not A (Cor. 6) Inevitable: no attractor to return to [9] ting Fine-tuning require- None: convergence to A is intrinsic Required: surrogate for absent absorbing set ment Stability under noise Guaranteed (Proposition 5) Not guaranteed: orbit can diverge Relaxation time τ = 1/α, tunable by parameter Undefined (no attractor) Goal-directedness Intrinsic: P(Φ) = ∇ · (D∇Φ) + T (Φ) Extrinsic: only via Latt
that intelligence is not sequential prediction but the possession D. Comparison with Other Formal AGI Definitions of a globally stable “pattern of coherence”. Several formal definitions of AGI have been proposed in the lit- B. Four Principles of MDEI-AGI erature. We compare Definition 3 with the three most prominent: Theorem 8 (Four Emergent Principles). Every system satisfying Definition 3 exhibits: Universal AIXI (Hutter). AIXI defines intelligence as the (P1) Self-Organisation. P(Φ) generates structured patterns in agent maximising expected reward over all computable environ- A via nonlinear diffusion, without memorising input–output ments, using a mixture of all computable priors [4]. This defi- pairs. Mechanism: Turing-type instabilities are suppressed nition is incomputable and relies on a fixed prior (Kolmogorov by H1; the attractor A carries genuinely emergent structure complexity). MDEI-AGI (Definition 3) does not require enumer- from (7) [8]. ating environments; it characterises intelligence through the geo- (P2) Stability. Every orbit with Φ0 ∈ H enters Nε (A ) in finite metric property of possessing a global attractor. The two frame- time T (Φ0 ). Proof: Lemma 1 and Theorem 2. works are complementary: AIXI maximises reward; MDEI-AGI (P3) Causalidad. A state Φ⋆ ∈ M is causally stable: small pertur- minimises cognitive energy (Lyapunov functional). bations return to Φ⋆ exponentially; no spurious associations are generated outside M . Proof: Lyapunov’s second method Integrated Information Theory (Tononi). IIT [7] defines at a local minimum of V . intelligence/consciousness as the quantity ΦIIT (integrated infor- (P4) Dimensional Compression. System complexity = dimF (A ), mation). Definition 3 does not measure information integration, not the parameter count. Three operators (A, V , P) vs. 175× but the attractor dimension dimF (A ) provides an upper bound 109 weights; law vs. table. Proof: Theorem 3 and the inertial on the number of “irreducible” degrees of freedom—a structural manifold. constraint consistent with IIT. C. Connection to Friston’s Free-Energy Principle There is a deep structural parallel between MDEI-AGI and Fris- Legg–Hutter universal intelligence. The universal intelli- ton’s free-energy framework [7]. In both frameworks the system gence measure [4] is the reward-weighted average over all en- minimises a functional trading prediction accuracy against model vironments. Like AIXI, it requires computability assumptions complexity. In MDEI this functional is V ; in Friston’s setting not needed in MDEI. However, our relaxation-time result (Corol- it is variational free energy. The gradient descent −∇Φ V corre- lary 6) gives an operational interpretation of intelligence speed: sponds to perception (belief update); the teleological pressure P the characteristic time τ = 1/α for the system to adapt to a new corresponds to action (environment update to match beliefs). The environment after a perturbation. attractor A is the set of minimum-surprise states—maximum The key advantage of Definition 3 is verifiability: conditions model evidence—the “dynamic pattern of coherence” of [2]. (G1)–(G4) can be checked directly from the operator structure, IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2026 7
(a) Spectrum of A = + I (b) diam( ) 2 K/( 1 ) (c) Analytical Entry-Time Formula without reference to rewards, environments, or priors. (analytical: k = k 2 2 + on [0, 1]) 3.0 (Corollary 3, K = 0.5, 1 = 2, analytical) = 0.4 T * (d0) = 1 log d0 ( = 0.20) 0.78 7 = 0.8 = 1.5 103 2.5 6 = 2.5 E. Falsifiability of the MDEI-AGI Definition 0.69 = 0.2
(log scale)
diam( ) upper bound 0.30
(Lyapunov rate) 5 2.0
T * (d0) = 1 logd0 = 0.5, k = k + 0.60 A scientific definition must be falsifiable. Definition 3 is falsifi- 102 = 1.0, k = k + = 2.0, k = k + 1.5 0.51 4
k+ Floor = 1.2 (H1 lower bound) 3 0.40 able in the following senses: 101 0.42
k= 1.0 2 0.33 1 (i) Attractor existence (G2): If one constructs an MDEI sys- 100 (A) >0 0.5 Note: diam( ) 1/ 0.60; independent of when C2 = 0 0.80 0.24 0 5 10 15 20 25 0.5 1.0 1.5 2.0 2.5 3.0 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 tem satisfying H1–H4 but can demonstrate that some orbit Mode index k (damping rate) Initial distance d0 = dist( 0, )
diverges to infinity, Theorem 2 would be refuted. Figure 5. [Mathematical simulation — analytical formulas only. No em- (ii) Dimension bound (G2, C2): If one finds that dimF (A ) > n⋆ pirical data.] (a) Eigenvalue spectrum µk = νk2 π 2 + α of A = −ν∆ + αI on for an MDEI system, formula (31) would be refuted. the unit interval [0, 1] (Dirichlet eigenvalues λk = k2 π 2 ), for ν ∈ {0.5, 1.0, 2.0}, (iii) Stability under noise (G2, Prop. 5): If a perturbation ∥η∥ < α = 1.2. All eigenvalues satisfy µk ≥ α = 1.2 > 0, confirming H1. This is a δ produces an orbit that permanently leaves Nε+Cδ (A ), q not a measurement. (b) Contour plot of the diameter bound formula evaluation,
Proposition 5 would be refuted. diam(A ) ≤ 2 K̃/(λ1 β̃ ) (Corollary 4; K̃ = 0.5, λ1 = π 2 , β̃ ≈ β ) as a function (iv) Attention incompatibility (C3): If an attention-based sys- of (α, β ). All values derived from the formula; not measured from any system. (c) Entry-time formula T ⋆ (d0 ) = α1 log dε0 (derived from Corollary 6; ε = 0.20) tem is demonstrated to possess a compact global attractor in for four declared values of α. H, Theorem 7 would be refuted. None of these falsifications has been observed, and our proofs guarantee they cannot occur under H1–H4. of the system: 1 X. F OUR P RINCIPLES : D ETAILED O PERATOR M ECHA - diam(A ) × |{z} τ ≲ . (29) | {z } β NISMS accuracy speed
Table 2 gives an operator-level account of each principle, its Relation (29) is an accuracy–speed tradeoff : for fixed β , in- mathematical mechanism on A , and a discriminating experi- creasing α (faster √ convergence) does not worsen accuracy, since mental test distinguishing MDEI from attention-based systems. diam(A ) ∝ 1/ α also decreases. Only the Lyapunov rate β governs the joint product. Table 3 gives design rules for three XI. S PECTRAL T HEORY AND PARAMETER D ESIGN operating regimes. This section analyses how the operator parameters ν, α, β , d0 , d1 determine the geometry of A and how they can be chosen by C. Spectral Dimension Estimate design to achieve prescribed performance. From (19), the dimension bound becomes explicit when we use the eigenvalue structure: A. Eigenvalue Structure of A n On a bounded domain Ω with Dirichlet boundary conditions, the Laplacian −∆ has a countably infinite family of eigenvalues ∑ µi ≤ −nα +CP n1/2 , (30) i=1 0 < λ1 ≤ λ2 ≤ · · · ↗ ∞ with corresponding orthonormal eigen- functions {ek }∞ k=1 forming a Hilbert basis of H. The eigenvalues where CP = ∥DP∥L∞ (A ) + ∇2 V L∞ (A ) . Setting the right side of A = −ν∆ + αI are therefore negative determines 2 µk = νλk + α, k = 1, 2, . . . (28) CP n⋆ ≤ + 1. (31) α2 with µk ≥ µ1 = νλ1 + α > α > 0. This explicit spectrum gives three important consequences: This shows dimF (A ) ≲ CP2 /α 2 : stronger dissipation (α large) or (i) Uniform lower bound: Re(σ (A)) ≥ α, confirming H1 with weaker teleological coupling (CP small) reduces the dimension, the quantitative rate α. confirming that MDEI scales by physics, not by data. (ii) Spectral gap: The gap λgap = ν(λ2 − λ1 ) determines the ex- Fig. 5 shows the eigenvalue spectrum, the diam(A ) control ponential rate of convergence to the inertial manifold M and surface, and the entry-time formula. must satisfy λgap > ∥DP∥L∞ (A ) for M to be Lipschitz [3]. XII. ISDM–TSM C OUPLING : A FFECTIVE –S EMANTIC (iii) Compactness: Since µk → ∞, the resolvent (A − λ I)−1 is DYNAMICS compact on H, guaranteeing that e−At is a compact operator A central novelty of MDEI is the coupling between the Internal for every t > 0. State Dynamics Model (ISDM) governing affective regulation For a cube Ω = [0, L]n , the Poincaré constant is λ1 = π 2 /L2 and the Teleological Semantics Model (TSM) governing seman- and the spectral gap is λ2 − λ1 = π 2 (4 − 1)/L2 = 3π 2 /L2 (for tic dynamics [1, 2]. This section analyses this coupling from the n = 1). Choosing ν large enough ensures the gap condition is attractor perspective. met even for large ∥DP∥, i.e., strong teleological coupling. A. Decomposition of the State Field B. Parameter Design Rules p The state field Φ ∈ H = L2 (Ω; R3 ) decomposes as Corollary 4 gives diam(A ) ≲ 1/ αβ , while Corollary 6 gives τ = 1/α. These two quantities define the operational envelope Φ = (Φaff , Φsem , Φtel ) ∈ R3 , (32) IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2026 8
Table 2. Four Principles of MDEI-AGI: Operator-Level Mechanisms and Tests
Principle MDEI Operator Mechanism on A Discriminating Test
P1: Self-Org. P(Φ) = ∇ · (D∇Φ) + Nonlinear diffusion generates emer- Remove 20% of input tokens: MDEI T (Φ) gent patterns; Turing instabilities ex- reorganises within A ; attention col- cluded by H1 lapses q P2: Stability diam(A ) ≤ 2 K̃/(λ1 β̃ ) All orbits absorbed into B0 , then at- Inject noise ∥η∥ < δ : MDEI stays in tracted to A (Theorem 2) Nε+Cδ (A ); attention drifts P3: Causali- M = {∇Φ V = 0, ∇2 V > Stable fixed points of −∇V generate Counterfactual queries: MDEI re- dad 0} causally stable responses; no halluci- sponds from M ; attention samples nations from non-minima outside M P4: Compres- dimF (A ) ≤ C(ν, α, β ) 3 operators (A, V , P) vs. 175 × 109 Scale N: dimF (A ) constant; atten- sion params; complexity = n⋆ , a physical tion cost ∝ N 2 constant
Table 3. Parameter Design Rules for MDEI Operation (b) Analytical Invariant Measure (a) ISDM-TSM Coupled Phase Portrait MDEI: stationary (0, 2/2Aeff) (compact support on ) Exact gradient flow of joint; a = 1.0, s = 0.8, = 0.4 Attention: diffusing Gaussian, variance t (no attractor) 16 , r 0.61 (formula) Attention p( , t = 10), Var = 2t 2 ( ) Attention p( , t = 30), Var = 2t Core (unique min.\ of joint) 14
Probability density p( ) (analytical) Attention p( , t = 100), Var = 2t Regime α β Effect 12 MDEI stationary p ( ) = (0, 2/2Aeff) 1
sem (semantic state) 10 High accuracy Large Large Small diam(A ), fast τ 0 8
Fast response Large Moderate Fast τ, moderate accuracy 6 1 4 Energy-efficient Moderate Large Small diam(A ), slower τ 2 2 0 2 1 0 1 2 3 2 1 0 1 2 3 aff (affective state) (state projection)
where Φaff encodes the affective (energetic) component, Φsem Figure 6. [Mathematical simulation — exact gradient flow and analytical encodes the semantic (cognitive) component, and Φtel encodes probability densities. No empirical data.] (a) Streamlines of the exact deter- the teleological orientation. The three-dimensional structure was ministic gradient flow Φ̇ = −∇Vjoint (Φ) of the joint Lyapunov functional (35) introduced in [1] as the minimal algebraic framework admitting (αa = 1.0, αs = 0.8, γ = 0.4, c1 = 0.25). Computed by numerically evaluating all four principles of Theorem 8. the exact gradient formula; not a simulation p of any cognitive or neural system. The analytical attractor radius rA = K/ min(αa , αs ) is marked. (b) Analytical B. Coupled Evolution Equations probability density functions derived from the theory (no sampling, no Monte Carlo): Blue — stationary (invariant) density of the Ornstein–Uhlenbeck process In the ISDM–TSM formulation [2], the components evolve as: dΦ = −Aeff Φ dt + σ dW : p⋆ (Φ) ∝ exp(−Aeff Φ2 /σ 2 ), with Aeff = α + 1 = 2.0, σ = 0.10. This is the exact invariant measure of the MDEI-OU model (Sec- dΦaff + αa Φaff + ∂Φa V − Pa dt = σa dWa , (33) tion 18, eq. (55)). Red family — marginal density of the attention random walk dΦsem + αs Φsem + ∂Φs V − Ps − γΦaff dt = σs dWs , (34) dΦ = JΦ dt + σ dW at times t = 10, 30, 100: p(Φ,t) = N (0, σ 2 t) (exact for- mula, not sampled). The expanding variance demonstrates Theorem 7. where γ > 0 is the affective–semantic coupling constant and (Wa ,Ws ) are independent Q-Wiener processes. The coupling term γΦaff in (34) encodes the principle that affective states modulate semantic processing—a key claim of the ISDM frame- Vjoint ensures coherence: states in Ajoint have bounded affective– work [2]. semantic misalignment, formalising the “pattern of coherence” C. Joint Lyapunov Functional and Attractor Coupling of [2]. The joint Lyapunov functional for the coupled system is Remark 5 (Lyapunov Construction of [2]). The joint func- tional (35) is precisely the Lyapunov construction introduced by γ Rudolph and Vieira [2] in the context of stability of ISDM–TSM Vjoint (Φ) = Vaff (Φaff ) + Vsem (Φsem ) + ∥Φaff − Φsem ∥2 , (35) 2 dynamics. The present paper embeds that result in the infinite- where the cross-term penalises affective–semantic misalignment. dimensional SPDE framework and proves that the corresponding The dissipativity condition (6) for Vjoint reads: attractor Ajoint is compact with finite fractal dimension.
d D. Phase Portrait and Invariant Measure E[Vjoint ] ≤ − min(αa , αs )E[Vjoint ] + Kjoint , (36) dt Fig. 6 shows the phase portrait of the coupled ISDM–TSM sys- which is satisfied as long as both damping rates αa , αs > 0. The tem on the projected plane (Φaff , Φsem ) and the invariant measure joint attractor Ajoint projects onto attractors for each subsys- on A compared with the diffuse distribution of a transformer- tem: πa (Ajoint ) ⊆ Aaff and πs (Ajoint ) ⊆ Asem . The cross-term in based system. IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2026 9
XIII. C ONTROL T HEORY P ERSPECTIVE : T UNING THE Table 4. MDEI-SPDE vs. Neural ODE: Structural Comparison ATTRACTOR From a systems-engineering perspective, the MDEI framework Property MDEI-SPDE Neural ODE can be viewed as a controlled dissipative system where the pa- Dissipativity (H1) Built-in (A = −ν∆ + αI) Learned (not guaranteed) rameters (α, β , ν, γ, d0 , d1 ) are design variables. This section analyses the problem of prescribing the attractor geometry. Stochastic driving Q-Wiener, trace-class (H4) Typically absent Global attractor Proven (Theorem 2) Depends on fθ A. Attractor Placement Problem Dimension bound n⋆ = ⌊CV ,P /α⌋ + 1 No general bound Definition 4 (Attractor Placement). Given a target diameter D⋆ > 0 and target relaxation time τ ⋆ > 0, find parameters (α, β ) Fine-tuning Not needed (attractor intrinsic) May be needed such that diam(A ) ≤ D⋆ and τ = 1/α ≤ τ ⋆ . Interpretability Operators A, V , P explicit fθ black-box From Corollary 4 and (??):
1 4K̃ Theorem 9 (Closed-Loop Stability). Let α(t) = α0 + α≥ , β≥ . (37) τ⋆ λ1 α(D⋆ )2 kα max(0, Vˆ (t) − V ⋆ ) as in (39) with α0 > 0, kα > 0, V ⋆ = 2R0 . Assume Vˆ (t) = E[V (Φ(t))] (exact observation). Then: These are linear constraints on (α, β ) and define a feasible (i) {t : E[V (Φ(t))] ≤ V ⋆ } is positively invariant; half-plane in parameter space. The minimum-energy solution (ii) For all t ≥ 0, E[V (Φ(t))] ≤ max(E[V (Φ0 )], V ⋆ ); (minimising α 2 + β 2 subject to (37)) can be found analytically: (iii) lim supt→∞ E[V (Φ(t))] ≤ V ⋆ .
1 4K̃τ ⋆ Proof. Define e(t) = (E[V (t)] − V ⋆ )+ ≥ 0. When e(t) > 0, α⋆ = , β⋆ = . (38) α(t) = α0 + kα e(t) and the closed-loop energy estimate (12) τ⋆ λ1 (D⋆ )2 gives B. Lyapunov-Based Feedback Control d The parameters α and β can be made time-varying adaptive E[V ] ≤ −(β̃ + kα e)E[V ] + K̃ ≤ −β̃ E[V ] + K̃ − kα e · V ⋆ . dt controls. Suppose we observe an estimate Vˆ (t) of V (Φ(t)) and (42) wish to maintain Vˆ (t) ≤ V ⋆ . Define the feedback law: Since e > 0 implies E[V ] > V ⋆ = 2R0 = 2K̃/β̃ , we have kα e · V ⋆ > 0, so dtd E[V ] ≤ −β̃ E[V ] + K̃, which is the same α(t) = α0 + kα max(0, Vˆ (t) − V ⋆ ), (39) estimate as in the open loop. This gives part (iii) by Gronwall. For part (i): if E[V (t0 )] ≤ V ⋆ , then e(t0 ) = 0, α(t0 ) = α0 , and where kα > 0 is a gain. When Vˆ exceeds V ⋆ , the damping dt E[V ] ≤ −β̃ V + K̃ = −β̃ · 2K̃/β̃ + K̃ = −K̃ ≤ 0, so E[V ] d ⋆ increases, pulling the orbit back toward B0 . The closed-loop cannot increase above V ⋆ . Part (ii) follows from combining (i) energy estimate (12) becomes: and the Gronwall bound. □ d Remark 6. Theorem 9 shows that the Lyapunov feedback pro- E[V ] ≤ − β + kα max(0, Vˆ − V ⋆ )/V E[V ] + K̃, (40) vides a provably safe mechanism for maintaining the orbit in dt the absorbing set B0 . The feedback requires only measuring showing that the effective decay rate increases whenever the E[V (Φ(t))] (a single scalar), not retraining on new data. system exceeds the target energy. This Lyapunov-based adaptive E. Comparison with Neural ODE Architectures damping is a principled alternative to ad-hoc fine-tuning. Neural ODEs [13] replace the discrete residual stream x(l+1) = C. Comparison with Fine-Tuning x(l) + fθ (x(l) , l) of transformers with a continuous-time ODE Fine-tuning an LLM can be viewed as an ad-hoc projection step: ẋ = fθ (x,t). Their dynamics can be written as given that the orbit has left the region of previous knowledge (no attractor to constrain it), fine-tuning attempts to return Φ dx = fθ (x,t), x(0) = x0 , (43) to a neighbourhood of the desired output distribution. This is dt fundamentally reactive and does not prevent future escape. By which has a global attractor if and only if fθ is dissipative: contrast, the Lyapunov feedback (39) is proactive: it prevents the ⟨ fθ (x,t) − fθ (y,t), x − y⟩ ≤ −c ∥x − y∥2 for some c > 0. Stan- orbit from leaving Nε (A ) in the first place, with no data-specific dard neural ODE architectures do not impose this condition; intervention required. The cost difference is dramatic: the network learns fθ freely, potentially violating dissipativity. MDEI, by contrast, builds dissipativity into the structure via H1: Fine-tuning cost ≫ Lyapunov feedback cost ∈ O(1). the term −AΦ in (4) guarantees c ≥ α > 0 uniformly, regardless of the learned components ∇Φ V and P. Table 4 compares the | {z } | {z } data+GPU hours parameter update (41) two architectures. D. Closed-Loop Stability Theorem F. Robustness to Distribution Shift We now state and prove that the Lyapunov-based adaptive feed- A further advantage is robustness to distribution shift: changes back (39) renders B0 positively invariant. in the input distribution correspond to changes in the noise term IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2026 10
Table 5. MDEI vs. Reservoir Computing / Echo State Networks C. Finite-Dimensional Projection as Readout The inertial manifold M (Section 6) provides a natural read- ⋆ Property MDEI-SPDE ESN / Reservoir out layer for MDEI: the projection Π : H → Rn onto the n⋆ - dimensional manifold extracts the order parameters, analogous Echo state property Guaranteed (α > Requires ρ(W ) < to the trained readout weights in an ESN. The difference is fun- 0, H1) 1 damental: in MDEI the readout dimension n⋆ is determined by Attractor dimen- dimF (A ) ≤ n⋆ Bounded by reser- the operator physics (formula (31)), whereas in an ESN it is cho- sion voir size N sen arbitrarily. Furthermore, the MDEI readout Π is determined Training None (attractor in- Readout only by the attractor geometry and requires no training data—only trinsic) knowledge of the operators A, V , and P. Stochastic driving Q-Wiener, H4 Typically absent XV. I NVARIANT M EASURES AND E RGODIC P ROPERTIES Infinite- Yes (H = No (finite RN ) A. Existence of Invariant Measures dimensional L2 (Ω; R3 )) Under H1–H4, the Markov semigroup {S(t)} is Feller: t 7→ Goal-directedness Intrinsic (P opera- Only via readout S(t)µ is weakly continuous for every probability measure µ on tor) H. By the Krylov–Bogoliubov theorem [6], the existence of a State space Hilbert space H RN , fixed N compact absorbing set (Lemma 1) implies the existence of at Interpretability Explicit A, V , P Random reservoir least one invariant probability measure µ ⋆ for S(t), supported on A: Z Z µ ⋆ (A ) = 1, f d(S(t)# µ ⋆ ) = f dµ ⋆ ∀ f ∈ Cb (H), t ≥ 0. Σ(Φ) dW in (4). By H4 and Proposition 5, as long as the new H H distribution satisfies (8), the orbit remains in Nε+Cδ (A ). In (45) contrast, transformer models without an attractor exhibit catas- B. Ergodicity and Time Averages trophic performance degradation under moderate distribution If the invariant measure µ ⋆ is unique (which holds under a suit- shift [9], consistent with the absence of H1. able irreducibility and strong Feller condition [6]), the system is ergodic: for µ ⋆ -almost every Φ0 , XIV. C ONNECTION TO R ESERVOIR C OMPUTING AND Z T E CHO S TATE N ETWORKS 1 Z T →∞ f (Φ(t)) dt −−−→ f dµ ⋆ , f ∈ Cb (H). (46) Reservoir computing [5] is a framework for processing time- T 0 H series data using a fixed, high-dimensional dynamical system This is the rigorous foundation for the empirical observation that (the reservoir) and training only a linear readout. Echo State MDEI produces stable long-run statistics: the time averages of Networks (ESNs) are the discrete-time version. The mathemat- any observable converge to their µ ⋆ -expectation, independent of ical condition for an ESN reservoir to function correctly is the the initial condition Φ0 (after the relaxation time τ). In contrast, echo state property: the effect of initial conditions fades expo- a system without an attractor has no invariant measure supported nentially, i.e., the network has a unique response to any input on a compact set; its long-run statistics depend on the entire sequence. This is precisely the property captured by our Corol- history of fine-tuning steps, producing the drift observed in [9]. lary 6: E[dist(Φ(t), A )] ≤ Ce−αt dist(Φ0 , A ). The relaxation C. Mixing and Correlation Decay time τ = 1/α is the MDEI counterpart of the spectral radius con- dition ρ(W ) < 1 in discrete ESNs (where ρ(W ) is the spectral A stronger property is exponential mixing: there exists γmix > 0 radius of the reservoir weight matrix). such that
A. Echo State Property as a Special Case of Corollary 3 E[ f (Φ(t))g(Φ(0))] − Eµ ⋆ [ f ]Eµ ⋆ [g] ≤ Ce−γmix t ∥ f ∥Cb ∥g∥Cb . In continuous time, the echo state property for an autonomous (47) reservoir ẋ = f (x, u) driven by input u(t) is: Exponential mixing follows from the spectral gap of the Kolmogorov operator (the generator of the semigroup in ∥x1 (t) − x2 (t)∥ → 0 as t → ∞ (44) L2 (H, µ ⋆ )) [6]. Physically, (47) means that correlations between past and future states decay exponentially with the mixing rate for any two initial conditions x1 (0) ̸= x2 (0) but the same input γmix : the system forgets its past at the rate γmix , which is related u. This is exactly the content of Corollary 6 restricted to the to but distinct from the relaxation rate α in Corollary 6. deterministic case (Σ ≡ 0): dist(Φ(t), A ) ≤ Ce−αt dist(Φ0 , A ). Remark 7 (Causalidad and Mixing). The Causalidad principle Thus the MDEI framework implies the echo state property, with (P3) of Theorem 8 is the deterministic analogue of ergodicity: an explicit rate α > 0 guaranteed by H1. stable fixed points Φ⋆ ∈ M generate responses that are indepen- dent of the transient history. Ergodicity (46) is the stochastic B. Contrasting Properties extension: long-run statistics are governed by µ ⋆ , not by initial Table 5 compares MDEI with classical reservoir computing along conditions. Both properties would be absent without the attractor the dimensions relevant to attractor theory. A. IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2026 11
XVI. D ISCUSSION teleological pressure P corresponds to active inference (minimis- A. Connections to Attractor Theory for PDEs ing surprise by changing the environment). The attractor A The existence and finite dimensionality of global attractors for corresponds to the set of free-energy minima: states of minimum dissipative PDEs is classical [16]. Ladyzhenskaya (1972) proved surprise, maximum model evidence. The core M = {∇V = 0} corresponds to the prior expectation states where the model is the first boundedness result for the 2D Navier–Stokes equations fully consistent with data. in a bounded domain. Constantin–Foias–Temam [3, 14] derived sharp dimension bounds via Lyapunov exponents: for the 2D This formal correspondence suggests that the MDEI-SPDE NSE, dimF (A ) ≤ c0 ν −3/2 ∥ f ∥3/2 G1/2 , where ν is viscosity, f a is a field-theoretic extension of Friston’s framework from finite- body force, and G the Grashof number. Robinson [5] extended dimensional mean-field dynamics to infinite-dimensional con- the theory to stochastic evolution equations. Our Theorems 2 tinuous fields—a generalisation that has not been previously formalised. and 3 instantiate this classical machinery for the specific MDEI operators, giving an explicit bound n⋆ = ⌊CV ,P /α⌋ + 1 that is E. Efficient Attention and Structural Limits the MDEI analogue of the Grashof-number estimate. The O(N 2 ) bottleneck has motivated numerous approximations: B. Stochastic Attractors linear attention [11] (O(N)), Nyströmformer, LSG attention, For SPDEs driven by Q-Wiener noise, the theory of random at- sparse attention [22], and Mamba (state-space models). How- tractors was developed by Crauel–Flandoli [18]. Debussche [15] ever, reducing computational complexity does not recover H1: an proved finite Hausdorff dimensionality of random attractors; O(N) approximation of a non-dissipative dynamics is still non- Langa–Robinson [12] extended this to fractal dimension esti- dissipative. The fundamental√issue is structural: the attention mates under assumptions analogous to our H1–H4. Our proof score matrix softmax(QK ⊤ / d) has spectrum bounded away follows their approach but incorporates the teleological opera- from −∞ but not bounded away from zero; no amount of ap- tor P, which couples semantic and affective dynamics follow- proximation fixes this. Hybrid continuous-depth approaches [13] ing [1, 2]. The key new element is the verification that P satisfies incorporating neural ODEs replace the attention residual stream the dissipativity condition (6) jointly with ∇Φ V ; this is guaran- with a differential equation, potentially restoring H1 if the ODE teed by the ISDM–TSM structural assumptions of [2]. is chosen to be dissipative. This is the direction most consistent with MDEI principles. C. Formal Analogy with Navier–Stokes The structural parallel between the MDEI-SPDE (4) and the F. Synergetics and Order Parameters incompressible Navier–Stokes equations (NSE) is illuminating. Haken’s synergetics [8] explains self-organisation via the slav- In both cases: ing principle: fast (stable) modes are enslaved to slow (order (i) A sectorial operator (−ν∆ in both) provides high-frequency parameter) modes, reducing effective dimensionality from the dissipation; full phase space to a finite-dimensional manifold. The MDEI (ii) A nonlinear term (∇Φ V in MDEI; (u · ∇)u in NSE) intro- inertial manifold M (Section 6) is the rigorous realisation of duces the essential nonlinearity; this principle: the n⋆ coordinates on M are the order parameters, (iii) A forcing term (P in MDEI; external force f in NSE) pre- and the remaining (infinite-dimensional) complement of M in H vents trivial collapse to zero; consists of the slaved modes that relax to M at rate λgap /2. The (iv) A compact global attractor exists in both cases, with dimen- pattern-formation results of Cross–Hohenberg [20] further sup- sion bounded by physical parameters. port this picture: dissipative nonlinear PDEs generically exhibit The difference is that the NSE nonlinearity (u · ∇)u is energy- finite-dimensional attractors whose structure is determined by preserving (it vanishes in the energy balance), whereas ∇Φ V the interplay of dissipation and driving forces. in MDEI is genuinely dissipative: it contributes positively to G. Limitations and Open Problems the Lyapunov decay, making the MDEI analysis cleaner and the (i) Explicit n⋆ for MDEI operators: Computing CV ,P = bounds sharper. ∇2 V L∞ (A ) + ∥DP∥L∞ (A ) for the specific ISDM–TSM op- D. Friston’s Free Energy: Formal Correspondence erators of [1, 2] requires knowledge of A , which is itself The free-energy principle [7] proposes that biological cognition the unknown. A self-consistent fixed-point argument may minimises variational free energy: resolve this. F(Φ, s) = KL[q(Φ) ∥ p(Φ|s)] − ln p(s), (48) (ii) Galerkin discretisation: Finite-element discretisations of | {z } | {z } (4) introduce discretisation error; convergence Ah → A as complexity accuracy mesh h → 0 requires a Galerkin analysis following [3], Chap- where q(Φ) is a variational density, p(Φ|s) the generative model, ter IV. and s sensory observations. In MDEI, the Lyapunov functional (iii) Adaptive control of diam(A ): The feedback law (39) pro- V plays the role of F: vides a starting point, but a rigorous stability proof for the closed-loop system (40) under the adaptive α(t) requires a V (Φ) ←→ F(Φ, s). (49) time-varying Lyapunov argument. The gradient descent −∇Φ V corresponds to perceptual infer- (iv) Biological validation: The prediction that biological cog- ence (minimising surprise by updating the internal model); the nition exhibits a finite-dimensional attractor is testable in IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2026 12
principle via dimensionality reduction of EEG/fMRI time rium, but as a dynamic pattern of coherence. The global attractor series. The methods of [3] for estimating dimF (A ) from A is that pattern, made rigorous by the present work. data (correlation dimension, Lyapunov exponents) could be XVIII. W ORKED E XAMPLE : E XPLICIT ATTRACTOR FOR applied to neural recordings. THE S CALAR MDEI-OU M ODEL (v) Hybrid transformer-MDEI architectures: Replacing the attention residual stream x 7→ x+Attention(x) with an MDEI- To make the abstract theory concrete, we work out the attractor type SPDE layer dx = −Ah x dt + fθ (x) dt + σ dW would completely for a scalar, linear version of (4). This is the MDEI- restore H1 while preserving the expressivity of transformer Ornstein–Uhlenbeck (MDEI-OU) model, which is the simplest architectures. non-trivial instance of the framework. (vi) Non-autonomous (input-driven) MDEI: The present anal- A. Model Specification ysis treats the SPDE as autonomous. Extending to non- Let H = R (scalar state) and consider autonomous forcing P(t, Φ) (representing time-varying in- puts) would require the theory of uniform attractors [3] or dΦ(t) = −aΦ(t) dt + σ dW (t), Φ(0) = Φ0 ∈ R, (50) pullback attractors [5]. where a > 0 is the effective damping (a = α + β0 combining A XVII. C ONCLUSIONS and ∇Φ V for a quadratic V (Φ) = β20 Φ2 ) and W (t) is a standard We have established seven main results, which we state precisely. Wiener process with σ > 0. Hypothesis H1 is satisfied with C1 (Theorem 2). Under H1–H4, the MDEI-SPDE (4) admits α = a; H2 holds for V (Φ) = β20 Φ2 since V (Φ) → ∞ as |Φ| → ∞; a compact global attractor A ⊂ H that is invariant, attracts every H3 is trivial (no coupling); H4 holds with C1 = σ 2 , C2 = 0. bounded set uniformly, and is the unique compact maximal B. Exact Solution invariant set of {S(t)}t≥0 . The exact solution is the classical Ornstein–Uhlenbeck process: C2 (Theorem 3). The fractal dimension satisfies dimF (A ) ≤ n⋆ = ⌊CV ,P /α⌋ + 1, where CV ,P = ∇2 V L∞ (A ) + ∥DP∥L∞ (A ) . Z t Φ(t) = e−at Φ0 + σ e−a(t−s) dW (s). (51) This bound is data-independent: it grows with coupling strength, 0 not with token count or parameter count. First explicit formula It is Gaussian with: for dimF (A ) in terms of all MDEI model parameters. C3 (Theorem 7). Attention dynamics violate H1. No compact E[Φ(t)] = e−at Φ0 , (52) absorbing set, no global attractor, no finite-dimensional iner- σ2 1 − e−2at . tial manifold exist. Catastrophic forgetting [9] and fine-tuning Var[Φ(t)] = (53) 2a dependence are structural, not incidental, consequences. C4 (Corollaries 4, 6; Prop. 5). The C. Absorbing Set: Explicit Formula q attractor geometry is explicitly computable: diam(A ) ≤ 2 K̃/(λ1 β̃ ), τ = 1/α, and From Lemma 1 with β̃ = a, K̃ = σ 2 /2: stability under perturbation ∥η∥ < δ holds with margin Cδ . s s C5 (Def. 3; Thm. 8). AGI is characterised geometrically as σ2 n β0 2 o h σ2 σ2 i R0 = , B0 = Φ ∈ R : Φ ≤ 2R0 = −2 ,2 . the possession of a global attractor of bounded fractal dimension. 2a 2 aβ0 aβ0 Four emergent principles (Self-Organisation, Stability, Causali- (54) dad, Dimensional Compression) follow necessarily. D. Global Attractor: Exact Description C6 (Section 15). MDEI systems possess an invariant measure As t → ∞, (53) gives Var[Φ(t)] → σ 2 /(2a), and the distribution µ ⋆ supported on A and exhibit exponential mixing with rate γmix . converges to the unique invariant measure: Long-run statistics are governed solely by µ ⋆ , not by history. σ2 This property is absent in transformer systems, which have no µ = N 0, ⋆ . (55) µ ⋆ supported on a compact set. 2a C7 (Section 13). The attractor geometry is prescribable: set- ting α ⋆ = 1/τ ⋆ and β ⋆ = 4K̃τ ⋆ /(λ1 (D⋆ )2 ) achieves any target The attractor is the support of µ ⋆ in the relevant functional sense; diameter D⋆ and relaxation time τ ⋆ . Lyapunov-based adaptive its “effective diameter” (two standard deviations) is damping (39) maintains the orbit in Nε (A ) at O(1) cost, replac- 2σ ing the data- and compute-intensive fine-tuning cycle entirely. diameff (A ) = √ . (56) a Conceptual summary. The transition from the discrete-token p Comparing with Corollary 4: (56) is precisely 2 K̃/a = paradigm to the MDEI continuous-field paradigm is the transi- p √ 2 σ 2 /(2a2 ) · a, consistent with the bound (differing only tion from systems lacking a global attractor to systems possess- by the Poincaré constant λ1 = 1 in the scalar case). ing one. Fine-tuning is the symptom; the absence of H1 is the cause. AGI is not the extrapolation of next-token prediction; it is E. Relaxation Time and Memory Loss the architectural commitment to dynamics that carry a compact, From (52), the relaxation time is τ = 1/a, exactly as in Corol- finite-dimensional, globally stable invariant set. In the words lary 6. After time t ≫ τ, the distribution of Φ(t) is µ ⋆ regardless of [2]: stability emerges not as convergence to a static equilib- of Φ0 : the system has completely forgotten its initial condition. IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2026 13
Table 6. Scalar MDEI-OU Model: Exact Attractor Quantities maps H into D(A1/2 ) = V = H01 (Ω; R3 ) for s < t. Specifically,
Quantity Formula Dependence C A1/2 e−A(t−s) ≤ , 0 < s < t, (59) p L (H) (t − s)1/2 Absorbing set radius σ 2 /(aβ0 ) ↑ σ , ↓ a, ↓ β0 √ Attractor std. dev. σ / 2a ↑ σ, ↓ a which is integrable in s ∈ [0,t). Therefore Φ2 (t) ∈ D(A1/2 ) = V √ almost surely. Since the embedding V ,→ H is compact (Sobolev– Effective diameter 2σ / a ↑ σ, ↓ a Rellich–Kondrachov), the image S2 (t)B is precompact in H. Relaxation time τ 1/a ↓a □ Invariant measure N (0, σ 2 /2a) Gaussian, compact effective support B. Invariance of A : Detailed Proof Fractal dimension ≤ ⌊β0 /a⌋ + 1 = 0 (det.), ≤ 1 (stoch.) Proof of S(t)A = A , Theorem 2(ii). The inclusion S(t)A ⊆ Memory loss time t ≫ 1/a Independent of Φ0 A follows from the invariance of the omega-limit: if Φ ∈ A = All entries derived from (a, σ , β0 ); no empirical data. ω(B0 ), then there exist tn → ∞ and Ψn ∈ B0 with S(tn )Ψn → Φ. By the semigroup property, S(tn + t)Ψn = S(t)(S(tn )Ψn ) → S(t)Φ, which is in ω(B0 ) = A . F. Fractal Dimension: Exact Value For the reverse inclusion A ⊆ S(t)A : let Φ ∈ A . By invari- For the scalar OU model, H = R, so A is a bounded interval ance of the omega-limit, Φ lies on a complete orbit {Φ(s) : s ∈ on R. Its fractal dimension is dimF (A ) = 0 (a point for the R} ⊂ A . In particular Φ(−t) ∈ A and S(t)Φ(−t) = Φ. Hence deterministic skeleton) or dimF (A ) ≤ 1 for the stochastic case. Φ ∈ S(t)A . □ From formula (31): n⋆ = ⌊CV ,P /a⌋ + 1 = ⌊β0 /a⌋ + 1, which C. Uniqueness of the Attractor equals 1 for β0 < a (subcritical damping), confirming that the Corollary 11 (Uniqueness). The global attractor A is unique attractor is zero-dimensional (a single point at the origin in the among compact invariant sets attracting bounded sets. deterministic case). G. Parameter Table for the Scalar Example Proof. Let A ′ be any other compact invariant set that attracts Table 6 summarises the complete quantitative description of the bounded sets. Since A ′ is bounded, dist(S(t)A ′ , A ) → 0, MDEI-OU attractor. All entries are derived analytically from the i.e., A ′ is attracted by A . Since S(t)A ′ = A ′ (invariance), model parameters (a, σ , β0 ). dist(A ′ , A ) = 0, so A ′ ⊆ A . Symmetrically A ⊆ A ′ . □
XIX. P ROOFS : C OMPACT D ECOMPOSITION AND I NVARI - ACKNOWLEDGEMENTS ANCE The author thanks Hans-Joachim Rudolph for foundational dis- This section gives the proof details for the compact decompo- cussions on the ISDM–TSM coupling [2] and the Valer Inteligên- sition and the invariance of A = ω(B0 ) that were sketched in cia Artificial group for support. Priority for the MDEI framework Section 4. and the joint Lyapunov construction resides with [1, 2]. A. Compact Decomposition: Detailed Proof I. F ULL I TÔ –LYAPUNOV D ERIVATION Proposition 10 (Compact Decomposition). Under H1–H3, for Starting from the mild solution (9), apply (2) to F = V . The each t > 0 and bounded B ⊂ H, the map S(t) : B → H decom- stochastic integral term has zero expectation (martingale). The poses as S(t) = S1 (t) + S2 (t) where: trace term satisfies (i) ∥S1 (t)∥L (H) ≤ Me−αt ; (ii) S2 (t)B is precompact in H for each t > 0. 1 2 ∑ ∇2 V Σek , Σek H ≤ C′ ∥Σ∥2L20 ≤ C′ (C1 +C2 V ), (60) k
Proof. Let Φ(t) = S(t)Φ0 solve (4). Define Φ1 (t) = e−At Φ0 by H4. Combined with (6): (the pure semigroup part) and Φ2 (t) = Φ(t) − Φ1 (t). Then Φ1 satisfies part (i) by (3). d For part (ii), Φ2 (t) satisfies the equation E[V ] ≤ −(β −C′C2 )E[V ] + K +C′C1 = −β̃ E[V ] + K̃. dt (61) ′C > 0 (by H4: C < β ) dΦ2 +AΦ2 dt = −∇Φ V (Φ)+P(Φ) dt +Σ(Φ) dW, Φ2 (0) = 0. Gronwall then gives (13) with β̃ = β −C 2 2 (57) and K̃ = K +C′C1 ≥ 0. □ By the variation-of-constants formula, II. P OINCARÉ I NEQUALITY AND S PECTRAL DATA Z t Z t A. Poincaré Inequality e−A(t−s) ∇Φ V −P (Φ(s)) ds+e−A(t−s) Σ(Φ(s)) dW (s). 1 Φ2 (t) = − 0 0 On H0 (Ω), the Poincaré inequality reads: (58) Since Φ(s) ∈ B0 for s ≥ T (B) (by Lemma 1), the integrands 1 are bounded in H. The analyticity of e−At implies that e−A(t−s) ∥Φ∥2L2 (Ω) ≤ ∥∇Φ∥2L2 (Ω) , ∀Φ ∈ H01 (Ω). (62) λ1 IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2026 14
For a hypercube Ω = [0, L]n (side L), the Dirichlet eigenvalues C. Hypothesis H3: Bounded Teleological Pressure of −∆ are P(Φ) = ∇ · (D(Φ)∇Φ) + T (Φ) satisfies H3 if: (a) D(Φ) is uni- formly bounded and elliptic (d0 ≤ D ≤ d1 ), which holds if D π2 2 k1 + · · · + kn2 , is Lipschitz and bounded away from zero; (b) T : H → H is a λk1 ,...,kn = 2 ki ∈ N, (63) L bounded operator, i.e., ∥T (Φ)∥H ≤ CT (1 + ∥Φ∥H ). The TSM so λ1 = nπ 2 /L2 (achieved at k1 = · · · = kn = 1). For a ball operator of [2] satisfies both conditions. ✓ Satisfied for all Ω = B(0, R), the first eigenvalue is λ1 = ( jn/2−1,1 /R)2 , where MDEI teleological operators considered in [1]. jν,1 is the first zero of the Bessel function Jν . D. Hypothesis H4: Subcritical Stochastic Noise B. Eigenvalues of A and Spectral Gap The key condition is 2C2 < 2β , i.e., the noise growth rate is The eigenvalues of A = −ν∆ + αI are strictly dominated by the Lyapunov decay rate. For additive noise (C2 = 0, Σ = const): H4 is automatically satisfied. For µk = νλk + α, k = 1, 2, . . . , (64) multiplicative noise with ∥Σ(Φ)∥ ≤ C ∥Φ∥: C2 = C2 and the condition becomes C2 < β , i.e., the noise must be subcritical with µk ≥ µ1 = νλ1 + α > α > 0. The spectral gap is relative to the dissipation. ✓ Satisfied when noise intensity is not too large relative to β . λgap = µ2 − µ1 = ν(λ2 − λ1 ). (65) IV. N OTATION S UMMARY For Ω = [0, L]3 (relevant for n = 3): λ1 = 3π 2 /L2 , λ2 = 6π 2 /L2 , so λgap = 3νπ 2 /L2 . Symbol Meaning
C. Role of the Spectral Gap in Inertial Manifolds H = L2 (Ω; R3 ) State Hilbert space The existence of a Lipschitz inertial manifold M of dimension V = H01 (Ω; R3 ) Energy (Sobolev) space n⋆ requires the spectral gap condition: Φ∈H MDEI state field A = −ν∆ + αI Dissipative operator (H1) µn⋆ +1 − µn⋆ > 2 ∥DP∥L∞ (A ) + ∇2 V L∞ (A ) . (66) V (Φ) Lyapunov functional (H2) P(Φ) = ∇ · (D∇Φ) + T Teleological pressure (H3) When (66) holds, every orbit is exponentially attracted to M at Σ(Φ) dW (t) Multiplicative Q-Wiener noise (H4) ⋆ rate (µn⋆ +1 − µn⋆ )/2, providing the factored dynamics on Rn {S(t)}t≥0 Markov semigroup on H described in Section 6. A = ω(B0 ) Global attractor (compact, invariant) B0 = {V ≤ 2R0 } Absorbing ball D. Coercivity of V via Poincaré Nε (A ) Stability neighbourhood From (5) and (62): M = {∇V = 0, ∇2 V > 0} Core manifold (equilibria) Z M Inertial manifold (dim n⋆ ) V (Φ) = 21 ∥∇Φ∥2 + F dx ≥ λ21 ∥Φ∥2 + c1 ∥Φ∥4L4 − c2 |Ω|. µ⋆ Invariant probability measure on A Ω (67) dimF (A ) Fractal (box-counting) dimension of A This confirms that V → ∞ as ∥Φ∥ → ∞, ensuring the existence diam(A ) Attractor diameter (intrinsic error) of a level set B0 = {V ≤ 2R0 } that is bounded in H. τ = 1/α Characteristic relaxation time R0 = K̃/β̃ Energy radius of B0 III. V ERIFICATION OF H YPOTHESES H1–H4 FOR THE n⋆ = ⌊CV ,P /α⌋ + 1 Dimension bound (Thm. 3) MDEI-SPDE λ1 First Dirichlet eigenvalue of −∆ For the reader’s convenience, we provide explicit conditions λgap Spectral gap of A under which each hypothesis is satisfied for the MDEI model. β̃ = β −C2 /2 > 0 Effective Lyapunov decay rate K̃ = K +C1 /2 ≥ 0 Effective energy forcing A. Hypothesis H1: Sectorial Dissipation γmix Exponential mixing rate A = −ν∆ + αI is sectorial for any ν > 0, α > 0 since: (a) α, ν > 0 Damping rate, diffusion coefficient −∆ is a non-negative self-adjoint operator on H01 (Ω); (b) αI β,K ≥ 0 Lyapunov rate, forcing constant shifts the spectrum by α, ensuring σ (A) ⊂ [α, ∞); (c) the re- d0 , d1 > 0 Ellipticity bounds on D(Φ) solvent estimate (λ I − A)−1 ≤ M/|λ − α| holds in the sector C1 ,C2 ≥ 0 Noise growth constants (H4) | arg(λ − α)| < π − δ for any δ ∈ (0, π) [3]. ✓ Always satisfied for any ν, α > 0. R EFERENCES B. Hypothesis H2: Lyapunov Coercivity [1] T. A. Vieira, “Formalização Algébrico-Dinâmica Apli- The dissipativity condition (6) holds if: F ∈ C2 (R3 ; R), F(Φ) ≥ cada ao Modelo MDEI: Da Tabela Vetorial ao Sistema c1 |Φ|4 −c2 , c1 > 0, and the nonlinearity satisfies the dissipativity Cognitivo-Afetivo,” ARACÊ, vol. 7, no. 7, pp. 40384– estimate F ′ (Φ) · Φ ≥ γ|Φ|4 −C for some γ > 0, C ≥ 0. This is 40397, 2025. DOI: 10.56238/arev7n7-300. satisfied by all polynomial F of degree ≥ 4 with positive leading [2] H.-J. Rudolph and T. A. Vieira, “The Mathematical Pat- coefficient. ✓ Satisfied for typical ISDM energy functions [1, 2]. tern of Coherence: Convergence between the Internal IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2026 15
State Dynamics Model and the Teleological Architec- Phys.65.851. ture,” PhilArchive, 2024. [Online]. Available: https: [21] G. Bekal, A. Pujari, and S. D. Kelly, “Continual //philarchive.org/rec/RUDTMP-3 learning with query-only attention,” arXiv preprint [3] R. Temam, Infinite-Dimensional Dynamical Systems in arXiv:2510.00365, 2024. Mechanics and Physics, 2nd ed. New York, NY, USA: [22] R. Child et al., “Generating long sequences with sparse Springer, 1997. transformers,” arXiv preprint arXiv:1904.10509, 2019. [4] J. K. Hale, Asymptotic Behavior of Dissipative Systems, Mathematical Surveys and Monographs, vol. 25. Provi- dence, RI, USA: Amer. Math. Soc., 1988. [5] J. C. Robinson, Infinite-Dimensional Dynamical Systems. Cambridge, U.K.: Cambridge Univ. Press, 2001. [6] G. Da Prato and J. Zabczyk, Stochastic Equations in Infinite Dimensions, 2nd ed. Cambridge, U.K.: Cambridge Univ. Press, 2014. DOI: 10.1017/CBO9781107295513. [7] K. Friston, “The free-energy principle: a unified brain theory?” Nature Rev. Neurosci., vol. 11, no. 2, pp. 127– 138, Feb. 2010. DOI: 10.1038/nrn2787. [8] H. Haken, Synergetics: An Introduction, 3rd ed. Berlin, Germany: Springer, 1983. [9] O. Y. L. Imanov, “Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-Tuning,” arXiv preprint arXiv:2601.18699, Jan. 2026. [Online]. Available: https://arxiv.org/abs/2601. 18699 [10] A. Vaswani et al., “Attention is all you need,” in Proc. NIPS, vol. 30, 2017, pp. 5998–6008. [11] A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret, “Transformers are RNNs: fast autoregressive transformers with linear attention,” in Proc. 37th ICML, PMLR vol. 119, 2020, pp. 5156–5165. DOI: arXiv:2006.16236. [12] J. A. Langa and J. C. Robinson, “A note on the fractal dimension of attractors of dissipative dynamical systems,” Nonlinear Anal., vol. 45, no. 2, pp. 207–217, 2001. [13] C. Li et al., “Mitigating catastrophic forgetting in lifelong learning: a hybrid architecture integrating neural ODEs with memory-augmented transformers,” Sci. Rep., vol. 15, 2025. DOI: 10.1038/s41598-025-31685-9. [14] P. Constantin and C. Foias, “Global Lyapunov exponents, Kaplan–Yorke formulas and the dimension of attractors for the 2D Navier–Stokes equations,” Commun. Pure Appl. Math., vol. 38, pp. 1–27, 1985. [15] A. Debussche, “Hausdorff dimension of a random invariant set,” J. Math. Pures Appl., vol. 77, pp. 967–988, 1998. [16] A. V. Babin and M. I. Vishik, Attractors of Evolution Equa- tions. Amsterdam, Netherlands: North-Holland, 1992. [17] I. Chueshov and I. Lasiecka, Von Kármán Evolution Equa- tions. New York, NY, USA: Springer, 2010. [18] H. Crauel and F. Flandoli, “Attractors for random dy- namical systems,” Probab. Theory Relat. Fields, vol. 100, pp. 365–393, 1994. [19] T. Wu et al., “Continual learning for large language models: a survey,” arXiv preprint arXiv:2402.01364, 2024. [20] M. C. Cross and P. C. Hohenberg, “Pattern forma- tion outside of equilibrium,” Rev. Mod. Phys., vol. 65, no. 3, pp. 851–1112, Jul. 1993. DOI: 10.1103/RevMod-