Research paper draft25 July 2026Not peer reviewed

Grounded Recurrent Control in Toy Partially Observable Worlds

A thesis-driven synthesis of recurrence, grounded valence, attention, world models, hierarchical access, maintenance, learned detours, and embodied predictive control.

Scope. This project investigates operational access consciousness: internal states made reportable, selectively available across modules, and causally available for control. It makes no claim of phenomenal experience or sentience.

This paper audits and synthesizes small control experiments that combine recurrent policies, signed outcome feedback, workspace-like routing, learned forward models, and model-predictive control.

The contribution is empirical and methodological, including an operational test of access-consciousness-like function rather than a claim about phenomenal experience. A five-seed delayed-preference POMDP isolates long-delay memory and grounded feedback; withheld navigation families test topology transfer and hidden-state dependence; engineered workspace assays test compressed report/control packets, selective broadcasting, hierarchical arbitration, learned four-context specialist selection, and causal intervention; Unity telemetry anchors predictive calibration and cross-runtime deployment. The results support a narrow thesis: temporal state, externally grounded outcome signals, selective information routing, and uncertainty-aware receding-horizon control can be composed into compact partially observable agents.

Keywords recurrent control · partial observability · access consciousness · machine consciousness · model-predictive control · workspace routing · zero-shot navigation · uncertainty · robustness · Unity

Key results at a glance

Demonstrated
0.820Accuracy at delay 28

Recurrent + grounded outcome feedback, five seeds (±0.158 SD)

98.25%Withheld-topology success

Five-seed continuous-ray benchmark across C-shape and zigzag families

41.7%Lower held-out prediction MAE

Unity telemetry calibration, 0.1267 → 0.0739

12 / 12Completed Unity courses

One recorded session; two incomplete fragments excluded from this count

1.00Workspace report accuracy

The promoted packet drove both grounded report and control across obstruction labels

6 / 6Stable resonant categories

Twenty-four seeds; 1.000 familiar retention, novel adaptation, and unknown rejection

96.39%Adaptive conductor success

Twenty seeds; 97.52% context-optimal routing across four frozen specialists

10 / 10Prospective red-rule confirmations

Isolated red pickups; 15 isolated blue controls had zero positive responses

256 / 256Masked causal-language matches

Two independent audits; 34.711% fewer visible tokens than matched JSON on seed 309

3 / 3Committed PGNW convergence

Three matched Unity pairs; passive reached the same threshold in one of three runs

0.9999999996Dynamic held-out posterior

Thirty-three post-admission updates after Gemma formulated a new live L1 candidate

Values are descriptive within the recorded assays. Seed counts, selection rules, incomplete episodes, engineered masks, and missing paired data materially limit interpretation.

A regulated division of labor

The supported architecture is not a monolithic “mind.” A recurrent controller integrates local rays, visible-goal direction, hunger, prior action, and prior reward. A body-clearance mask removes invalid moves. When a target is hidden, temporal state supports exploration; once sensor-visible, a calibrated ensemble forward model supports short-horizon planning. Separate synthetic assays test a hierarchy in which local workspaces compress specialist state, a master workspace arbitrates conflict, and regional sub-masters reduce routing load at larger scales.

01Local observationrays · goal · hunger · prior outcome
02Recurrent statehistory-dependent exploration
03aRouting gateselective shared packet
03bPredictive branchsensor-grounded MPC
04Bounded actionbody-clearance mask · re-anchor

Experimental sequence. Minimal delayed-outcome POMDP → withheld navigation topologies → workspace routing assays → Unity telemetry calibration → fixed and adaptive MPC → recorded Unity deployment → synthetic robustness sweeps.

The full thesis, not only the Unity capstone

Converging toy assays

The Functional Ego thesis is a claim about regulated composition. No single score or module constitutes the result. Earlier symbolic and neural experiments test distinct functional roles; the learned foraging and Unity work then test whether those principles survive contact with a partially observable body and environment.

01Integration is not a scoreboard

Recurrence supplies temporal coupling, but random feedback and excessive cross-talk show that more integration can reduce grounding.

02Valence needs grounding and boundaries

Externally derived progress can orient control; directly writable positive valence produces wireheading.

03Imagination must answer to reality

Ungrounded internal simulation degrades behavior; prediction checks, uncertainty limits, and sensory re-anchoring make lookahead useful.

04Attention allocates integration

Stable moments use cheap local control. Surprise, disagreement, and prediction error open broader workspace access.

05Report must share a cause with control

Self-representation matters when the same compressed state changes both explicit report and downstream action.

06External correction must be independent

Grounded and complementary observers help; echo peers can amplify confidence without adding knowledge.

07Hierarchy compresses conflict

Local, master, and regional workspaces can reduce routing load, while slow propagation turns hierarchy into bureaucracy.

08Useful intelligence lives in routing

Causal credit determines which specialist controls action in context, rather than rewarding every module uniformly.

09Integrated systems need maintenance

Active repair extends operation; offline down-selection more strongly restores recurrent separability after saturation.

10Learning tests the synthesis

Reward-grounded foraging, memory-dependent detours, hidden-state interventions, and Unity transfer test the stack beyond symbolic rules.

Defended synthesis

A Functional Ego is regulated access: which state becomes available, what it controls, how reality corrects it, and how the system remains functional over time.

Full treatment. The expanded 6,900-word paper maps each proposition to its experiments, counter-results, and evidentiary boundary. The Unity controller is the embodied capstone, not the whole thesis.

Recurrent memory + grounded outcome feedback

Demonstrated

A hidden preferred action reverses midway through each episode. The only evidence is a signed outcome pulse after each choice, followed by a variable distractor delay. The feedforward policy receives an eight-frame window; the recurrent policy uses a closely matched 32-unit GRU.

Finding

At delay 28, recurrent + outcome reached 0.820 accuracy. State reset, outcome zeroing, and outcome shuffling each returned accuracy to 0.500; sign reversal reduced it to 0.180.

Preferred-action accuracy across temporal delays.
Figure 1Preferred-action accuracy across temporal delays.Shading is ±1 across-seed standard deviation, not a confidence interval. Five training seeds; training stability remains unresolved.

Boundary. This demonstrates use of temporal state and signed task feedback in one assay. It does not establish recurrent superiority in general or biological affect.

Hierarchical selective workspace routing

Engineered toy system

Conditional-coupling and workspace-lift assays test compressed shared packets, routing cost, downstream intervention, reuse across obstruction labels, and susceptibility to false broadcast. The promoted packet reduces specialist state to five action-relevant fields: intent, problem, strategy, feeling, and confidence. Movement, memory, valence, and explicit report all read the same representation rather than receiving every raw local signal.

Hard-threshold routing retained 0.953 late accuracy with mean coupling 0.067, versus 0.976 accuracy and coupling 1.0 for always-on control. A forced packet accelerated escape but also induced a false alarm.

Level 01Local specialists

Sensorimotor, valence, memory, and context modules process local signals without broadcasting every raw state.

Level 02Master workspace

Confidence, tension, surprise, and disagreement become reportable packets for conflict arbitration.

Level 03Regional sub-masters

A further compression tier protects one global master from overload as specialist count grows.

0.714early post-shift accuracyhierarchical master
15steps to recoveryversus 16 monolithic
0.858workspace efficiencyversus 0.851 monolithic

The result is conditional, not “more levels are always better.” A deliberately slow hierarchy recovered in 29 steps with 0.781 efficiency. The assays support compressed multi-level arbitration when reduced channel load outweighs added propagation delay.

Supported

Designed local, master, and regional levels expose measurable tradeoffs among compression, conflict resolution, routing load, and delay.

Boundary

The hierarchy and routers are substantially hand-designed and evaluated in synthetic tasks. This supports engineered operational access, not phenomenal awareness.

Phase-coherent workspace binding

Software timing mechanism

Inspired by Llinas's thalamocortical synchrony hypothesis, this assay asks a narrower engineering question: can shared phase organize distributed features into one reportable control packet? Color, shape, motion, and valence modules receive identical content and emit the same number of events in every condition. The workspace receives no object identity label, so temporal coincidence is its only binding cue.

1.000coherent binding
0.312same frequency, private phase
0.302mixed-frequency binding
0.000valence phase-shift accuracy
Identical distributed content under coherent, jittered, unlocked, mixed-frequency, phase-shifted, and asynchronous timing.
Figure 3Identical distributed content under coherent, jittered, unlocked, mixed-frequency, phase-shifted, and asynchronous timing.1,600 matched trials per condition. A half-cycle shift to only the valence stream produced systematic false binding while preserving packet content and event count.

Coherent 20 Hz-labelled timing also reached 1.000 accuracy. The result therefore supports shared phase as a causal binding and routing coordinate in this explicit architecture, not a privileged role for exactly 40 Hz in software. It does not simulate neurons, thalamocortical anatomy, electromagnetic fields, biological gamma, or phenomenal consciousness.

Synchronization and turn-taking emerge under different routing demands

A second assay removes the assigned phase relationships and trains only 12 context-conditioned phase offsets. One context rewards six-way coincidence. The other requires two internally coherent packets to share a capacity-limited bus without overlapping. Payloads, amplitudes, carrier label, and bus structure remain fixed.

0.029initial routing utility
1.000learned routing utility
0.049phase-scrambled utility
0.814πpacket phase separation
Gradient-learned synchrony for feature binding and phase division for competing packets.
Figure 4Gradient-learned synchrony for feature binding and phase division for competing packets.Twenty-four matched seeds. Exact restoration and common global rotation preserved 1.000 utility; frequency mismatch reduced it to 0.143.

Random phase scrambling preserved every payload but collapsed performance; restoring the learned offsets recovered it exactly. A common rotation of all modules had no effect, identifying relative phase rather than absolute clock position as the causal variable. The no-bottleneck control retained variable inter-packet timing, so channel contention is what selected phase division.

Primary sources. Llinas and Ribary (1993), Llinas et al. (1998), Llinas et al. (2002), and Fries (2005). The Consciousness Atlas served as a discovery map, not as evidence that the theory is true.

Adaptive resonance stabilizes reportable categories

Engineered ART analogue

Grossberg's Adaptive Resonance Theory links bottom-up evidence with learned top-down expectations. In this assay, a sensory pattern proposes a category, its template must pass a vigilance match, mismatch resets the candidate and continues search, and successful resonance supplies one shared packet to learning, action, valence, and symbolic report.

1.000familiar report retention
1.000novel action adaptation
1.000unknown-pattern rejection
0.250reset-lesion familiar accuracy
Stable category learning, mismatch reset, and forced false-resonance intervention.
Figure 5Stable category learning, mismatch reset, and forced false-resonance intervention.Twenty-four seeds. Latest-sample overwrite reduced familiar category-report retention to 0.510; removing reset conflated six categories into one.

Holding sensory input fixed while forcing a poorly matching category switched both action and report on every eligible trial. This supplies a second causal route to the project's operational access claim: selectively stabilized internal content is jointly available to report and control. The false-resonance result also preserves the central warning that causal availability does not guarantee truth.

Primary sources. Grossberg (1999) and Grossberg (2017). The software implements a compact match-vigilance-reset mechanism, not Grossberg's full biological architecture, and it provides no measurement of phenomenal experience.

A learned conductor arbitrates four control regimes

Adaptive routing · synthetic

A Bach-inspired conductor receives noisy embodied-style signals without a context label. It learns from reward to allocate four frozen specialists: recurrent exploration on clear terrain, episodic ART playback for familiar hidden goals, predictive MPC for visible targets, and stable fallback during genuine physical wedges.

96.39%learned-conductor success
97.52%context-optimal routing
76.78%strongest fixed policy
20 / 20seeds beating best fixed policy
Adaptive specialist allocation, routing confusion, targeted lesions, and boundary entropy.
Figure 6Adaptive specialist allocation, routing confusion, targeted lesions, and boundary entropy.Twenty independent training seeds and 48,000 held-out evaluation steps. True context labels were withheld from the conductor.

The gate approached oracle utility while remaining stable inside contexts: unnecessary handoffs averaged 0.40 per 100 steps. Routing entropy rose briefly from 0.491 bits during steady operation to 0.947 bits at context boundaries. Scrambling the gate reduced optimal routing to 7.65%.

Targeted lesions produced matching deficits. Blocking recurrent control hurt clear terrain most; blocking ART hurt familiar hidden goals; blocking MPC hurt visible targets; and blocking fallback hurt wedges. This selective damage is stronger evidence for dynamic arbitration than a single-context "always choose ART" result.

Boundary. This is a synthetic contextual-bandit benchmark over engineered specialists and signals. Its checkpoint is Python-only and has not yet demonstrated continuous four-way arbitration in Unity, terrain episodic recall, or phenomenal consciousness.

A multi-theory access-consciousness suite

Functional · bounded

Five theory-inspired experiment families test distinct computational mechanisms rather than treating any theory's philosophical conclusion as established. Together they form converging evidence for engineered access mechanisms, with explicit negative controls and a standing boundary around phenomenal experience.

Llinas-inspired timingCausal temporal binding

Shared relative phase organized distributed features into one report/control packet. Both 20 Hz and 40 Hz worked, so coordination rather than a privileged frequency carried the result.

Pockett-inspired field codeSpatial-temporal distribution

Coupled spatial and phase coordinates bound information, while a matched random distributed code showed that redundancy, not physical wave geometry, was sufficient for fault tolerance in this assay.

Doyle-inspired playbackValence-bound precedent

Episodic retrieval was 11.6 times faster per step than MPC but less successful, supporting a cheap intermediate controller rather than a replacement for prospective planning.

Criticality-inspired regulationAdaptive gain tradeoff

Recurrent gain controlled a memory-sensitivity-stability frontier. Distinct functions peaked at different gains, and adaptive regulation did not significantly beat the best fixed-gain control.

Mathematical structuralismReport-proxy geometry

Report relations corresponded to causal state and behavior, survived coordinate relabeling, collapsed under report shuffling, and responded predictably to upstream interventions.

Defensible synthesis

The repository is a reproducible, multi-theory experimental suite for engineered access-consciousness mechanisms. It demonstrates causally active temporal binding, distributed spatial-temporal coding, valence-bound episodic retrieval, adaptive recurrent-gain regulation, and an implementation-consistent relational geometry of reportable experience proxies.

The organizational-invariance assay sharpens that synthesis. Symbolic, dense-vector, and message-passing realizations agreed at 1.000 on ordinary trajectories, interventions, and novel counterfactual inputs. An observational replay clone also matched ordinary outputs at 1.000, but failed every causally effective intervention and all 318 changed-output novel trials.

Causal organization across three software encodings compared with observational replay.
Figure 7Causal organization across three software encodings compared with observational replay.Twenty-four seeds. This is software-representation invariance inside one Python process, not physical-substrate invariance.

The relational-structure assay then measured full causal state, upstream memory and grounded valence, explicit report, and behavior as separate spaces. Upstream-state/report distance correlation reached 0.786, and upstream displacement predicted report change at AUC 0.813 versus 0.509 shuffled and 0.500 replay controls.

Relational geometry of causal state, behavior, and reportable experience proxies.
Figure 8Relational geometry of causal state, behavior, and reportable experience proxies.The high full-state correspondence is partly definitional; the upstream intervention result is the stronger causal test.

Boundary. These experiments support a functional access-consciousness architecture while remaining neutral about phenomenal experience. They do not verify the source theories wholesale or identify report proxies with qualia.

Operational access consciousness

Engineered · operational

The lab uses access consciousness in a deliberately functional sense: selected internal information enters a shared workspace, becomes available to report and multiple downstream systems, and can be shown by intervention to alter what the agent both reports and does. This is a testable architectural property, separate from claims about subjective experience.

1.00report accuracygrounded global packet
11.0–11.5steps to escapetree, rock, and mushroom pockets
19.4–22.2private-module stepswithout shared generalized report
01 · Reportability

Promoted workspace packets can generate explicit self-reports grounded in the same state used for control.

02 · Hierarchical availability

Local workspaces preserve specialist processing while compressed master packets become available to movement, memory, valence, and report.

03 · Causal control

Forced workspace interventions measurably shift action distributions and escape behavior.

04 · Flexible reuse

A generalized obstruction packet transfers across tree, rock, and mushroom-pocket labels.

05 · Falsifiability

False broadcasts can steer behavior incorrectly, showing causal power without guaranteeing truth.

06 · Resonant stability

Bottom-up/top-down match and mismatch reset preserve reportable categories while admitting novel ones.

The self-report is therefore not treated as decorative narration. In the workspace-lift and Ego Lens assays, perturbing the shared state changes both the action distribution and the report distribution. That shared causal source is the project’s evidence for grounded self-report and operational awareness-like access.

Bounded claim

The system demonstrates an engineered operational access-consciousness profile: reportable state with shared, intervention-sensitive access to downstream control.

Interpretability boundary. ego_lens_lab.py is a Jacobian-Lens-inspired explicit attribution assay, not a transformer Jacobian lens. It measures whether perturbing explicit workspace and drive variables changes action and report together; that causal alignment supports the operational claim but does not establish phenomenology.

Python-to-Unity deployment + MPC

Cross-runtime simulation

A Python controller receives local body/world telemetry over UDP and returns Unity actions. Only predictive heads were calibrated on a chronological Unity split; policy weights and recurrent memory were frozen. Held-out transition MAE fell from 0.1267 to 0.0739. In 108 matched Python courses, calibrated fixed MPC reached 95.37% success versus 90.74% for the raw recurrent baseline.

4,879usable Unity transitions
731chronological test transitions
29.74 smean completed course time
35 / 36adaptive stochastic MPC successes

A visibility-gated controller completed two cycles of six primitive Unity courses. Recurrent control handled hidden-goal exploration; MPC engaged after sensor grounding. The same file contains two short unsuccessful fragments excluded from the completed-episode count. A later 36-episode Python benchmark found adaptive stochastic MPC matched fixed MPC at 35/36 while reducing mean steps from 88.1 to 82.4.

Boundary. “12/12” describes completed episodes in one session, not all parsed fragments. The controller comparison was sequential, not randomized; adaptive MPC has not yet been validated live on the course suite.

Neuro-symbolic Tiny Scientist in Unity

Verified in synthetic Unity

A passive episode analyzer reduced frozen embodied telemetry to red/blue pickup outcomes. Gemma 3 1B then selected a causal candidate from that grounded summary; a formal compiler, rather than the language model, bound the retest action to the selected cause and derived the logical contradiction. The resulting candidate was held at zero action authority while a new Unity run tested it prospectively.

10 / 10isolated red responses
15 / 15isolated blue controls
10.890 smean red-response latency
0 / 10red counterexamples

In the pre-registered full-stack run, all ten isolated red pickups preceded a positive internal-pressure response (mean change +0.3373) after 10.890 seconds. None of the fifteen isolated blue controls produced a positive response. The restored navigation stack was active during collection: orbit recovery, AIR route guidance, resource memory, systemic routing, and bounded adaptive GNW improved access to the relevant encounters; the passive scientific GNW observer remained causally disconnected from movement.

Bounded claim

The demonstrated result is a neuro-symbolic system-level loop: grounded episode selection, language-mediated candidate formulation, formal falsifier compilation, and prospective verification in a programmed synthetic metabolism.

Boundary. Gemma did not read raw telemetry or autonomously choose the intervention. The red-to-pressure contingency is programmed in this Unity environment, so this does not establish unaided language-model abduction, natural-world causal discovery, sentience, or a rule with live motor authority. The verified rule is eligible only for future passive shadow production memory.

From a fixed hypothesis pool to live dynamic formation

A later phase replaced the fixed five-candidate ontology with a dynamic admission path. PGNW first gathered a mixed discovery set while its pool contained only spontaneous-probe and no-tested-cause baselines. Without pausing the Unity control loop, the local Gemma 3 1B adapter then formulated the compact record L1 c red k blue e + t 10.0 q 0.5. The formal layer checked its role binding, direction, delay, and vocabulary before admitting a new red-causes-probe-rise candidate at prior 0.20.

4discovery observations frozen before admission
6held-out updates to posterior 0.95
33strictly post-admission updates
20L1 output tokens

In corrected seed 154, proposal generation began at 171.232 seconds and admission completed at 178.113 seconds. Discovery observations were never reused as verification evidence. Six later clean observations raised the new candidate above 0.95 at 548.356 seconds; 33 held-out updates ended at posterior 0.999999999584. Background generation took 1.686 seconds, while the largest observed telemetry interval was 0.936 seconds. The run recorded zero critical-hunger exposure, respawns, or survival failures.

Closed embodied science loop

Discovery evidence → language-mediated hypothesis formation → formal admission → PGNW experiment selection → held-out embodied verification.

Before dynamic formation, a three-pair manipulation check tested whether PGNW authority changed evidence acquisition rather than merely observing passive wandering. Committed guidance reached posterior 0.95 in all three seeds; passive reached it in one, only 20.2 seconds before the 1,200-second limit. Censoring misses at the limit gave mean times of 633.1 versus 1,193.3 seconds. Committed control also produced eight clean red observations versus two and reduced the discarded-window fraction from 49.2% to 27.4%.

Claim boundary. Dynamic formation currently operates inside a symmetric finite red/blue/probe grammar; it is not open-ended ontology invention. Seed 154 is one successful corrected live replication, and the three committed/passive pairs are a manipulation check rather than a population-level significance result. Commitment, observation isolation, and immediate retreat were bundled, and committed mode logged 21 stuck events versus 17 passive.

A mechanistic diagnosis resolves the observed Pareto trade-off

A matched Gemma 3 1B LoRA experiment then asked whether a compact labeled causal language could retain JSON-level variable binding with fewer visible tokens. On a frozen seed-205 set of 128 held-out tables, JSON scored 128/128; L1 scored 121/128 while reducing mean visible tokens from 282.367 to 184.430. A layerwise probe and causal head ablation localized four failures to comparison- role binding; three others were grammar or serialization errors.

256 / 256masked L1 exact matches
128 / 128matched JSON and L1
−34.711%L1 visible tokens
+2.55%mask overhead vs greedy L1

The repair restricts one-path greedy decoding to a symmetric set of valid L1 continuations. It requires distinct observed cause and comparison roles but does not identify the correct causal assignment. Masked L1 scored 128/128 on seed 309, then 128/128 on untouched seed 311. On the matched seed-309 audit, JSON and masked L1 both scored 128/128 while mean visible tokens fell from 282.328 to 184.328—a 34.711% reduction. The mask added only 2.55% elapsed time versus ordinary L1.

Claim boundary. This is zero observed defects across 256 frozen new-seed cases, not universal zero-defect precision. The 49.15% observed speed advantage over JSON is descriptive because adapter runs were sequential. Broader causal distributions and embodied replication remain future work.

Syntax alone does not explain the repair

A stricter control separated generic grammar enforcement from the semantic distinct-role invariant. On untouched seed 313, a syntax-only decoder repaired three malformed records but retained four repeated cause/comparison bindings. Adding the distinct-role constraint repaired those four remaining cases, with no regressions relative to syntax-only decoding.

Untouched seed-313 constrained-decoding control
DecoderExact matchesWhat it enforces
Ordinary L1121 / 128No output constraint
Syntax-only mask124 / 128Valid L1 record grammar
Syntax + distinct roles128 / 128Grammar plus non-repeated cause/comparison roles

The same four-case semantic separation appeared in the matched seed-309 diagnostic. The confirmatory seed-313 syntax-to-role comparison was four repairs and zero regressions; its exact two-sided McNemar p-value was 0.125, so that run alone is underpowered. The result supports a bounded conclusion: grammar constraints explain part of the gain, while the engineered role invariant adds repeatable value beyond syntax. It does not show unaided perfect model binding.

Biologically inspired regulation under synthetic stress

Operational analogues

The Functional Ego stack borrows organizing ideas from biological regulation without claiming to reproduce a nervous system. Its named variables are deliberately simplified control signals: they alter reward sensitivity, urgency, sensory gain, workspace promotion, metabolic pressure, and maintenance timing. Their value is experimental separation: each signal can be swept, ablated, or intervened on while behavior and self-report are recorded.

ValenceGrounded outcome

Signed progress and cost orient action. Directly writable positive valence exposes a wireheading failure mode.

Dopamine-likeReward modulation

Mushroom reward temporarily raises the normalized signal before it decays toward an adjustable baseline.

Norepinephrine-likeUrgency and arousal

Modulates response intensity and how quickly obstruction pressure escalates into escape behavior.

Acetylcholine-likeSensory focus

Raises trust in grounded sensory evidence while suppressing disruptive internally recurrent noise.

Calcium-likeExcitability gate

Controls how readily weak signals are promoted. High gain can improve sensitivity or amplify noise into false workspace reports.

HungerMetabolic pressure

Accumulates without reward, expands food sensing, and shifts control toward exploration and survival-relevant targets.

Fatigue and sleepMaintenance regulation

Fatigue schedules active repair, successor handoff, or sleep-like offline down-selection rather than constant visible shutdown.

Grounding governorMetacognitive stabilization

Reality checks, sensory re-anchoring, hunger protection, and reliability monitoring limit runaway internal promotion.

A 4×4 sweep varied internal noise and the calcium-like excitability gate across 80 seeded episodes per cell. Survival fell from 0.6375 at noise 0.35 / calcium 1.0 to 0.0 in the strongest combined perturbation cells. At maximum noise and calcium, nine engineered stabilizers were then compared across 100 replicates.

0.00unregulated survivalmaximum noise + calcium
0.94sensory-focus survivalstrongest single stabilizer
0.74next-stack survivalwith zero false promotions
0.88offline-sleep accuracyno-sleep late accuracy: 0.25
Engineered stabilizers under maximum synthetic noise and calcium-like excitability pressure.
Figure 9Engineered stabilizers under maximum synthetic noise and calcium-like excitability pressure.Sensory focus preserved survival most strongly but retained false promotions; the next-generation stack traded some survival for stricter grounding.

Biological boundary. These mechanisms are biologically inspired computational analogues, not quantitative models of neurotransmitters, ion channels, psychiatric conditions, drug response, or subjective affect. The experiments support architectural control claims only within their defined simulations.

From spatial trajectories to musical time

Architecture transfer

The same recurrent design principles were adapted from embodied navigation to procedural symbolic music. The music policies were retrained on pitch and rhythm objectives; this is portability of architecture and learning method, not reuse of the navigation checkpoint.

0.520recurrent delayed-motif accuracyfeedforward: 0.380
0.592recurrent final-return accuracyfeedforward: 0.454
0.352final return after memory resetdown from 0.592
0.576learned rhythm objectiveweighted random: 0.473
PitchRoot · scale · octave
TimeBPM · learned or stochastic rhythm
MemoryMotif recall · novelty · reset ablation
OutputIAC bus · DAW · MIDI instrument

Compatibility. Apple Silicon and macOS 13+; produces MIDI rather than audio. Route it to the macOS IAC Driver or another MIDI destination. This research build is ad-hoc signed but not Apple-notarized, so first launch may require Control-click → Open.

Boundary. Recurrent memory improved delayed motif and final-return measures, but did not win every metric: the feedforward model had higher immediate next-note accuracy. Learned rhythm was optimized under engineered musical rewards, not discovered without an objective.

Demonstrations and hypotheses are not interchangeable

A

Demonstrated in the recorded systems

  1. Recurrent state carries signed task outcome beyond an eight-frame feedforward context in the delayed-preference assay.
  2. Memory is causally necessary for the trained recurrent policies in specific detour families, especially zigzag gates.
  3. The recurrent design transfers at the architecture level to symbolic music after domain-specific retraining; navigation weights are not reused.
  4. A compressed workspace packet carrying intent, problem, strategy, feeling, and confidence jointly alters report and downstream control.
  5. A Python policy and calibrated predictive controller execute through a Unity UDP bridge; this is cross-runtime simulation, not sim-to-real transfer.
  6. Biologically inspired regulatory analogues expose distinct operational roles for reward, urgency, sensory gain, excitability, hunger, fatigue, and sleep-like maintenance in defined toy assays.
  7. Synthetic perturbation sweeps expose defined failure regions, and engineered governors partially restore function within that simulator.
  8. Related control roles are implemented across binary transition systems, symbolic routers, learned neural policies, Unity, and retrained symbolic MIDI tasks.
  9. ART-like vigilance and mismatch reset stabilize learned categories while one resonant packet supplies both symbolic report and behavioral control.
  10. A reward-trained conductor dynamically routes clear terrain, familiar hidden goals, visible targets, and physical wedges to recurrent, episodic, predictive, and fallback specialists.
  11. Three independent software encodings preserve ordinary, intervention, and counterfactual behavior while an observational replay clone fails causal tests.
  12. Upstream memory and grounded-valence changes predict changes in the relational geometry of explicit report and behavior.
  13. A bounded neuro-symbolic Tiny Scientist can select a causal candidate from grounded episode summaries, use formal code for falsifier binding, and prospectively verify that candidate in the synthetic Unity metabolism.
  14. A mechanism-guided finite-state decoder removed the compact causal language's observed binding and serialization failures across two independent 128-case audits without encoding the correct causal assignment.
  15. A live Unity loop collected discovery evidence, asynchronously asked Gemma to formulate an L1 causal candidate, admitted it into a null-only hypothesis pool, selected later experiments through PGNW, and verified it only from post-admission observations.
H

Hypotheses consistent with the evidence

  1. Recurrent memory and grounded outcome feedback may be complementary primitives for partially observable control.
  2. Dynamic workspace routing may be more efficient than constant global coupling when conflict is sparse.
  3. A controller may benefit from memory-driven exploration when goals are hidden and predictive optimization when targets are sensor-grounded.
  4. Ensemble disagreement may be a useful operational signal for truncating imagined rollouts.
  5. These motifs may be substrate-independent at the computational level; this is a design hypothesis, not evidence about phenomenology.
Read the full 47-claim evidence ledger

What is established, and what comes next

The present artifact supports compact recurrent control, memory-dependent transfer, predictive navigation, and engineered operational access consciousness in toy simulated systems. The next milestone is to broaden and independently replicate those results, not to stretch them into claims of phenomenal consciousness or AGI.

Learned routing
Replace more hand-designed promotion and safety logic with learned, ablated routing while preserving report-control alignment.
Broader transfer
Replicate the initial MIDI architecture transfer and test additional temporal domains without assuming that navigation competence transfers automatically.
Independent replication
Freeze protocols, preregister primary metrics, preserve exact code-output provenance, and add confidence intervals across more seeds.
Phenomenal boundary
Operational access is a functional engineering claim. The experiments provide no measurement of subjective experience or sentience.
Embodiment scope
The included asset-free Unity course is reproducible simulation; physical-world sim-to-real transfer remains future work.

Read, audit, and reproduce

Tiny Consciousness Lab. “Grounded Recurrent Control in Toy Partially Observable Worlds.” Research paper draft, 20 July 2026.