This paper audits and synthesizes small control experiments that combine recurrent policies, signed outcome feedback, workspace-like routing, learned forward models, and model-predictive control.
The contribution is empirical and methodological, including an operational test of access-consciousness-like function rather than a claim about phenomenal experience. A five-seed delayed-preference POMDP isolates long-delay memory and grounded feedback; withheld navigation families test topology transfer and hidden-state dependence; engineered workspace assays test compressed report/control packets, selective broadcasting, hierarchical arbitration, learned four-context specialist selection, and causal intervention; Unity telemetry anchors predictive calibration and cross-runtime deployment. The results support a narrow thesis: temporal state, externally grounded outcome signals, selective information routing, and uncertainty-aware receding-horizon control can be composed into compact partially observable agents.
Keywords recurrent control · partial observability · access consciousness · machine consciousness · model-predictive control · workspace routing · zero-shot navigation · uncertainty · robustness · Unity
Key results at a glance
Recurrent + grounded outcome feedback, five seeds (±0.158 SD)
Five-seed continuous-ray benchmark across C-shape and zigzag families
Unity telemetry calibration, 0.1267 → 0.0739
One recorded session; two incomplete fragments excluded from this count
The promoted packet drove both grounded report and control across obstruction labels
Twenty-four seeds; 1.000 familiar retention, novel adaptation, and unknown rejection
Twenty seeds; 97.52% context-optimal routing across four frozen specialists
Isolated red pickups; 15 isolated blue controls had zero positive responses
Two independent audits; 34.711% fewer visible tokens than matched JSON on seed 309
Three matched Unity pairs; passive reached the same threshold in one of three runs
Thirty-three post-admission updates after Gemma formulated a new live L1 candidate
Values are descriptive within the recorded assays. Seed counts, selection rules, incomplete episodes, engineered masks, and missing paired data materially limit interpretation.
A regulated division of labor
The supported architecture is not a monolithic “mind.” A recurrent controller integrates local rays, visible-goal direction, hunger, prior action, and prior reward. A body-clearance mask removes invalid moves. When a target is hidden, temporal state supports exploration; once sensor-visible, a calibrated ensemble forward model supports short-horizon planning. Separate synthetic assays test a hierarchy in which local workspaces compress specialist state, a master workspace arbitrates conflict, and regional sub-masters reduce routing load at larger scales.
Experimental sequence. Minimal delayed-outcome POMDP → withheld navigation topologies → workspace routing assays → Unity telemetry calibration → fixed and adaptive MPC → recorded Unity deployment → synthetic robustness sweeps.
The full thesis, not only the Unity capstone
The Functional Ego thesis is a claim about regulated composition. No single score or module constitutes the result. Earlier symbolic and neural experiments test distinct functional roles; the learned foraging and Unity work then test whether those principles survive contact with a partially observable body and environment.
Recurrence supplies temporal coupling, but random feedback and excessive cross-talk show that more integration can reduce grounding.
Externally derived progress can orient control; directly writable positive valence produces wireheading.
Ungrounded internal simulation degrades behavior; prediction checks, uncertainty limits, and sensory re-anchoring make lookahead useful.
Stable moments use cheap local control. Surprise, disagreement, and prediction error open broader workspace access.
Self-representation matters when the same compressed state changes both explicit report and downstream action.
Grounded and complementary observers help; echo peers can amplify confidence without adding knowledge.
Local, master, and regional workspaces can reduce routing load, while slow propagation turns hierarchy into bureaucracy.
Causal credit determines which specialist controls action in context, rather than rewarding every module uniformly.
Active repair extends operation; offline down-selection more strongly restores recurrent separability after saturation.
Reward-grounded foraging, memory-dependent detours, hidden-state interventions, and Unity transfer test the stack beyond symbolic rules.
A Functional Ego is regulated access: which state becomes available, what it controls, how reality corrects it, and how the system remains functional over time.
Full treatment. The expanded 6,900-word paper maps each proposition to its experiments, counter-results, and evidentiary boundary. The Unity controller is the embodied capstone, not the whole thesis.
Recurrent memory + grounded outcome feedback
A hidden preferred action reverses midway through each episode. The only evidence is a signed outcome pulse after each choice, followed by a variable distractor delay. The feedforward policy receives an eight-frame window; the recurrent policy uses a closely matched 32-unit GRU.
At delay 28, recurrent + outcome reached 0.820 accuracy. State reset, outcome zeroing, and outcome shuffling each returned accuracy to 0.500; sign reversal reduced it to 0.180.

Boundary. This demonstrates use of temporal state and signed task feedback in one assay. It does not establish recurrent superiority in general or biological affect.
Hierarchical selective workspace routing
Conditional-coupling and workspace-lift assays test compressed shared packets, routing cost, downstream intervention, reuse across obstruction labels, and susceptibility to false broadcast. The promoted packet reduces specialist state to five action-relevant fields: intent, problem, strategy, feeling, and confidence. Movement, memory, valence, and explicit report all read the same representation rather than receiving every raw local signal.
Hard-threshold routing retained 0.953 late accuracy with mean coupling 0.067, versus 0.976 accuracy and coupling 1.0 for always-on control. A forced packet accelerated escape but also induced a false alarm.
Sensorimotor, valence, memory, and context modules process local signals without broadcasting every raw state.
Confidence, tension, surprise, and disagreement become reportable packets for conflict arbitration.
A further compression tier protects one global master from overload as specialist count grows.
The result is conditional, not “more levels are always better.” A deliberately slow hierarchy recovered in 29 steps with 0.781 efficiency. The assays support compressed multi-level arbitration when reduced channel load outweighs added propagation delay.
Designed local, master, and regional levels expose measurable tradeoffs among compression, conflict resolution, routing load, and delay.
The hierarchy and routers are substantially hand-designed and evaluated in synthetic tasks. This supports engineered operational access, not phenomenal awareness.
Phase-coherent workspace binding
Inspired by Llinas's thalamocortical synchrony hypothesis, this assay asks a narrower engineering question: can shared phase organize distributed features into one reportable control packet? Color, shape, motion, and valence modules receive identical content and emit the same number of events in every condition. The workspace receives no object identity label, so temporal coincidence is its only binding cue.

Coherent 20 Hz-labelled timing also reached 1.000 accuracy. The result therefore supports shared phase as a causal binding and routing coordinate in this explicit architecture, not a privileged role for exactly 40 Hz in software. It does not simulate neurons, thalamocortical anatomy, electromagnetic fields, biological gamma, or phenomenal consciousness.
Synchronization and turn-taking emerge under different routing demands
A second assay removes the assigned phase relationships and trains only 12 context-conditioned phase offsets. One context rewards six-way coincidence. The other requires two internally coherent packets to share a capacity-limited bus without overlapping. Payloads, amplitudes, carrier label, and bus structure remain fixed.

Random phase scrambling preserved every payload but collapsed performance; restoring the learned offsets recovered it exactly. A common rotation of all modules had no effect, identifying relative phase rather than absolute clock position as the causal variable. The no-bottleneck control retained variable inter-packet timing, so channel contention is what selected phase division.
Primary sources. Llinas and Ribary (1993), Llinas et al. (1998), Llinas et al. (2002), and Fries (2005). The Consciousness Atlas served as a discovery map, not as evidence that the theory is true.
Adaptive resonance stabilizes reportable categories
Grossberg's Adaptive Resonance Theory links bottom-up evidence with learned top-down expectations. In this assay, a sensory pattern proposes a category, its template must pass a vigilance match, mismatch resets the candidate and continues search, and successful resonance supplies one shared packet to learning, action, valence, and symbolic report.

Holding sensory input fixed while forcing a poorly matching category switched both action and report on every eligible trial. This supplies a second causal route to the project's operational access claim: selectively stabilized internal content is jointly available to report and control. The false-resonance result also preserves the central warning that causal availability does not guarantee truth.
Primary sources. Grossberg (1999) and Grossberg (2017). The software implements a compact match-vigilance-reset mechanism, not Grossberg's full biological architecture, and it provides no measurement of phenomenal experience.
A learned conductor arbitrates four control regimes
A Bach-inspired conductor receives noisy embodied-style signals without a context label. It learns from reward to allocate four frozen specialists: recurrent exploration on clear terrain, episodic ART playback for familiar hidden goals, predictive MPC for visible targets, and stable fallback during genuine physical wedges.

The gate approached oracle utility while remaining stable inside contexts: unnecessary handoffs averaged 0.40 per 100 steps. Routing entropy rose briefly from 0.491 bits during steady operation to 0.947 bits at context boundaries. Scrambling the gate reduced optimal routing to 7.65%.
Targeted lesions produced matching deficits. Blocking recurrent control hurt clear terrain most; blocking ART hurt familiar hidden goals; blocking MPC hurt visible targets; and blocking fallback hurt wedges. This selective damage is stronger evidence for dynamic arbitration than a single-context "always choose ART" result.
Boundary. This is a synthetic contextual-bandit benchmark over engineered specialists and signals. Its checkpoint is Python-only and has not yet demonstrated continuous four-way arbitration in Unity, terrain episodic recall, or phenomenal consciousness.
A multi-theory access-consciousness suite
Five theory-inspired experiment families test distinct computational mechanisms rather than treating any theory's philosophical conclusion as established. Together they form converging evidence for engineered access mechanisms, with explicit negative controls and a standing boundary around phenomenal experience.
Shared relative phase organized distributed features into one report/control packet. Both 20 Hz and 40 Hz worked, so coordination rather than a privileged frequency carried the result.
Coupled spatial and phase coordinates bound information, while a matched random distributed code showed that redundancy, not physical wave geometry, was sufficient for fault tolerance in this assay.
Episodic retrieval was 11.6 times faster per step than MPC but less successful, supporting a cheap intermediate controller rather than a replacement for prospective planning.
Recurrent gain controlled a memory-sensitivity-stability frontier. Distinct functions peaked at different gains, and adaptive regulation did not significantly beat the best fixed-gain control.
Report relations corresponded to causal state and behavior, survived coordinate relabeling, collapsed under report shuffling, and responded predictably to upstream interventions.
The repository is a reproducible, multi-theory experimental suite for engineered access-consciousness mechanisms. It demonstrates causally active temporal binding, distributed spatial-temporal coding, valence-bound episodic retrieval, adaptive recurrent-gain regulation, and an implementation-consistent relational geometry of reportable experience proxies.
The organizational-invariance assay sharpens that synthesis. Symbolic, dense-vector, and message-passing realizations agreed at 1.000 on ordinary trajectories, interventions, and novel counterfactual inputs. An observational replay clone also matched ordinary outputs at 1.000, but failed every causally effective intervention and all 318 changed-output novel trials.

The relational-structure assay then measured full causal state, upstream memory and grounded valence, explicit report, and behavior as separate spaces. Upstream-state/report distance correlation reached 0.786, and upstream displacement predicted report change at AUC 0.813 versus 0.509 shuffled and 0.500 replay controls.

Boundary. These experiments support a functional access-consciousness architecture while remaining neutral about phenomenal experience. They do not verify the source theories wholesale or identify report proxies with qualia.
Operational access consciousness
The lab uses access consciousness in a deliberately functional sense: selected internal information enters a shared workspace, becomes available to report and multiple downstream systems, and can be shown by intervention to alter what the agent both reports and does. This is a testable architectural property, separate from claims about subjective experience.
Promoted workspace packets can generate explicit self-reports grounded in the same state used for control.
Local workspaces preserve specialist processing while compressed master packets become available to movement, memory, valence, and report.
Forced workspace interventions measurably shift action distributions and escape behavior.
A generalized obstruction packet transfers across tree, rock, and mushroom-pocket labels.
False broadcasts can steer behavior incorrectly, showing causal power without guaranteeing truth.
Bottom-up/top-down match and mismatch reset preserve reportable categories while admitting novel ones.
The self-report is therefore not treated as decorative narration. In the workspace-lift and Ego Lens assays, perturbing the shared state changes both the action distribution and the report distribution. That shared causal source is the project’s evidence for grounded self-report and operational awareness-like access.
The system demonstrates an engineered operational access-consciousness profile: reportable state with shared, intervention-sensitive access to downstream control.
Interpretability boundary. ego_lens_lab.py is a Jacobian-Lens-inspired explicit attribution assay, not a transformer Jacobian lens. It measures whether perturbing explicit workspace and drive variables changes action and report together; that causal alignment supports the operational claim but does not establish phenomenology.
Python-to-Unity deployment + MPC
A Python controller receives local body/world telemetry over UDP and returns Unity actions. Only predictive heads were calibrated on a chronological Unity split; policy weights and recurrent memory were frozen. Held-out transition MAE fell from 0.1267 to 0.0739. In 108 matched Python courses, calibrated fixed MPC reached 95.37% success versus 90.74% for the raw recurrent baseline.
A visibility-gated controller completed two cycles of six primitive Unity courses. Recurrent control handled hidden-goal exploration; MPC engaged after sensor grounding. The same file contains two short unsuccessful fragments excluded from the completed-episode count. A later 36-episode Python benchmark found adaptive stochastic MPC matched fixed MPC at 35/36 while reducing mean steps from 88.1 to 82.4.
Boundary. “12/12” describes completed episodes in one session, not all parsed fragments. The controller comparison was sequential, not randomized; adaptive MPC has not yet been validated live on the course suite.
Neuro-symbolic Tiny Scientist in Unity
A passive episode analyzer reduced frozen embodied telemetry to red/blue pickup outcomes. Gemma 3 1B then selected a causal candidate from that grounded summary; a formal compiler, rather than the language model, bound the retest action to the selected cause and derived the logical contradiction. The resulting candidate was held at zero action authority while a new Unity run tested it prospectively.
In the pre-registered full-stack run, all ten isolated red pickups preceded a positive internal-pressure response (mean change +0.3373) after 10.890 seconds. None of the fifteen isolated blue controls produced a positive response. The restored navigation stack was active during collection: orbit recovery, AIR route guidance, resource memory, systemic routing, and bounded adaptive GNW improved access to the relevant encounters; the passive scientific GNW observer remained causally disconnected from movement.
The demonstrated result is a neuro-symbolic system-level loop: grounded episode selection, language-mediated candidate formulation, formal falsifier compilation, and prospective verification in a programmed synthetic metabolism.
Boundary. Gemma did not read raw telemetry or autonomously choose the intervention. The red-to-pressure contingency is programmed in this Unity environment, so this does not establish unaided language-model abduction, natural-world causal discovery, sentience, or a rule with live motor authority. The verified rule is eligible only for future passive shadow production memory.
From a fixed hypothesis pool to live dynamic formation
A later phase replaced the fixed five-candidate ontology with a dynamic admission path. PGNW first gathered a mixed discovery set while its pool contained only spontaneous-probe and no-tested-cause baselines. Without pausing the Unity control loop, the local Gemma 3 1B adapter then formulated the compact record L1 c red k blue e + t 10.0 q 0.5. The formal layer checked its role binding, direction, delay, and vocabulary before admitting a new red-causes-probe-rise candidate at prior 0.20.
In corrected seed 154, proposal generation began at 171.232 seconds and admission completed at 178.113 seconds. Discovery observations were never reused as verification evidence. Six later clean observations raised the new candidate above 0.95 at 548.356 seconds; 33 held-out updates ended at posterior 0.999999999584. Background generation took 1.686 seconds, while the largest observed telemetry interval was 0.936 seconds. The run recorded zero critical-hunger exposure, respawns, or survival failures.
Discovery evidence → language-mediated hypothesis formation → formal admission → PGNW experiment selection → held-out embodied verification.
Before dynamic formation, a three-pair manipulation check tested whether PGNW authority changed evidence acquisition rather than merely observing passive wandering. Committed guidance reached posterior 0.95 in all three seeds; passive reached it in one, only 20.2 seconds before the 1,200-second limit. Censoring misses at the limit gave mean times of 633.1 versus 1,193.3 seconds. Committed control also produced eight clean red observations versus two and reduced the discarded-window fraction from 49.2% to 27.4%.
Claim boundary. Dynamic formation currently operates inside a symmetric finite red/blue/probe grammar; it is not open-ended ontology invention. Seed 154 is one successful corrected live replication, and the three committed/passive pairs are a manipulation check rather than a population-level significance result. Commitment, observation isolation, and immediate retreat were bundled, and committed mode logged 21 stuck events versus 17 passive.
A mechanistic diagnosis resolves the observed Pareto trade-off
A matched Gemma 3 1B LoRA experiment then asked whether a compact labeled causal language could retain JSON-level variable binding with fewer visible tokens. On a frozen seed-205 set of 128 held-out tables, JSON scored 128/128; L1 scored 121/128 while reducing mean visible tokens from 282.367 to 184.430. A layerwise probe and causal head ablation localized four failures to comparison- role binding; three others were grammar or serialization errors.
The repair restricts one-path greedy decoding to a symmetric set of valid L1 continuations. It requires distinct observed cause and comparison roles but does not identify the correct causal assignment. Masked L1 scored 128/128 on seed 309, then 128/128 on untouched seed 311. On the matched seed-309 audit, JSON and masked L1 both scored 128/128 while mean visible tokens fell from 282.328 to 184.328—a 34.711% reduction. The mask added only 2.55% elapsed time versus ordinary L1.
Claim boundary. This is zero observed defects across 256 frozen new-seed cases, not universal zero-defect precision. The 49.15% observed speed advantage over JSON is descriptive because adapter runs were sequential. Broader causal distributions and embodied replication remain future work.
Syntax alone does not explain the repair
A stricter control separated generic grammar enforcement from the semantic distinct-role invariant. On untouched seed 313, a syntax-only decoder repaired three malformed records but retained four repeated cause/comparison bindings. Adding the distinct-role constraint repaired those four remaining cases, with no regressions relative to syntax-only decoding.
| Decoder | Exact matches | What it enforces |
|---|---|---|
| Ordinary L1 | 121 / 128 | No output constraint |
| Syntax-only mask | 124 / 128 | Valid L1 record grammar |
| Syntax + distinct roles | 128 / 128 | Grammar plus non-repeated cause/comparison roles |
The same four-case semantic separation appeared in the matched seed-309 diagnostic. The confirmatory seed-313 syntax-to-role comparison was four repairs and zero regressions; its exact two-sided McNemar p-value was 0.125, so that run alone is underpowered. The result supports a bounded conclusion: grammar constraints explain part of the gain, while the engineered role invariant adds repeatable value beyond syntax. It does not show unaided perfect model binding.
Biologically inspired regulation under synthetic stress
The Functional Ego stack borrows organizing ideas from biological regulation without claiming to reproduce a nervous system. Its named variables are deliberately simplified control signals: they alter reward sensitivity, urgency, sensory gain, workspace promotion, metabolic pressure, and maintenance timing. Their value is experimental separation: each signal can be swept, ablated, or intervened on while behavior and self-report are recorded.
Signed progress and cost orient action. Directly writable positive valence exposes a wireheading failure mode.
Mushroom reward temporarily raises the normalized signal before it decays toward an adjustable baseline.
Modulates response intensity and how quickly obstruction pressure escalates into escape behavior.
Raises trust in grounded sensory evidence while suppressing disruptive internally recurrent noise.
Controls how readily weak signals are promoted. High gain can improve sensitivity or amplify noise into false workspace reports.
Accumulates without reward, expands food sensing, and shifts control toward exploration and survival-relevant targets.
Fatigue schedules active repair, successor handoff, or sleep-like offline down-selection rather than constant visible shutdown.
Reality checks, sensory re-anchoring, hunger protection, and reliability monitoring limit runaway internal promotion.
A 4×4 sweep varied internal noise and the calcium-like excitability gate across 80 seeded episodes per cell. Survival fell from 0.6375 at noise 0.35 / calcium 1.0 to 0.0 in the strongest combined perturbation cells. At maximum noise and calcium, nine engineered stabilizers were then compared across 100 replicates.

Biological boundary. These mechanisms are biologically inspired computational analogues, not quantitative models of neurotransmitters, ion channels, psychiatric conditions, drug response, or subjective affect. The experiments support architectural control claims only within their defined simulations.
From spatial trajectories to musical time
The same recurrent design principles were adapted from embodied navigation to procedural symbolic music. The music policies were retrained on pitch and rhythm objectives; this is portability of architecture and learning method, not reuse of the navigation checkpoint.
Compatibility. Apple Silicon and macOS 13+; produces MIDI rather than audio. Route it to the macOS IAC Driver or another MIDI destination. This research build is ad-hoc signed but not Apple-notarized, so first launch may require Control-click → Open.
Boundary. Recurrent memory improved delayed motif and final-return measures, but did not win every metric: the feedforward model had higher immediate next-note accuracy. Learned rhythm was optimized under engineered musical rewards, not discovered without an objective.
Demonstrations and hypotheses are not interchangeable
Demonstrated in the recorded systems
- Recurrent state carries signed task outcome beyond an eight-frame feedforward context in the delayed-preference assay.
- Memory is causally necessary for the trained recurrent policies in specific detour families, especially zigzag gates.
- The recurrent design transfers at the architecture level to symbolic music after domain-specific retraining; navigation weights are not reused.
- A compressed workspace packet carrying intent, problem, strategy, feeling, and confidence jointly alters report and downstream control.
- A Python policy and calibrated predictive controller execute through a Unity UDP bridge; this is cross-runtime simulation, not sim-to-real transfer.
- Biologically inspired regulatory analogues expose distinct operational roles for reward, urgency, sensory gain, excitability, hunger, fatigue, and sleep-like maintenance in defined toy assays.
- Synthetic perturbation sweeps expose defined failure regions, and engineered governors partially restore function within that simulator.
- Related control roles are implemented across binary transition systems, symbolic routers, learned neural policies, Unity, and retrained symbolic MIDI tasks.
- ART-like vigilance and mismatch reset stabilize learned categories while one resonant packet supplies both symbolic report and behavioral control.
- A reward-trained conductor dynamically routes clear terrain, familiar hidden goals, visible targets, and physical wedges to recurrent, episodic, predictive, and fallback specialists.
- Three independent software encodings preserve ordinary, intervention, and counterfactual behavior while an observational replay clone fails causal tests.
- Upstream memory and grounded-valence changes predict changes in the relational geometry of explicit report and behavior.
- A bounded neuro-symbolic Tiny Scientist can select a causal candidate from grounded episode summaries, use formal code for falsifier binding, and prospectively verify that candidate in the synthetic Unity metabolism.
- A mechanism-guided finite-state decoder removed the compact causal language's observed binding and serialization failures across two independent 128-case audits without encoding the correct causal assignment.
- A live Unity loop collected discovery evidence, asynchronously asked Gemma to formulate an L1 causal candidate, admitted it into a null-only hypothesis pool, selected later experiments through PGNW, and verified it only from post-admission observations.
Hypotheses consistent with the evidence
- Recurrent memory and grounded outcome feedback may be complementary primitives for partially observable control.
- Dynamic workspace routing may be more efficient than constant global coupling when conflict is sparse.
- A controller may benefit from memory-driven exploration when goals are hidden and predictive optimization when targets are sensor-grounded.
- Ensemble disagreement may be a useful operational signal for truncating imagined rollouts.
- These motifs may be substrate-independent at the computational level; this is a design hypothesis, not evidence about phenomenology.
What is established, and what comes next
The present artifact supports compact recurrent control, memory-dependent transfer, predictive navigation, and engineered operational access consciousness in toy simulated systems. The next milestone is to broaden and independently replicate those results, not to stretch them into claims of phenomenal consciousness or AGI.
- Learned routing
- Replace more hand-designed promotion and safety logic with learned, ablated routing while preserving report-control alignment.
- Broader transfer
- Replicate the initial MIDI architecture transfer and test additional temporal domains without assuming that navigation competence transfers automatically.
- Independent replication
- Freeze protocols, preregister primary metrics, preserve exact code-output provenance, and add confidence intervals across more seeds.
- Phenomenal boundary
- Operational access is a functional engineering claim. The experiments provide no measurement of subjective experience or sentience.
- Embodiment scope
- The included asset-free Unity course is reproducible simulation; physical-world sim-to-real transfer remains future work.
Read, audit, and reproduce
Tiny Consciousness Lab. “Grounded Recurrent Control in Toy Partially Observable Worlds.” Research paper draft, 20 July 2026.
