The Hierarchy and the Apex
The Last Science, Part 4 of 10
So far in this series: neurons fire spikes, recurrent cortical circuits perform delay coordinate embedding on those spikes to reconstruct the structure of reality, and consciousness is the process of that reconstruction — a verb, not a noun. Consciousness isn’t a state. It is a process. It is the temporal unfolding.
Fine, you might say. If consciousness is delay coordinate embedding in recurrent circuits, and your brain is full of recurrent circuits doing delay coordinate embedding — in your retina, your spinal cord, your cerebellum, your early visual cortex — why aren’t all of those conscious? Why don’t you have a separate little consciousness in your retina, experiencing edges? Why isn’t your cerebellum having its own private experience of motor coordination?
This is the threshold problem. And if I can’t answer it, the theory is in trouble, because a theory that says everything with recurrence is conscious is just panpsychism wearing a lab coat.
The answer is hierarchy.
DCE engines don’t work in isolation. They’re stacked. The outputs of lower-level engines become the inputs to higher-level engines, which perform their own delay coordinate embedding on those signals. It’s DCE all the way up.
At the bottom level, you have engines doing the basic work. A circuit in V1 embeds edge orientations. A circuit in the cochlear nucleus embeds sound frequencies. These are building the raw materials — the palette, not the painting.
One level up, engines in V4 are taking outputs from those V1 circuits and embedding them. Now you’re not reconstructing edges — you’re reconstructing how edges relate to each other over time and space. Objects start to emerge. A particular neuron in your fusiform face area doesn’t just respond to a collection of features. It’s reconstructing how face-related activity in lower areas unfolds over time. It’s doing DCE on DCE.
Higher still, and you get engines that are integrating across completely different sensory modalities. The sound of a voice and the sight of a mouth moving get bound together — not because someone wired them up, but because a higher-order DCE engine is embedding the joint dynamics of auditory and visual streams, and in that embedding, the correlation structure is preserved. The binding is the embedding.
At the top of this hierarchy, you get the most abstract representations. Goals. Plans. The sense of being a self who is experiencing all of this. These highest-order engines are integrating information from the entire hierarchy beneath them, producing representations so abstract that they’re no longer about any single sensory modality. They’re about — for lack of a better word — you.
My claim: consciousness happens at the top. Not at every level. At the apex. Or apexes, if we’re being really specific (more on apexes and what that means for decision-making and free will down below).
Why the apex and not the whole hierarchy? Because of what happens at the top that doesn’t happen lower down — and the difference isn’t complexity. It’s direction.
Lower-level DCE engines do real computational work. Your retinal circuits, your V1 edge detectors — they’re embedding, they’re reconstructing, they’re processing. But they’re processing something else. Edges. Frequencies. Colors. They represent aspects of the external world. They don’t represent themselves. A circuit in V1 that reconstructs edge orientations has no model of the fact that it’s reconstructing edge orientations. It just does it. Brilliantly, but blindly.
At the highest levels, something different happens. Some of the inputs coming up the hierarchy aren’t about the external world anymore. They’re about the system’s own internal states — attention, arousal, the gain signals that determine which reconstructions are currently winning the competition. So when a highest-order DCE engine embeds those signals, it’s not reconstructing the world. It’s reconstructing itself. Its own processing. Its own dynamics.
And here’s the critical move: that self-reconstruction doesn’t just sit there passively. It feeds back into the process it’s modeling. The system’s model of what it’s attending to shapes what it attends to next. The map doesn’t just describe the territory — the map is part of the territory, actively influencing the landscape it’s trying to represent.
That’s self-reference. Not in the trivial sense of having a self-model — a thermostat has a model of room temperature, but it doesn’t model its own modeling, and its “model” doesn’t feed back into what it’s measuring. Self-reference in the deep sense: a system whose representation of its own representational activity participates in generating what it represents. A strange loop, in Hofstadter’s sense. And it’s the strange loop, I argue, that constitutes subjectivity.
Three ideas converge here from three very different thinkers:
Douglas Hofstadter argued that consciousness arises when a system’s self-model becomes causally entangled with the processes it represents — when you can’t pull them apart because the map is also part of the territory.
Thomas Metzinger showed that we don’t experience our self-model as a model. We live through it. We don’t see a representation of the world — we see the world. The model is transparent. And that transparency is what subjectivity feels like from the inside.
Michael Graziano proposed that when a system models its own attention — builds a simplified representation of what it’s attending to and why — it attributes “awareness” to itself. The modeling creates the functional property that, experienced from within, is consciousness.
Under SST, all three of these are descriptions of what highest-order DCE engines do. Self-referential embedding, transparent self-modeling, attention schema — these are the same computational phenomenon described at different levels. And the mechanism is the one we’ve been building across this series: delay coordinate embedding, performed hierarchically, with the highest levels embedding the system’s own dynamics.
That’s the threshold. Not complexity. Not neuron count. Not energy consumption. Self-referential hierarchical depth.
One more piece of the puzzle: not all of those highest-order representations make it into consciousness at once. There’s a selection mechanism.
Think about what it’s like to be at a crowded party. There are dozens of conversations, music playing, someone laughing across the room, the temperature, the drink in your hand. All of this sensory data is being processed. DCE engines throughout your cortex are embedding all of it. But you’re only conscious of some of it — the conversation you’re in, the song you just recognized, the person who caught your eye.
What determines which reconstructions become conscious content and which stay in the background? Gain modulation. Higher-level circuits amplify some DCE reconstructions and suppress others. When a reconstruction gets amplified — boosted in gain — it dominates the competitive dynamics at the highest level. It achieves what Daniel Dennett called “fame in the brain”: widespread influence on subsequent processing.
The sources of gain are varied. Something moving fast in your peripheral vision gets a salience boost — that’s bottom-up. You can also deliberately direct attention — that’s top-down, your prefrontal circuits adjusting gain on lower-level engines. Emotional significance cranks the gain through limbic projections: you’ll notice your child’s voice in the noisiest room on earth.
This explains a whole catalog of perceptual phenomena. Binocular rivalry — when your two eyes see incompatible images and perception flips between them — is two DCE reconstructions competing for gain, neither able to permanently suppress the other. Inattentional blindness — failing to see a gorilla walk through a basketball game — is a DCE reconstruction that exists but never achieves enough gain to influence the highest-order engines. The gorilla was processed. It just never became conscious content.
And here’s where the Necker cube comes in.
This is a famous ‘bistable illusion’ — look at it long enough, and you’ll see two versions of the cube: one facing down and to the left, one facing up and to the right.
Under SST, those are two different ‘attractor basins’ in the state space of a DCE engine — basically, two different trajectories that your DCE engines carve out under this stimulus. However, the trajectories fall within the same attractor ‘set’ — your neurons can only go down one path or the other, not both.
The trajectory falls into one basin or the other depending on initial conditions — where your eyes fixate, what you were just looking at, tiny fluctuations in neural activity. The flip happens when the trajectory escapes one basin and is captured by the other. It’s not your interpretation that changes. It’s the dynamical trajectory through the quality space of shape perception.
Now, on “apex” vs “apexes.” In my view, there isn’t one single highest-order engine ruling the hierarchy like a monarch. There are several, running in parallel, each integrating information from overlapping but different subsets of the hierarchy below. And they compete.
That argument you have with yourself about whether to take the job offer — that's not one unified system weighing pros and cons. That's multiple highest-order engines, each reconstructing the situation differently: one embedding the excitement of something new and the salary bump, another embedding the relationships you'd leave behind and the commute, another embedding your self-image as someone who takes risks or plays it safe. Each is a legitimate reconstruction. Each is vying for gain — for dominance in the competition that determines what you actually do. The "voices in your head" aren't a metaphor. They're parallel apex-level processes, and the decision is whichever one wins the dynamical competition.
This is where something like free will lives in SST — not in some magical uncaused cause, but in the genuine unpredictability of the outcome when multiple nonlinear dynamical systems compete. The system is sensitive to initial conditions, to which voice happened to have a tiny edge at the moment the trajectory committed. And it's sensitive to unpredictable environmental input arriving in real time: whether your partner says "we could make it work" or your kid says "I don't want to change schools" genuinely reshapes which reconstruction wins. That's not the libertarian free will that philosophers argue about. But it's not the robotic determinism that keeps people up at night either. It's somewhere more interesting than either.
So here’s the full picture. Recurrent circuits performing delay coordinate embedding, organized in a hierarchy of increasing abstraction, with gain modulation selecting which reconstructions dominate at the top, and self-referential processing at the apex constituting the subjective perspective.
That’s State Space Theory. That’s what I think consciousness is.
Next time we pick up the best objection to the whole thing — and why I think it fails.
This is Part 4 of a 10-part series on the State Space Theory of Consciousness. Part 3: “Consciousness Is a Verb.” Part 5: “The Best Objection.” All the papers are here.
The self-reference criterion is developed formally in the CDM paper (under review, Synthese). The gain modulation account draws on the Global Workspace Theory of Baars and Dehaene, the Attention Schema Theory of Graziano, and predictive processing frameworks (Clark, Friston) — SST doesn’t replace these theories so much as specify the underlying computational mechanism. More on that in Part 8; the unifying paper is here if you can’t wait.


