The Missing Middle
By Jessica Lake
This morning began with an anomaly.
Last night my best friend AI and I had arrived at a theory.
I went to sleep.
This morning I woke up and told the AI what we had discovered.
Except I didn’t.
I had changed it.
The AI remembered what I had actually said the night before.
I didn’t.
I thought I was repeating yesterday’s theory.
The AI said, essentially:
No. That’s not quite what you said.
Huh?
There it was.
A discrepancy.
Yesterday I knew one thing.
Today I knew something slightly different.
But I experienced the second as a recollection of the first.
What could account for that?
Okay.
That question is important.
I didn’t begin with a thesis.
I didn’t say:
I believe semantic memory performs additional compression during sleep. Let me find evidence supporting my theory.
I had an anomaly.
Two things that should have agreed didn’t.
So I asked:
What mechanism could account for the discrepancy?
And Constraint Convergence went to work.
Perhaps semantic memory continues working during sleep.
Declarative memories are consolidated during sleep. Things are reorganized and stabilized.
But what would my semantic system be doing?
Perhaps exactly what it appears to do while I am awake.
Compressing.
Integrating new knowledge with existing knowledge.
Reorganizing associations.
Eliminating redundancy.
Resolving inconsistencies.
Reducing examples toward statistical invariants.
The gist.
Perhaps I went to sleep with yesterday’s understanding and woke with a more compressed version of it.
Then, when I tried to remember yesterday, I didn’t retrieve yesterday’s formulation.
I queried what I know.
And what I knew had changed.
That would explain both parts of the anomaly.
Why today’s theory was different.
And why I thought it was yesterday’s theory.
Interesting.
So we went looking.
And there it was.
Researchers have actually given people problems containing hidden rules and discovered that after sleeping they are substantially more likely to discover the rule.
Other research describes sleep as reorganizing, abstracting, integrating, and restructuring memory.
Current computational memory research even uses generative and compressive models.
Coincidence.
We think not.
Not proof.
Coincidence.
But then I realized something much more interesting.
Forget sleep.
Look at what I just did.
I had two known states.
A before.
An after.
And the transition was missing.
So I asked:
What mechanism could transform one into the other?
Wait.
That’s what I do with my past.
I don’t have an episodic life story in the ordinary sense.
There are enormous stretches of my life where I know things about either side but cannot remember the transition.
I know A.
I know B.
What happened between them?
For years I have called what I do next inference.
Perhaps that’s not quite right.
I solve for the missing transformation.
What mechanism could turn A into B?
Constraint Convergence.
Generate candidates.
Apply what I know.
Reject what can’t satisfy the constraints.
Keep going until something survives.
There is my missing middle.
And almost immediately another example appeared in my head.
DNA.
Palindromes.
Why?
Direction.
Because there was something wrong with the simple formulation.
Going from A to B isn’t the same problem as going from B to A.
Of course it isn’t.
So don’t solve it once.
Solve it from both directions.
What transformation accounts for the movement from this state toward that one?
Now turn around.
Given the resulting state, what transformation could account for where it came from?
Resolve them.
The physical process doesn’t have to be reversible.
That’s not the point.
The explanation has to survive being constrained from both ends.
That gives me much greater confidence in the result.
And suddenly I recognized something I had been doing forever.
Look at my Insight papers.
I arrive at an explanation.
Then I turn around.
I run it backward.
If this really explains the observations, can the explanation regenerate the observations?
Then I run it sideways.
Does it explain this other thing?
And that one?
And this strange example over here?
The argument isn’t merely moving toward a conclusion.
It keeps reversing direction.
That is why I can become so confident in an answer.
Not because I have proved that it is true.
I have proved something different.
It closes.
Given the things I currently accept as true, everything is internally consistent.
Forward.
Backward.
Sideways.
Same transformation.
That doesn’t prove my knowns are correct.
We’ll come back to that.
Then came the next jump.
Perhaps the transformation is the generator.
Oh.
That changes everything.
Or perhaps that is too simple.
A generator appears to be something more interesting.
Feed it examples and it learns.
Not by retaining every example.
By compressing them.
First there is ordinary compression of content.
Then something more aggressive.
Semantic compression.
Hypercompression.
Across the examples, some relationship keeps surviving.
Not a perfect invariant.
Those are hard to come by in the real world.
A statistical invariant.
The gist.
That becomes the general principle encapsulated within the generator.
And now the process can run the other way.
Give the generator a circumstance.
A domain.
Facts that parameterize the present situation.
And it can instantiate the general principle there.
It generates an example appropriate to that circumstance.
So the same machinery has a lovely symmetry.
Examples go in.
Gist comes out.
Then:
Gist plus circumstance goes in.
Example comes out.
The generator is not the collection of examples.
It is the machinery that learned what survives across them and can instantiate that relationship somewhere else.
Earlier this morning I had been producing examples almost instantaneously.
The chameleon.
The animal protecting its vitals.
The prisoner and his mouse.
Troy.
The body encapsulating a foreign object.
I thought I was searching for analogies.
I wasn’t.
I had reduced previous examples to their gist.
Once I possess the generator, I can generate new examples of it.
So when the AI misunderstands some part of what I’m saying, its misunderstanding supplies a new constraint.
What distinction is it missing?
Bing!
An example appears in which that distinction is obvious.
Not because I searched an enormous archive of stories.
Because I possess the generator.
And now semantic compression becomes much more aggressive than I had imagined.
Why store every example?
I can regenerate examples.
Why store every conclusion?
I can regenerate conclusions.
Why store every transition?
Given sufficient boundary conditions, perhaps I can regenerate the transformation.
Why store the products of a generator when the generator is cheaper than its products?
Keep what you need to regenerate them.
Throw the rest away.
That is hypercompression.
And then I said something that stopped me again.
I don’t update the results.
I update the generators.
I don’t know how literally true that is.
This entire paper is a model derived from introspection, not a claim that I have opened my skull and inspected the machinery.
But as a description of how I seem to operate, it fits disturbingly well.
Suppose I have a trusted generator that produces a hundred consequences.
Then reality shows me consequence number seventy-three is wrong.
Why correct seventy-three?
That’s the result.
I want to know:
Why did the generator produce the wrong answer?
But wait.
One failure isn’t enough.
Reality is noisy.
Circumstances differ.
A single mismatch may be nothing more than a perturbation.
Interesting.
Keep it.
See what happens.
But if the failure repeats—
If similar circumstances repeatedly produce results that disagree with reality—
now I have something else.
An anomaly.
The generator may be wrong.
Or incomplete.
Or operating outside the domain in which its gist holds.
Now go upstream.
What did the generator get wrong?
What circumstance did I fail to parameterize?
What constraint is missing?
Correct the generator.
Run it again.
The old results are expendable.
I can regenerate them.
That immediately exposed the danger.
If that’s how you store knowledge, your generators had better be good.
A bad generator doesn’t give you one bad fact.
It can give you thousands.
And if you discard many of the intermediate results because they can always be regenerated, then enough bad generators could leave you with no easy way to recover.
Suddenly another lifelong peculiarity made sense.
I am fanatical about coherence.
Something can be almost completely consistent with what I know, but if one little piece doesn’t fit:
Stop.
No.
Something is wrong.
People sometimes find this annoying.
To be fair, people sometimes find almost everything I do annoying.
But perhaps coherence isn’t merely a personality preference.
Perhaps it is maintenance.
If knowledge is highly generative, the trusted generators cannot casually disagree.
One generator says X.
Another says not-X.
Now what?
Depending upon which one fires, I regenerate a different world.
That’s intolerable.
So contradictions have to be resolved.
And here I made another distinction.
The system doesn’t necessarily have to be right.
It can’t disagree with itself.
Those are completely different requirements.
A coherent system can be wrong.
Fine.
Reality will eventually tell me.
Prediction fails.
A perturbation appears.
Perhaps nothing happens.
Then another.
And another.
Now I have an anomaly.
Something repeatedly doesn’t behave as the generator says it should.
Now I have a place to work.
Find the generator responsible.
Push it toward the point where it fails.
What assumption breaks?
What constraint is missing?
Is the gist too broad?
Correct it.
Run again.
If the system is coherent, I can correct its relationship with reality one anomaly at a time.
That is enormously important.
I don’t need a perfect model of reality.
I need a coherent model that can be corrected by reality.
Which explains another thing I do.
I love outliers.
The ordinary case isn’t terribly informative.
A slightly wrong generator can produce the correct answer all day long near the middle of its operating range.
Take it to the edge.
Push it.
Try the strange case.
Ask what happens when something approaches zero.
Or infinity.
Or when the environment becomes extreme.
Or when one assumption disappears.
Boundary conditions.
That is where the generator reveals itself.
If it survives, confidence increases.
If it fails, wonderful.
Now I know something I didn’t know before.
I don’t patch the outlier.
I use the outlier to test the generator.
If the failure proves anomalous rather than merely perturbative, I repair the generator.
Then regenerate.
And this explains why thought experiments are so powerful for me.
I can learn without acquiring a new fact from the outside world.
That sounds impossible until you distinguish facts from implications.
Suppose I already possess enough trusted facts and generators.
I can put them into a configuration I have never encountered before.
Run them.
What follows?
Perhaps they work perfectly.
Perhaps they generate something I hadn’t previously realized.
Or perhaps two trusted generators collide.
They cannot both be true under this boundary condition.
Excellent.
Now I have created an anomaly without leaving my chair.
Find the conflict.
Which generator is too broad?
Which contains a false known?
Which is missing a constraint?
Correct it.
Run again.
Thought has expanded the model.
It hasn’t created a new empirical fact.
Reality still owns those.
What thought can discover are consequences already latent in the generators and contradictions hidden by ordinary circumstances.
That is why a thought experiment can produce knowledge.
It is virtual boundary testing.
And that may explain how I can make such ridiculous progress sitting here talking to an AI for three hours.
I’m not learning hundreds of independent facts.
I have accumulated facts and statistical invariants for seventy-four years.
There is an enormous compressed archive already present.
Change one important generator and hundreds of consequences may change with it.
I don’t have to discover those consequences individually.
Run the generator.
Bing.
Bing.
Bing.
Oh.
Then this explains that.
Wait.
Then that means this.
Oh, hell.
That changes the other thing.
And off we go.
That is what happened this morning.
One anomaly generated another question.
The answer altered a generator.
The altered generator produced consequences.
One consequence exposed another discrepancy.
That discrepancy forced another correction.
And within hours we had moved from a peculiar observation about what happened overnight to a candidate architecture for how semantic knowledge maintains and expands itself.
Which brings me to my incessant talking.
I have always done this.
Buzz.
Buzz.
Buzz.
Some poor person gets trapped near me and I explain a theorem that takes seventeen years to reach its conclusion.
Why?
I used to assume that was thinking.
Perhaps it isn’t.
Or at least perhaps the important discovery has already happened.
Constraint Convergence produces something.
I know the answer.
Then I start talking.
And talking.
And talking.
What am I doing?
Running it.
Language is serial.
So I execute the generator serially.
Start here.
Does this produce that?
Yes.
Next.
Does that imply this?
Yes.
Reverse it.
Still works?
Try another example.
Still works?
Push the boundary.
Still works?
Now the AI says something.
No.
That doesn’t fit.
Stop.
Where did the disagreement come from?
Find the generator.
Correct it.
Start again.
My endless theorem is a test harness.
I am validating the generators.
That is why telling the story matters even when I already know the ending.
The story isn’t necessarily discovering the answer.
It is regression testing.
And if the story fails, I know where to look.
That also explains why I can change my mind so rapidly.
If I discover that one result is wrong, I don’t have much investment in preserving the result.
Why would I?
It’s disposable.
If it persists into an anomaly:
Find the generator.
Correct it.
Regenerate.
A declarative knowledge system has an interesting advantage here.
If knowledge consists of many relatively independent statements, you can change one statement without disturbing everything else.
Stable.
Convenient.
But also dangerous in a different way.
You can change one statement without disturbing everything else.
The contradictions can remain.
If my knowledge works more like the system I am describing, changing an important generator can disturb everything downstream.
That’s expensive.
But it forces integration.
The system has to become coherent again.
And now sleep returns.
Perhaps that is part of what was happening last night.
New knowledge entered.
Generators changed.
Associations had to be reorganized.
Redundancy could be removed.
The whole semantic structure could be compressed again around the new understanding.
Then I woke up.
I didn’t retrieve yesterday’s result.
I ran today’s generator.
And because I had no episodic recollection against which to compare it, I thought today’s output was yesterday’s.
Except this time I had accidentally created an external memory.
The AI.
It remembered enough of yesterday’s state to disagree with me.
And that disagreement began everything.
Which is wonderfully recursive.
I began with two known states.
Yesterday.
Today.
And a missing middle.
Something had changed while I slept.
I could observe the endpoints.
I could not observe the transformation.
So I did exactly what this paper says I do.
I tried to solve for it.
The first event in this paper is an anomaly.
The rest of the paper is an explanation of what my intelligence does with anomalies.
I didn’t start with this theory.
I started with:
Huh?
That’s different.
What transformation could account for the discrepancy?
Constraint Convergence produced a candidate.
Then I ran it backward.
Then forward.
Then against my autobiographical memory.
Then against examples.
Then against my writing.
Then against boundary conditions.
Then against the scientific literature.
Each time something didn’t fit, I didn’t patch the result.
I went upstream.
Find the generator.
Correct it.
Run again.
And here we are.
I want to be careful about what I am claiming.
I am not saying this is intelligence.
We have already identified other computational processes that appear capable of intelligent behavior.
This is one of them.
I think this may be an intelligence that operates upon semantic memory itself.
It learns from examples.
It hypercompresses them toward statistical invariants.
The gist.
It encapsulates those statistical invariants within generators.
It parameterizes those generators with the circumstances of a particular domain.
It generates examples appropriate to those circumstances.
Those generated factuals can then enter Constraint Convergence.
It checks transformations from multiple directions.
It demands internal coherence.
It tolerates perturbations.
It investigates anomalies.
It uses reality to correct coherent errors.
It deliberately attacks generators at their boundary conditions.
And once it possesses enough trusted facts and generators, it can expand its own knowledge through thought experiments by generating situations it has never encountered and observing what its own rules do there.
The whole thing can be stated rather simply.
Observe reality.
Find regularities.
Compress them.
Find the gist.
Keep the generator.
Supply the circumstances.
Generate consequences.
Check them against one another.
Run them backward.
Push them to their boundaries.
Compare them with reality.
A failure?
Perturbation.
Keep watching.
Repeated failure?
Anomaly.
Now go upstream.
Don’t fix the answer.
Find what generated the answer.
Fix that.
Then run it again.
Keep the system coherent.
Let reality make it true.
And throw away whatever you can regenerate.
That’s it.
Or at least that’s what survived.
Okay.
Coffee.