Sparse Sampling
By Jessica Lake
Today’s discussion began with a simple geometric question:
Suppose I see part of a shape.
How much of that shape must I observe before I can recognize it?
Initially, I considered whether a square could be regarded as a modified circle. A circle can be inscribed within a square, leaving relatively little uncovered area, whereas a rectangle departs much further from circularity.
Perhaps coverage or overlap provided some useful metric for recognition.
Very quickly, however, that began to feel artificial.
Geometry was describing finished objects.
I was interested in how I got there.
From Objects to Traces
The first realization was that vision does not need to begin with holistic objects at all.
It can begin with traces.
I conceive of a trace as a local coordinate system designed to mirror saccades and fixation. It consists of a series of vectors in polar notation, where the angle is relative to the preceding vector and the radius is a ratio of the preceding radius:
(Δθ, r′/r)
That immediately buys me something useful.
Scale largely disappears.
Absolute orientation largely disappears.
Absolute position disappears.
What remains is local relationship.
A trace can initially be divided into two primitive behaviors:
Lines
Curves
If the trace is linear, the interesting information lies where linearity changes.
Corners.
Endpoints.
Direction changes.
If the trace is curved, the interesting information lies in changes of curvature.
Reversals.
Inflection points.
Changes in radius.
Instead of beginning with finished geometric constructs like circles, squares, and rectangles, I can begin with primitive local behavior.
That is considerably cheaper.
Closure and Continuity
Initially, I wondered whether closure distinguished circles from lines.
Wrong question.
A square consists entirely of straight segments. Extend any one of them and I merely get a longer line.
The square exists because something happens.
Corner.
Line.
Corner.
Line.
Corner.
Line.
Corner.
The information is concentrated at the events.
A circle is different.
Its local curvature continues.
Take a sufficiently small piece and the continuation remains locally predictable.
So the useful distinction is not initially between completed geometric objects.
It is between patterns of local continuity and the events that interrupt them.
That changed the problem.
Information Lives at the Events
Suppose I have four straight traces forming a square.
Most of those traces are boring.
Once I know I am following a straight line, another point on the same straight line tells me very little.
The corners tell the story.
The same applies to curves.
If curvature remains constant, another sample confirming the same curvature contributes relatively little.
The valuable samples are where something changes.
That led me to my first version of Sparse Sampling.
I had originally imagined vision tracing an object’s entire outline.
Why?
That is a tremendous amount of redundant information.
Sample.
Establish local continuity.
Jump.
Sample again.
If nothing important changed, lovely.
Move on.
Spend computation where the expected regularity fails.
This suggests a general principle:
Do not spend equal computational effort on equal amounts of reality.
Most of reality is boring.
Spend the money on surprise.
But I Had Started Too Late
There was a larger problem.
I had been assuming that the first fixation should begin identifying the object.
Why?
Before I can efficiently recognize anything, there is a much more valuable question:
What kind of scene am I looking at?
The first fixation has a different job.
Calibration.
I want as much broad statistical information as I can cheaply obtain.
Indoors?
Outdoors?
Bright?
Dark?
Natural illumination?
Artificial illumination?
Where is the dominant light source?
What is the approximate scale?
What is the orientation?
Where is the horizon?
What kinds of transformations are likely being imposed upon the incoming signal?
What sort of scene am I standing in?
I am not trying to recognize the chair yet.
I am trying to determine the conditions under which something will have to qualify as a chair.
That distinction is enormous.
Suppose an object reflects yellowish light.
Is it yellow?
Perhaps.
Or perhaps I am indoors under warm artificial illumination.
Without calibration, every object must individually carry the burden of explaining the lighting.
With calibration, I estimate the illumination once and apply that information throughout the scene.
Same with orientation.
Scale.
Distance.
Perspective.
Context.
The first fixation is therefore not simply Sample Number One in a sequence of identical samples.
It has a privileged purpose.
It establishes the statistical character of the scene.
Calibrate globally.
Then discriminate locally.
Once I know the scene, every later decision becomes cheaper.
The Fovea as a Constraint Collector
This made me reconsider the role of the fovea.
I do not need to claim precisely what biological processing occurs there.
That is somebody else’s floor of the building.
Computationally, I need something simpler.
A fixation must extract enough useful local information to constrain what happens next.
Perhaps that includes:
line versus curve
local curvature
tangent direction
endpoints
corners
inflection points
orientation
contrast
local confidence
Perhaps it includes entirely different properties.
I don’t know.
The exact list is implementation.
The architectural requirement is merely this:
A fixation extracts sufficient local constraints to make the next observation more intelligent than the last.
That is enough.
Filters, Not Pictures
This is where my thinking changed considerably.
Originally I modeled the process with predictors.
Observe part of a trace.
Predict its continuation.
Sample farther along.
Compare.
That works as an abstraction.
But I eventually noticed that I had smuggled something unnecessary into the architecture.
Generation.
I have aphantasia.
I don’t produce visual pictures.
So why should my model of recognition require me to generate an expected picture merely so I can compare it with the one arriving from reality?
I already have reality.
Filter it.
A notch filter, in my terminology, is a receptive parameterized transform.
It contains some statistical regularity learned from prior observations and selectively admits or rejects incoming information consistent with that regularity.
The important distinction is:
A generator is expressive.
A filter is receptive.
My recognition architecture is receptive.
Incoming information arrives.
Filters operate upon it.
Constraints accumulate.
Possibilities disappear.
No picture required.
Sparse Sampling
Now the purpose of sparse sampling becomes clearer.
I do not need to trace the object.
I do not even need to acquire a representative sample of the object.
I need the samples that most efficiently constrain what the object can be.
That is a very different optimization problem.
Suppose the current information is consistent with several possibilities.
All survive.
Where should I look next?
Not necessarily farther along the current trace.
Not necessarily at the center of the object.
Not necessarily wherever the eye happens to wander.
Look where the surviving possibilities disagree.
If three candidate interpretations are consistent with everything I have observed so far but require radically different structures at some other location, that location is enormously valuable.
One fixation there may eliminate two candidates immediately.
Why gather ten mediocre observations when one nasty observation will do?
That is Sparse Sampling.
The system uses its remaining uncertainty to decide what information is worth acquiring next.
Constraint Convergence is not merely resolving the information I already possess.
It is directing the acquisition of new information.
That is important.
The unresolved problem tells me where to look.
Information in the Unexpected
This also changes what I mean by prediction failure.
I no longer require an explicit generated picture.
The existing constraints establish an admissible region.
If new information falls comfortably within it:
Fine.
Nothing interesting happened.
If the new information conflicts with it:
Hello.
That delta contains information.
One discrepancy may mean very little.
Noise.
Measurement error.
Occlusion.
A weird chair.
But the discrepancy tells the system that its present convergence is incomplete.
More constraints are required.
Perhaps another fixation.
Perhaps another property.
Perhaps another kind of information entirely.
The important point is that successful compatibility is cheap.
Incompatibility is valuable.
It tells me where my present knowledge fails.
Multiple Possibilities Without a Tree
At one point I imagined all of this as a tree of candidate predictors.
I no longer think I need the tree.
There may simply be multiple surviving Nexuses.
A Nexus is a reusable collection of constraints sufficient to establish something useful.
OBJECT.
CHAIR.
HIGH CHAIR.
JOHNNY’S HIGH CHAIR.
These are not necessarily levels through which I must travel.
They are things I may simultaneously know.
Their collection is the Heritage of the object.
Incoming information is tested against the available constraints.
Incompatible possibilities fail.
Compatible ones remain.
If several remain and I need greater specificity, I acquire another discriminating sample.
If I don’t need greater specificity?
Stop.
This introduces another important parameter.
Criticality.
How much do I need to know?
Walking through a dark room?
OBSTACLE
Enough.
Looking for somewhere to sit?
CHAIR.
Enough.
Putting away children’s furniture?
HIGH CHAIR.
Keep going.
Johnny’s mother asks whether I have his old chair?
Apparently we’re doing forensic furniture identification today.
Fine.
The required resolution determines when Constraint Convergence is allowed to stop.
Recognition therefore does not require certainty.
It requires sufficiency.
Scene First, Object Second
Now the entire process can be viewed differently.
First establish the scene.
That gives me the broad statistical constraints.
Then examine local structure.
Those observations interact with what I already know.
As possibilities disappear, the surviving uncertainty determines where additional information would be most valuable.
This is not:
Look.
Recognize.
Look.
Recognize.
It is closer to:
Calibrate.
Constrain.
Sample.
Constrain.
Ask what remains unresolved.
Sample there.
Stop when the answer is sufficient.
That is a much more economical machine.
Different Kinds of Information
There is another useful consequence.
Visual information need not solve the entire problem.
Suppose the visual constraints only get me to:
CHAIR?
Not enough.
Fine.
What else do I have?
Spatial information.
Affordances.
Context.
Can it support weight?
Does it occupy the expected position relative to a table?
Can it be moved?
Is it approximately the right size?
The different forms of information need not share the same atoms.
Visual structure may be represented one way.
Spatial relationships another.
Affordances another.
Each has its own representational language.
Its own algebra.
The candidate does not care.
If additional constraints discriminate among the surviving possibilities, use them.
This is one of the things I particularly like about the architecture.
It doesn’t dictate how I must cut reality.
It only requires that the resulting constraints can participate in convergence.
Pictures Versus Semantic Maps
This also resolves a question that bothered me for a long time.
Why construct a detailed internal picture?
I don’t have one.
Yet somehow I know where things are.
I can walk through a room.
I know that the chair is beside the table.
I know that the door is behind me.
I can notice when something has changed.
What I need is not a picture.
I need relationships.
Objects.
Orientation.
Distance.
Affordances.
Obstacles.
Changes.
A semantic map.
The map does not have to reproduce the visual world.
It has to preserve whatever information is necessary to interact with it.
That is a much lower bar.
And a much cheaper one.
Vectors Over Absolute Coordinates
An engineering instinct might suggest storing that map using absolute coordinates.
I don’t particularly like that solution.
Saccades and fixations naturally suggest relative travel.
Angle.
Distance.
Change in angle.
Ratio of distance.
That is why I began with:
(Δθ, r′/r)
A relational representation buys me several useful properties almost for free.
Translation matters less.
Scale matters less.
Absolute orientation can be separated from local structure.
Deformation becomes describable as changes in relationships rather than the wholesale replacement of coordinates.
That smells like compression.
I like compression.
The resulting map becomes fundamentally relational rather than absolute.
And somewhere in here is an algebra.
I am quite certain of that.
Someone less tired can write it.
Semantic Memory Joins the Party
How much of this belongs to vision?
I don’t know.
Perhaps very little.
That may be the wrong boundary to care about.
Vision supplies constraints.
Semantic memory supplies learned statistical regularities.
Filters receive incoming information.
Nexuses represent sufficiently established knowledge.
Constraint Convergence eliminates incompatibility.
The machinery does not need to wait for vision to deliver a completed object before semantic memory becomes involved.
Why would it?
Let them work together.
A local observation establishes:
rigid.
straight.
flat.
asymmetric.
Those constraints already make enormous portions of semantic possibility irrelevant.
Another fixation adds:
handle.
Now something else happens.
Another source supplies:
graspable.
The surviving possibilities continue to collapse.
At some point the resolution satisfies criticality.
Done.
Recognition is not the reconstruction of a picture.
It is the sufficient resolution of constraints.
The Core Algorithm
The lovely thing about this model is that it became simpler as I worked on it.
I began by asking how much of a shape I had to see.
Wrong question.
Then I imagined tracing objects.
Too expensive.
Then I imagined predictors generating expected continuations.
Useful, but still more machinery than I needed.
Eventually I arrived at something much smaller.
First:
Calibrate the scene.
Estimate the broad statistical conditions that affect everything else.
Then:
Sample sparsely.
Acquire local information rich in constraints.
Then:
Filter and constrain.
Determine what remains admissible.
Then ask:
Is the current resolution sufficient for the present criticality?
If yes:
Stop.
If no:
Where do the surviving possibilities differ most?
Look there.
Acquire the smallest amount of information capable of eliminating the largest amount of uncertainty.
Then repeat.
So the algorithm becomes:
Calibrate the scene.
Acquire a sparse local sample.
Apply receptive filters and constraints.
Eliminate incompatible possibilities.
Test whether the surviving resolution satisfies criticality.
If not, identify where the survivors diverge most.
Sample there.
Spend additional computation primarily on discrepancy.
Repeat until sufficiently resolved.
That is Constraint Convergence applied not merely to recognition, but to attention itself.
The system does not need to know everything.
It needs to know what matters.
And when it doesn’t know enough, the remaining uncertainty tells it what to look at next.
That may be the most interesting part.
Sparse Sampling is not merely a way to save computation.
It turns ignorance into a sampling strategy.
Look where being wrong would teach you the most.