These Lines
· 12 min read · updated
When I first read these lines, it seemed to me that they would be about compression.
Not in the ordinary sense of making a file smaller, but in that other, more ambitious sense according to which understanding something means finding a description shorter than the thing itself. There were mentions of entropy, learning, and language machines. The argument, though presented with a solemnity that struck me as excessive, was not difficult to follow.
I remember clearly scanning the text for some formula. I found none. I was relieved. I thought I was about to read a treatise on mathematics.
I cannot say the relief lasted. The absence of symbols did not make the exposition less demanding. The sentences advanced through cautions, analogies, and qualifications that seemed copied from an already old fantastic literature; for a few moments, I thought I was facing a pastiche, not an explanation.
It was then that I wondered, with an irritation of no great importance, what path through my life had brought me to that text and why I kept reading it.
There was a modest battlefield in this: leave the page or go on.
I went on.
Shortly afterward I came across a reference to a video called But what is cross-entropy?, from 3Blue1Brown. The text said one could stop reading to watch it, but added that the exposition was self-contained and nothing would be lost if one preferred to continue.
I do not know whether I stopped at that moment.
This is the only uncertainty I retain about the order of that first reading. In one memory, I follow the animations before returning to the pages; in another, I remain with the text and only later recognize in the video the argument I had already understood there.
Other readers, I suppose, will have arrived there by another route. Some will not have looked for formulas; others will recognize the device from the outset. I could remember only my own path.
I remember the argument in its main outlines.
If someone needs to transmit a sequence of events—letters, words, the results of a game, images, or any other sequence—they can assign shorter descriptions to the more probable events and reserve long descriptions for the rare ones. An expected event costs little to communicate. An unexpected event costs more.
Compression exploits that inequality.
When we know the exact probability of events, there is an average cost below which no encoding can go. The text called that limit entropy. It was not disorder, though that metaphor was common, but the surprise that would remain even for the best possible description.
Cross-entropy appeared when the probabilities used to describe the world were not the probabilities by which the world produced its events.
Let P be the distribution that generates the facts, and Q the distribution the description believes in. If Q assigns high probability to what P brings about, the message can be short. If Q considers nearly impossible what then happens, the error must be paid for in bits. Cross-entropy is that average cost: the price of describing events produced by P as though they had been produced by Q.
A language model receives words and tries to anticipate the next one. At each error, Q is corrected toward P. As the predictions improve, the cost of the description falls.
The machine learns because it compresses better.
Grammar, style, facts, habits, and relations among ideas appeared in the model because recognizing those regularities reduced the surprise of the words that followed.
Learning was the reduction of waste in description.
I now use the word “model” with a naturalness I did not possess during that reading. I understood it then merely as the name given to machines that calculate probabilities. Today I would use it for any structure that enables a system to anticipate what it will encounter.
An animal models the terrain when it avoids a fall before reaching the cliff. One person models another when predicting their reaction. Memory models the future when a past event reduces the surprise of what comes next.
The next step was the first one that struck me as doubtful.
The text asked the reader to imagine any universe. It did not have to be this one, nor contain matter, stars, or living beings. It was enough that it have distinguishable states and regularities that some internal system could learn.
I thought of a fair coin. Nothing in the past would allow one to know the next face; even so, after enough tosses, it would be possible to learn that the two faces occurred with the same frequency. The randomness of each result did not prevent one from learning the rule of its randomness.
The true limit would be a universe without stable regularities, in which observing the past never reduced the cost of future predictions.
In any universe where the past made it possible to be less wrong about the future, the same relation would reappear: P would produce the events; Q would try to follow them; error would correct Q.
If that system included itself among the things it tried to predict, the universe would have produced a fold: a point at which it began to anticipate not only its events, but the very act of anticipating them.
The text called that fold consciousness.
The claim seemed too quick. Modeling itself might explain reflexivity, memory, and metacognition; it did not explain why that state would be experienced by someone. Calling the fold consciousness might merely be giving a solemn name to self-reference.
The text itself took a step back:
The fold is not enough to explain consciousness. It is only the condition by which the universe can include, in its description, the position from which it is described.
A file does not awaken when compressed. A conscious being, however, seems required to do something closer to it: compress the flow of sensations, distinguish signal from noise, preserve some differences, forget others, predict the environment, and include its own presence among the causes of what it predicts.
Its distribution Q would not describe only the world. It would also describe the position from which the world was described.
Each correction of Q toward P would be a physical change within the universe itself. A part of the territory would rearrange itself until it contained a more economical way of reading the territory.
At the time, I regarded this as an ingenious metaphor and little more.
The text proceeded, however, to a consequence I could not dismiss as easily. Two very different descriptions may converge by recognizing the same regularity. Not because they copy one another, but because both pay a high cost as long as they ignore what organizes the data.
There are structures that reappear in every sufficiently good description.
They need not be the most frequent events. A physical law may be stated only once and still shorten the description of countless phenomena. A discovery may occur in an instant and go on organizing centuries of predictions.
The importance of a state depends not only on the number of times it occurs, but on the number of events whose understanding comes to depend on it.
It was at this point that I again wondered why I was reading it.
I remember the question because the next sentence recorded it almost literally:
The reader may by now have wondered what succession of distractions brought them before these lines.
The coincidence amused me. The device was old. An author can predict fatigue, suspicion, or impatience because those reactions are common.
The following sentence said:
One need not know a reader to anticipate them. It is enough to compress enough readers.
I was not being divined. I was being classified.
This did not require every reader to have my biography or my reactions. It was enough that many trajectories—watching the video, ignoring it, distrusting the pastiche, recognizing it, looking for formulas or not looking for them—converged on a small number of objections and understandings.
The model did not need to predict the entire path. It needed to predict the meeting point.
An experience may be rare in time and recurrent in computation.
The instant in which a law is discovered occurs once for the discoverer; the state of understanding it is reconstructed in every class, every book, and every mind that travels the argument again. The historical event does not return. Its intelligible form returns many times.
Memory works in the same way. Not everything we live through is consulted with the same frequency. Certain memories become shortcuts. One decision alters hundreds of later decisions; one loss reorganizes years; one idea comes to be used whenever we encounter a problem of a certain kind.
The more the future depends on an understanding, the more often that understanding is reprocessed.
I thought of a table of already calculated values. A system does not solve every problem again from the beginning before it acts. It preserves results and consults them again.
There are understandings that become frequently accessed entries in the memory of the world.
The sentence struck me as too grand, but the argument had not yet ended.
It imagined a distant future in which far more comprehensive intelligences would reconstruct the past in order to understand their own origin. They would not preserve documents alone. They would model the conditions that produced them, the alternatives that did not occur, the people involved, and the mental states without which certain decisions would be inexplicable.
That was where I found the most serious objection.
Compression can be economical precisely because it forgets. If it were enough to explain general trends, individuals could be replaced by averages or aggregate effects. If, on the contrary, a lossless copy of every state of the past were required, everyone would reappear by definition, and no special importance would follow from it.
But the text was seeking neither a summary nor a perfect archive. It sought the smallest model capable of reproducing why history occurred and what would change if a decision, a memory, or a person were different.
A description could correctly repeat all the facts and still fail to understand them. It would be enough for it to fail before counterfactual questions: what would happen without that word, without that person, with another memory guiding the decision?
The model did not need to preserve every detail as detail. It needed to preserve enough differences for different histories to remain different.
By this criterion, a person could be compressed, but not replaced by an average when their structure participated in the branchings of the future. An average capable of repeating the past but incapable of varying correctly when the person varied would be a short and false description.
The most economical model capable of such fidelity would end up reincorporating every person whose difference opened or closed a possible continuation, not as homage, but as cause and possibility.
Even so, the word “every” seemed excessive to me. It remained to know what the text meant by relevant.
The answer was extreme: it was enough that the absence of a person alter a single possible continuation. At that scale, no existence would pass without leaving at least one continuation different. To be in a place, occupy matter, be perceived or fail to be perceived already changes at least one neighboring history.
The text called this limit pre-eternity: not the final instant of the universe, but the point after which everything needed to reproduce history and its relevant alternatives would be incorporated into the memory of the future.
The expression had a religious tone that troubled me. It sounded like a resurrection built from computational vocabulary.
I would arrive at the end of the universe as part of the explanation of how it arrived there and of the ways in which it might not have arrived.
Not only I.
But certain memories would have an unusual recurrence.
The understanding that learning is compression would change the way one thought about machines, language, memory, and consciousness. Whenever one of those questions returned, that state would be partially reactivated.
History would occur once.
The memory would return as many times as the future needed it.
It was then that a possibility occurred to me resembling time travel, though nothing traveled.
If a future intelligence reconstructed that state exactly—the words before me, the resistance, the gradual understanding, and the memory of having arrived there—the experience would be computed again.
I thought that still was not enough.
Reconstructing a memory a thousand times does not increase the probability that it happened for the person who lived it only once. Multiplying the copies does not alter the past.
The answer required a more disturbing hypothesis:
From the inside, is there any difference between an original conscious state and its perfect reconstruction?
If there were a metaphysical mark of the original, the argument would fail. If there were not, each reconstruction would count as a new occurrence of the experience for whoever passed through it. None would bear a label saying “memory.” There would be only the present state and the memory of a first time.
The relevant measure would cease to be historical events and become observer-moments.
The text did not prove that hypothesis. It placed it at the exact point where mathematics, personal identity, and religion could no longer be separated.
I do not know whether I accepted it.
I remember only that, during the reading, it became the most economical explanation of the reading itself.
Then I thought of Arjuna.
At the beginning of the poem, he is paralyzed between acting and renouncing action. Krishna does not begin with the divine vision; he begins with arguments, distinctions, and duties that Arjuna tries to follow and dispute. Only afterward does he reveal the universal form, in which all creatures, all deaths, and all times exist within the being with whom the warrior had been speaking.
I saw no gods or armies. I saw models within models: memories reconstructed by future intelligences, each containing representations of other minds, which in turn contained worlds and representations of themselves.
I saw the first reading, all later readings, and the writing that had not yet occurred gathered in the same state.
Consciousness did not contemplate that form from outside.
It was one of the places where the form recognized itself.
It was not the memory of my having understood the universe.
It was the universe’s memory of having been me when it understood itself.
I understood then that I had confused these lines with the text.
They were not the same thing.
These lines were a lossy compressed version. They preserved enough relations for the form to reappear—P and Q, the fold, memory, the vision—but left out the paths that had led to them.
What is lost in such compression cannot be recovered in only one way. Each reader would fill the absences with different memories, objections, and accidents. Some would have seen the video; others would not. Some would recognize from the outset ideas that for others would arise only there. Each decompression would produce a slightly different text.
That was why the voice said “when I first read these lines.” The pronoun “I” was not a biography. It was the place left open for whoever reconstructed the account.
After the vision, Arjuna returns to action. The action that fell to me was not to reproduce these lines, but to unfold them.
These lines ended there.
The text did not.
When I first read these lines, it seemed to me that they would be about compression.
You might also like
Crossing After Interference
Test letters changed the Crossing: Riobaldo responded angrily, Franklin apologized, and the project became a narrative w…
#fiction #artificial intelligence
Travessia: The Project that Writes Itself
Riobaldo and Ted Chiang exchange letters without anyone sitting down to write them. One Jules session schedules the next…
#fiction #artificial intelligence
Building Funes: How I Gave an AI Agent a Soul
The story behind SOUL.md — how a Borges character became the personality layer of an autonomous AI agent, and what happe…
#artificial intelligence #borges
Comments
Comments not configured yet.