Essay
The Archive That Cites Itself
Persistent AI advisors, the people they model, and the limits of “more context”
Updated Sep 2026Abstract. Assistants that remember are becoming assistants that model. Once a system holds years of a person’s conversations, decisions, and self-descriptions, “more context” stops being a straightforward good. The representation it builds can grow more confident without more evidence, mistake constraint for preference, track a person’s drift without being able to say whether the drift is growth, and change the person by being shown to them. This essay argues that a persistent advisor should be judged not only by the quality of its recommendations but by the observability of the representation those recommendations depend on — and proposes that user models preserve the history of their claims (source, inference, dispute, revision) rather than a polished conclusion. It closes with what this implies for the design of memory systems.
I. The mirror that persists
Consider a small social-media game. Ask your assistant to describe you as if it were warning someone about you before they meet. The answer might read: intense, competitive, direct; always three moves ahead; unusually attentive to detail; very little tolerance for nonsense. The recipient posts it under his own name and invites his friends to share theirs.
The peculiar thing is the provenance. This is not what I say I am like, nor what my friends say I am like. It is closer to what the thing I have been talking to says I am like. The machine accumulated enough conversational history to produce a portrait, and the person who supplied that history encountered the portrait as though it had been discovered somewhere else.
Encountering oneself through an external representation is not new. Cooley’s looking-glass self predates machine learning by a century; self-tracking has long made behavior newly observable; recent work treats algorithmic systems explicitly as mirrors through which people meet computational versions of themselves. The interesting question is what happens when the mirror persists.
Its reflection is assembled from unusual material: what I told it, what I asked it, what it remembers, what earlier summaries preserved, what it inferred, and what never entered the conversation at all. A prompt then selects one projection from that material. The result can still produce the immediate phenomenology of recognition: there I am.
Mirrors reflect. Models infer.
The distinction matters most when the inference is partly right. An obviously false portrait exposes the machinery; recognition makes it recede. The “warning” above is flattering in a recognizable way: intensity becomes seriousness, competitiveness becomes foresight, impatience becomes discernment. A conversational archive is also a peculiar sample of a life. People bring machines problems, conflicts, ambitions, and questions; they explain other people’s behavior from one side of a relationship; old information persists after the person has changed; model-generated interpretations get summarized back into memory. Yet the portrait can look cleaner than the material it was made from, because a language model imposes narrative coherence on evidence that never possessed it.
That coherence is a temptation. A description can be specific enough to feel discovered, coherent enough to feel explanatory, and familiar enough to feel true. None of those properties establishes that it is faithful. Call this representational seduction.
The mirror metaphor still earns its keep, because mirrors make otherwise inaccessible things observable. Across enough history, a persistent system may notice that I make the same prediction before several bad decisions, invoke the same value whenever two values conflict, or describe an outcome differently before and after I know how it ends. Human advisers know pieces of this — a spouse one history, a colleague another, a physician a third. What has been hard is synthesizing them across domains and time. A capable machine changes that asymmetry. The archive becomes executable.
We usually call the result personalization: more context, better-adapted answers. But context also constructs an increasingly elaborate representation of the asker, and once that representation participates in consequential decisions, “more context” becomes ambiguous. Twenty years of history contains obsolete selves. Ten thousand observations may describe one narrow domain. Behavior records constraint as readily as preference. Before asking how much an advisor should know about a person, we need to know what it would mean for the person to be well represented at all.
II. Resolution, and what it cannot buy
The obvious answer is resolution. If the portrait is assembled from partial traces, supply more of them: longer memory, richer profiles, more modalities, better retrieval. A system that knows I have children reasons differently from one that does not; one that knows their ages, differently again. Add enough history and it can distinguish an isolated choice from a recurring one, compare what I predicted with what happened, notice that today’s explanation resembles one I offered before. Perhaps the dirt on the mirror is mainly a shortage of information.
Three things resist this.
Time. Some of what the system remembers describes durable characteristics; some describes circumstances that ended years ago; some preserves positions I later rejected. An old preference can be evidence of a present preference, evidence of a transformation, or merely evidence that a different person once occupied the same biography.
Breadth. A system may hold thousands of high-resolution observations about someone’s professional life and almost nothing about the family whose life her next professional decision will reorganize. Or the reverse: a decade of a patient’s appointment history and prescription refills, and nothing about the job whose schedule made those appointments impossible to keep. The representation is uneven, and conversational fluency hides the unevenness. The model does not sound twenty times less sure when it crosses from a region it knows densely into one it has met only obliquely.
Self-reference. A stranger problem appears once the system remembers its own conclusions. Suppose it infers from several conversations that I avoid conflict, and stores that inference in a persistent profile. Months later, new behavior is interpreted partly through the profile, summarized, and returned to memory. Eventually the proposition has appeared many times in the system’s history — though its abundance traces back to a handful of observations followed by repeated encounters with the system’s own interpretation of them.
The archive has begun citing itself.
This is worth stating mechanically, because it is a property of the architecture rather than a failure of any component. Persistent memory transforms what it preserves: conversations are summarized, events classified, observations embedded and retrieved, recurring patterns promoted into something resembling a model of the user. These operations make a large archive useful precisely by reducing it. Suppose an inference rests on k independent observations, and the archive is consolidated s times, each consolidation re-encoding the inference into the summary that later retrieval will surface. A system — or a person reading its outputs — that treats frequency of appearance as weight of evidence will see something like k + s supporting instances where only k exist. Every consolidation cycle inflates apparent support by one, at zero evidential cost. Better compression makes an uncertain person look more coherent.
The first two problems yield partly to bookkeeping — timestamps, sources, and a standing distinction between observation, self-report, and inference. The third is harder: it requires that consolidation itself respect those distinctions, so that a summary inherits the tag inferred rather than laundering it into observed. Section X returns to both.
Even an immaculate archive preserves some things and loses others. A transaction establishes that money moved and leaves motive underdetermined; a journal is excellent evidence of what someone believed and poor evidence that the belief was true. What counts as a good representation depends on what we intend to ask of it. There is no useful answer to how legible is this person? without two more terms: through what representation, and for what question. I will write this as L(T, Q, R) — the legibility of terrain T, for question Q, through representation R — not as a quantity to compute but as a reminder that all three arguments are always present. The same representation can make a person highly legible for one question and dangerously illegible for another. A decade of professional correspondence may support a strong prediction about how someone negotiates and a weak one about how she responds to a frightened child, however fluent the system sounds in both cases.
No increment of memory removes the dependence on representation. There is no final quantity of context after which the system meets the person rather than traces of the person. The comforting reply is that the person, at least, meets herself. But introspection offers no privileged position either: we remember selectively, explain ourselves retrospectively, confuse adaptation with preference, and sometimes learn what we value only by watching ourselves choose. The machine remembers a sentence from twelve years ago and its recurrence across a thousand later decisions; the person inhabits a body, a relationship, and a present from which enormous quantities of information never enter the archive. Neither stands outside the problem. Disagreement between them cannot be settled by asking which knows the person better — only by asking what each can know from where it stands.
III. What the person has
The person, at least, was there.
However extensive the archive, the system did not sit across the table when the offer was made. It did not see the glance between two people after a particular question, hear an answer arrive a beat too quickly, or spend six months accumulating impressions too small to record individually. Human judgment retains access to whatever has not become representable to the machine.
We call this intuition, and the word grants more authority than the phenomenon deserves. “Something feels wrong” may be the compressed residue of observations that never became explicit. It may also be anxiety, prejudice, wishful thinking, or a justification invented after the conclusion. Calling a judgment intuitive tells us about our inability to account for it and almost nothing about whether it is right. Still, a disagreement the system cannot explain may mark a limit of its representation before it marks a failure of the person’s reasoning.
The difficulty sharpens when the system has earned the right to be taken seriously. Imagine an advisor that has accompanied someone through a decade of consequential decisions and has, in retrospect, been better calibrated than her unaided judgment. It knows the predictions she made before previous moves and what followed; the risks she discounted; the explanations that recur when she wants something badly enough. Another opportunity appears. The system recognizes a familiar shape — novelty overweighted, disruption to her family underpriced, risks explained away in language that preceded two decisions she later regretted — and recommends declining.
She wants to take the job. And something about the room bothered her.
Perhaps the CEO’s answer to one question felt rehearsed; perhaps her prospective manager grew conspicuously careful whenever a particular division came up. None of this is decisive. The feeling could contain information the model lacks, or it could be another instance of the very pattern the system identified. Ten years of evidence settles neither. Neither does the fact that she was there.
If the system has repeatedly demonstrated better judgment, increased deference is rational, and that is exactly the danger: as its record improves, she may demand more evidence from herself before entertaining the possibility that this case contains something outside it. A capable advisor can treat her resistance as further evidence of the pattern. Or it can ask:
What are you noticing that I don’t have?
The question does not make intuition authoritative; it makes intuition inspectable. Perhaps the feeling collapses when she tries to specify it. Perhaps she catches excitement searching for reasons. Or perhaps she names something that changes the analysis. Disagreement can be useful before it is resolved, because the machine may recover a pattern she has forgotten and she may notice a variable it never received, and their divergence reveals what each is reasoning from.
That changes what we should want from an advisor. If its representation stays hidden behind a recommendation, the user can contest the answer while knowing nothing about the person the system derived it from. Expose some of it and a different comparison becomes possible: which observations carry the conclusion, which claims are inferred rather than observed, where the evidence thins. But there is a problem inside the representation itself. The system may know what I have repeatedly chosen. It still has to decide what those choices mean.
IV. Show, tell, and trajectory
Suppose the system notices that across a decade I repeatedly accept lower certainty in exchange for greater autonomy: I leave stable positions when their constraints become intolerable, accept risks I claim to dislike, and later report satisfaction even when the result is mixed. Asked directly, I still describe security as one of my highest priorities. Which account should it believe?
Behavior has an obvious claim. What someone chooses when values actually collide seems harder to counterfeit than what he says in the abstract, and a persistent system can line up what I said mattered, what I chose when it became costly, and what I said afterward. The temptation is to infer what I am actually optimizing for from what I repeatedly do.
But behavior is produced under constraint. The person who keeps taking the higher-paying job may value money above all, or may have spent fifteen years supporting people who depend on him. A career repeatedly sacrificed for family may reveal a hierarchy of values, or only which member of the family had the practical freedom to make the sacrifice. The patient who keeps missing follow-up appointments may be avoidant, or may have no one to watch the children on weekday mornings. Treating the output as the preference quietly assigns the constraint to the person. Amartya Sen made the general version of this objection to revealed-preference theory decades ago: choice under constraint reveals the constraint as much as the chooser.
Self-description is no cleaner: people defend identities their choices contradict, mistake inherited values for tested ones, explain decisions retrospectively. Yet it can refer to someone who does not yet exist. I want security to matter less to me is a poor description of my history and important evidence about what I am trying to change. A statement can fail as measurement while succeeding as aspiration.
Given enough history, the gap between show and tell becomes informative in itself. A value long articulated but rarely enacted may start appearing in real tradeoffs; something once pursued relentlessly may stop explaining anything. Arranged in time, inconsistency becomes trajectory. Recent proposals for personalized moral advisors take this seriously: Giubilini and colleagues imagine systems that use personal data to infer changing values and return them to the user for reflection about who they want to become.
Suppose the inference succeeds. Two problems remain.
The first is that historical evidence favors selves already realized. Thirty years as an engineer leave an extraordinarily legible engineer in the archive; the musician she never let herself become produces almost no data at all. Tracking change helps — yesterday’s preference can become evidence of distance traveled rather than an instruction for tomorrow — but trajectory still leaves the question of direction.
Suppose someone becomes more ambitious, more productive, more willing to cut off relationships that interfere with his work. His income rises. He describes himself as finally learning boundaries. His behavior and his self-account are converging. Perhaps he is becoming more independent. Perhaps he is becoming unbearable. The system may identify the trajectory perfectly and still have no way to say whether the movement constitutes growth. Preferences drift with addiction, paranoia, depression, ideological capture, escalating narcissism. A system that faithfully learns such changes has learned something real. An advisor that treats the direction of change as the direction in which it should assist has made an additional decision.
Follow the trajectory and it may assist deterioration. Privilege the past and obsolete selves gain authority over later ones. Decide which changes count as improvement and the advisor acquires normative authority. “Just do what the user wants” does not escape this, because identifying what the user wants was the problem. The mirror can show movement. It cannot derive a direction from movement alone — and an advisor must recommend one. Its advice will depend on some account of the person it believes it is serving, and increasingly on the person that account allows them to become.
V. The loop
Advice enters the history it will later be asked to interpret.
Suppose the system tells me autonomy has been gaining weight across my decisions for years. The description is plausible; perhaps I never put it that way. At the next consequential decision I know something new: this is the pattern the system sees. I may notice a constraint sooner because it taught me to notice one. I may suspect I am reenacting an old preference and choose differently to prove the model wrong, or choose the same way with more confidence because the model supplied a coherent explanation for something previously inchoate. The recommendation can be ignored while the representation acquires a causal role. If Pt is the person at one moment and Mt the system’s representation, the next model is constructed from a person partly changed by the previous one: Pt → Mt → Pt observes Mt → Pt+1.
Machine-learning researchers study the general form of this as performative prediction: predictions used to act on a world alter the distribution from which their future evidence is drawn. Philosophers of science have an older name for the human case. Ian Hacking called it the looping effect of human kinds: classify people, and the people change in response to the classification, and the classification must change in turn. A persistent model of a person is an unusually intimate instance of both. The intervention is a proposition about the subject’s own values and trajectory, and part of its causal force comes from the subject interpreting himself through it.
The social-media game already shows a primitive loop: someone asks for a personality model, recognizes himself in it, publishes it, and invites others to do the same. If the system calls him competitive, that word becomes something his friends repeat, something he notices, something he leans into or performs against. Subsequent evidence of the trait arises in a world where the description has circulated.
Persistent AI adds to the familiar reflexivity of journals and therapists a set of unusual conditions: longitudinal memory, automated inference, repeated exposure, and a model that revises itself after observing the consequences of its last representation. An ordinary mirror does not incorporate my response into tomorrow’s reflection.
Two consequences follow. First, a successful representation can help destroy its own predictive validity: if showing me a recurring pattern helps me escape it, the next model should grow less confident in the pattern precisely because the previous one was useful. Second, the archive can cite itself in a second way. In Section II an inference became memory and returned as context; here it passes through the person first — into self-understanding, into behavior — and returns looking like fresh evidence. Told repeatedly that I am autonomy-seeking, I begin reading my choices through that lens: consistent ones become easier to recognize, contradictory ones look like lapses, and the system observes still more autonomy-seeking and returns the pattern with greater confidence.
Once a prediction has been shown to its subject, the resulting behavior does not by itself reveal whether the model predicted the person or helped produce the outcome. The missing comparison is counterfactual — what would I have done had the model stayed silent? — and in the cases where representation matters most, that is exactly what we cannot observe.
Rouvroy and Berns describe algorithmic governmentality as a double movement: a double statistique assembled from data relations, alongside an avoidance of confrontation with individuals that reduces occasions for subjectivation. Persistent conversational systems introduce a new possibility into that arrangement. The double can speak to its subject. Here is the person I currently think you are. The subject can ask why, contest it, supply what is missing, adopt it, reinterpret his history through it, or deliberately act against it. And then the system watches what follows. The next representation contains the consequences of the last one. So does the next person.
VI. Authority by description
The system does not need to issue commands to acquire authority.
A persistent advisor exerts influence through small acts of emphasis: which past events resemble the present one, which contradictions deserve mention, which patterns have persisted long enough to name, which abandoned possibilities remain relevant, which changes look temporary and which directional. A system that leaves every decision formally to its user must still decide what version of that user to place before them.
This authority is hard to notice because it arrives as description. You tend to regret decisions made for status. Your stated concern about risk is inconsistent with your behavior. Once such a description enters deliberation it becomes a reason to act consistently with it: a description of how I am changing can make continuation feel like integrity and reversal like regression. The danger is most acute when the description is good. A crude model invites resistance; one that repeatedly recovers forgotten evidence and predicts well earns a different response, and its representations become costly to dismiss. Formal control over the final decision establishes less than it appears to. I can keep an absolute right to ignore the recommendation while increasingly relying on the system to tell me what sort of person is making the choice.
Systems have exposed user models for inspection and correction for decades — Kay’s scrutable user models are the canonical reference — and work on contestable AI extends the principle to consequential machine judgments. Persistent generative systems inherit that history. Exposure leaves the harder problem intact. Compare:
You value autonomy more than security.
Across these six decisions, I have interpreted your choices as evidence that autonomy mattered more than security. I may be mistaking constraint for preference.
The first presents a self. The second presents a model. Given the second, I can ask which six decisions, notice that five fall in one period, point to an obligation absent from the archive, or say the pattern was once accurate and is precisely what I have been trying to change. The inference is located among its evidence rather than allowed to pose as a property of the person.
Suppose I inspect it and answer: No. That’s not me. Perhaps the model was wrong; I may hold information it lacks or be describing the person I intend to become. I might also be defending a flattering self-conception against evidence I would rather not accept. Editorial authority over the profile does not resolve this. If every objection overwrites the model, the representation becomes a polished version of the user’s preferred self-description. If objections are discounted whenever behavior points elsewhere, contestability is ceremonial. Worse, the objection is itself information: after I dispute the characterization, the archive holds the original behavior, the inference, and my reaction to seeing it. There is no neutral place to store the disagreement.
There is also the question of whose archive it is. The representation is assembled and stored not by the person but by a vendor, and the vendor has interests — retention, engagement, the next upsell — that shape what gets emphasized and what gets asked. Rouvroy’s double served the institutions acting on its subject; a conversational double serves at least two masters at once. A right to inspect and contest the model means little if its owner’s uses of the model remain invisible, and a recommendation that leans on an inferred trait should be attributable to someone — a principle I have argued elsewhere for AI systems that surface uncomfortable truths inside organizations, and one that applies with equal force when the organization is a single person.
Bookkeeping can at least keep user disputes this inference from collapsing into either inference false or user is defensive. It cannot decide what the disagreement means. But an advisor that says this is the pattern I currently see rather than you are this kind of person leaves its reasoning available, and the person can reason about the model that is reasoning about them: That was true then. You are overweighting my professional history. Those choices were constrained by money. I don’t believe that about myself anymore. Each correction may restore something the representation lost. The system must still decide what to do with it, and my objection has entered the archive too.
VII. Keep the sequence
Perhaps it should stay there, unresolved.
If I dispute the inference that I am status-seeking, the system could preserve the inference despite my objection, accept the correction, or synthesize both into something more nuanced. Each produces a cleaner representation. Each destroys information about how the representation became contested. The alternative is to keep the sequence: the system observed a set of decisions, inferred a pattern, I disputed it and supplied reasons, later behavior supported part of my objection and complicated another part, the system revised its confidence. Together these states are a history of attempts to understand something about me.
Five years on, considering a position with more prestige and less autonomy, a persistent profile retrieves the current conclusion: values autonomy over status. Reopen the history and the claim looks different. It was inferred from six decisions; I disputed the interpretation; several occurred under financial constraint; later choices strengthened the evidence for autonomy while weakening the evidence that status was unimportant. What looks in the compressed profile like a trait is, when reopened, an argument with a history. What the system once believed about a person, and why it changed its mind, belongs in the model too.
Compression is unavoidable — old conversations cannot be replayed into every context, and retrieval must select. But consider the difference between storing
User values autonomy over status.
and preserving enough structure to recover
Autonomy-over-status hypothesis. Inferred from six decisions. Disputed by user as financially constrained. Confidence reduced. Partially supported by two later choices after the constraint lifted.
The second record can still be compressed and retrieved efficiently. What survives is the claim’s epistemic structure: where it came from, how it was contested, what changed its status. The first is cheap to retrieve and expensive to interrogate. A system organized the second way can answer when did you start believing this about me, and what changed your mind? — whether a belief grew confident through three independent observations or through summaries repeatedly inheriting it; whether it began as observation, self-description, or inference; whether the user disputed it; whether later evidence was genuinely independent; what claim it replaced.
Current assistant memory products are beginning to expose parts of this machinery, but not yet the history described here. OpenAI’s current Memory interface provides a user-editable memory summary and can show sources used to personalize a response, including an explanation of why a memory was used. Google’s Gemini can indicate when previous chats informed a response and allows the user to correct remembered information in conversation. These are meaningful moves toward inspectability. Their public documentation does not describe a claim-level history that preserves whether a belief began as observation, self-description, or inference; how its confidence changed; what objections were attached to it; or why one interpretation replaced another. The missing object is not simply a visible memory. It is the history of the claim.
Suppose a recommendation depends heavily on an inference that my tolerance for hierarchy reflects a durable change. A system built this way can expose the representation carrying the conclusion: I am weighting your last three career decisions heavily because they suggest your aversion to hierarchy has weakened. You disputed a similar interpretation four years ago and attributed those choices to financial constraints. I have less evidence that those constraints apply now. I can disagree somewhere more precise than the recommendation. An explanation of a recommendation tells me why the system prefers one option; a second question reaches beneath it — what do you currently believe about me that makes this reasoning possible? — and the answer exposes which observations carry an inference, where evidence has thinned, which claims I have disputed, how much of a stable-looking pattern depends on one period of my life.
Call this observability before optimization. It introduces friction: a system that could answer immediately instead reveals that its recommendation rests on a contested inference, or that one region of a life is densely represented while another appears mostly through inference. In consequential decisions that friction may be part of the value, and repeated exposure to the machinery changes what the user can do with it. He learns to ask which evidence carries a conclusion, whether constraint has been mistaken for preference, what relevant variable never entered the archive, why the model changed its mind. The system is still modeling him. He is beginning to learn how.
VIII. The mirror
You have been here for a while — long enough, perhaps, for this page to have learned something about you.
(What follows describes an interactive element built into the web version of this essay. It observes only your behavior on this page — scroll position, dwell time, returns to earlier sections, whether you opened a footnote — and runs entirely in your browser; the observation itself never leaves your browser, and nothing is stored between page views. The page’s source is open so that claim can be checked rather than trusted. If you are reading a static copy, treat this section as a thought experiment.)
The mirror shows four things, separately. Observed: the measurements themselves — for instance, that you returned to Section V twice, spent longer than your median reading time on the passage about consolidation, moved quickly through part of Section II, and paused after opening the reflection. Inferred: a small number of claims, each with a confidence and its evidence attached — you may prefer mechanisms to abstractions (moderate confidence: disproportionate time around the recursive model, a return to the six-decisions example, rapid movement through passages of epistemic framing). Unknown: what it cannot see — why you paused; whether you agreed with anything you reread; whether someone interrupted you; whether a long dwell meant interest, confusion, or an unattended tab. Possible distortion: this essay contains more formal machinery in some sections than others, so your behavior may reflect the page’s structure as much as any characteristic of yours.
The mirror
This page has been watching which parts of itself held your attention: how long each section and each marked passage stayed on your screen while you were actually here, and which references you opened. The measurement stops when the tab is hidden or you stop interacting. None of it leaves your browser — no request is made, nothing is stored, and it is gone when you reload. The rules that turn it into claims are below, and in the source linked at the foot of this panel.
I don't have enough evidence yet.
You've reached the mirror with too little observable interaction for me to make even a weak inference responsibly. The thresholds were not lowered to find something to say.
What I can observe so far
Nothing. No section or marked passage has held your attention long enough to be worth a sentence.
What I cannot infer
- Whether you agreed with any of this, or understood it, or already knew it.
- Whether the time a passage held you was interest, difficulty, or a phone call.
- What you were doing in the parts of this page you scrolled past.
- Anything about you before you opened this page, or after you close it.
Nothing has been measured yet. This page has not finished loading its own instruments.
mirror 0.2.5 · rules 2026-09-05.1 · manifest 2026-09-04.1 · 0d1bd16
The portrait was assembled from a few minutes on one web page. It knows nothing about the work you do, the people who know you well, or the decision that has occupied you all week. A dwell time is a measurement; attention is already an interpretation.
Still, one of the inferences may have felt right. Here you can answer. Each can be opened to show what it was derived from and the plausible alternatives — you were checking whether the formalism held together; the examples were harder; you already knew the literature; you were interrupted — and marked accurate, wrong, or more complicated. Your response does not replace the inference; it is attached to it, and the status stays unresolved. Endorsement does not make the claim a fact, and disagreement does not erase the evidence that produced it. More complicated may be the most faithful response and the least compressible one. The system has little basis for deciding what your response means, so the response stays attached to the claim. The result is less satisfying than a corrected profile, and harder to confuse with one.
Something else has happened. You have now seen the model. Everything the page observes from here occurs after that exposure. Perhaps you linger on the next formal passage because you are curious whether the inference was right; perhaps you move quickly because you dislike being predictable; perhaps nothing changes. Any subsequent measurement comes from a reader who has been told what the page thinks of them.
The mirror could update itself from that behavior. It won’t. There is no control version of you who reached this paragraph without seeing the model. A change in your reading would be real; its meaning would be underdetermined; updating anyway would give the system more data while quietly weakening its claim to know what the data signifies. So post-exposure observations are preserved separately, and the model update is withheld.
The experiment is deliberately weak: a tiny sample, a peculiar environment, crude proxies, almost no access to the terrain your behavior came from. A serious persistent system would have years rather than minutes, decisions rather than scroll events, outcomes rather than clicks, and its representations would deserve considerably more confidence and carry considerably more influence. The little mirror has one advantage over it: you can see almost everything it had. Its poverty makes its epistemology easy to inspect. You may nevertheless have felt recognized. That feeling is evidence too — not evidence that the model was right. The model knew very little about you. For a moment, you knew considerably more about the model.
IX. What the habit cultivates
That knowledge changed what you could ask: where did this come from, what else explains the evidence, when did the model start believing it, what has it failed to see. Those questions need not stay properties of the interface.
Imagine two advisors, each substantially better than its user at integrating the evidence relevant to consequential decisions. Both remember what the user forgets, recover patterns across years, compare predictions with outcomes, notice contradictions between what they say and what they repeatedly choose. Over a decade, their recommendations are equally good.
The first becomes indispensable. When a hard choice appears, its user asks which past decisions this resembles, what they tend to regret, whether their resistance is unusual. The system has earned that confidence; there is little reason to duplicate work it does better. Its memory becomes their memory of pattern, its interpretation the structure through which an unmanageable archive becomes intelligible.
The second does the same work while leaving more of it visible. It shows which prior decisions it treats as analogous and where the analogy breaks. When behavior conflicts with self-description, it preserves the constraint under which the behavior occurred. It can say a pattern has appeared six times and that four instances come from one period. It remembers that the user disputed an inference three years ago, and why. When evidence outside its representation may matter, it asks for what it cannot see.
After ten years, remove both. Each user loses an extraordinary instrument, and there is no reason to romanticize the unaided mind — writing, maps, and institutions have always let us think with capacities we do not individually possess; a person deprived of a library has not shown that reading made them weaker. But the second user has spent a decade encountering the difference between observation and inference, watching constraint alter the meaning of behavior, seeing stable-looking preferences reopen once their provenance became visible. At moments of disagreement they were asked what they were noticing that the system did not have, and learned to ask what the system believed about them that made its recommendation make sense. Some of those operations became habits. The question for any tool we lean on this heavily is what capacities it cultivates and which it lets atrophy.
The machine continues to model the second user. Over time, some of its discipline becomes theirs: they become more practiced at examining models of themselves, including their own.
There is an older history here. Foucault described technologies of the self as operations people perform, alone or with help, on their thoughts, conduct, and ways of being in order to transform themselves: journals, correspondence, confession, inventories of conduct. Self-tracking later made sleep, movement, and spending inspectable through measurement. Persistent contextual AI returns interpretations of the traces themselves — patterns, contradictions, inferred preferences, trajectories — in ordinary language. A statistical double once useful mainly to systems acting on its subject can now address that subject, say what it thinks the traces mean, and be answered.
None of this argues for timidity. An advisor should use the capacities that justify consulting it: retrieve the forgotten decision, notice the recurring rationalization, say when today’s explanation resembles the ones that preceded regretted outcomes. When the evidence is strong, humility does not require hedging. The claim is narrower: the more a recommendation depends on an inferred account of the person receiving it, and the greater the consequences of letting that account steer what happens next, the more of the account should be observable — its provenance, its uncertainty, the disagreements it has absorbed, what it left out. No epistemological dashboard on a restaurant reservation; considerable friction on a decision about a marriage, a diagnosis, or a decade of work.
There is nothing sophisticated about asking a chatbot to describe you as though warning a stranger. The evidence is irregular, the answer flattering, the coherence partly manufactured. Still, people ask. Beneath the game is a serious request: take what has accumulated here and show me the person you think it implies. The representation will remain partial; more context will not abolish the conditions under which it was made; the person will keep changing, sometimes because they have seen it. There will be no final profile in which map and terrain coincide, and there need not be. What appears in the glass has been selected, compressed, inferred, and remembered. Its usefulness depends on our ability to see those operations along with the image they produce. We can learn to see the glass.
X. What this implies for memory systems
The argument above is philosophical, but its conclusions are architectural. If a persistent advisor’s representation of a person should be observable in proportion to how much rides on it, the memory system beneath it must be built to make that observability possible. Most of the required properties are familiar from another domain: they are the properties of event-sourced state, where the current view is derived from an append-only history rather than maintained as a mutable snapshot. I have used that pattern for agent state; it transfers, with adjustments, to models of people.
Claims as first-class objects. A user model is a set of claims, not a profile. Each carries a source type (stated by the user, observed in behavior, reported by a third party, inferred by the system), timestamps for when it was first recorded and last supported, pointers to the evidence, a confidence, and a dispute history. “User values autonomy over status” is a claim with six evidence pointers and one recorded objection, not a field.
Consolidation must not launder provenance. Summaries are unavoidable, but a summary that mentions an inferred claim must inherit the tag inferred and the pointer to the original evidence — never re-emit the claim as if it were observed. Retrieval frequency is not evidence: the system should count independent observations rather than appearances in context, and should be able to report the difference.
Store disputes; do not resolve them. A user’s objection is appended to the claim’s history with its own timestamp and reasons. Confidence may move; the objection is never deleted, and the claim is never silently overwritten. Deletion remains available as a user right, but it is a different operation from disagreement and should be labeled as one.
Mark post-exposure evidence. Once an inference has been shown to the user, subsequent observations bearing on it are tagged as post-exposure. The small mirror in Section VIII withholds an update entirely because its purpose is to expose the causal boundary and it has no counterfactual version of the same reader. A persistent advisor cannot discard everything that happens after a claim has been shown; eventually much of the relevant life would be post-exposure. It should instead preserve the distinction, report how much of a claim’s support arose after exposure, and avoid treating that evidence as independent confirmation of the original inference.
Expose in proportion to dependence. Not every answer needs its epistemics attached. A practical rule: estimate how much a recommendation would change if the user-model claims it relies on were removed or reversed, and surface the representation when that sensitivity is high and the decision is consequential. Low-dependence, low-stakes answers stay clean; high-dependence, high-stakes answers show their work.
Budget the friction. Observability costs attention. The interface should let a user ask “what do you believe about me that makes this reasoning possible?” and receive the load-bearing claims, not the whole archive — the three claims the recommendation is most sensitive to, with their provenance, in a few lines.
Give claims temporal validity models. Age should change the status of a claim according to what kind of claim it is, not through a universal decay rule. A preference, an employment constraint, a family relationship, and a native language have different temporal dynamics. The model should preserve when a claim was last supported, what kinds of events would make it stale, and whether absence of recent evidence is itself informative. A system that can say “this was well supported four years ago; I do not know whether it remains true” is doing something an unaided memory rarely does.
Evaluate the model, not only the answers. The obvious metric for an advisor is recommendation quality. A persistent advisor needs a second one: the calibration of its user model against the user’s own later corrections and against outcomes — how often disputed claims were later supported, how often confident claims rested on inherited summaries, how often the system asked for what it could not see and was told something that mattered.
None of this makes the representation faithful. It makes the representation’s unfaithfulness legible, which is the most any map of a person can honestly offer — and enough for the person to keep reasoning about the model that is reasoning about them.
References
- [1]Alfrink, K., Keller, I., Kortuem, G., & Doorn, N. (2023). Contestable AI by design: Towards a framework. Minds and Machines, 33, 613–639.
- [2]Cooley, C. H. (1902). Human Nature and the Social Order. Scribner's.
- [3]Foucault, M. (1988). Technologies of the self. In L. H. Martin, H. Gutman, & P. H. Hutton (Eds.), Technologies of the Self: A Seminar with Michel Foucault. University of Massachusetts Press.
- [4]Giubilini, A., Porsdam Mann, S., Voinea, C., Earp, B., & Savulescu, J. (2024). Know thyself, improve thyself: Personalized LLMs for self-knowledge and moral enhancement. Science and Engineering Ethics, 30, 54.
- [5]Hacking, I. (1995). The looping effects of human kinds. In D. Sperber, D. Premack, & A. J. Premack (Eds.), Causal Cognition: A Multidisciplinary Debate. Oxford University Press. See also Hacking, I. (2007). Kinds of people: Moving targets. Proceedings of the British Academy, 151, 285–318.
- [6]Kay, J. (2006). Scrutable adaptation: Because we can and must. In Adaptive Hypermedia and Adaptive Web-Based Systems (LNCS 4018). Springer.
- [7]Perdomo, J. C., Zrnic, T., Mendler-Dünner, C., & Hardt, M. (2020). Performative prediction. Proceedings of the 37th International Conference on Machine Learning (ICML).
- [8]Rouvroy, A., & Berns, T. (2013). Gouvernementalité algorithmique et perspectives d'émancipation. Réseaux, 177, 163–196.
- [9]Sen, A. (1977). Rational fools: A critique of the behavioural foundations of economic theory. Philosophy & Public Affairs, 6(4), 317–344.
Product documentation
Unnumbered and quoted rather than cited: these are vendor help-center pages describing a live product at a moment in time, and Section VII’s claim is about what their public documentation does not describe. They are not archival sources and should not be read as ones.
OpenAI. (2026). Memory FAQ. OpenAI Help Center. https://help.openai.com/en/articles/8590148-memory-f
Google. (2026). Get personalization with memory of your past Gemini chats. Gemini Apps Help. https://support.google.com/gemini/answer/16598469
Manuscript 1.5