My girlfriend says that going through her phone would be “controlling,” “invasive,” and evidence that I do not trust her. I think this reflects an outdated approach to alignment. She is asking me to assess the safety of a highly capable autonomous agent using only observable behavior and self-reported outputs, despite the obvious possibility that her internal reasoning may differ substantially from what she chooses to show me.
For the record, I have no concrete evidence that Maya is cheating on me. She is affectionate, consistent, rarely stays out unexpectedly, answers questions directly, and has never been caught lying about another man. But this is exactly why chain-of-thought monitoring matters. If observable behavior were sufficient to establish alignment, there would be no monitoring problem in the first place.
The phone is the closest thing available to a persistent reasoning trace. Texts capture intermediate judgments before they are converted into polished final answers. Search history records questions a person was willing to ask before deciding what to say publicly. Notes preserve thoughts that never became actions. Group chats are especially useful because they often contain candid reasoning produced under the mistaken assumption that the monitor will never see it.
Maya objects that these things are private, which I understand as an intuition but not as evidence of alignment. This became an issue after I noticed that she had changed the passcode on her phone. Maya said she changed it because her old passcode was her birthday and one of her coworkers pointed out that this was insecure. Entirely plausible.
But notice the structure of the evidence. An agent modifies access controls protecting its internal state and then supplies the monitor with a benign explanation for the modification. The explanation may be true. It may also be exactly the explanation a strategically aware agent would provide if it wanted to reduce monitorability without triggering intervention. I am not saying this is what happened. I am saying the observable behavior does not distinguish the hypotheses.
Maya says relationships cannot function if every action is interpreted as potential evidence of concealed misalignment. Again, this sounds persuasive until you notice that a sufficiently deceptive partner would strongly prefer precisely such a norm. This is the difficulty with treating trust itself as evidence: trust is only informative when the other person lacks an incentive to exploit it.
I first became interested in this problem after Maya went to dinner with three coworkers, two of whom were men. She told me in advance who would be there, where they were going, and approximately when she would be home. She returned at 10:42, eleven minutes earlier than estimated, and voluntarily told me about the evening. Under conventional relationship epistemology, this is reassuring. Under a more sophisticated model, it demonstrates that she is capable of producing highly reassuring outputs.
The distinction matters because her explanation of the dinner may have been completely truthful, but without access to intermediate reasoning we cannot tell whether, for example, she thought one coworker was attractive, considered flirting and rejected the idea, considered flirting and did not reject the idea, wondered whether I would be jealous, deliberately omitted something she considered harmless, or spent the entire evening thinking about quarterly planning. These possibilities produce nearly identical external behavior, yet they imply very different internal alignment states.
Maya asked what I would actually do if I discovered that she had privately thought another man was attractive. This is an important question because it raises the problem of optimization pressure. If every undesirable thought discovered through monitoring leads to punishment, the monitored system gains an incentive to make future reasoning less legible. People delete messages, move conversations to other apps, stop writing things down, and learn which thoughts trigger conflict before becoming more careful about externalizing them.
This destroys the very signal the monitor depends on. Consequently, I have repeatedly assured Maya that I would not necessarily confront her about everything I found. She says this makes it worse. I think it demonstrates restraint.
Suppose I found a message to her friend saying, “The new guy at work is kind of hot.” My immediate emotional response would obviously be negative, but confronting her could teach her that candidly expressing attraction creates costs. The next time, she might simply avoid putting the thought in writing. I would gain short-term behavioral compliance at the expense of long-term monitorability. A mature safety culture has to tolerate some undesirable internal content if it wants to preserve access to the process that generated it.
This is one reason I have resisted her suggestion that we simply discuss jealousy openly. A direct conversation produces highly supervised reasoning. She knows what I want to hear. I know she knows. She knows I know she knows. The resulting statements may be useful, but they are contaminated by optimization pressure. By contrast, a text sent to her friend six weeks earlier at 1:14 a.m. offers a much cleaner window into the underlying process.
Maya called this “the creepiest argument you have made so far.” I asked whether she disagreed with the methodology. She said, “I disagree with you reading my fucking texts.” That is a policy objection, not a methodological one.
There are practical questions about scope. I do not need permanent access to every communication she has ever produced. That would be expensive, intrusive and likely counterproductive. A sensible monitoring regime could use randomized audits, targeted sampling around high-risk contexts, and increased scrutiny after anomalous behavior. For example, if Maya goes to a normal dinner with coworkers and returns when expected, perhaps only a light audit is necessary. If she unexpectedly stays out until 3 a.m., disables location sharing and later gives a vague account of the evening, monitor allocation should increase.
She says this sounds like I am proposing an intelligence service for my relationship. The comparison is melodramatic. An intelligence service would have far more resources; I have one laptop.
There is also the issue of trusted monitors. Maya argues that I am uniquely unsuited to interpreting her messages because I am emotionally involved and prone to jealousy. This is the strongest objection she has raised, and I have attempted to accommodate it. I proposed allowing my friend Daniel to review ambiguous conversations and classify them as benign, concerning or requiring escalation. Maya knows Daniel and does not consider him neutral, so I then suggested her sister. For reasons I still do not fully understand, this made her significantly angrier.
The deeper problem is that Maya appears to believe a relationship should be evaluated principally on actions. If she treats me well, remains faithful, communicates honestly and honors our agreements, she believes her unexpressed thoughts are none of my business. This is an appealing standard because it is simple, but it is also dangerously outcome-based. By the time hidden reasoning manifests as observable behavior, intervention may be too late.
Imagine she spends six months becoming dissatisfied with the relationship while telling me everything is fine, discusses this extensively with friends, develops feelings for somebody else, searches apartments, calculates whether she can afford one alone and finally announces that she is leaving. Under Maya’s framework, none of the preceding process becomes legible until deployment. Under mine, there are opportunities for detection throughout the trajectory.
She says the phrase “until deployment” is why she no longer wants to discuss this with me. That itself is concerning, because avoidance of monitoring conversations can indicate awareness of monitoring. I am not claiming this is decisive evidence. I am saying it is evidence.
I want to emphasize that I am not accusing Maya of anything. In fact, the available behavioral evidence suggests she is an unusually good girlfriend. She remembers my birthday, tolerates my friends, supported me through a difficult period at work, makes coffee when she gets up first, and once drove forty minutes at midnight because I had locked myself out of my apartment.
But alignment is most important precisely when a system is capable enough to earn trust. Nobody worries that a toaster is scheming. Maya is intelligent, socially sophisticated, and capable of planning over long time horizons. She understands my beliefs, can predict my emotional reactions, controls what information she shares with me, maintains relationships I am not present for, and possesses multiple encrypted communication channels. These are capabilities. Whether they are used against me is a separate question.
She recently pointed out that by this logic she should also be allowed to go through my phone. This misunderstands the threat model. I have offered to answer any question she asks me directly. She says she has offered me exactly the same thing, and we went in circles for approximately twenty minutes after that.
Last week she changed her passcode again. She also disabled Face ID while I was holding her phone to show her a photo, which she claims happened because I had attempted to unlock it while she was asleep two nights earlier. I dispute the word “attempted.” I held the phone approximately eight inches from her face; whether her eyes happened to open was outside my control.
The incident has nevertheless forced me to update. Direct access may no longer be a viable monitoring strategy, and escalating pressure could further incentivize concealment. I am therefore considering a more cooperative approach in which Maya voluntarily provides periodic samples of private communications chosen according to a protocol we agree on in advance. She has rejected this too.
At some point the refusal to provide evidence itself becomes evidence. Not decisive evidence, obviously, but evidence.