Last week my girlfriend asked whether I loved her, and I said 82%.

She stared at me for a few seconds and asked what that was supposed to mean. I told her it meant I assigned an 82% probability to the proposition that what I currently feel toward her should be classified as love, given my prior relationships, my behavior over the past year, and the fact that “love” itself is not an especially well-defined category. She asked why I could not just say yes. I said I could, but it would communicate a level of certainty I did not actually have.

I do not think this is as strange as she does. People are willing to think probabilistically about almost everything except their own feelings. We do not say a startup will definitely succeed because the founders seem smart, or that a medication definitely works because the first trial looks promising. We update on evidence. Romantic relationships are at least as complicated as either of those things, yet the accepted convention is that you compress your entire posterior into yes or no and then repeat whichever answer is more socially useful.

I started keeping track about four months into our relationship, when I was at 74%. The low starting number was mostly a base-rate correction. I had previously believed I was in love three times and, looking back, I think only one of those cases really qualifies. One was mostly sexual novelty plus intermittent reinforcement. Another was an intense preference for not being single combined with the fact that she lived close to my office. My subjective certainty had been extremely high in both cases, which is exactly why I no longer trust subjective certainty by itself.

When my current girlfriend first told me she loved me, I told her I was very bullish. She did not appreciate that phrasing, but the number continued to rise. I noticed that she was usually the first person I wanted to tell things to. I changed travel plans to spend more time with her. I became invested in outcomes in her life that did not benefit me directly. I met her mother twice, including once when there was no holiday requiring it. By our first anniversary I was at 86%.

The problem began when she found the spreadsheet.

I had not hidden it. It was in a folder called personal_metrics, and I simply did not expect her to open it. What upset her most was not that I had quantified the relationship, but that the number had declined from 86% to 82%. I explained that this was ordinary updating. We had been fighting more. Sex frequency was down. I had caught myself enjoying a work trip partly because she was not there. None of those observations proved anything by themselves, but refusing to update on them because the conclusion was emotionally unpleasant would have been straightforward motivated reasoning.

She asked whether I had a column for sex frequency. I said yes, which turned into a different argument.

The larger disagreement is that she wants “I love you” to do two jobs at once. She wants it to describe my internal state, but she also wants it to reassure her, reinforce the relationship, and make future commitment more likely. Those are all reasonable goals, but they are not the same goal. If I say “I love you” when my confidence is 82%, she hears something closer to certainty. I then know that she has incorporated that certainty into her model of the relationship, and future decisions are being made using bad information.

She says this is not how normal people use the sentence. I know that. Normal people are frequently miscalibrated.

I tried explaining that 82% is actually very high. If a prediction market put an outcome at 82%, nobody would call the market undecided. If there were an 82% chance of rain, you would bring an umbrella. She said she did not want to be compared to rain. I did not think that was the important part of the analogy, but I stopped there.

The calibration issue matters because my romantic forecasting record is objectively bad. Looking through old journals, I found at least five statements equivalent to “I know this is the person I am going to marry,” distributed across four women. That is not a subtle error. If I had made public forecasts with that level of confidence and that level of performance, nobody in a serious forecasting community would take my probabilities at face value.

So once a month I record estimates for several propositions: whether we will still be together in a year, whether we will eventually marry, whether I currently love her, whether she currently loves me, and whether a neutral observer with full information would describe the relationship as healthy. I briefly tracked the probability that she was angry with me at any given moment, but the number moved too quickly to be useful and seemed to become self-fulfilling whenever she saw me updating it.

The exercise has improved my calibration considerably. Six months ago, for instance, I put the probability that we would still be together in a year at 91%. It is now 73%. She found this alarming even after I explained that six months had elapsed and therefore the forecast was answering a different question. She said that the fact I thought this distinction would make her feel better was itself concerning.

What I find difficult is that she seems to actively prefer overconfidence when it concerns us. She wants me to say “I love you” without qualification, “I want to be with you forever” despite obvious uncertainty around that claim, and “of course we’ll work this out” during arguments where neither of us has introduced meaningful new evidence since the previous seven times we had the same argument. I understand why those statements are comforting. I just do not understand why comfort should be allowed to determine the reported probability.

She says being in a relationship with me feels like being on a quarterly earnings call. I think that is overstated, although I did once describe a bad month as “relationship headwinds,” and in retrospect that probably did not help.

There are also genuine methodological problems I am still working through. My estimate is partly based on my own behavior, but my behavior changes in response to believing that I love her, which makes some of the evidence endogenous. If I buy her flowers because I think a loving boyfriend should buy flowers and then count the flowers as evidence that I am a loving boyfriend, I am effectively grading my own test. When I corrected for this, my estimate dropped to 79%.

I have not told her that part.

Jealousy is similarly difficult to interpret. Some jealousy is evidence of attachment, but too much may indicate possessiveness rather than love. Distress at the thought of a breakup is also ambiguous because a breakup would mean losing not just her but our apartment, mutual friends, travel plans, regular sex, and approximately $11,000 already spent on vacations booked around both of our schedules. Introspection does not cleanly separate these variables.

She thinks the desire to separate them is itself proof that I do not understand love. I think the opposite. If something matters, that is usually an argument for measuring it more carefully. We measure blood pressure because health matters. We calculate load tolerances because bridges matter. Nobody tells an engineer that if he truly cared about the bridge, he would stop asking how much weight it can hold.

A few nights ago she asked me again whether I loved her. By then my internal estimate was still hovering around 79%, with a fairly wide confidence interval because of the recent fighting. I knew what answer she wanted, and for once I decided that preserving epistemic precision was probably not the highest-value intervention available to me.

I said yes.

She hugged me, relaxed immediately, and told me she loved me too. We had a genuinely good evening for the first time in several days.

That was new evidence.

I am back to 83%.