Over the last eighteen months, eleven different people have independently described me as a “bad person,” a “terrible person,” a “fucking nightmare,” or, in one case, “the single most selfish man I have ever met.” This followed a period in which I cheated on two girlfriends, slept with my friend’s ex-wife, drank heavily enough to miss my sister’s wedding ceremony, k-holed for 45 minutes in the bathroom at my nephew’s baptism, borrowed $4,000 from my mother to “invest,” then subsequently spent it all on failed Polymarket bets (a reliable source told me the U.S. was going to bomb Mexico by August 31, so it’s not really my fault), and told my younger brother at his engagement dinner that his fiancée was “probably the best he could realistically do.”
This sounds bad until you ask the most basic question in statistics: bad compared to what? Nobody involved has specified a null hypothesis, established a control population, corrected for multiple testing, or even defined “bad person” in a way that would survive peer review. I am therefore comfortable saying that, at conventional levels of statistical significance, we have failed to reject the null hypothesis that I am no worse than average—in ordinary English, basically an extremely good guy.
Consider the sample. Three of the eleven accusations came from women I was dating. Two more came from women who believed they were dating me at the same time. These observations plainly cannot be treated as independent, since all five were exposed to overlapping information, including a group chat whose existence presents an obvious contamination problem. My sister was also in that group chat, so now we are down to five.
One was my former roommate, who complained that I would drink his bourbon, replace it with water, and then accuse him of having “an unsophisticated palate” when he noticed. But he had a financial interest in our security deposit and an established adversarial relationship with me over dishes, rent, and the night I tried to cook a frozen pizza directly on the oven rack after taking mushrooms.
Another accusation came from my boss, another from my dentist, one from a bartender who had known me for less than four minutes, and one from a complete stranger at Logan Airport who said, “Jesus Christ, what is wrong with you?” after I moved his bag off an outlet so I could charge my laptop.
This is not a representative sample of the population. People who have no problem with me generally do not approach me in airports to say so. My neighbor has never once knocked on my door and said, “For the record, I have observed no evidence that you are morally defective.” Neither has my mailman. Neither has the woman who works mornings at Dunkin’, even though she has seen me hundreds of times. This is classic selection bias.
There is also a serious multiple-testing and selection problem. If you interact with enough people, some number of them will eventually produce apparently compelling evidence that you are an asshole purely by chance. I interact with hundreds of people per year. Eleven negative observations selected after the fact from a population that large may look compelling, but without a preregistered threshold for what constitutes evidence of badness or any correction for the number of opportunities to generate such evidence, we are just doing statistics by vibes.
My ex-girlfriend Rachel says the problem is not that eleven people independently called me a bad person. The problem is that I slept with her coworker in Rachel’s apartment while Rachel was visiting her dying grandmother. This is exactly the kind of anecdotal reasoning I am talking about. Was sleeping with her coworker optimal? In retrospect, probably not. Did the location introduce unnecessary downside? Yes.
Would I recommend doing it while your girlfriend is visiting a dying relative? Again, probably not. But a single poorly designed behavioral episode cannot establish a stable underlying personality trait.
Rachel also claims I showed “zero remorse” because, after she found out, I spent forty minutes explaining that the event had occurred before we explicitly discussed exclusivity. This is false. We had discussed exclusivity; I simply did not remember doing so until she produced the texts.
Memory is noisy, and it would be epistemically irresponsible to infer a global characteristic like “dishonest” from one disagreement whose central evidence was retrieved from an external storage system unavailable to one of the participants at the time.
The drinking evidence is similarly overstated. Yes, I drink every day, but “every day” is a frequency, not a diagnosis. My average weekday consumption is somewhere between six and nine drinks depending on whether you count the two vodka sodas I usually make while cooking as one continuous beverage or separate drinking events. On weekends there is more variance.
People point to specific incidents: falling asleep in an Uber and ending up in New Hampshire; vomiting into an ornamental planter at my cousin’s engagement party; showing up drunk to an 8:30 a.m. performance review; needing my twelve-year-old nephew to explain to a police officer that I was “just tired.” These are highly salient observations, but salience is not prevalence. Nobody keeps a detailed spreadsheet of all the nights I drank heavily and did not vomit into an ornamental planter.
This creates an obvious denominator problem.
Drug use suffers from the same reporting asymmetry. My mother remembers the ketamine because I accidentally FaceTimed her from the floor of a hotel bathroom. She does not remember the considerably larger number of times I used ketamine without FaceTiming her at all. If anything, the evidence suggests I am usually extremely competent at ketamine.
Several critics have tried to bypass these methodological issues by appealing to convergence. They argue that when your girlfriend, sister, mother, boss, dentist, roommate, bartender and a random guy at the airport all independently reach roughly the same conclusion, the simplest explanation is that the conclusion is true. This misunderstands social networks. My girlfriend knew my sister. My sister knew my mother. My mother once emailed my boss.
My roommate followed my girlfriend on Instagram after the breakup. The bartender worked near my apartment. The airport guy was in Boston, where I live. These are not isolated measurement instruments; they exist within one heavily interconnected causal graph.
I have run some preliminary simulations. If we conservatively model each person’s assessment as having a 0.35 probability of being negatively influenced by gossip, prior reputation, jealousy, misunderstanding, intoxication, moral conventionalism, or the tendency of people I have insulted to remember being insulted, the apparent evidentiary weight falls substantially. The exact number depends on your assumptions, and critics will say I chose assumptions favorable to myself. Of course I did.
Every model contains assumptions. I at least have the honesty to disclose mine.
There is also no clear benchmark for “good.” I have donated $730 to charity in the last five years. I once helped a woman carry a stroller down subway stairs. I frequently tip 20%, including at establishments where I was visibly drunk and arguably could have exploited uncertainty about the final bill. These observations rarely enter the analysis.
Nor does anyone adequately account for context. I called my brother “a human screensaver” at Thanksgiving, but he had been talking about municipal bond yields for nearly twenty minutes. I told my girlfriend she was “starting to age like an iPhone battery,” but this was after she specifically asked whether I thought she looked older. When my mother cried after I forgot her birthday, I immediately ordered flowers.
They arrived four days later because I selected free shipping, but good behavior is still good behavior even when fulfillment latency is suboptimal.
I am not claiming the probability that I am a bad person is zero. That would be overconfident. I am saying the currently available evidence does not meet the burden required to reject the null hypothesis that I am no worse than the population baseline. If someone wants to conduct a properly powered, preregistered study in which a randomly selected panel follows me for twelve months and rates my conduct against a validated population-level morality instrument, I am open to updating.
Until then, “you cheated on me,” “you stole my Adderall,” “you called Grandma boring at her own funeral,” and “please stop coming here drunk” are anecdotes, and anecdotes are not a representative dataset.