I sent our promotion committee a document titled “Unresolved Questions About Lila's Judgment” three days before her review. The questions concerned an incident from two years ago, a client she had already made whole, and a version of our team's budget that she had never seen. I did not copy Lila. When she found out, she called it sabotage. That word assumes a unitary agent with stable preferences, which is the very assumption shard theory asks us to investigate. My current best model is that a promotion-seeking shard generated the document, while the part of me writing this post objects to several of its choices.

The shard's training history is fairly easy to reconstruct. Early in my career I learned that being helpful in meetings produced approval. Later I learned that finding risks other people had missed produced more approval, especially when the risks involved someone above me. Those rewards need not create a global objective to ruin Lila's life. They can produce a local policy: in contexts resembling promotion review, notice potential downside in another candidate and make it legible to the evaluator. I can observe the policy activating without endorsing its goal. The document's careful formatting is particularly diagnostic; I only use that table style when I want something taken seriously.

Lila points out that I wrote the document, attached the old client email, chose the recipients, and sent it from my account at 11:48 p.m. I grant all four observations. The question is whether they are best explained by a single enduring preference for professional harm or by a learned subroutine pursuing a local proxy for esteem. This matters because an intervention aimed at the wrong level will fail. Her proposed intervention is that I tell the committee the questions were misleading and withdraw from consideration. That would remove the immediate reward, but I worry it might train the shard to hide its work next time.

I have instead proposed a three-week interpretability sprint. I would log the moments when I feel compelled to turn a coworker's mistake into a document and ask two volunteers to rate the internal cues. Lila is the obvious domain expert, given her recent firsthand exposure, although she has declined. Our manager also declined and said that submitting a 14-page protocol to explain a three-page accusation was making the original problem worse. I would have preferred a written explanation of which page introduced the regression.

There is a substantive issue in Lila's complaint that I don't want to evade. My document said she had “unresolved judgment concerns,” even though the client incident had been resolved, and omitted that the budget discrepancy came from Finance. I knew those facts when I hit send. I remember hesitating over each one. The hesitation is exactly the kind of conflict a simple selfish-agent story misses: one learned tendency had access to the correction, while another selected what would be salient to the committee. We should be wary of a moral vocabulary that collapses these internal processes before we can study them.

The committee has postponed both promotions. Several coworkers now avoid discussing mistakes around me, which limits the data for the sprint. On Monday Lila asked, in front of them, whether the part of me that knows the truth could send a correction before Friday. I said I would need to see whether the part that controls the email account agreed.