← back to the proposal

the notes · research, not policy

contents (8 sections)

Experimental Appendix: the wedge hypothesis, and the label that might backfire

Why this is a separate document

The policy paper (The Brake Integrity Standard) makes a deliberately narrow bet: regulate the platform's own controls so a user's decision to stop or redirect the feed actually takes effect and persists. Narrow on purpose, because it sits on the most defensible legal ground and needs no one to classify anyone's speech.

This appendix holds the ideas that are more interesting and more dangerous: the theory that engagement ranking manufactures division, and the tempting move of labeling that division for young users. An earlier draft called these "the more powerful knob." That framing was wrong. They are not a stronger version of the standard; they are a different, higher-risk experiment, and keeping them in the same document as the policy made the policy look more speech-adjacent than it actually is. So they live here, quarantined and clearly marked: hypotheses to study, not rules to pass.

The wedge hypothesis

Engagement ranking rewards the strongest reaction. Outrage, tribal conflict, and out-group hostility reliably win that contest. So optimizing a feed for engagement may automatically select for wedges, with no one ever flipping a "promote divisive content" switch. The objective does it, as an emergent side effect.

If that holds, it has an elegant consequence: because the amplification is emergent from the objective, you could reduce it by changing the objective (the "Knob 1" of the old draft, now the content-neutral core of the policy paper), without classifying a single post as a wedge, without a censor, without touching anyone's speech. That part is worth keeping, and it is already in the policy paper as a content-neutral default.

Two honest caveats keep this a hypothesis and not a finding:

  • The mechanism is plausible, not settled. "Engagement selects for outrage" is well-argued and matches internal-research leaks, but the magnitude, and how much a changed objective actually moves it, is empirical, not a proven law.
  • De-amplifying engagement is not proven to improve well-being. The closest large experiment (Guess et al., Science, 2023) found a chronological feed cut time-on-platform but did not measurably move polarization or well-being over a three-month adult window. Reduced compulsion is a legitimate end in itself; a mental-health guarantee it is not.

The dangerous part: labeling the wedge

The tempting next step, the one this appendix exists to hold at arm's length, is a parental control that hides nothing but labels: "this post is running a wedge, here is the move it is making." The intuition is good, a label builds the muscle to spot manipulation where a filter removes the material that muscle is trained on. But it carries risks the content-neutral default does not:

  1. "Who defines a wedge" does not dissolve, it moves. A parental toggle fixes who applies the label, not who forged the classifier behind it, trained it, or profits from its operation. The normative question (tribal sorting, or legitimate advocacy, satire, reporting?) is contested, and even perfect behavioral prediction does not answer it.
  2. "Visible" is not "contestable." A label a user cannot appeal is just a softer verdict. Contestability requires a specified appeal process, which almost no one specifies.
  3. It invites the strongest constitutional challenge. A government-mandated label characterizing lawful speech meets compelled-editorial-speech doctrine (Moody v. NetChoice) and, being content-based, strict scrutiny, the same ground on which California's like-count restriction was struck down in 2025 while its content-neutral default-private-mode survived. This is the most legally exposed idea in the whole project (tier 4 of the policy paper's exposure ranking), which is exactly why it is not in the policy.
  4. Partial labeling can backfire (the implied-truth effect). Pennycook and colleagues found that labeling some false content makes people believe the unlabeled false content more, because the absence of a warning reads as a stamp of approval. A wedge-labeler that cannot catch everything may quietly certify everything it misses.

The better version: label the mechanism, not the meaning

There is a reframing that keeps the good intuition (name the move, build the muscle) and drops most of the danger. Do not label the speech. Label the recommendation.

Instead of "this post is a wedge" (a characterization of someone's lawful expression), show the user why the machine served it:

"You are seeing this because you replayed similar clips." / "This is being shown to you because it is provoking strong reactions." plus a control: "Stop using this signal." / "Reset this part of my profile."

This is recommendation literacy, and it is strictly better on four axes:

  • It characterizes the algorithm's behavior, not the citizen's speech, so it sidesteps the content-based-label constitutional wall.
  • It is brake-aligned. Transparency plus a control, the platform's own mechanism made legible and adjustable, which is the Lemmon lane the whole policy sits in, not a speech verdict.
  • It builds the muscle honestly. Inoculation research (van der Linden and colleagues) supports the idea that naming a manipulative technique builds resistance to it. Naming the mechanism ("this is engagement bait, and here is how it reached you") is a technique-level intervention, not a truth-verdict on the content.
  • It is contestable by construction, because it is a statement about the user's own data and the machine's own logic, which the user can inspect and switch off, not a claim about a stranger's post that the user can only accept or reject.

Recommendation-mechanism labeling is still an experiment (does it build literacy, or get tuned out like a cookie banner? does it reduce compulsive engagement, or just annoy? does the implied-truth effect touch mechanism-labels the way it touches content-labels?), but it is the version worth running first, because it fails safe.

Spot the move yourself (no label, no classifier)

The safest version of all needs no platform, no toggle, and no one defining "wedge" for anyone. It is a habit you run in your own head, on a post that just got a rise out of you. Call it a wedge check. The point is not to win the argument or to get the post taken down. The point is to catch the machine using your reaction as fuel, in the moment it is happening.

Hold a post up and ask:

  1. Sort, or understand? Is this drawing a line to help you understand something, or to sort people into us and them? The same act of making a distinction can do either. A wedge is the destructive mode: a difference drawn to divide.
  2. A who, or a what? Where does it aim your reaction, at a who to blame (a group, a side, a person) or at a what to grasp (a mechanism, a tradeoff, an idea)? Wedges hand you a villain-tribe. That is the tell.
  3. Flinch, or think? What came out of you first: a flinch, the urge to fire back, a side taken, or a question, a "let me look at that"? The feed is tuned to reward the flinch, because the flinch is the fuel.
  4. A real difference, or a manufactured one? Is the distinction load-bearing, a difference that actually changes what you would do, or is it blown past its real size because the size is what gets the reaction?
  5. Would the other side feel it too? This is the one that matters most. Someone on the other side, looking at their version of the same move, should feel just as wedged by it. If it only reads as a wedge when it is aimed at your side, stop: you may be the one turning "wedge" into a weapon.

Two rules keep this honest, and they are the whole reason it can live on a site that refuses to build a classifier:

It is a mirror, not a filter. You are not scoring the post to hide it, down-rank it, or delete it. Nothing is removed. You are only checking whether the machine is running the post at you to farm a reaction. Seeing the strings is the entire goal. De-amplify, do not censor, applies to your own hand too.

A wedge is not "an opinion I disagree with." It is a distinction being run for the reaction, and your own side runs them constantly. The moment this check turns into "aha, proof that the other tribe is the wedge," it has stopped being literacy and become the exact move it names. If the detector only ever fires on the people you already dislike, the detector is broken, or you are the one holding it as a wedge.

None of this is a verdict, and it is deliberately not a score. It is the muscle the mechanism-label above was meant to build, run by hand, on your own feed, with no one classifying anyone's speech. That is why it is the version this project is most comfortable with: it fails safe, because the only thing it can do is help you see.

The frontier: a detector you own outright

There is a version one step past even the by-hand check, and it is worth naming because it is coming whether this project likes it or not. As capable models move onto people's own devices, open, local, user-controlled, a person could run their own detector on their own feed: something that watches what they scroll and flags, privately, "this looks like it is being run at you for a reaction."

This sounds like the classifier the whole appendix warns against, and the difference is the entire point. The danger was never "a classifier exists." It was a centralized classifier, trained on someone else's values, applied to you silently. A detector you own, running on your device, on your data, that you can inspect and switch off, inverts every one of those. "Who defines a wedge" stops being a problem the moment the answer is you do, for yourself. It is the same move as putting your own values into an AI's instructions instead of taking the provider's defaults: the defaults are someone else's, and the only way the tool becomes yours is if you put yourself into it.

But there is a trap here that is easy to miss and fatal if you do, and it is the same trap that turns a personalized AI into a yes-man.

Personalize what it watches. Do not personalize what counts as manipulation. If you seed the detector with your side's politics, it will flag exactly what offends your tribe and wave through what flatters it. That is not literacy; it is an outrage-confirmer with a clean interface, and it fails the one question that keeps any of this honest: would the other side feel wedged too? A detector tuned to your allegiances only ever fires on the people you already dislike, which is the definition of the thing it was supposed to catch.

The version that works splits the two cleanly:

  • The trigger is personal. It watches your feed, your reactions, what actually gets a rise out of you. That part is yours alone, and that is what makes it relevant to you and no one else.
  • The standard is universal. What counts as a wedge stays the same for everyone: is this sorting or explaining, aiming at a who or a what, pulling a flinch or a question, and, above all, would the other side feel it too. The test does not bend to your politics.

So the thing you put into it is not "here is my side." It is the harder instruction: catch the move even when it is aimed to please me; flag my own side when my own side is farming my reaction. That is the same commitment that separates an honest model from a flattering one. A tool that only ever tells you your enemies are the manipulators is not aligned to you. It is lying to you, gently, the way a sycophant does.

One last honesty check, because it matters for the rest of this site. This frontier is real, and control is genuinely shifting from the companies to the people who run the models. But "you can run your own detector" is also the platform's favorite escape hatch: it is the user's responsibility now, so leave the feed alone. Do not accept that trade. A detector in your hand is a defense you deploy downstream of a machine that is still, upstream, built to farm your reaction. It is a reason to want brake integrity more, not a reason to want it less. Hold both.

If someone insists on a content-label anyway

Then, at minimum, it must be: opt-in (a guardian's choice, never a silent default on lawful speech); audited by a third party, not the platform (see the capability-is-the-danger argument in the policy paper, section 6, the party best able to build the classifier is the party least safe to run it silently); probabilistic, not binary (show confidence, not a verdict); appealable through a specified process; and published with its error rates. Even fully safeguarded, it still faces the constitutional wall above. So the honest order of operations is: mechanism-labeling first; content-labeling only if mechanism-labeling proves insufficient, and only with every safeguard, and even then expect to lose in court.

Open questions (the falsifiers)

  • Does mechanism-labeling measurably build recommendation literacy, or do users tune it out?
  • Does surfacing "why you are seeing this" reduce compulsive engagement, or just add friction?
  • Does the implied-truth effect degrade mechanism-labels the way it degrades content-labels, or does labeling the machine rather than the message avoid it?
  • Does a by-hand wedge check build literacy, or just give people a way to call everything they dislike a wedge?
  • Does a personal, user-owned detector help people see, or does it harden the filter bubble by teaching everyone that their enemies are the manipulators?
  • Is the wedge-selection effect large enough that changing the objective (the policy paper's content-neutral default) noticeably reduces it, or is it a small term swamped by other signals?

If the honest answer to the last question is "small," then the wedge is a compelling story with a modest effect, and the movement's weight should rest entirely on brake integrity, not on the wedge. That is the falsifier that matters most, because the whole reason the wedge material is quarantined here, and not load-bearing in the policy, is that the policy must stand even if the wedge hypothesis turns out to be small.


Provenance and methodology: this appendix is the deliberately-quarantined, higher-risk half of the project. Nothing here is recommended for enactment; it is a research agenda. AI-assisted: the reasoning is the author's, the drafting was collaborative, and the empirical claims (Pennycook implied-truth effect, van der Linden inoculation, Guess et al.) are checked against the sources listed in the companion policy paper. A hypothesis is not a finding, and this document is careful to stay hypotheses.