
Features
Part of Platform moderation: a four-layer, six-stage framework
Testing a platform moderation policy change against the evidence
Platform moderation policy example: follow a fictional community testing a harassment rule, reviewer guidance, notices, appeals, metrics, and rollback limits.
What to take away
- This fictional case separates a policy problem from a staffing or detection problem.
- The pilot defines harassment by conduct, target, repetition, and context.
- Review quality, notice quality, appeal outcomes, and user exposure are measured separately.
- Reversals are treated as feedback, not proof that appeals or initial enforcement alone failed.
- The pilot has a rollback trigger for language-specific error.
Northline Forum is fictional. It hosts neighborhood groups in English, Spanish, and Vietnamese. During a local election, moderators receive repeated reports about users posting another member's workplace and urging others to contact the employer.
The old rule says, "No harassment or inappropriate conduct." Reviewers disagree about whether public workplace information is allowed, whether one post counts, and whether political criticism changes the result. That is the vague-rule failure in miniature: adjectives instead of conduct, targets, and exceptions.
Baseline review
The team samples 240 closed reports from the prior month, stratified by language and outcome. Two reviewers independently recode each case using the old rule. Independent recoding is borrowed method; the forum rule-change pilot runs the same two-reviewer baseline.
Baseline moderation measures
- 68%Reviewer agreement
- 31%Notices naming behavior
- 19 hoursMedian closure
- 16Reversals after appeal
| Baseline measure | Fictional result |
|---|---|
| Reviewer agreement on violation or no violation | 68% |
| Notices naming a specific behavior | 31% |
| Appeals completed | 44 |
| Decisions reversed after appeal | 16 |
| Median report closure | 19 hours |
| Cases involving off-platform contact requests | 27 |
These invented figures do not show how real services perform.
The revised rule
The pilot prohibits directing or encouraging others to contact a private person's employer, school, household, or family to punish participation in a forum dispute. It covers a single call to action when the likely effect is targeted pressure. It excludes providing an official public contact for a government office in a discussion of that office's duties.
The policy adds examples in all three languages and separates identifying information from the call to contact. A reviewer can remove the contact details, restrict the organizer, preserve evidence, and assess credible threats through a higher-risk route. Routing by severity follows the pipeline's triage stage; a doxing case and a scope dispute never share a queue.
Review and notice design
Reviewers receive ten paired examples, a context checklist, and an escalation channel. The notice must identify the post, the call-to-contact rule, the audience-directed phrase, the action, duration, and appeal route.
An appeal goes to a reviewer who did not make the original decision. The reviewer can restore the post, remove a strike, reverse a feature limit, and record the reason for correction.
Why reversals matter
A European Commission release about DSA moderation appeals and reversals reports aggregate figures for internal challenges under the EU framework and says many appealed decisions were reversed. Northline uses the release only to support tracking reversals as a meaningful system measure. Its legal obligations and fictional numbers are not inferred from the EU totals.
Reversal can reveal a rule problem, missing context, reviewer error, system error, or new evidence. The code matters more than the count alone.
Record the reason code for every changed decision.
Four-week pilot measures
The team records:
- report volume by language and source
- time to first safety action and final outcome
- notice completeness
- appeal access, completion, outcome, and time
- reversal reason
- repeat attempts to repost contact details
- affected users' exposure reports
- reviewer workload and escalation use
It does not use raw removal volume as the success target.
Australia's eSafety Commissioner report on mandatory transparency notice findings notes that time to a moderation outcome can help services assess the efficacy of trust and safety systems and track change. Northline therefore keeps timing, but pairs it with accuracy, reversal, and exposure measures so speed does not reward careless closure.
Fictional pilot results
Baseline
- Independent reviewer agreement
- 68%
- Notices naming specific behavior
- 31%
- Median report closure
- 19 hours
- Appeals completed
- 44
- Appeal reversals
- 16
- Repeat repost attempts
- Not measured
Pilot
- Independent reviewer agreement
- 87%
- Notices naming specific behavior
- 92%
- Median report closure
- 13 hours
- Appeals completed
- 38
- Appeal reversals
- 7
- Repeat repost attempts
- 11
The English and Spanish samples meet the team's 80% agreement threshold. Vietnamese reviewer agreement is 71% across a smaller sample, below the rollback trigger.
Rollback and correction
Northline pauses automated routing for Vietnamese reports under the new rule. It does not abandon the safety action. A bilingual senior reviewer handles those cases while the team checks translation, examples, and queue labels with paid community advisers. The pause treats routing and action as separable, which is the model comparison's core distinction.
The team also finds that seven removed posts quoted the prohibited call to criticize it. It adds a counterspeech example and restores four posts on appeal; three remain restricted because they repeated the contact details to a larger audience without redaction.
What the case can support
The pilot can support a local decision to continue the revised rule in two languages and repair the third workflow. It cannot show long-term deterrence, prove that all affected users felt safer, or compare Northline with a real platform.
The strongest result is procedural: the new rule, examples, notices, and independent appeal produced a record detailed enough to find where the system still failed.
Common questions
Is Northline Forum real?
No. The service, cases, measurements, and results are fictional.
Why did the team pause only one language workflow?
The predefined accuracy trigger failed there. A targeted pause preserves useful parts of the pilot while addressing the measured gap.
Does a lower reversal count prove better moderation?
No. Appeal access, case mix, user behavior, and reviewer standards can all affect reversals.
Why preserve evidence after removal?
Secure preservation can support appeals, threat assessment, legal duties, and pattern analysis while the material is no longer public.
Why measure notice completeness?
Specific notices help users understand the decision, change behavior, and submit relevant context when the decision is wrong.
What would end the pilot entirely?
Severe unaddressed safety harm, sustained low review agreement, inaccessible appeals, or a rule that cannot distinguish criticism from targeting would trigger full rollback.







