Content Moderation Policy
Effective: 2026-07-06
This document describes how moderation actually works on Debategle: what is automated, what humans review, what the systems can and cannot see, and what happens when something is flagged. We would rather be honest about limitations than pretend to perfection.
The four layers
- Pattern filters on debate text. Every submitted argument is checked against a curated blocklist of slurs, hate speech, and severe abuse, including common evasion spellings. Matches block the message and issue a strike instantly.
- AI classification of submitted arguments. After delivery, each argument is reviewed by an AI classifier tuned for a debate platform: heated disagreement and offensive opinions about topics are expected and not flagged. It looks for sexual solicitation, grooming, doxxing, credible threats, harassment, scams, and encouragement of self-harm. High-confidence serious violations issue a retroactive strike.
- Voice caption moderation. In video debates, your browser generates live captions of your speech. Captions pass through our servers on the way to your opponent and are screened by the same pattern filters. Banned speech is not relayed and issues a strike.
- Video frame analysis. While video is active, your browser periodically samples a downscaled frame from your own camera and submits it for AI safety classification (nudity, sexual activity, violence, weapons displayed as threats, self-harm, drug use, covered or fake cameras). Frames are processed transiently and are never stored.
Thresholds and actions
Automated systems act on confidence scores. Low-confidence signals do nothing. Mid-confidence signals warn. High-confidence signals strike, and for explicit sexual content on camera, end the video session immediately. Repeated warnings escalate to termination. Strikes follow the three-strike ladder described in the Community Rules; bans escalate with repeat offenses (24 hours, 72 hours, 7 days, permanent). Severe video violations and credible underage reports lead to account suspension pending human review.
Human review
Humans review user reports (triaged by severity), suspensions, appeals, and borderline automated decisions. Moderators see the debate transcript, the report, and the moderation event log (category, confidence, action). They do not see your video or hear your voice, because we never record either.
Known limitations
- Video is peer-to-peer: our servers never receive the media stream. Frame sampling runs in your browser, which means a technically sophisticated bad actor could tamper with it. User reports remain the backstop, and tampering is a bannable offense.
- No classifier is perfect. False positives happen; that is what warnings, human review, and appeals are for. False negatives happen too; that is what reports are for.
- Caption moderation sees text, not tone. Threats delivered in a calm voice with clean words can evade automated detection. Report them.
Record keeping
Every automated moderation decision is logged with its source, category, confidence, and action taken. No media is stored with these events. Logs are retained for 24 months and feed the transparency metrics we intend to publish.