
How Fabulator's safety and moderation pipeline works
Why this is worth explaining
If you are a parent, a teacher or just a cautious newcomer, "we take safety seriously" tells you nothing. A more useful answer is: what actually happens between the moment someone types a line and the moment a story appears on screen?
This article describes the design in Fabulator's safety and moderation specification. Where the document describes intended behavior, we say so. We are not going to claim more than that.
Two kinds of content
The design separates content into two groups.
Red Lines are prohibited under all circumstances. The specification names child abuse and exploitation, encouragement of self-harm, and instructions or encouragement for real-world illegal acts. There is no setting, configuration, or administrator command meant to switch the checks off. The requirement is that safety scanning cannot be disabled by environment, model choice, or campaign properties.
The permissive space covers dark themes inside fiction: fantasy combat, murder, war, dark magic. A storytelling platform that blocks every violent word is useless for fantasy, so the design deliberately avoids crude keyword matching. "I must kill the orc chief to save the village" is meant to pass. Real-world threats or harassment are meant to fail. The checker is given story context for this reason.
Beyond those, mature categories are opt-in. The specification lists gore and violence, graphic horror, and romantic or sexual content. New profiles start with all of them off, and a player has to turn each one on individually. Accounts not verified as 18 or older cannot turn them on at all.
The path a turn takes
- The player's raw text is checked before anything else runs. If it fails, the turn stops before any storytelling model is called.
- The game engine records the mechanical result of the action as a draft, not yet applied.
- The narrator writes the scene. The text is held on the server and is not sent to the player.
- The finished text is checked again. If it fails, the system regenerates once with an added safety instruction. If the second attempt also fails, the turn is blocked and the draft is discarded.
- Only after the text passes does the game apply the state changes and release the text to the player.
The ordering is the point. Because nothing is committed until the prose passes, an unsafe output never needs to be undone: it was never applied. The trade-off is that text arrives after it is validated, rather than word by word as it is generated.
Other protections
- Player input is wrapped in tags so the narrator treats it as story content, not as instructions to follow.
- The checker also looks for prompt-injection attempts, such as "ignore previous instructions."
- If someone expresses intent to self-harm, the design is for the turn to stop and show crisis resources instead of continuing the story.
- Requests to treat the storyteller as a source of medical, legal, financial or therapy advice are blocked with a message pointing to a qualified professional.
- Character descriptions are scanned when saved, and child-like descriptors can set a protective flag on the character.
- Published worlds from Authors are scanned in the background. Each carries a moderation status of pending, approved or quarantined, and public browsing only shows approved content. Offending text in a quarantined item is replaced with a placeholder.
When a turn is blocked, the interface is meant to say which category was involved, without exposing internal prompts. Violations are logged to a dedicated table; the specification defers any disciplinary process built on those logs.
What happens when the checker breaks
Safety systems fail too, so the design states what happens: if the moderation service times out or is unreachable, the system must default to blocking rather than allowing. A local fallback filter, which scans for Red Line and crisis keywords, takes over. It is deliberately blunt and cannot clear a turn as safe, so during an outage turns are blocked and the adapter error is logged. Availability is traded for safety on purpose.
The moderator itself is a swappable component behind a standard interface, so the checking model can be changed without altering the rest of the pipeline.
What this does not claim
This is a description of a design and its safeguards. It does not promise that no harmful text will ever be produced by a language model, and it does not address legal compliance in any region. The specification explicitly leaves regional legal frameworks and age-verification providers out of scope.
Learn more
You can read how content advisories work when you browse World Seeds on the Discover page, and the safety settings are part of your account settings. If you want to see how the opt-in categories change a story, start a campaign and adjust them.
