X's Multi-Step Reply Spam Filter: What Its Code Shows (2026)
X added a filter on September 4, 2026 for chains of self-replies under accounts with 1,000+ followers. What its public code checks, exempts and does.

X's open-source code contains a filter, added on September 4, 2026, that targets chains of self-replies under posts from accounts with at least 1,000 followers. When an AI model flags at least 2 posts of the same author in such a chain, X's code gives each flagged post a reply-spam label and a reply ranking score of 0.
What did X add to its open-source code on September 4, 2026?
X's public repository xai-org/x-algorithm gained a new reply-spam flow, multi_step_reply_spam, on September 4, 2026. The commit added four files in grox/flows/reply_spam/:
plan_multi_step_reply_spam.pychains 4 steps: an eligibility filter, media hydration, spam detection, and a write step.task_multi_step_reply_spam.pyholds the filter, the detection step and the write step.classifier_multi_step_reply_spam.pybuilds the thread the model reads, parses its answer and applies 3 safeguards.state_multi_step_reply_spam.pystores the result: a list of flagged post IDs and a reason.
X's README describes the grox/ folder as code that runs as posts are published, with classifiers for categories such as spam. The new plan is registered with the same stream generators as X's other reply-ranking flows.
The task file was edited again on September 8, 2026: its write step now calls one publish function instead of two separate writes. The value it writes is still 0.0.
Which replies can the multi-step reply spam filter target?
Free, no login, results in seconds.
The filter (TaskMultiStepReplySpamFilter) selects a reply only when all of these are true:
- The post is a reply with ancestors, and none of them was deleted.
- Its direct parent is a post by the same author: the account is replying to itself.
- It sits at least 2 levels below the root post.
- Its author is not the author of the root post.
- The root post's author has at least 1,000 followers (
FOLLOWER_COUNT_THRESHOLD_FOR_SPAM_DETECTION = 1000). Below that, the code skips the reply with the reasonlow_blast_radius.
The smallest qualifying shape is a post by another account, a reply from the account being evaluated, and a second reply from that same account to its own first reply. The filter never looks at the replier's own follower count.
A reply whose direct parent belongs to someone else is skipped (parent_not_same_author). So is any reply under the replier's own post (same_user_reply_as_root).
Which accounts does the filter exempt?
X's multi-step reply spam filter skips 5 kinds of cases before any model runs:
- Replies written by the Grok or Gork accounts.
- Authors marked as high page rank (version 2), checked with
is_high_page_rank_v2_user. - Authors carrying a grey badge, checked with
is_grey_badge_user. - Replies under a root author with fewer than 1,000 followers.
- Threads where an ancestor post was deleted, or where a post has no author record.
These files do not define "high page rank" or "grey badge".
How does the model decide what is spam?
X's code shows what the model is given and what happens after it answers, but not the criteria it applies. The scorer is bound to the model identifier oai-gemma4-26b-2, set in constants.py.
The model reads:
- the profile of the reply's author (name, username, bio and profile location);
- the thread as numbered posts, where Post 0 is the root post, tagged "never flag", and the last post is tagged "the reply being evaluated";
- for each post, the handle, follower count, display name, text, URLs and up to 2 media items;
- the root author's bio, added to the root post.
A thread with more than 10 ancestors is cut to the first 5 and the last 5 (MAX_THREAD_SIZE = 10). The model answers with JSON: the indexes of the posts it judges spam, and a reason.
The prompt itself is not public. prompts.py loads a template named multi_step_reply_spam_system.j2, which is missing from the repository, and the file's own comment says prompts are excluded to reduce gameability. X's README lists Grox prompts among the files it does not publish.
Which safeguards follow the model's answer?
Three checks run in this order, and each can erase the flags:
- If the newest reply is not among the flagged posts, all flags are dropped.
- Flagged posts written by other authors are removed.
- If exactly 1 flagged post remains, it is dropped.
A penalty therefore needs at least 2 flagged posts from the same author, the newest reply included.
What happens to a reply the filter flags?
For each flagged post in the thread, X's code applies a reply-spam safety label and publishes a reply ranking score of 0.0. The write step (TaskWriteMultiStepReplySpamReplyRanking) is disabled outside production through DisableTaskForNonProd.
For each flagged post, the code:
- applies the safety label
RiskyHighVizReply, unless the author is high page rank, carries a grey badge or is a test user (a credibility score of 62 or more counts as high page rank when the score exists); - publishes a reply ranking score of 0.0 with the model's reason, cut to its last 500 characters.
The main reply-ranking scorer in the same folder rejects any score outside a 0 to 3 rubric, so 0 is the lowest value on that scale. Once a post holds a 0, the standard reply-ranking write step skips any new score that is not lower than the stored one.
These files act on individual posts. They do not suspend the account or label it as a whole.
What the published files do not show
The repository does not show how a conversation view uses the score or the label: a search for the score's name finds it only in the reply_spam folder. The label name RiskyHighVizReply does not appear anywhere in the repository's under-the-hood folder, including its list of labels, so the published code shows no Under the Hood entry for it.
How can you tell whether your replies are affected?
To tell whether your replies are affected, look at where they sit in the conversation. Reply deboosting means your replies are pushed to the bottom of a conversation or hidden behind Show more replies.
Shadowban Radar defines the test from the outside. Reply deboosting is detected when X serves a reply in the conversation read in chronological order but leaves it out of the default view of that same conversation. The reply deboosting page explains how to check by hand.
A free check on Shadowban Radar runs this test on any public handle, with no login. It shows the position of a reply, not the cause. A reply ranked low can come from several signals, so the check cannot say whether this filter flagged it.
A reply that is missing from the conversation altogether is a different restriction, a ghost ban.
What does the code suggest doing about it?
X has not published the model's criteria, so these points follow only from the filter's conditions.
- Say it once. The filter acts only when at least 2 posts of the same author are flagged, and a single reply under someone else's post cannot qualify. Merging several short replies into one avoids the chain the filter looks for.
- Put extra detail in your own post. A follow-up, a link or a longer explanation can go in your own post or thread: replies under your own post are never evaluated (
same_user_reply_as_root). - Answer people, not yourself. A reply whose direct parent is another account's post is outside the filter (
parent_not_same_author). - A chain is judged together. The model reads the whole thread at once, and every flagged post of the same author gets its own score of 0.
The reply_spam files contain no step that lifts a label or raises a 0 once written.
Sources
Read on October 7, 2026, in the xai-org/x-algorithm repository at commit 78460ca. The 4 filter files were added in commit 9b0dc31 on September 4, 2026.
- plan_multi_step_reply_spam.py: the 4 steps and their order.
- task_multi_step_reply_spam.py: eligibility, detection and write steps.
- classifier_multi_step_reply_spam.py: the thread given to the model and the 3 safeguards.
- state_multi_step_reply_spam.py: the stored result.
- task_write.py: the label function and the standard reply-ranking write.
- README.md: the
grox/description and the list of unpublished files. - Commit 9b0dc31: September 4, 2026.
Shadowban Radar is not affiliated with X Corp. For how X's own report relates to a live check, see X's Under the Hood.


