·

7 min read

Instagram’s Hidden Words Filter: What It Catches by Default

Instagram's Hidden Words Filter: What It Catches by Default

Someone turns on Instagram’s Hidden Words filter the day before a launch, a casting call, or a pile on they can already see coming, and then has to guess what it actually screens out. The settings screen never lists what it is checking against, and the announcements that would explain it sit on a host that would not yield readable text in this research session. This piece states plainly what could and could not be confirmed, what a manager still has to type in themselves, and where the honest limit sits.

What Hidden Words is, according to Instagram

Instagram’s Hidden Words feature is described by the company as a tool that filters offensive words, phrases and emojis out of comments and DM requests, sorting them into a folder the account holder never has to open. [EVIDENCE NEEDED: Instagram’s own announcement post, fetched and quoted directly, establishing the August 2021 global rollout date and the exact language used to describe what the feature filters]. The general shape of the feature, a comment and DM triage tool aimed at abuse and spam rather than a general content moderation system, is consistent with how third parties who covered the 2021 rollout describe it, but this piece could not confirm Instagram’s own wording against a readable copy of its own announcement.

That framing matters for a practitioner because it tells you what the feature was built to do before you ask what it currently catches. It is something an account holder switches on, not something Instagram polices centrally on their behalf. That much is consistent across every description of the feature this research turned up, even without a verified quote from Instagram’s own post.

What the default list covers, and what Instagram says about updating it

Instagram has repeatedly signalled that Hidden Words is not a static list. [EVIDENCE NEEDED: Instagram’s own wording confirming that the company periodically expands the list of offensive words, hashtags and emojis filtered automatically, and states it will keep updating that list]. What can be said without a verified quote is narrower: no source available in this research, Instagram’s own or otherwise, publishes the literal words, phrases, or emojis on the default list. What gets described publicly are categories, the general shape of what gets caught, and periodic notes on how that shape has changed.

The clearest documented example of that pattern is a follow up update covering improvements to Hidden Words rather than the list itself. [EVIDENCE NEEDED: Instagram’s own announcement, fetched and quoted directly, describing improved detection of intentional misspellings such as character substitutions, expanded language coverage naming specific languages, and new filtering for scam or spam message requests]. Even a confirmed version of that update would not tell a moderator which words are on the list. It would tell them the list got better at catching disguised versions of words and at handling more languages, which is a different kind of information and a narrower promise than it might sound.

Default filter versus custom word list: where the line actually sits

Instagram is understood to let account holders build a custom list of additional words, phrases and emojis on top of whatever the default filter already catches. [EVIDENCE NEEDED: Instagram’s exact wording confirming that everyone can build a custom list with additional words, phrases and emojis they want to hide, quoted directly rather than paraphrased]. The word additional, if that is in fact the term used, would be doing real work in that sentence: a custom list layered on top of a default only makes sense if the default does not cover everything.

The table below separates what is generally described as automatic from what is generally described as something the account holder has to build, pending direct confirmation of Instagram’s own wording for each row.

Layer What is documented Who controls it
Default filter Offensive words, phrases and emojis, described only as categories rather than a list; spammy or low quality DM requests. [EVIDENCE NEEDED: confirmation from a readable Instagram source of the specific categories and any documented detection improvements] Instagram, undocumented in specific terms
Custom list Additional words, phrases and emojis the account holder adds, described as something everyone can build on top of the default. [EVIDENCE NEEDED: Instagram’s exact wording] The account holder, entirely
Coded terms, targeted phrases, slur variants specific to one account’s harassment Not confirmed as covered by the default in any source read for this piece Relies on the account holder having added it to the custom list

Because the default categories are described in general terms rather than as an itemised list, a moderator cannot know from the outside whether a specific slur variant, a coded term, or a phrase unique to one account’s harassment pattern is inside that default set. The only thing that can be stated with confidence is that a custom list exists at all, which by itself implies the default is not exhaustive.

What has reportedly changed since 2021, and why an old answer is not a current one

  • August 2021: [EVIDENCE NEEDED: Instagram’s own dated announcement confirming the global rollout of Hidden Words, its filtering of offensive words, phrases and emojis into a folder for comments, and its filtering of spammy or low quality DM requests].
  • October 2022: [EVIDENCE NEEDED: Instagram’s own dated announcement confirming intentional misspelling detection, added language support, expanded scam and spam message request filtering, extension to Story replies, and testing of automatic default on for Creator accounts specifically].
  • After the most recent confirmed update: [EVIDENCE NEEDED: a dated Instagram announcement describing any further change to Hidden Words’ scope]. A search of Instagram’s own announcements listing and Help Center did not turn up a newer dated update during this research, but that absence is not confirmation that none exists.

A manager reading any description of Hidden Words that cites only a single launch date is reading a description that may already be out of date on scope, since the feature has reportedly changed at least once since it launched. Before acting on anything in this piece, it is worth checking Instagram’s own announcements page directly for anything newer, and reading it in a way that yields actual article text rather than a rendered shell.

Setting up comment and DM filtering without assuming the default is enough

Instagram is understood to treat this as three separate controls, not one moderation system: a default Hidden Words filter, a custom word, phrase and emoji list layered on top of it, and account level tools such as Restrict, block, and report that operate independently of what gets filtered into the Hidden Folder. Treating these as one setting is how a manager ends up assuming coverage that was never promised. [EVIDENCE NEEDED: Instagram’s own description confirming that the default filter and the custom list both feed the same Hidden Folder and Hidden Requests folder as two distinct inputs].

The honest limit is this: because no readable Instagram source in this research publishes the literal default word list, no one outside the company, including this publication, can say with certainty whether a specific slur variant, coded term, or targeted phrase is or is not caught before anyone adds it manually. What is defensible without a verified quote is narrower and still useful. A custom list is understood to exist specifically because the default does not catch everything, which means the responsible default for any account running comments or DMs at volume is to build that list rather than to assume the out of the box filter is doing that job already.

Get more of this in your inbox

RecurPost Blog covers the platform mechanics behind the accounts you run, dated and sourced back to the platforms themselves wherever a source can actually be read. Follow along for the next teardown of what a feature actually does versus what it is described as doing.

FAQ

Does Instagram publish the exact list of words Hidden Words blocks by default?

No. Every description of the default scope, Instagram’s own and otherwise, uses categories, offensive words, phrases, emojis, and spammy or low quality DM requests, rather than an itemised list. This piece could not locate any source, readable or otherwise, that publishes the literal words filtered by default.

If I turn on Hidden Words, do I still need a custom list?

Yes, on the available evidence. A custom list of additional words, phrases and emojis is described as something every account holder can build on top of the default filter, which only makes sense if the default does not cover every case, including coded slurs, targeted phrases, or terms specific to one account’s harassment pattern. [EVIDENCE NEEDED: Instagram’s own exact wording confirming this, quoted directly rather than paraphrased].

Dinesh Agarwal Avatar