Somewhere in your feedback channel, someone has probably repeated a claim about how X ranks posts and cited “the algorithm they open sourced” as the reason. Chances are neither of you has opened the repository. The 2023 announcement is the thing everyone remembers. The repository it pointed to is a different object, one that keeps changing underneath that memory, and the two have drifted apart in a way that is easy to check and rarely is.
What shipped in 2023, and what has actually changed since
The repository is github.com/twitter/the-algorithm, created on 2023-03-27 according to GitHub’s own repository metadata. It is not a one-time dump a company posted and abandoned, which matters because most secondhand coverage of “the X algorithm” treats it as a frozen artifact from launch week. Queried against the GitHub commits API at the time of this research, the main branch had taken 31 commits in total since creation. That is a moving number by definition: anyone re-running the same query after X pushes again will get a higher one, and should re-check before citing it.
The most recent of those 31 commits is dated 2025-09-03, with the message “update for-you recommendations code,” and GitHub records it as pushed to the repository on 2025-09-08. Counting from that push to this piece’s publication date is 385 days, a little over twelve and a half months. Compare that cadence to the volume of secondhand explainer content produced about “how the X algorithm works” in the same period, almost none of which appears to have opened the diff.
What the code documents, component by component
The repository’s own README lays out its architecture as a table, not prose, and the safest way to describe the system is to reproduce that table rather than paraphrase the 2023 blog post about it. This is what the README stated when fetched directly for this piece, on 2026-09-23:
| Component | File path | What the README says it does |
|---|---|---|
| SimClusters | src/scala/com/twitter/simclusters_v2/README.md | Community detection and sparse embeddings into those communities |
| TwHIN | the-algorithm-ml, projects/twhin/README.md | Dense knowledge graph embeddings for Users and Posts |
| real-graph | src/scala/com/twitter/interaction_graph/README.md | Model to predict the likelihood of an X User interacting with another User |
| tweepcred | src/scala/com/twitter/graph/batch/job/tweepcred/README | Page-Rank algorithm for calculating X User reputation |
| light-ranker | src/python/twitter/deepbird/…/earlybird/README.md | Light Ranker model used by search index (Earlybird) to rank posts |
| heavy-ranker | the-algorithm-ml, projects/home/recap/README.md | Neural network for ranking candidate posts, one of the main signals used to select timeline posts post candidate sourcing |
| home-mixer | home-mixer/README.md | Main service used to construct and serve the Home Timeline, built on product-mixer |
| visibility-filters | visibilitylib/README.md | Filters content for legal compliance, quality, trust and revenue protection, via hard-filtering, visible treatments and coarse-grained downranking |
| pushservice light and heavy rankers | pushservice/README.md; pushservice/src/main/python/models/light_ranking and heavy_ranking | Light ranker pre-selects candidates for notifications; heavy ranker is a multi-task model predicting open and engagement probability for a sent notification |
Every description in that table is the README’s own wording at its own named file path, not a gloss on the original announcement. Two entries, TwHIN and heavy-ranker, live in a second repository, the-algorithm-ml, which the main README links out to directly rather than duplicating.
What X said in 2023, and what it has not said since
The public record behind the release is two dated posts, both from 31 March 2023: the engineering blog post titled “Twitter’s Recommendation Algorithm,” and a companion post on the company blog, “A new era of transparency for Twitter,” which frames the code release as part of a broader openness push and states plainly that Twitter made “the decision not to release training data or model weights associated with the Twitter algorithm at this point.”
Checking the engineering blog’s own “Open source” tag listing, as archived on 2026-01-05, the newest entry under that tag is still that same 31 March 2023 post. Nothing about ranking has been added to that topic listing since. This describes the public blog record as archived and checked on that date, not X’s internal engineering activity, which this piece has no way to see. An attempt to re-check the live blog.x.com listing directly on 2026-09-23 was blocked by an automated bot challenge that returned no readable page, so the archived snapshot is the most recent version this research could actually read. A reader citing this finding later should try the live page again first.
The specific claims worth checking before you repeat them
Most of what circulates as folklore about X’s ranking traces back to the 2023 announcement rewritten in someone else’s words, not to a specific file. Before repeating any of the following as fact, the method is the same each time: find the exact file that would have to say it, read what it actually says, and drop the claim if no such file exists.
- “Verified accounts get an automatic ranking boost.” The README’s table names no such boost. If a verified-status feature exists, it would need to be hydrated somewhere in home-mixer’s scoring pipeline and then weighted in a specific parameter object. Absent a named file and line showing that weight, this stays [EVIDENCE NEEDED].
- “Getting reported or blocked enough times quietly reduces your reach platform-wide.” The closest real component is visibility-filters, at visibilitylib/README.md. Its own README states that “part of the code has been removed and is not ready to be shared yet” and that “the remaining part of the code needs further review.” A component whose own maintainers say part of it was withheld cannot be read as confirming or denying a specific mechanism it doesn’t fully contain. [EVIDENCE NEEDED].
- “Replies that keep a thread going with the original author get ranked higher than one-off replies.” This is checkable in principle, since the heavy-ranker documentation names discrete predicted engagement types rather than one generic “engagement” score. Whether one of those types is actually weighted higher in the live scoring configuration, and by how much, is a separate question from whether the type exists, and the two should not be conflated.
None of this is a claim that these mechanisms don’t exist. It’s a claim that stating them as confirmed requires a path, not a memory of a 2023 headline.
Reading one component yourself, start to finish
Here is the walk any reader can repeat in under five minutes. Start at the main README’s architecture table and find the heavy-ranker row. It links out to a second repository, the-algorithm-ml, at the path projects/home/recap/README.md. That file, read directly, describes the heavy ranker as a parallel MaskNet model producing several separate predicted-engagement probabilities, named things like “scored_tweets_model_weight_fav”, “scored_tweets_model_weight_reply”, and “scored_tweets_model_weight_negative_feedback_v2”, which are then combined into one score by a weighted sum. The same file states that the specific weight values it lists are “current” as of a date written into the file itself, April 5, 2023, and it links to one exact location for the live version of those weights: home-mixer/server/src/main/scala/com/twitter/home_mixer/product/scored_tweets/param/ScoredTweetsParam.scala, line 84.
Fetching that exact file today, 2026-09-23, it runs 881 lines and none of the named weight parameters the linked README describes appear in it anymore; line 84 now sits inside an unrelated content-exploration parameter block. That doesn’t prove the weighting logic left the codebase, it may have moved elsewhere in a repository this size. It does prove the specific cross-repository citation a reader would have followed to find “the current weights” no longer resolves to weights at all, three and a half years after the file it points to was written. Anyone repeating those 2023 figures as X’s current ranking weights is citing a page that has since moved.
What this changes about reading the next platform statement
The finding here is narrow and specific to this one repository: 31 commits and a fresh push in September 2025, against an engineering blog whose own open-source topic has recorded nothing new about ranking since March 2023, and a cross-repository citation that pointed at real weights in 2023 and points at something else now. That gap is not evidence of bad faith. It’s what happens by default when an artifact is maintained on an engineering cadence and a statement about that artifact is published once, on a communications cadence, and neither team goes back to reconcile them.
The practical lesson is about the evidence, not about what to do differently when posting: a platform’s code and its PR copy about that code are two different kinds of record on two different update schedules, and folklore fills the gap between them when nobody checks which one they’re citing. This piece stops at what is documented and what is or isn’t still said about it, deliberately, with nothing to add about posting strategy. The same method applies to how LinkedIn’s own engineering posts describe its feed ranking, to what Instagram has actually published about its ranking signals, and what it hasn’t, and to what a platform’s earnings call can tell you about ranking, and what it can’t, for a fuller kit of primary sources to check before repeating the next claim.
Before you cite this again
If a claim about X’s ranking is worth repeating in a deck or a client email, it’s worth the extra two minutes to open the file it supposedly comes from. The repository, the commit history, and the linked component READMEs are all public and all dated. Check the specific path before the specific claim goes out under your name.
FAQ
Is the code at github.com/twitter/the-algorithm the exact code ranking posts on X today?
The repository itself does not claim real-time parity with production: the heavy-ranker documentation notes there “may be small differences from the production model,” and no file asserts it is a live mirror. It is released under the AGPL-3.0 license and has continued to receive commits, most recently in September 2025, but what those commits mean for production is not something the public documentation states either way.
Has X explained any changes to how ranking works since the 2023 open-source release?
Checked against the engineering blog’s own “Open source” topic listing, as archived on 2026-01-05, the most recent entry under that topic is still the original 31 March 2023 recommendation-algorithm post. That describes the public blog record as archived and checked on that date, not X’s internal process, and it should be re-checked against the live listing before being repeated as current.
Sources
- github.com/twitter/the-algorithm, repository and README, fetched 2026-09-23
- GitHub API, repository metadata (created_at, pushed_at, license), fetched 2026-09-23
- GitHub API, commits endpoint (commit count via pagination, most recent commit), fetched 2026-09-23
- the-algorithm-ml, heavy ranker README, fetched 2026-09-23
- home-mixer, ScoredTweetsParam.scala, fetched 2026-09-23
- visibility-filters README (visibilitylib), fetched 2026-09-23
- “Twitter’s Recommendation Algorithm,” engineering blog, archived snapshot, fetched 2026-09-23
- “A new era of transparency for Twitter,” company blog, archived snapshot, fetched 2026-09-23
- Engineering blog, “Open source” tag listing, archived snapshot dated 2026-01-05, fetched 2026-09-23



