x algorithmopen sourcealgorithmic transparencymethodprovenance18 أغسطس 2026

X نشرت خوارزميتها. قرأناها، ارتكبنا خطأً واحدًا، وقد اكتشف المصدر ذلك.

إصدار خاص، خارج الدورة المعتادة. ليست مجموعة بيانات - بل قطعة من الكود المصدري يقتبسها نصف الإنترنت عن ظهر قلب. قرأنا جميع المكونات الـ 123، وربطنا كل رقم بسطره، وأصدرنا ادعاءً واحدًا دحضه مصدرنا الخاص بعد ثلاثة أيام.

X published its algorithm. We read it, we got one thing wrong, and the source caught us.

A special edition, off the usual cycle. Not a data cluster this time — a piece of source code that half the internet is quoting from memory. We read all 123 components, published a page that links every number to its line, and shipped one claim that our own source refuted three days later. Here is what the code says, and what that mistake taught us about the method we sell.

On 13 August 2026, X open-sourced the ranking code behind its For You feed, under Apache-2.0. Within seventy-two hours, the numbers were everywhere: a repost is worth twenty likes, a profile click twelve, one report cancels four hundred and sixty-eight likes.

Almost none of those figures describe the code that runs today.

They are not inventions. They are leftovers — true for a Scala pipeline retired years ago, quoted as if they described the Rust one that replaced it. A number with no date and no line of code is not a fact. It is a rumour with a decimal point.

So we did the unglamorous thing: we read the repository, line by line, and built the page that was missing.

d-nvest.com/x-algorithm — free, no account, nine languages.

1. What the code actually says

190 parameters that X can change without shipping code. 24 constants compiled into the binary. 123 components in the ranking pipeline. Every one of those numbers on our page links to the exact line it was read from, pinned to a single commit — and, since this week, unfolds the actual lines of Rust in place, so you never have to take our word for a figure.

Four of the claims we checked, and what the source says instead:

"A repost is worth 20 likes." It is worth two. `RetweetWeight` is 1.0 against `FavoriteWeight` at 0.5. No published version of this algorithm ever used 20 for a repost — not even the 2023 one people are quoting.

"A profile click is worth 12 likes." Today it is worth nothing at all: the weight is zero. The 12 was real, for a composite signal that no longer exists in the tree.

"One report cancels 468 likes." The arithmetic is right and the conclusion is wrong. Those weights multiply a predicted probability, not a count of actions. X states a report is thousands of times rarer than a like — the two terms were never the same size to begin with. Dividing one weight by another gives no exchange rate.

"The blue check buys reach." Nothing in the published code multiplies a score by subscription status, and the 2023 boost is gone from the tree. The honest limit, which we print on the page: the final ranking runs through a learned model whose weights are not published. A correlation the model learned on its own cannot be ruled out by reading code. We say so, because it is true.

And the finding almost nobody mentions: the heaviest positive signal in the published code is not the repost, the reply or the like. It is someone copying your link to share it elsewhere. The 20 that everyone quotes is real. It is just attached to a completely different action.

2. The claim we shipped, and that our own source refuted

Here is the part a marketing team would cut.

Our page went live saying, as its headline finding, that no person and no organisation is named anywhere in that repository — no account id, no handle, no hardcoded list. We put a full-screen zero on the carousel.

It was false.

At the exact commit we had pinned, `home-mixer/filters/brazil_2026_election_filter.rs` carries 665 account ids written into the source, each annotated with its handle. The file says so itself: "usernames are included for transparency." The filter removes those accounts' posts from For You — including reposts, quotes and thread ancestors — unless the viewer already follows them. X ships it deliberately, in accordance with Brazilian electoral law, and documents it in its own README.

Why we got it wrong is the only interesting part. The sentence was not extracted from anything. It was written by hand. It was true of the subset we had read — the visibility-filtering rules — and then quietly generalised to "anywhere in this repository." A generalisation never re-verifies itself.

Worse: a test in our own suite pinned the mistake in place. It checked that the phrase contained certain words. It guaranteed the sentence was there, never that it was true. A test can lock a lie.

We corrected it the same day. That claim is now computed from a traversal of the repository instead of asserted about it, the page shows the accounts, and the number moves if the source moves.

The source refuted us before anyone else did. That is exactly what a source is for — and it is the whole argument for working this way.

3. The country restriction that has no country in it

Once we started looking properly, the geography turned out to be the opposite of what everyone assumes.

In the entire repository, exactly one country has a list of its own. Two other country lists are hard-coded — 16 countries where sensitive content is withheld from anyone who has not stated an age, and 18 countries the model treats individually — and every other country on earth falls into a catch-all: not blocked, indistinct.

The per-post takedown data — the thing people mean when they say "withheld in country X" — is not in the repository at all. Those rules compare the viewer's country to codes carried by each post at runtime, and that data is not published. A map colouring in "censored" countries from this code would be inventing them. Ours does not.

And the one country-specific filter that is written into the code has a property we did not expect: no country condition appears anywhere in it, nor around the line that switches it on. It is registered unconditionally in the pipeline's filter chain.

We publish that with its limit attached, because the limit is the honest part: which pipeline serves which market may be decided in a layer X does not publish. What we can say is what the published code contains — and it contains no country test.

4. What gets removed is reach, not speech

The silencing machinery exists, it is readable, and it is narrower than the word "shadow ban" suggests.

32 takedown rules ship in the clear, acting on 27 label types. Two of them hide a post from non-followers only — your followers still see you. That distinction is in the code, and it disappears every time someone reaches for the phrase that gets clicks.

Everywhere except that one filter, the rules act on a label attached at runtime, and the labelling itself is not published. The code shows what can be done. Never to whom it was done.

5. Why a data company spent a week on this

Because it is the work we do for a living, done in public, on a source everyone already has an opinion about.

Our job is taking something opaque — a register, a filing, a dataset, a repository — and making it checkable line by line. The part that takes discipline is not finding the answers. It is saying out loud where the answers stop, and correcting yourself in public when the source says you were wrong.

The page watches itself, too. A job compares our pinned commit to X's every morning; when they diverge, the page says it is behind rather than silently serving stale figures, and the regeneration lands on a branch waiting for a human. A page that calls out other people's stale numbers does not get to go stale itself.

→ Read it: d-nvest.com/x-algorithm

Get the next analysis

One deep-dive per edition on where valuable data is hiding — the evidence, the sources, and who would pay for it. No noise.

One email per edition. Unsubscribe any time. We never share your address.

From the marketplace

Explore live data opportunities

Browse datasets by sector & use-case
Found this useful? Share it

d-nvest turns the data assets behind these deals into scored, actionable opportunities.

Explore the pipeline →