Telegram RegisterThe public register of Telegram

Calibration · measured against reality

The measured accuracy of our own methods

Every number on this register is measured, estimated, or inferred, and every page already says which. This page states, in numbers, how often each one is actually right — checked against held-out truth, against a hand-check of live Telegram pages, or against an independent dataset — including the two places it has been visibly wrong. Nothing below is asserted from the pipeline’s own confidence in itself; every figure is either read live from this database, this run, or sourced from a specific dated hand-check on a named file, cited in place.

Channel creation dates: two methods, scored against reality

How we measure already explains the four measured sources and the one estimate. This section is the scoring, not the method: how often each one is actually right, re-checked against the live database rather than quoted from memory.

The estimate was withdrawn for one day, on 2026-08-09, and restored. A check reported 3.6% coverage; it had matched channels to a reference dataset by username, and every pair it compared was a reassigned handle — the reference dataset’s ids run near 1.2 billion where this register holds ids near 4.4 billion for the same name, and Telegram ids only increase. It was comparing a new channel against a dead one’s birthday. Re-measured id-joined, the figures below are what replaced it. Full account: /methodology#created-dates.

The id_interpolation estimate: how often the published range contains the truth

Two independent held-out checks, both re-run against the live database on 10 August 2026, both read-only, both left runnable so either can be re-checked rather than taken on trust.

CheckWhat it held outScoredTruth inside the published band
Pooled random halfA random half of 489,741 pooled anchor ids, removed from every source before the curve was refit — so nothing predicts itself.244,870 channels82.8% (202,848/244,870)
Whole source held outTGDataset (120,979 ids) purged entirely from the anchor pool — never seen at fit time — then scored as unseen truth. The strictest test available: not a random split of one pool, but whether the method generalises to a dataset it has never seen at all.120,979 channels79.5% (96,125/120,979)

Corpus-weighted median error of the point estimate, pooled check: 97 days. Across the 24 finer id bands the pooled script actually scores, coverage ranged from 78.4% (its single worst band) to 91.4% (its best) — every band at or above 78%, none catastrophically under the 80% target. Sources: /tmp/coverage_fine.py, run on web07 against the live database, 2026-08-10 and /tmp/coverage_extval.py, run on web07 against the live database, 2026-08-10.

The fitting pipeline grades its own point estimate continuously, not just in a one-off script. Its most recent run, 12 August 2026, fit on 720,219 anchors (382,804 from our own first-post readings, 264,777 + 120,979 from the two third-party datasets, 54from Telegram’s own MTProto record). Treating each measured source as briefly unseen and dating it from the rest: tgdataset — median error 182 days over 60,489 held-out anchors, first_message — median error 44 days over 191,402 held-out anchors, and tg_graph_dataset — median error 119 days over 132,388 held-out anchors.

The observational method: first_message against an independent date

first_message— the timestamp of message id 1, the “channel created” service message — is a measurement, not a model. Checked here against posted_at of msg_id=1 (the "channel created" service message) vs. the same channel’s date in an independent, third-party-measured dataset, joined on the durable numeric id only:

Compared againstChannels held by bothMedian differenceAgree within a day
ext_tg_channel3,6480 days100%
tgdataset9970 days99.9%

Read live from the register, this run, joined on the durable numeric channel id only — never on username, for the same reason the estimate above is joined on id: a handle can belong to a different channel today than it did when a third-party dataset was collected.

What the detectors get right, and what they do not

Three worker processes write observations onto channel pages — clones.py, erratio.py, velocity.py — and a fourth, avatars.py, writes findings that are collected but deliberately never rendered anywhere. Every one of them is severity 0 in the database: not “low risk”, but ungraded, because until a detector’s precision is hand-checked it is not entitled to a grade. What an observation is, and is not.

Duplicate content: the match is reliable, the direction is not

45 of the highest-confidence flagged pairs were hand-checked by fetching the live t.me pages their own evidence names, and re-deciding from the page.

Text match wrong0/45The two live posts really were near-identical text, every time.
Attribution reading wrong18/45The live page showed a forward header the database had no record of — naming this pair’s other side, or a third channel.
Confirmed a genuine unattributed copy27/45No forward header on either live page.

Strict precision: 60% (27/45), and every single error is an attribution miss, not a text miss — the cause is dated exactly, in clones.py’s own header: header capture only began reliably on 2026-08-06, so posts published before that date can show “no attribution recorded” when the channel did, in fact, credit its source. That is why the site states the duplication plainly and the direction as a reading, not a fact — see the Observations section of any affected channel page.

30,479 channels currently carry an active duplicate-content observation.

Shared avatars: the picture match is reliable, “impersonation” mostly is not

60 pairs opened live and compared by eye: 58 really were the same picture — a corpus-weighted image-claim precision of 95.2%. But the relationship behind a shared picture is, overwhelmingly, not one channel faking another.

Arm hand-checkedSampledActually impersonation-shapedPrecision
Every shared-avatar pair, undifferentiated30413.3%
avatar_impersonation_candidate — same avatar, similar name, decisive size gap301653.3%

26 of 30 undifferentiated shared-avatar pairs turned out to be one operator running a regional or topical network off a single brand asset — legitimate, and the common case. The name+size filter roughly quadruples the precision, from 13.3% to 53.3%, and 53.3% is still not a majority you would call a verdict — 9 of 30 in that same sample were clearly legitimate and 5 were ambiguous.

This is why avatar_impersonation_candidate is collected and never rendered on any channel page. 16,433 channels carry it, active, right now (98,019 more carry the broader avatar_shared finding it is drawn from). lib/flags.ts’s own rule is that an unrecognised flag key renders nothing at all, and this register has deliberately not written the sentence that would let it render: at roughly half precision, the honest sentence is “maybe”, and this site does not publish maybe as a finding about a named channel.

Views moving without reactions: too rare to grade at all

velocity.py has written 141 of these observations corpus-wide (117still active). That is small enough that no hand-check of it has been run — the honest statement about this detector’s precision is that it has not been measured, and this page does not manufacture a percentage to fill the gap. More.

Values we hold and refuse to print as exact

t.me rounds to three significant figures above 1,000 — a view count, a reaction count, a channel’s lifetime photo/video/link counters. Two columns exist for the sole purpose of carrying that rounding forward rather than hiding it: post_metric.views_approx and channel_snapshot.counters_approx. A large count here is not a weakness in the data — it is a direct measure of how much of what this register holds it declines to overstate the precision of.

Of 102,043,436 view readings this register holds, 42,392,413 (41.5%) arrived already rounded by Telegram and are flagged as such; 48,985,128 (48.0%) are exact, below 1,000; 10,665,895 (10.5%) were captured before this flag existed and their precision is simply unknown — not assumed exact.

Of the register’s 2,803,777 live entries, 436,066 (15.6%) carry lifetime photo/video/link counters that arrived short-form and are marked rounded; 1,525,561 (54.4%) carry exact ones; 842,150 (30.0%) have never had these counters captured at all — a different state from a measured zero, and rendered differently on every page that shows them (why).

What fraction of the register carries what

Four questions, asked of the same population every other whole-register figure on this site uses — live channels and groups, user accounts always excluded (kind <> 3, never published, see above).

CarriesEntriesShare of the register
A measured creation date592,13621.1%
An estimated creation-date range2,181,30177.8%
Neither — no creation-date claim at all30,4001.08%
A topic classification52,2401.86%
An active observation35,5651.27%

Denominator: 2,803,837live, non-user entries — the same figure the homepage publishes as “Indexed”.

“2+ measurements” needs a caveat, so here it is. 2,786,321 entries (99.4%) hold 2 or more subscriber readings — but readings from different surfaces never confirm each other (why), so a channel touched once by the sweep and once by a promotion already has two rows with no time depth behind them. Narrower and more honest: only 1,806,368 (64.4%) span so much as a day, and only 910,789 (32.5%) show an actual subscriber-count change between readings. None yet span the fourteen days the primary indexing pathrequires — consistent with that page’s own “path 1 passed zero” count.

A discovery method that did not work, published anyway

Not every method measured here is one this register uses. The Wayback Machine’s CDX index was tried as a source of Telegram handles the crawl had not otherwise found. First real run: 200,000 archived URLs walked, of which 13 were valid Telegram handles — a yield of 0.0065%. Continued to 540,000 rows across 108 pages: 139 handles, cumulative. The cause, measured rather than assumed: the t.me keyspace in Wayback’s index is dominated by malformed archived URLs — broken hyperlinks someone else typed — which sort before genuine handles and have to be walked through first. For contrast, a directory sitemap tried the same night yielded a 52% new-handle rate for free. The method is kept, disabled, rather than deleted, specifically so this finding can be re-checked instead of taken on trust.

What this page leaves out

Not every method on this register has been graded. velocity.py’s view/reaction decoupling detector has written too few observations corpus-wide to support a hand-checked precision figure, and none is published above for it — a manufactured percentage would be less honest than none. The engagement-rate detector (erratio.py) publishes its own comparison population in full, rather than a single precision number, at the cohort baselines. Where a figure from this page’s original brief could not be reproduced against the live database or a written record in the repository, it was left out rather than quoted from memory — this page states only what it could check itself.

Reproducing these figures

Every live figure above is a plain read against this register’s own database — no figure on this page is licensed, surveyed, or taken from a vendor. The two held-out coverage checks for the creation-date estimate and the two detector hand-checks are dated, sourced to a specific file, and left runnable on web07 rather than archived as a one-off claim. Disagree with something here the same way you would dispute any other figure on this site: how to dispute an observation or a number.