Calibration · measured against reality
The measured accuracy of our own methods
Every number on this register is measured, estimated, or inferred, and every page already says which. This page states, in numbers, how often each one is actually right — checked against held-out truth, against a hand-check of live Telegram pages, or against an independent dataset — including the two places it has been visibly wrong. Nothing below is asserted from the pipeline’s own confidence in itself; every figure is either read live from this database, this run, or sourced from a specific dated hand-check on a named file, cited in place.
Channel creation dates: two methods, scored against reality
How we measure already explains the four measured sources and the one estimate. This section is the scoring, not the method: how often each one is actually right, re-checked against the live database rather than quoted from memory.
The estimate was withdrawn for one day, on 2026-08-09, and restored. A check reported 3.6% coverage; it had matched channels to a reference dataset by username, and every pair it compared was a reassigned handle — the reference dataset’s ids run near 1.2 billion where this register holds ids near 4.4 billion for the same name, and Telegram ids only increase. It was comparing a new channel against a dead one’s birthday. Re-measured id-joined, the figures below are what replaced it. Full account: /methodology#created-dates.
The id_interpolation estimate: how often the published range contains the truth
Two independent held-out checks, both re-run against the live database on 10 August 2026, both read-only, both left runnable so either can be re-checked rather than taken on trust.
Corpus-weighted median error of the point estimate, pooled check: 97 days. Across the 24 finer id bands the pooled script actually scores, coverage ranged from 78.4% (its single worst band) to 91.4% (its best) — every band at or above 78%, none catastrophically under the 80% target. Sources: /tmp/coverage_fine.py, run on web07 against the live database, 2026-08-10 and /tmp/coverage_extval.py, run on web07 against the live database, 2026-08-10.
The fitting pipeline grades its own point estimate continuously, not just in a one-off script. Its most recent run, 12 August 2026, fit on 720,219 anchors (382,804 from our own first-post readings, 264,777 + 120,979 from the two third-party datasets, 54from Telegram’s own MTProto record). Treating each measured source as briefly unseen and dating it from the rest: tgdataset — median error 182 days over 60,489 held-out anchors, first_message — median error 44 days over 191,402 held-out anchors, and tg_graph_dataset — median error 119 days over 132,388 held-out anchors.
The observational method: first_message against an independent date
first_message— the timestamp of message id 1, the “channel created” service message — is a measurement, not a model. Checked here against posted_at of msg_id=1 (the "channel created" service message) vs. the same channel’s date in an independent, third-party-measured dataset, joined on the durable numeric id only:
Read live from the register, this run, joined on the durable numeric channel id only — never on username, for the same reason the estimate above is joined on id: a handle can belong to a different channel today than it did when a third-party dataset was collected.
What the detectors get right, and what they do not
Three worker processes write observations onto channel pages — clones.py, erratio.py, velocity.py — and a fourth, avatars.py, writes findings that are collected but deliberately never rendered anywhere. Every one of them is severity 0 in the database: not “low risk”, but ungraded, because until a detector’s precision is hand-checked it is not entitled to a grade. What an observation is, and is not.
Duplicate content: the match is reliable, the direction is not
45 of the highest-confidence flagged pairs were hand-checked by fetching the live t.me pages their own evidence names, and re-deciding from the page.
Strict precision: 60% (27/45), and every single error is an attribution miss, not a text miss — the cause is dated exactly, in clones.py’s own header: header capture only began reliably on 2026-08-06, so posts published before that date can show “no attribution recorded” when the channel did, in fact, credit its source. That is why the site states the duplication plainly and the direction as a reading, not a fact — see the Observations section of any affected channel page.
30,479 channels currently carry an active duplicate-content observation.
Shared avatars: the picture match is reliable, “impersonation” mostly is not
60 pairs opened live and compared by eye: 58 really were the same picture — a corpus-weighted image-claim precision of 95.2%. But the relationship behind a shared picture is, overwhelmingly, not one channel faking another.
26 of 30 undifferentiated shared-avatar pairs turned out to be one operator running a regional or topical network off a single brand asset — legitimate, and the common case. The name+size filter roughly quadruples the precision, from 13.3% to 53.3%, and 53.3% is still not a majority you would call a verdict — 9 of 30 in that same sample were clearly legitimate and 5 were ambiguous.
This is why avatar_impersonation_candidate is collected and never rendered on any channel page. 16,433 channels carry it, active, right now (98,019 more carry the broader avatar_shared finding it is drawn from). lib/flags.ts’s own rule is that an unrecognised flag key renders nothing at all, and this register has deliberately not written the sentence that would let it render: at roughly half precision, the honest sentence is “maybe”, and this site does not publish maybe as a finding about a named channel.
Views moving without reactions: too rare to grade at all
velocity.py has written 141 of these observations corpus-wide (117still active). That is small enough that no hand-check of it has been run — the honest statement about this detector’s precision is that it has not been measured, and this page does not manufacture a percentage to fill the gap. More.
Values we hold and refuse to print as exact
t.me rounds to three significant figures above 1,000 — a view count, a reaction count, a channel’s lifetime photo/video/link counters. Two columns exist for the sole purpose of carrying that rounding forward rather than hiding it: post_metric.views_approx and channel_snapshot.counters_approx. A large count here is not a weakness in the data — it is a direct measure of how much of what this register holds it declines to overstate the precision of.
Of 102,043,436 view readings this register holds, 42,392,413 (41.5%) arrived already rounded by Telegram and are flagged as such; 48,985,128 (48.0%) are exact, below 1,000; 10,665,895 (10.5%) were captured before this flag existed and their precision is simply unknown — not assumed exact.
Of the register’s 2,803,777 live entries, 436,066 (15.6%) carry lifetime photo/video/link counters that arrived short-form and are marked rounded; 1,525,561 (54.4%) carry exact ones; 842,150 (30.0%) have never had these counters captured at all — a different state from a measured zero, and rendered differently on every page that shows them (why).
What fraction of the register carries what
Four questions, asked of the same population every other whole-register figure on this site uses — live channels and groups, user accounts always excluded (kind <> 3, never published, see above).
“2+ measurements” needs a caveat, so here it is. 2,786,321 entries (99.4%) hold 2 or more subscriber readings — but readings from different surfaces never confirm each other (why), so a channel touched once by the sweep and once by a promotion already has two rows with no time depth behind them. Narrower and more honest: only 1,806,368 (64.4%) span so much as a day, and only 910,789 (32.5%) show an actual subscriber-count change between readings. None yet span the fourteen days the primary indexing pathrequires — consistent with that page’s own “path 1 passed zero” count.
A discovery method that did not work, published anyway
Not every method measured here is one this register uses. The Wayback Machine’s CDX index was tried as a source of Telegram handles the crawl had not otherwise found. First real run: 200,000 archived URLs walked, of which 13 were valid Telegram handles — a yield of 0.0065%. Continued to 540,000 rows across 108 pages: 139 handles, cumulative. The cause, measured rather than assumed: the t.me keyspace in Wayback’s index is dominated by malformed archived URLs — broken hyperlinks someone else typed — which sort before genuine handles and have to be walked through first. For contrast, a directory sitemap tried the same night yielded a 52% new-handle rate for free. The method is kept, disabled, rather than deleted, specifically so this finding can be re-checked instead of taken on trust.
What this page leaves out
Not every method on this register has been graded. velocity.py’s view/reaction decoupling detector has written too few observations corpus-wide to support a hand-checked precision figure, and none is published above for it — a manufactured percentage would be less honest than none. The engagement-rate detector (erratio.py) publishes its own comparison population in full, rather than a single precision number, at the cohort baselines. Where a figure from this page’s original brief could not be reproduced against the live database or a written record in the repository, it was left out rather than quoted from memory — this page states only what it could check itself.
Reproducing these figures
Every live figure above is a plain read against this register’s own database — no figure on this page is licensed, surveyed, or taken from a vendor. The two held-out coverage checks for the creation-date estimate and the two detector hand-checks are dated, sourced to a specific file, and left runnable on web07 rather than archived as a one-off claim. Disagree with something here the same way you would dispute any other figure on this site: how to dispute an observation or a number.