Исследователи Tencent и Университета Цинхуа представили CALM — языковую модель, которая генерирует текст не по одному токену, а сразу целыми фрагментами. Автоэнкодер сжимает четыре токена в один непрерывный вектор, а затем восстанавливает исходный текст с точностью более 99,9%. В результате модели требуется в четыре раза меньше последовательных шагов генерации. В экспериментах CALM с 371 млн параметров показала рез…

Channel
Data Portal | DS & ML
@LLMScience
On this record: Growth · Engagement · What this channel posts · Advertising · Posts · Citations · Cite this entry
8,376subscribers
+20 since we began measuring on 9 August 2026
Risers and fallers across the register · movement among entries of 3,162–10,000.
Register entry
| Telegram ID | -1002177572398 |
|---|---|
| Type | Channel |
| Username | @LLMScience |
| Created | Between 1 June 2024 and 30 September 2024— estimated from Telegram’s id allocation, not measured. How this range is calculated. |
| First recorded | 9 August 2026 |
| Last confirmed live | 12 August 2026 |
| Measurements held | 3 |
| Confirmed unchanged | 1 time, most recently 12 August 2026 |
| On Telegram | t.me/LLMScience |
Growth
| Measured (UTC) | Subscribers | Change |
|---|---|---|
| 12 Aug 2026, 17:16 | 8,376 | +21 |
| 10 Aug 2026, 01:21 | 8,355 | -1 |
| 9 Aug 2026, 18:16 | 8,356 | first reading |
Engagement
33 posts held, back to 4 August 2026 — the reader has not yet reached the start of this channel’s public history, so older posts may sit further back, unread. Read across 6 pagesof Telegram’s post history, 20 posts per page.
- ERR · 30 days
- 7.77%
- avg views ÷ 8,376 subscribers
- Avg views / post
- 651
- 33 posts measured
- Reaction rate
- —
- this channel exposes no reaction counts
- Posts in window
- 33
- of 33 held
ERR is average views per post over the last 30 days divided by subscribers, the definition TGStat uses, so this figure is comparable with the one you will see elsewhere. It falls structurally as a channel grows: a high ERR on a small channel and a low one on a large channel describe reach mathematics, not quality. We publish the figure and the sample it came from and pass no verdict on it.
ER is defined industry-wide as (forwards + reactions + comments) ÷ views— note the denominator is views, not subscribers. Telegram’s public web preview carries views and reactions but not forward or comment counts, so the reaction rate above is the reactions term only and is therefore a floor: the true ER for this channel is higher by an amount we have not measured and will not estimate.
| Window | Rolling 30 days · latest post in window 12 August 2026 |
|---|---|
| Posts held | 33 (4 August 2026 – 12 August 2026) |
| Views total | 21,490 |
| Reactions total | — |
| Forwards / comments | not exposed by the public surface — not measured, not estimated |
| Readings taken | 12 Aug 2026, 16:28 UTC |
Views are a single reading per post, taken at the time above. A post published in the last day or two is still accumulating views, which pulls the 30-day average down slightly. That is a property of the standard definition rather than a fault in it, so we keep the definition rather than “correcting” the number into something nobody can reproduce.
Precision. Telegram publishes view counts on its public widget in short form — 8.12K, 3.7M — so any reading at or above 1,000 reaches us rounded to three significant figures, and only counts below 1,000 are exact. Averages and rates derived from them are shown to the same precision rather than to the unit: a figure like 3,701,250 would assert digits nobody measured.
Reaction counts are published per emoji and rounded the same way, so a total below 1,000 is exact and a larger one is a sum that may carry a rounded component from each emoji above 1,000. Because it is a sum, it does not look rounded — read a large reaction total as three significant figures per contributing emoji rather than as the figure it prints.
What this channel posts
- Video runtime
- 36s
- Average length
- 36s
Measured directly from 1 video with a duration reading, out of the posts we hold for this channel — not this channel’s whole posting history, only the sample this register has actually read. An exact reading to the second, taken from the post itself rather than from Telegram’s own rounded chrome, so it carries no ≈ mark.
Advertising
- Ad load
- 3.03%
- 1 of 33 posts carry an ad marker
- Regulatory tokens
- 1
- posts carrying an erid · 1 distinct token
- Median views · ads
- 528
- over 1 measured post
- Median views · rest
- 726
- over 32 measured posts
An ad marker, not a judgement about a post. A post is counted here because it carries one of two explicit markings: an erid token, which Russian law has required on paid placements since 2022 and which is issued against a specific advertising contract, or a #реклама / #ad hashtag in the body, which is the channel declaring it itself. The first is documentary; the second is a self-declaration and is weaker. No classifier reads the text and decides — nothing on this site guesses that a post is an advertisement.
This is a floor, and it can only ever be a floor.A channel that runs paid placements without marking them produces no marker for us to count, and an unmarked ad is indistinguishable from an ordinary post on the public surface. The ad load above therefore means “the share of posts that declared themselves”, never “the share of posts that were paid for”. A low figure is not evidence of a channel that runs few ads.
Both figures are medians, and no ratio between them is published. Each is a view reading that actually occurred on a post, picked by percentile_disc rather than averaged, so one viral post cannot move it and no interpolated value is invented between two readings. The sample on one side is under five posts, which is too thin to compare. The two figures are shown side by side with the count behind each, and deliberately not divided into a headline like “ads get x% fewer views” — an arithmetic that is easy to print and, at this sample size, means nothing.
| erid | Posts | First seen | Last seen |
|---|---|---|---|
| 2VtzqvfcpBD | 1 | 11 August 2026 | 11 August 2026 |
A token repeated across several posts is one advertising contract placed more than once, which is what the identifier is for. The strings are reproduced exactly as they appeared in the post or in its click-through URL and are not validated against any registry — we record the marker a channel published, and whether it resolves to a real contract is a question for the register that issued it.
Measured over the 33 most recent posts we hold, published 4 August 2026 to 12 August 2026. Views are the latest single reading held for each post, and any reading at or above 1,000 is rounded by Telegram to three significant figures.
Recent posts
Как научить рекомендации думать не о следующем лайке, а о всей сессии? Обычно рекомендательные системы оптимизируют ближайший сигнал: клик, лайк или дослушивание. Но каждое действие меняет состояние пользователя, поэтому локально хорошая рекомендация не обязательно улучшает всю сессию. В новой работе AI VK Research рекомендации формулируются как задача обучения с подкреплением. Модель рассматривает пользователя как…
Исследователи впервые в больших масштабах извлекли настоящие скрытые рассуждения из закрытых моделей OpenAI, Anthropic и Google. Затем они использовали эти данные для изучения Kimi K3, GLM-5.2, DeepSeek и других открытых моделей. OpenAI, Anthropic и Google всё реже показывают пользователям полные цепочки рассуждений. Итоговый ответ можно скопировать, но гораздо ценнее понять, как именно модель к нему пришла. Если …
Теорема Байеса — одна из фундаментальных концепций в Data Science. Но мне понадобилось два года, чтобы по-настоящему понять её важность. За 2 минуты расскажу самое полезное, что узнал за последние два года изучения байесовской статистики. Поехали. 1. Предыстория Работа «An Essay towards solving a Problem in the Doctrine of Chances» была опубликована в 1763 году, через два года после смерти Томаса Байеса. В ней Б…
«Analysis of Functions of a Single Variable» — бесплатный учебник по вещественному и комплексному анализу. В книге разбираются вещественные и комплексные числа, последовательности, пределы, бесконечные ряды, функции, непрерывность, равномерная сходимость, дифференцирование, разложения Тейлора, интегрирование, длина дуги и контурные интегралы. Также выводятся ключевые результаты: теорема Грина, теорема Коши, теорема…
Вот как, по-хорошему, нужно преподавать машинное обучение. Наблюдать за тем, как модели реально учатся, невероятно залипательно. 🤩 http://ml-visualized.com
«Introduction to Probability» — отличный бесплатный учебник MIT по теории вероятностей. В книге разбираются вероятностные модели, условная вероятность, независимость, формула Байеса, дискретные и непрерывные случайные величины, распределения вероятностей, математическое ожидание, дисперсия, ковариация, корреляция, условное ожидание, преобразования и свёртки, процессы Бернулли и Пуассона, цепи Маркова, закон больших …
«Patterns, Predictions, and Actions» — бесплатная книга по математическим основам машинного обучения. Она сочетает понятное изложение с математикой, на которой построен современный ML. В книге разбираются предсказание, обучение с учителем, оптимизация, обобщающая способность моделей, глубокое обучение, датасеты, причинность, последовательное принятие решений и обучение с подкреплением. Отдельно авторы объясняют, ка…
Что скрывается за монетизацией поиска и реков в Авито 👀 Рассчитать ожидаемую выручку сложнее, чем кажется: нельзя просто умножить вероятность на ставку. На практике всё упирается в данные, метрики и бизнес-ограничения. Если вам интересно, как работает наше ранжирование, приходите на Хабр почитать и оставить комментарии.
Google DeepMind утверждает, что привычный RAG упирается в фундаментальный предел. Последние три года стандартный ответ почти на любую задачу с памятью или данными для ИИ выглядел одинаково: разбить данные на части, положить их в векторную базу и искать по эмбеддингам. Обычно предполагается, что если поиск работает плохо, проблема лишь в качестве модели: больше данных, больше параметров, лучше обучение — и всё стане…
Сильная новая работа от Meta. Обычно законы масштабирования предполагают, что размер модели и объём данных влияют на ошибку независимо друг от друга. В этой работе авторы вводят Skaling law — закон, который связывает ёмкость модели и данные через единый параметр взаимодействия. Дополнительный член снижает среднюю абсолютную процентную ошибку примерно в 1,5–3 раза как при интерполяции, так и при экстраполяции. Сам…
«The Mathematics of Bitcoin» — короткая работа, которая разбирает Bitcoin с математической точки зрения. В ней используются теория вероятностей, случайные процессы, мартингалы, комбинаторика и специальные функции, чтобы исследовать механизмы протокола Bitcoin. В частности, авторы разбирают вероятность двойной траты, прибыльность майнинга, генерацию блоков, стратегии майнеров и устойчивость протокола. Если хочется …
Showing the 12 most recent of 33 posts we hold for @LLMScience. View and reaction counts are the latest single reading for each post, not a live figure, and a recent post is still accumulating both. A view count marked ≈ was rounded by Telegram before we ever saw it — t.me prints views in full below 1,000 and to three significant figures above, so ≈1,200,000 means somewhere between 1,150,000 and 1,249,999. Unmarked counts are exact. Text is reproduced from the public post preview and truncated for length.
Citation-graph rank
Citation-graph rank — 813,719 of 1,160,990entries in the measured graph. A weighted position computed from the forward and mention edges below — republished posts weigh more than named mentions — and recomputed periodically, over the whole graph. Published only as this ordinal position, never as a score: a position is a fact, and a score printed beside one channel’s name would read as a verdict this register does not make. The two counts beneath stay separate for the same reason mentions are never summed with forwards anywhere else on this page — a named-by count costs nothing to manufacture. The top 100 by this measure, or how it is computed.
Forward network
Republished by
Channels on the register that have forwarded this channel's posts into their own feed.
Built only from forwarded posts we have actually read, on both sides. Coverage is early and deliberately incomplete: a missing link means we have not read the post that would prove it, never that the relationship does not exist. Counts are distinct forwarded posts observed, so they only ever go up as we read more.
Mentions
Names
Channels on the register whose handles appear in this channel's posts.
A mention is a weaker signal than a forward and is counted separately for that reason — naming a channel is not republishing it, and a handle in a post body is easy to place deliberately. The post counts beside each row below are distinct posts in which the handle appeared, from posts we have read on both sides — the “Named by N registered channels” figure above is a different count, of distinct NAMING CHANNELS rather than posts, and is not the sum of the rows under it.
Cite this entry
A live page changes as we take new readings, so a citation should name the measurement it is based on, not just the URL. The line below cites the subscriber count as measured 12 August 2026 — this entry's latest reading, not the date you are reading this.
“Data Portal | DS & ML” (@LLMScience), 8,376 subscribers as measured 12 August 2026. Telegram Register, tgregister.com/channel/LLMScience.
Full measurement history, CC BY 4.0. Every reading this register holds for this entry, not just the latest one, as a dated, downloadable record: CSV · JSON. Free to use with attribution to tgregister.com. Each file carries its own generation timestamp, which is the figure to cite for exactly when the data was retrieved.