https://www.reddit.com/r/homeassistant/comments/1vhofep/hacked_and_debloated_an_echo_dot_2_local_llm/

Channel
Speech Technology
@speechtech
On this record: Growth · Engagement · Posts · Citations · Cite this entry
1,689subscribers
+4 since we began measuring on 6 August 2026
Risers and fallers across the register · movement among entries of 1,000–3,162.
Register entry
| Telegram ID | -1001472248479 |
|---|---|
| Type | Channel |
| Username | @speechtech |
| Created | 2 April 2020 — measured — cross-checked against a third-party dataset (TGDataset) |
| First recorded | 6 August 2026 |
| Last confirmed live | 12 August 2026 |
| Measurements held | 3 |
| Confirmed unchanged | 2 times, most recently 12 August 2026 |
| On Telegram | t.me/speechtech |
Growth
| Measured (UTC) | Subscribers | Change |
|---|---|---|
| 10 Aug 2026, 02:41 | 1,689 | +6 |
| 6 Aug 2026, 21:20 | 1,683 | -2 |
| 6 Aug 2026, 13:51 | 1,685 | first reading |
Engagement
22 posts held, back to 13 July 2026 — the reader has not yet reached the start of this channel’s public history, so older posts may sit further back, unread. Read across 2 pagesof Telegram’s post history, 20 posts per page.
- ERR · 30 days
- 53.4%
- avg views ÷ 1,689 subscribers
- Avg views / post
- 903
- 21 posts measured
- Reaction rate
- —
- this channel exposes no reaction counts
- Posts in window
- 21
- of 22 held
ERR is average views per post over the last 30 days divided by subscribers, the definition TGStat uses, so this figure is comparable with the one you will see elsewhere. It falls structurally as a channel grows: a high ERR on a small channel and a low one on a large channel describe reach mathematics, not quality. We publish the figure and the sample it came from and pass no verdict on it.
ER is defined industry-wide as (forwards + reactions + comments) ÷ views— note the denominator is views, not subscribers. Telegram’s public web preview carries views and reactions but not forward or comment counts, so the reaction rate above is the reactions term only and is therefore a floor: the true ER for this channel is higher by an amount we have not measured and will not estimate.
| Window | Rolling 30 days · latest post in window 7 August 2026 |
|---|---|
| Posts held | 22 (13 July 2026 – 7 August 2026) |
| Views total | 18,957 |
| Reactions total | — |
| Forwards / comments | not exposed by the public surface — not measured, not estimated |
| Readings taken | 7 Aug 2026, 17:44 UTC |
Views are a single reading per post, taken at the time above. A post published in the last day or two is still accumulating views, which pulls the 30-day average down slightly. That is a property of the standard definition rather than a fault in it, so we keep the definition rather than “correcting” the number into something nobody can reproduce.
Precision. Telegram publishes view counts on its public widget in short form — 8.12K, 3.7M — so any reading at or above 1,000 reaches us rounded to three significant figures, and only counts below 1,000 are exact. Averages and rates derived from them are shown to the same precision rather than to the unit: a figure like 3,701,250 would assert digits nobody measured.
Reaction counts are published per emoji and rounded the same way, so a total below 1,000 is exact and a larger one is a sum that may carry a rounded component from each emoji above 1,000. Because it is a sum, it does not look rounded — read a large reaction total as three significant figures per contributing emoji rather than as the figure it prints.
Recent posts
There is a big interest in full duplex as I see, here is a nice collection of papers https://github.com/Ruiqi-Yan/Awesome-Full-Duplex-SDM
https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B NVIDIA NemotronLabs VoiceChat is a 11B end-to-end, real-time speech full duplex (FD) model for conversational AI that jointly performs streaming speech understanding and speech generation [1, 2]. Unlike traditional cascaded stacks (ASR → LLM → TTS), this model achieves full duplex, real-time, seamless voice interaction in one unified architecture, elimi…
Things move on in OpenAI as well. Interesting that voice model is separate. And no turn detector anymore. https://x.com/OpenAI/status/2084378415818579975 Lots of interesting technical details, from realtime inference, to dynamic compaction, to WebRTC optimization. https://openai.com/index/continuous-voice-interaction-with-gpt-live/
One more https://github.com/AnXMuy/AgenticASR
https://huckiyang.github.io/voice-memory/ from NVIDIA https://arxiv.org/abs/2607.26410 Voice Memory for Agentic Speech Recognition Chao-Han Huck Yang, Zih-Ching Chen, Piotr Zelasko, Zhehuai Chen, Jagadeesh Balam, Boris Ginsburg We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain this http URL and decides per utterance whether…
https://huggingface.co/nyralabs/CrisperWhisper2.0_large Most speech-to-text systems never actually decide whether to write down what was said or what was meant. They inherit that choice from their training data and apply it inconsistently. CrisperWhisper 2.0 makes it an explicit, controllable choice. One recording, two transcripts: Verbatim, exactly what was said, in one consistent format: [um] so we we need to, to …
Some recent Uzbek things https://huggingface.co/datasets/k2speech/FeruzaSpeech - single speaker 40 hours TTS dataset https://huggingface.co/collections/navai-uz/navai-whisper-collection - recently trained Whisper models from Navai https://navai.pro https://huggingface.co/instinct-org/collections - some loosely organized data https://huggingface.co/datasets/OvozifyLabs/asr_evaluate_set - evaluation dataset with Te…
Everyone builds self-improvement loops in LLMs, I wonder how they could look like in ASR/TTS. Not many publications on that yet.
We compared three LALM judges against a calibrated human panel across 15 dimensions of speech quality. The LALMs tracked humans closely on relevance, answer quality, and instruction following—what was said—but were much less reliable on naturalness, emotion, pronunciation, and overall feel—how it was said. https://research.withdavid.ai/blog/lalm-as-judge-vs-hitl
Interesting math on speech LLM https://arxiv.org/abs/2604.08003v1 Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs Yuan Xie, Jiaqi Song, Guang Qiu, Xianliang Wang, Ming Lei, Jie Gao, Jie Wu Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a dominant paradigm. Although recent LLM-based ASR models have shown promisin…
Interesting project https://github.com/Xiaobin-Rong/unipase UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations Xiaobin Rong, Zheng Wang, Yushi Wang, Jun Gao, Jing Lu Universal speech enhancement (USE) aims to restore speech signals from diverse distortions across multiple sampling rates. We propose UniPASE, an extension of the low-hallucination PASE framework tail…
Showing the 12 most recent of 22 posts we hold for @speechtech. View and reaction counts are the latest single reading for each post, not a live figure, and a recent post is still accumulating both. A view count marked ≈ was rounded by Telegram before we ever saw it — t.me prints views in full below 1,000 and to three significant figures above, so ≈1,200,000 means somewhere between 1,150,000 and 1,249,999. Unmarked counts are exact. Text is reproduced from the public post preview and truncated for length.
Citation-graph rank
Citation-graph rank — 261,445 of 1,160,990entries in the measured graph. A weighted position computed from the forward and mention edges below — republished posts weigh more than named mentions — and recomputed periodically, over the whole graph. Published only as this ordinal position, never as a score: a position is a fact, and a score printed beside one channel’s name would read as a verdict this register does not make. The two counts beneath stay separate for the same reason mentions are never summed with forwards anywhere else on this page — a named-by count costs nothing to manufacture. The top 100 by this measure, or how it is computed.
Forward network
Republished by
Channels on the register that have forwarded this channel's posts into their own feed.
Built only from forwarded posts we have actually read, on both sides. Coverage is early and deliberately incomplete: a missing link means we have not read the post that would prove it, never that the relationship does not exist. Counts are distinct forwarded posts observed, so they only ever go up as we read more.
Cite this entry
A live page changes as we take new readings, so a citation should name the measurement it is based on, not just the URL. The line below cites the subscriber count as measured 10 August 2026 — this entry's latest reading, not the date you are reading this.
“Speech Technology” (@speechtech), 1,689 subscribers as measured 10 August 2026. Telegram Register, tgregister.com/channel/speechtech.
Full measurement history, CC BY 4.0. Every reading this register holds for this entry, not just the latest one, as a dated, downloadable record: CSV · JSON. Free to use with attribution to tgregister.com. Each file carries its own generation timestamp, which is the figure to cite for exactly when the data was retrieved.