PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents This work introduces a benchmark to test if personal AI agents really become better over time by using saved experience across sessions. It compares agents with memory on and off across many task sequences, and shows improvement happens but is uneven and not always through the expected save-retrieve-update process. The paper a…

Channel
Boring AI News
@boring_ai_news
On this record: Growth · Engagement · Posts · Cite this entry
24subscribers
+3 since we began measuring on 7 August 2026
Risers and fallers across the register · movement among entries of Under 1,000.
Register entry
| Telegram ID | -1002129615196 |
|---|---|
| Type | Channel |
| Username | @boring_ai_news |
| Created | Between 1 November 2023 and 31 May 2024 — estimated from Telegram’s id allocation, not measured. How this range is calculated. |
| First recorded | 11 August 2026 |
| Last confirmed live | 7 September 2026 |
| Measurements held | 5 |
| Confirmed unchanged | 1 time, most recently 7 September 2026 |
| On Telegram | t.me/boring_ai_news |
Growth
| Measured (UTC) | Subscribers | Change |
|---|---|---|
| 7 Sept 2026, 16:42 | 24 | +1 |
| 23 Aug 2026, 02:02 | 23 | +1 |
| 14 Aug 2026, 19:47 | 22 | +1 |
| 11 Aug 2026, 14:01 | 21 | no change |
| 7 Aug 2026, 20:28 | 21 | first reading |
Engagement
20 posts held, back to 7 August 2026 — the reader has not yet reached the start of this channel’s public history, so older posts may sit further back, unread. Read across 1 page of Telegram’s post history, 20 posts per page.
- ERR · 30 days
- 20.8%
- avg views ÷ 24 subscribers
- Avg views / post
- 5.0
- 5 posts measured
- Reaction rate
- —
- this channel exposes no reaction counts
- Posts in window
- 5
- of 20 held
ERR is average views per post over the last 30 days divided by subscribers, the definition TGStat uses, so this figure is comparable with the one you will see elsewhere. It falls structurally as a channel grows: a high ERR on a small channel and a low one on a large channel describe reach mathematics, not quality. We publish the figure and the sample it came from and pass no verdict on it.
ER is defined industry-wide as (forwards + reactions + comments) ÷ views — note the denominator is views, not subscribers. Telegram’s public web preview carries views and reactions but not forward or comment counts, so the reaction rate above is the reactions term only and is therefore a floor: the true ER for this channel is higher by an amount we have not measured and will not estimate.
| Window | Rolling 30 days · latest post in window 10 August 2026 |
|---|---|
| Posts held | 20 (7 August 2026 – 10 August 2026) |
| Views total | 25 |
| Reactions total | — |
| Forwards / comments | not exposed by the public surface — not measured, not estimated |
| Readings taken | 11 Aug 2026, 14:01 UTC |
Views are a single reading per post, taken at the time above. A post published in the last day or two is still accumulating views, which pulls the 30-day average down slightly. That is a property of the standard definition rather than a fault in it, so we keep the definition rather than “correcting” the number into something nobody can reproduce.
Precision. Telegram publishes view counts on its public widget in short form — 8.12K, 3.7M — so any reading at or above 1,000 reaches us rounded to three significant figures, and only counts below 1,000 are exact. Averages and rates derived from them are shown to the same precision rather than to the unit: a figure like 3,701,250 would assert digits nobody measured.
Reaction counts are published per emoji and rounded the same way, so a total below 1,000 is exact and a larger one is a sum that may carry a rounded component from each emoji above 1,000. Because it is a sum, it does not look rounded — read a large reaction total as three significant figures per contributing emoji rather than as the figure it prints.
Recent posts
Prime Agent: A self-improving RLM agent A self-improving coding harness built on Recursive Language Models and Continual Harness, with a persistent REPL that gives programmatic access to history, sub-agents, tools, and state for long-horizon coding, evaluation, and research workflows. https://www.primeintellect.ai/blog/prime-agent #learning
SOTA alignment assessments don't strongly update us against misalignment Frontier labs' alignment assessments provide much weaker evidence against misalignment than system cards suggest, because the covert-capability evaluations behind those reliability claims are themselves unreliable. https://blog.redwoodresearch.org/p/sota-alignment-assessments-dont-strongly #learning
Introducing Flex: Let the Model Write the Code Flex uses model-generated code and prompt rewrites inside a sandboxed interpreter to produce faster, lower-cost programs. https://www.cmpnd.ai/blog/let-the-model-write-the-code.html #learning
The next chapter of our AI momentum New AI updates in products including Gmail and Maps add stronger personalization and contextual understanding, aiming to improve user experience and support Google's position in the tech industry. https://blog.google/company-news/inside-google/message-ceo/next-chapter-ai-momentum #news
Qwen-Image-3.0-Pro Qwen-Image-3.0-Pro generates complex layouts such as newspapers, storyboards, menus, and exam papers in a single pass. https://www.qwencloud.com/models/qwen-image-3.0-pro #models
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement This work turns open-ended LLM tasks into proxy game-like settings where rewards can be checked automatically, without human judges or reward models. In their SpyRL setup, agents solve the same task with different information and vote to find a preset spy, so the vote gives self-verifiable reward signals. Tes…
LoopX: Stateful Control Plane for Long-Running AI Agents LoopX is a lightweight, provider-neutral state kernel for managing long-running AI agent work across tools like Codex and Claude Code. It keeps goals, todos, gates, evidence, quota, and handoffs stable across bounded turns, making agent workflows more reviewable, restartable, and easier to govern. https://github.com/huangruiteng/loopx #code
ByteDance SeedRealtime A native audio-visual model that handles continuous video, audio, and text while generating speech in real time. https://seed.bytedance.com/en/SeedRealtime #learning
Science One Framework: A verifiable autonomous research framework via Chain-of-Evidence Science One Framework and CoE Audit provide an autonomous research system with zero phantom references and fully verifiable results, outperforming baselines that hallucinate up to 21% of citations. https://research.google/blog/science-one-framework-a-verifiable-autonomous-research-framework-via-chain-of-evidence #learning
Mistral Releases Shieldstral, a 3B Multimodal Safety Classifier Mistral has released Shieldstral, a 3B Apache 2.0 open-weights multimodal safety classifier that takes plain-language policies at inference time, unifies text and image safety evaluation without retraining, and matches or outperforms guard models up to 7x its size while running on a single 16GB NVIDIA GPU. https://mistral.ai/news/shieldstral #news #le…
NVIDIA's Real-Time Full-Duplex Voice Model NemotronLabs VoiceChat is an 11B end-to-end speech model with a single architecture for streaming understanding, speech generation, and tool calling. https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B #models
Showing the 12 most recent of 20 posts we hold for @boring_ai_news. View and reaction counts are the latest single reading for each post, not a live figure, and a recent post is still accumulating both. A view count marked ≈ was rounded by Telegram before we ever saw it — t.me prints views in full below 1,000 and to three significant figures above, so ≈1,200,000 means somewhere between 1,150,000 and 1,249,999. Unmarked counts are exact. Text is reproduced from the public post preview and truncated for length.
Cite this entry
A live page changes as we take new readings, so a citation should name the measurement it is based on, not just the URL. The line below cites the subscriber count as measured 7 September 2026 — this entry's latest reading, not the date you are reading this.
“Boring AI News” (@boring_ai_news), 24 subscribers as measured 7 September 2026. Telegram Register, tgregister.com/channel/boring_ai_news.
Full measurement history, CC BY 4.0. Every reading this register holds for this entry, not just the latest one, as a dated, downloadable record: CSV · JSON. Free to use with attribution to tgregister.com. Each file carries its own generation timestamp, which is the figure to cite for exactly when the data was retrieved.