6 Aug 2026, 07:59 UTC91 views1 reactionsread 7 August 2026 Photo
EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions
📅 Publication Date: Jun 22, 2026
📑 Paper: https://arxiv.org/pdf/2606.23654.pdf
🔗 Code: https://github.com/huggingface
📝 Description:
EnterpriseClawBench presents a benchmark for enterprise agents based on real-world sessions with 852 reproducible tasks, emphasizing comprehensive evaluation metrics beyond single performance scores.
❤1
4 Aug 2026, 08:04 UTC152 viewsread 7 August 2026 Photo
OpenRath: Session-Centered Runtime State for Agent Systems
📅 Publication Date: Jun 17, 2026
📑 Paper: https://arxiv.org/pdf/2606.19409.pdf
🔗 Code: https://github.com/huggingface
📝 Description:
OpenRath introduces a PyTorch-like programming model for multi-agent systems using Session as a central runtime abstraction that enables explicit fork, merge, and replay operations while recording comprehensive execution s…
2 Aug 2026, 09:59 UTC213 views1 reactionsread 7 August 2026 Photo
MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision
📅 Publication Date: Jun 15, 2026
📑 Paper: https://arxiv.org/pdf/2606.17162.pdf
🔗 Code: https://github.com/huggingface
📝 Description:
MemSlides presents a hierarchical memory framework for personalized presentation agents that separates long-term user profiles, working memory for session const…
❤1
1 Aug 2026, 10:00 UTC218 views2 reactionsread 7 August 2026 Photo
Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs
📅 Publication Date: May 7, 2026
📑 Paper: https://arxiv.org/pdf/2606.27378.pdf
🔗 Code: N/A
📝 Description:
An axiomatic evaluation framework reveals systematic failures in latent thought representations of LLMs across multiple reasoning tasks, demonstrating that current representations fail to satisfy fundamental functional axioms consisten…
❤2
30 Jul 2026, 07:02 UTC264 viewsread 7 August 2026 EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions
📅 Publication Date: Jun 22, 2026
📑 Paper: https://arxiv.org/pdf/2606.23654.pdf
🔗 Code: https://github.com/huggingface
📝 Description:
EnterpriseClawBench presents a benchmark for enterprise agents based on real-world sessions with 852 reproducible tasks, emphasizing comprehensive evaluation metrics beyond single performance scores.
28 Jul 2026, 08:35 UTC311 views1 reactionsread 7 August 2026 Photo
Heterogeneous Scientific Foundation Model Collaboration
📅 Publication Date: Apr 30, 2026
📑 Paper: https://arxiv.org/pdf/2604.27351.pdf
🔗 Code: https://github.com/Violet24K/Eywa
📝 Description:
Eywa is a heterogeneous agentic framework that extends language-centric systems to scientific foundation models by integrating domain-specific models with language-based reasoning interfaces for improved performance across …
🔥1
26 Jul 2026, 08:03 UTC310 viewsread 7 August 2026 OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents
📅 Publication Date: May 6, 2026
📑 Paper: https://arxiv.org/pdf/2605.05185.pdf
🔗 Code: https://github.com/shawn0728/OpenSearch-VL
📝 Description:
OpenSearch-VL presents an open-source framework for training advanced multimodal search agents using reinforcement learning, featuring specialized data curation, diverse tool environments, and a novel tr…
24 Jul 2026, 08:31 UTC329 views1 reactionsread 7 August 2026 Photo
WorldOlympiad: Can Your World Model Survive a Triathlon?
📅 Publication Date: Jun 9, 2026
📑 Paper: https://arxiv.org/pdf/2606.11129
💻 Project Page: https://alibaba-damo-academy.github.io/WorldOlympiad/
📝 Description:
The paper introduces WorldOlympiad, a comprehensive benchmark for evaluating video-based world models. The problem with current generative models is that they often focus on visual quality, but lack …
❤1
22 Jul 2026, 07:45 UTC327 viewsread 7 August 2026 Qwen-AgentWorld: Language World Models for General Agents
📅 Publication Date: Jun 23, 2026
📑 Paper: https://arxiv.org/pdf/2606.24597.pdf
🔗 Code: https://github.com/huggingface
📝 Description:
Language-based world models enable agentic environment simulation across multiple domains and enhance general agent performance through scalable simulation and improved downstream task performance.
18 Jul 2026, 11:01 UTC413 views1 reactionsread 7 August 2026 Photo
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
📅 Publication Date: Apr 30, 2026
📑 Paper: https://arxiv.org/pdf/2604.28196.pdf
🔗 Code: https://github.com/H-EmbodVis/HERMESV2
📝 Description:
HERMES++ combines 3D scene understanding and future geometry prediction through BEV representation, LLM-enhanced queries, temporal linking, and joint geometric optimization for autonom…
👍1
16 Jul 2026, 09:01 UTC412 views3 reactionsread 7 August 2026 Photo
ABot-Earth 0.5: Generative 3D Earth Model
📅 Publication Date: Jun 8, 2026
📑 Paper: https://arxiv.org/pdf/2606.09967
💻 Project Page: https://abot-earth.amap.com/
📝 Description:
The paper presents ABot-Earth 0.5, a generative framework that creates realistic 3D environments from satellite imagery. The problem addressed is the need for large-scale 3D reconstruction, which is currently expensive and technically chal…
❤2👏1
14 Jul 2026, 06:45 UTC401 viewsread 7 August 2026 OpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty Trajectories
📅 Publication Date: May 5, 2026
📑 Paper: https://arxiv.org/pdf/2605.04036.pdf
🔗 Code: https://github.com/PolarSeeker/OpenSeeker
📝 Description:
A simple supervised fine-tuning approach achieves state-of-the-art performance in deep search capabilities using minimal data, outperforming complex industrial pipelines a…
Showing the 12 most recent of 20 posts we hold for @data_science_research_papers. View and reaction counts are the latest single reading for each post, not a live figure, and a recent post is still accumulating both. A view count marked ≈ was rounded by Telegram before we ever saw it — t.me prints views in full below 1,000 and to three significant figures above, so ≈1,200,000 means somewhere between 1,150,000 and 1,249,999. Unmarked counts are exact. Text is reproduced from the public post preview and truncated for length.