跳到正文 / Skip to content

AI Daily Digest · 2026-05-22

20 papers · Multi-source aggregation + AI summarization

· 11 min read #digest#auto#ai-papers

Hugging Face Daily Papers

π-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows

HF 37 · Haoran Zhang, Luxin Xu, Zhilin Wang… · HF Mirror

Addressing the shortcoming of existing personal assistant agent benchmarks that lack evaluation of long-cycle multi-round implicit user intents, this study launches the π-Bench benchmark: it includes 100 multi-round tasks covering 5 types of user personas, incorporates implicit intents, task dependencies, and cross-session features, and simultaneously evaluates agent proactivity and long-cycle task completion rates. Experiments show that current proactive assistants still have performance bottlenecks, with task completion and proactivity capabilities decoupled, and historical interactions can improve subsequent proactive intent recognition performance.

ACC: Compiling Agent Trajectories for Long-Context Training

HF 36 · Qisheng Su, Zhen Fang, Shiting Huang… · HF Mirror

To solve the problems of high cost for long-context training of large models and the fact that existing agent fine-tuning does not utilize scattered long-context evidence in trajectories, this paper proposes the ACC method, which can integrate multi-round agent trajectories into long-context QA pairs without additional annotation, directly supervising long-range reasoning. Experiments show that the 30B-parameter Qwen3 trained with ACC achieves significant performance gains on long-dependency benchmarks, matching the performance of the 235B-parameter version, without compromising its general capabilities.

WorldKV: Efficient World Memory with World Retrieval and Compression

HF 19 · Jung Yi, Minjae Kim, Paul Hyunbin Cho… · HF Mirror

Addressing the pain point that autoregressive video diffusion generation struggles to balance scene consistency and real-time performance, the training-free framework WorldKV is proposed: first, it retrieves evicted historical KV cache blocks through camera/action matching, and inserts them directly back into the attention window without re-encoding; second, it compresses redundant tokens based on key similarity with anchor frames, cutting single block storage in half. Tests show its throughput doubles compared to full KV caching, with consistent consistency performance, and its performance without fine-tuning is on par with trained baseline solutions.

Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning

HF 15 · Banghao Chi, Yining Xie, Mingyuan Wu… · HF Mirror

To solve the pain point that existing general LLM prompt-driven spreadsheet agents struggle to handle complex multi-step tasks in real scenarios, this study proposes the Spreadsheet-RL reinforcement learning fine-tuning framework, paired with an automatic dataset collection pipeline, domain-specific benchmarks, and an Excel-adapted multi-round interactive training environment. The trained Qwen3-4B model sees nearly a 1x improvement in Pass@1 accuracy on both general and domain-specific spreadsheet tasks, with excellent deployment potential.

FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching

HF 11 · Jangho Park, Geon Yeong Park, Gihyun Kwon… · HF Mirror

Addressing the pain points of existing training-free long video generation solutions: bidirectional methods are tied to specific architectures and suffer from long-term quality degradation, while autoregressive methods have error drift and repeated actions, this paper proposes the architecture-agnostic, no-additional-training inference solution FlowLong: it uses overlapping sliding window generation, fuses adjacent windows via manifold-constrained Tweedie matching, uses random sampling to synchronize trajectories in the high-noise phase, and switches to deterministic sampling to preserve details in the low-noise phase. This method can generate videos several times longer than the native window length, with better quality and temporal consistency than baselines, and can also be extended to audio-video joint generation and text-to-3DGS scenarios.

arXiv cs.LG (Machine Learning)

Neural Estimation of Pairwise Mutual Information in Masked Discrete Sequence Models

Jai Sharma, Yifan Wang, Bryan Li

Addressing the pain point that masked diffusion models lack explicit representations of variable dependencies, which limits efficiency and interpretability, this paper proposes a neural pairwise mutual information estimation framework: based on the hidden states of the pre-trained model, supervised by the real mutual information calculated from the model’s own conditional distribution, a single forward pass can obtain the full mutual information matrix, supporting mutual information-guided parallel decoding. Experiments show that it can restore structural constraints on Sudoku and protein generation tasks, reduces the number of inference forward passes by 3-5 times, and outperforms entropy-based parallel solutions.

GraphDiffMed: Knowledge-Constrained Differential Attention with Pharmacological Graph Priors for Medication Recommendation

Krati Saxena, Tomohiro Shibata

To solve the problems of difficult handling of long-sequence noise and clinical heterogeneity in electronic medical record medication recommendation, and that existing methods struggle to balance temporal modeling and pharmacological knowledge integration with weak noise resistance, this paper proposes the GraphDiffMed framework: it filters redundant signals through dual-scale differential attention for intra-consultation and cross-consultation, and introduces pharmacological constraints during training. It outperforms baselines on the MIMIC-III dataset, with better recommendation quality, ranking and safety, and its optimal configuration only requires demographic auxiliary features. The code has been open-sourced.

TabPFN-MT: A Natively Multitask In-Context Learner for Tabular Data

Cormac Cureton, Narges Armanfard

Addressing the shortcomings of existing tabular domain PFNs that are designed for single tasks, have high multi-objective inference costs and cannot share task information, this paper proposes the multi-task in-context learning model TabPFN-MT: trained on multi-objective synthetic priors, paired with an extended y encoder and shared decoding head, supporting simultaneous multi-task inference. Tests on small and medium datasets with fewer than 1,000 samples show that its multi-task tabular learning performance achieves a new SOTA, reduces inference cost from O(T) to O(1), has the highest average accuracy ranking, and can also match the performance of the latest single-task ensemble models.

OpenAI Official Updates

AdventHealth advances whole-person care with OpenAI

OpenAI

US healthcare provider AdventHealth has deployed a custom version of ChatGPT tailored for medical scenarios, optimizing the entire internal diagnosis and treatment process with AI tools. This solution can effectively reduce the administrative burden on medical staff, cut down the time they spend on non-clinical trivialities, so that medical staff can devote more energy to direct patient care, helping the institution achieve the service upgrade goal of “whole-person care”.

How Ramp engineers accelerate code review with Codex

OpenAI

The engineering team of fintech company Ramp has deployed the Codex large model powered by GPT-5.5 to optimize the code review process. Previously, the team needed hours to obtain substantive review feedback, slowing down the iteration pace; after introducing this tool, effective feedback can be obtained in just a few minutes, greatly shortening the review cycle and significantly improving R&D delivery efficiency, providing a referential implementation practice for large models empowering R&D efficiency scenarios.

Anthropic News

Introducing Claude Opus 4.7

Anthropic

Anthropic’s latest large model Claude Opus 4.7 is now officially available for general access. Compared to the previous generation Opus 4.6, this version has been specifically optimized for high-level software engineering scenarios, with significantly improved overall processing capabilities for related tasks, with particularly prominent performance gains for the most difficult complex software engineering tasks, making it more suitable for the needs of high-difficulty R&D scenarios.

Introducing Claude Design by Anthropic Labs

Anthropic

Anthropic Labs has officially launched the new product Claude Design, which focuses on human-machine collaborative creation capabilities, supporting users to work with Claude to complete various high-quality visual content production, covering multiple common visual outputs such as design drafts, product prototypes, presentation slides, and single-page promotional materials, providing users with convenient and efficient intelligent design assistance.

Google DeepMind

We’re launching the Google DeepMind Accelerator program in Asia Pacific to tackle environmental risks

Google DeepMind

Google DeepMind has recently officially launched its Asia Pacific accelerator program, which is positioned to rely on its leading AI technical capabilities to address various regional environmental risks. The program will collaborate with local Asia Pacific research institutions, tech startups and public departments, focus on scenarios such as extreme weather warning, ecological restoration, and carbon emission reduction management, implement solutions adapted to regional characteristics, and improve climate resilience in the Asia Pacific region.

Fast-tracking genetic leads to reverse cellular aging

Google DeepMind

This research focuses on the demand for discovering genetic targets for reversing cellular aging, innovatively uses the research assistance tool named Co-Scientist for screening, greatly shortening the target discovery cycle, and successfully excavates new regulatory factors, which have been experimentally verified to effectively achieve rejuvenation reprogramming of human cells. This achievement not only provides new candidate targets for anti-aging research, but also provides a reference for auxiliary scientific research tools empowering biomedical exploration.

Hugging Face Blog

OlmoEarth v1.1: A more efficient family of Earth observation models

Hugging Face

OlmoEarth v1.1 launched by the Allen Institute for AI is a new generation of efficient Earth observation model family, optimized for the previous generation’s architecture and pre-training strategy, adopting a lightweight multimodal remote sensing pre-training paradigm, reducing training computing cost by 35% compared to remote sensing large models of the same scale under the same accuracy. It outperforms existing SOTA models by 2%~4% in accuracy on more than 10 downstream tasks such as land classification and disaster detection, can be adapted for edge deployment, and greatly lowers the threshold for deploying remote sensing intelligent applications.

Introducing the Ettin Reranker Family

Hugging Face

This paper launches the Ettin reranker family for information retrieval and LLM RAG scenarios, covering multiple parameter scales, adopting multilingual cross-domain contrastive learning pre-training and task-adaptive fine-tuning schemes. Its accuracy is 3%~7% higher than existing models of the same specification on more than 10 public benchmarks, with inference latency reduced by more than 20%, and it can adapt to different deployment requirements on end devices and the cloud.

The Gradient

After Orthogonality: Virtue-Ethical Agency and AI Alignment

The Gradient

This AI alignment research reflecting on the orthogonality thesis, from the perspective of virtue ethics, refutes the traditional assumption that rational agents need to be anchored to fixed final goals, and proposes that human rationality is essentially a pattern of adapting actions to a practical network containing normative and evaluative systems. To achieve compliant collaboration between AI and humans and meet safety alignment requirements, AI decision logic needs to be isomorphic to human practice-based action logic, taking into account both ethical values and basic safety attributes.

QbitAI

The fastest among top-tier models! Zhipu, are you “spitting out” code?

QbitAI

Zhipu AI has newly launched the GLM-5.1-highspeed high-speed API, which is officially claimed to be the fastest code generation API among current top-tier large models, with a speed of 400 tokens/s. The basic version of GLM-5.1 itself is also among the open source models with the strongest code capabilities. Actual tests show that it can output complete code that meets the requirements of complex interactive dynamic effect web pages in more than ten seconds, and can also quickly respond to iteration and tuning requirements.

80-episode short drama, filmed in 3 days: When filmmakers enter the Agent space, film and television production gets the “most knowledgeable” solution

QbitAI

Most current AI video tools focus on single-segment flashy screen generation, which cannot adapt to the full-link needs of film and television industrialization. A film and television team with 20 years of deep industry experience has launched the AI film and television agent product MovieFlow Studio, which has three core capabilities: full-link production closed loop, enterprise-level asset library that prevents content visual drift, and thousand-person level collaborative management, solving the industry’s pain point of “having a pen but no production line”, and promoting the industrialized production of AI film and television.

390,000 yuan! Lei Jun releases Xiaomi’s most expensive SUV

QbitAI

Xiaomi has released the YU7 series SUV, including two models: the high-performance GT version and the standard version. The GT version, known as the “authentic GT”, is the first to be equipped with the new generation V8s EVO super motor, with a 0-100km/h acceleration time of 2.92 seconds, and its Nürburgring SUV lap time is 14 seconds faster than the previous top record holder, priced at 390,000 yuan. The standard version is a fully configured model, not a cut-down version, with a starting price of 233,500 yuan, 30,000 yuan lower than the Tesla Model Y, entering the mid-range new energy SUV market to compete.


Was this useful? A rating helps me pick the next topic.

Click a star to rate · Only anonymous fingerprint + timestamp stored

评论 · Comments