跳到正文 / Skip to content

AI Daily Highlights · 2026-09-28

17 papers · Multi-source aggregation + AI summaries

TL;DR · Today’s highlights in 30 seconds
  • DeepMind launches Gemini 3.8 Live with real-time digital humans, while Anthropic and OpenAI announce new technical and deployment progress in quick succession
  • Hugging Face releases multiple LLM optimization technologies, covering attention mechanisms, tool calling RL, multimodal acceleration and other areas
  • China’s quantum AI track sees a Tsinghua-founded startup team, desktop quantum computing products, and the anonymous Yutu model tops two performance rankings
🤖 Model Releases🔬 Technical Breakthroughs💼 Commercial Deployment⚛️ Quantum AI⚡ Performance Optimization

Hugging Face Daily Papers

FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

HF ★ 31 · Hongyang Du, Yunfei Xie, Junjie Ye… · HF Mirror

Aiming at the problems of performance trade-off between deep and shallow layers when selecting encoder layers corresponding to shared latent space for Representational Autoencoders (RAE), and the reconstruction-generation gap caused by fixed heuristic fusion, this paper proposes the FuseReg regularization method, which trains by randomly sampling subsets of encoder layers and penalizes cross-layer inconsistency sensitivity. This method does not require modifying the pretrained encoder, can improve reconstruction PSNR, reduces gFID of unguided image generation by up to 29%, and effectively narrows the performance gap between the two types of tasks.

InternW0-Δ: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data

HF ★ 11 · Xingyu Miao, Zizun Li, Baole Fang… · HF Mirror

For general robot manipulation tasks, this paper proposes InternW0-Δ, a unified world action model, which adopts a hybrid Transformer framework to integrate multiple types of pretrained priors, and is equipped with a causal imprint module to reduce inference overhead; after pretraining on a self-built multi-source heterogeneous open-source robot corpus of more than 20,000 hours, the model outperforms existing methods in both simulation and real-machine benchmark tests, and relevant resources will be open-sourced as required.

Game Arena: Strategic LLM Evaluation in Competitive Environments

HF ★ 3 · Bovard Doerschuk-Tiberi, Yao Yan, Justin Chiu… · HF Mirror

This research launches Kaggle Game Arena, an open and scalable LLM competitive evaluation platform. Different from traditional static evaluation benchmarks, the platform uses three types of competitive games covering perfect/imperfect information and multi-player scenarios: chess, poker, and Werewolf, which can dynamically evaluate LLM capabilities such as strategic planning and adaptability under uncertainty, avoid the performance saturation problem of static evaluation, have reproducible and highly transparent evaluation, and can also adapt to new game types.

Block Sparse Attention with Log-Linear Complexity

HF ★ 1 · Bohao Tang, Zhen Qin, Yuqi Pan… · HF Mirror

Aiming at the pain point that the quadratic complexity of self-attention limits the long context expansion of LLMs, and the block selection of traditional block sparse attention still has quadratic complexity, this paper proposes the PISA pyramid block sparse attention mechanism: it screens candidate blocks through a coarse-to-fine hierarchical LogSumExp, reduces the complexity to O(NlogN), and is equipped with hardware-adapted Triton operators. Its commonsense reasoning performance is comparable to the baseline, and its performance on retrieval tasks is better.

SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL

HF ★ 1 · Yan Zhan, Shaobo Liu, Qiunan Liu… · HF Mirror

Aiming at the problem of cross-segment credit misattribution and gradient noise interfering with tool decision-making in on-policy reinforcement learning such as GRPO in tool calling scenarios, this paper proposes the SLCA-GRPO framework: first build a pattern-guided LLM simulator as the training base, and match execution and preference advantages to tool and text generation segments respectively through segment-locked credit allocation combined with hierarchical rewards. Under the 7B base model, its performance is better than multiple baselines, and the accuracy on τ²-Bench is improved by up to 9.15 percentage points, with faster convergence and lower tool calling costs.

OpenAI

Proaction boosts sales 60% and saves 75+ hours with Codex

OpenAI

The publicly released deployment results show that enterprise service vendor Proaction uses three AI tools: Codex, GPT-Live-1, and GPT-6 Astra to intelligently improve the efficiency of the entire process of product construction, operation, and sales of its modern fleet management business, finally achieving a 60% increase in sales, while saving a total of more than 75 hours of operational labor hours, providing a measured reference for the deployment benefits of LLMs in ToB vertical business scenarios.

Two years of OpenAI Academy

OpenAI

This is a commemorative disclosure for the second anniversary of the operation of OpenAI Academy, a skill inclusive project under OpenAI. The core function of the project is to promote and teach practical AI-related skills to the public. On the occasion of the second anniversary, the project will further expand its service coverage, reach more diverse community groups, promote AI skill inclusion, and lower the threshold for AI technology learning for different groups.

Anthropic News

Claude discovers a novel enzyme system

Anthropic

A newly established life science laboratory announced its initial research results: during the research process, a Claude agent discovered a completely new enzyme system that has not been reported before. At present, the specific physiological function and mechanism of action of this system have not been clearly analyzed. This discovery expands the boundary of human understanding of the enzyme family, and provides a new research target for subsequent basic enzymology research and synthetic biology application exploration.

Partnering with Accenture on embedded evaluation

Anthropic

AI company Anthropic announced a partnership with Accenture to carry out third-party independent evaluation of cutting-edge AI. This cooperation is part of the “embedded evaluators” governance commitment previously proposed by Anthropic. To build a suitable evaluation capability system, the two parties agreed to invest no less than US$1 billion each in this field in the next five years, and work together to consolidate the relevant foundation for cutting-edge AI risk assessment.

Google DeepMind

Introducing Gemini 3.8 Live with Live Avatar

Google DeepMind

Google officially launched the real-time interactive version of Gemini 3.8 Live, with the core new Live Avatar real-time virtual human function. This version adopts an optimized end-to-end low-latency multimodal scheduling framework, paired with a dedicated generation model that integrates lip sync and micro-expression fitting, which can synchronously respond to voice and text input and output anthropomorphic virtual image interaction, with end-to-end latency lower than 200ms, supporting smooth multi-turn conversations, and the interaction naturalness is greatly improved compared to the previous generation, which can be adapted to multiple scenarios such as virtual customer service and online companionship.

Advancing Private AI Compute with secure, server-side memory

Google DeepMind

This paper focuses on the pain point of private AI computing privacy protection in personal AI scenarios, and proposes an innovative solution of introducing secure server-side private memory into the private AI computing system. This solution not only retains the advantages of cloud computing power, but also blocks the risk of personal sensitive data leakage in the AI computing link, taking into account both AI service operation efficiency and user privacy security, providing a more reliable architectural support for the deployment of personal private AI.

Hugging Face Blog

Accelerating vision-language models with LFM2.5-VL-DSpark

Hugging Face

At present, only the title of the paper is provided, and the full text of the abstract is not pasted, so translation and extraction work cannot be carried out. Please supplement the complete text information of this paper’s abstract, and I will highlight the core methods and conclusions as required to generate a streamlined Chinese summary of about 120 words for you.

How to Use NVIDIA Warp and MjWarp to Accelerate Robotics Simulation and Learning Workflows

Hugging Face

This article introduces the implementation method of NVIDIA’s high-performance GPU parallel computing framework Warp, and the MjWarp interface adapted to the MuJoCo physics engine: by offloading all links such as simulation sampling, gradient calculation, and strategy optimization to GPU for parallel execution, without major reconstruction of existing robotics code pipelines, the efficiency of reinforcement learning training, motion planning and other tasks can be increased several times, greatly shortening the R&D iteration cycle.

Lil’Log

Harness Engineering for Self-Improvement

Lilian Weng

This article sorts out the evolution of the concept of Recursive Self-Improvement (RSI): In 1965, scholar I.J. Good first proposed the core idea, referring to a superintelligent system with performance far exceeding that of humans that can independently design better models to achieve self-iteration; in 2008, Eliezer Yudkowsky clarified that its feedback logic is that AI relies on existing intelligence to optimize its own cognitive bottom layer. The study points out that current AI RSI can be realized either by directly rewriting its own weights, or expanded to optimize the entire training process.

QbitAI

A “Tsinghua Dream Team” emerges in quantum AI entrepreneurship: valued at 1 billion RMB, transforming the underlying layer of LLMs with quantum technology

QbitAI

The core team of Fermi Universe, China’s first quantum-enabled AI (Q4AI) startup, is mostly from Tsinghua University, and has completed a 100 million RMB seed round of financing, with a post-investment valuation of 1 billion RMB, the largest seed round in the same field in China. It has launched the world’s first full-link quantum-enhanced LLM FermiQLLM 1.0, which embeds quantum methods into the full link of LLMs, does not rely on immature quantum hardware, can be deployed on existing computing power, has significantly improved inference performance compared with traditional LLMs of the same parameters, and is suitable for industrial deployment.

Quantum computing goes to the desktop! The “small box” implements end-to-end operation, with data never leaving the local environment throughout the process

QbitAI

Unitary Quantum, incubated by Shanghai Jiao Tong University, recently released UnitarySpark, the world’s first heterogeneous computing hardware base, and simultaneously launched the public beta version of UnitaryLab 2.5, a natural language-driven quantum scientific computing platform. Both have built-in AI Agent, users do not need quantum professional background, describe requirements in plain language to automatically schedule quantum and classical computing power to complete calculations, support local deployment to ensure data security, and greatly lower the threshold for using quantum computing.

Fast and powerful! The anonymous Yutu model tops two rankings, full record of actual Coding test

QbitAI

The recently popular anonymous AI model “Space Bunny (Yutu)” has topped the daily call volume rankings of both OpenRouter and OpenCode. Many overseas users have expressed disbelief, generally reporting that its output quality is excellent, it runs fast, and its performance exceeds Astra. QbitAI specially conducted an actual test of its coding ability, assigning the task of generating an interactive cube-style Temple of Heaven using Three.js, which was completed in less than 2 hours and met the standards, with performance exceeding expectations.

Was this useful? A rating helps me pick the next topic.

Click a star to rate · Only anonymous fingerprint + timestamp stored

评论 · Comments