Daily AI Highlights · 2026-07-13
18 papers · Multi-source aggregation + AI-generated summaries
- Hugging Face and the academic community released multiple cutting-edge AI research results covering panoramic generation, video model learning, LLM generalization and other fields
- OpenAI, Anthropic, and Google DeepMind successively announced business progress including model implementation, industry cooperation, and new product launches
- New developments emerged in China’s AI track, including humanoid world models, LLMs running on consumer-grade hardware, and the rising embodied data sector
Hugging Face Daily Papers
PanoWorld: Real-World Panoramic Generation
HF ★ 3 · Haoyuan Li, Dizhe Zhang, Yuemei Zhou… · HF Mirror
To address the long-term memory bottleneck of panoramic world models, this study leverages the rotational equivariance of omnidirectional representations to propose the PanoWorld model. It simplifies camera trajectory modeling through dense panoramic ray conditions and a geometry-aware memory enhancement module, optimized with a three-stage training process. The research also builds the large-scale World360 evaluation dataset containing real aerial photography and simulation data. Experiments show its performance is significantly better than similar solutions, and relevant resources will be open-sourced.
Video Generation Models are General-Purpose Vision Learners
HF ★ 3 · Letian Wang, Chuhan Zhang, Rishabh Kabra… · HF Mirror
To meet the R&D demand for general vision models, this study proposes that large-scale text-to-video generation can serve as a powerful pretraining paradigm, based on which GenCeption is built: it uses a pretrained video diffusion backbone and can complete multiple types of visual tasks scheduled via text instructions. It achieves SOTA on multiple tasks including depth estimation and segmentation, outperforms similar pretraining solutions, improves data efficiency by up to 500 times, and also has out-of-distribution generalization capabilities, proving that video generation is an important development path for general visual intelligence.
From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models
HF ★ 3 · Zanyi Wang, Xin Lin, Haodong Li… · HF Mirror
To solve the redundancy problem of existing text-to-image models when applying the generation paradigm for dense prediction, this paper proposes the ReChannel solution: it retains the input-side VAE encoder of the pretrained DiT, freezes DiT and only performs fine-tuning with LoRA, removes the original generation decoder, and adds a shared linear head with only 33K parameters to directly output task-native pixel results. It achieves SOTA on 3 out of 6 dense prediction tasks, with comparable performance on the rest. With the same 4B parameters, it has higher accuracy and is 2.48 times faster than existing generative solutions.
Trust Region Policy Distillation
HF ★ 3 · Zhengpeng Xie, Li Lyna Zhang, Zeke Xie… · HF Mirror
To address the pain points of unstable training and high gradient variance in existing On-policy Distillation (OPD), this paper proposes the Trust Region Policy Distillation method TOP-D, which realizes a stable training paradigm by dynamically constructing proximal teachers. It theoretically proves that it can control gradient variance, has global convergence and a monotonic improvement boundary, and has no additional computational overhead. Experiments show that its training stability, sample efficiency, and final performance on mathematical reasoning tasks are significantly better, making it an ideal alternative to OPD.
Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning
HF ★ 1 · Lu Dai, Ziyang Rao, Yili Wang… · HF Mirror
To address the “knowledge-application gap” problem that LLMs can quickly memorize new knowledge during fine-tuning but cannot apply it to downstream inference, this paper proposes a self-patching intervention technology to track the propagation dynamics of knowledge inside the model, verifying that the problem stems from knowledge circuit misalignment: memorized representations are not routed to the computationally effective layer. The simple heuristic strategy designed based on this finding can repair 58%~75% of generalization loss, and the conclusion is verified to be robust through cross-domain experiments.
OpenAI
How Deutsche Telekom is rewiring telecommunications with AI
OpenAI
This article introduces that Deutsche Telekom is cooperating with OpenAI to promote its transformation into an AI-native telecom operator. Currently, AI technology has covered four core business scenarios: optimizing customer service quality and efficiency, streamlining employee workflows, improving the intelligence level of network operation and maintenance, and exploring the next generation of voice service forms. This practice provides a referable implementation paradigm for the digital and intelligent upgrading of the telecom industry.
GPT-5.6 is now the preferred model in Microsoft 365 Copilot
OpenAI
Microsoft has now designated GPT-5.6 as the preferred LLM for Microsoft 365 Copilot. This model has stronger AI capabilities, can cover functional requirements of all office scenarios including Word, Excel, PowerPoint, intelligent chat, and collaborative office, can provide users with higher-quality AI assistance, effectively speed up office processing, and improve the overall quality of various office outputs.
Anthropic News
Inviting hard questions
Anthropic
The research team launched the public interaction project “Inviting hard questions”, which aims to respond to public concerns about AI development and eliminate technical information gaps. It openly collects the most challenging and most concerned difficult questions in the field of artificial intelligence from the whole society. The team also promises that the full process of subsequent research responding to the questions will have public work paths and details, ensuring the process is transparent and traceable.
Redeploying Claude Fable 5
Anthropic
Anthropic announced that after the relevant export controls are officially lifted, it will re-deploy the Claude Fable 5 LLM from July 1. The launched version has completed two core security upgrades: first, it has iterated a more complete network security protection mechanism, and second, it has added an exclusive jailbreak risk prevention and control framework adapted to industrial scenarios, and the overall security and compliance has been greatly improved compared with the old version.
Google DeepMind
Google DeepMind and A24 announce first-of-its-kind research partnership
Google DeepMind
Top AI R&D institution Google DeepMind and well-known independent film and television label A24 have reached the first cross-industry research cooperation in the industry. The two parties will explore scenarios such as film and television creative generation and production process optimization, focusing on exploring feasible paths for collaboration between AI and creators, expanding the implementation boundary of AI in the entertainment and creative industry, and exploring a new development model that balances content production efficiency and artistic originality.
Start building with Nano Banana 2 Lite and Gemini Omni Flash
Google DeepMind
It is detected that you have not pasted the corresponding abstract text, and the summary is made based on the core content of this public development document: This guide is for embedded AI developers, introducing the rapid deployment solution of the low-power open-source development board Banana Pi Nano Banana 2 Lite adapted to Google’s lightweight edge LLM Gemini Omni Flash, with supporting implementation examples for scenarios such as offline voice and lightweight visual recognition, which can greatly reduce the development threshold of edge generative AI applications, and meet the AI implementation needs of low-computing embedded devices.
Hugging Face Blog
Profiling in PyTorch (Part 3): Attention is all you profile
Hugging Face
This is the third part of the PyTorch performance profiling series, focusing on the performance diagnosis of Transformer’s core attention operator. Based on PyTorch’s built-in profiler tool, the article provides a practical solution to disassemble attention bottlenecks from three dimensions: computing latency, video memory usage, and memory access efficiency, which can accurately locate computing redundancy and resource waste points, and provide quantitative references for attention optimization in LLM training and inference stages.
Data for Agents
Hugging Face
Hello, you have only provided the title of this post “Data for Agents” so far, and the core content of the abstract is missing, so the requirements for translation and key point extraction cannot be fulfilled. Please supplement and upload the full English abstract of this paper, and I will highlight the core methods and conclusions as required to generate a concise summary of about 120 words.
The Gradient
After Orthogonality: Virtue-Ethical Agency and AI Alignment
The Gradient
This paper challenges the orthogonality hypothesis in the AI alignment field from the perspective of virtue ethics, proposing that human rational action is not oriented towards a fixed ultimate goal, but follows the logic of a practical network composed of actions, evaluation criteria, resources, etc. To achieve AI safety and compliance and effective collaboration with humans, it is necessary to abandon the idea of setting fixed goals for AI, and make its decision logic match the practical action logic of humans.
Lil’Log
Harness Engineering for Self-Improvement
Lilian Weng
The concept of Recursive Self-Improvement (RSI) was first proposed by scholar I. J. Good in 1965, referring to the core feature of superintelligent machines that can surpass all human intelligent activities and iteratively design better systems. In 2008, Eliezer Yudkowsky clearly defined it as the feedback loop where AI uses its existing intelligence to optimize its own cognitive architecture. In the current AI context, this loop includes both the model directly rewriting its own weights, and can broadly refer to the model optimizing its own training pipeline.
QbitAI
1998-born HIT Professor Founded Startup to Build Humanoid Dexterous Manipulation World Model
QbitAI
The team of Shuo Yang, a tenured professor at Harbin Institute of Technology (Shenzhen) born in 1998, has overcome three major technical nodes: tactile data collection, low-cost augmentation, and tactile world model, forming a complete chain of tactile-enabled embodied intelligence. The Poxiao Intelligence he founded starts from tactile data, builds full-stack capabilities from data, model to control, creates a world model for full-body dexterous manipulation of humanoid robots, and promotes robots to shift from “seeing the world” to “touching and operating the world”.
NVIDIA RTX Spark Real Machine Appears at Bilibili World! CPU and GPU are Directly Soldered Together, Laptop Runs 120B LLM
QbitAI
NVIDIA debuted the laptop equipped with the RTX Spark chip at Bilibili World. The chip adopts a design of Blackwell GPU + 20-core Grace CPU directly connected via NVLink-C2C, equipped with 1P computing power and 128GB unified memory, can run a 120B parameter LLM locally, supports a million-token context window, and is paired with OpenShell to ensure privacy, meeting the needs of gaming, AI creation and personal agent use.
Nearly 100 Players Enter Embodied Data: 4.47 Billion Yuan Financing in One Year, Who Can Really Make Money by “Selling Data”?
QbitAI
The large gap in embodied intelligence data has spurred a boom in data collection. According to QbitAI statistics, there are a total of 97 related players in China. In the past year, the total financing of independent service providers that only focus on data was 4.47 billion yuan, far lower than the financing amount of embodied model enterprises in the same period. There are various current data collection methods, even mobilizing ordinary users to participate in daily life, but the threshold for high-quality data collection is high. The mainstream technologies are divided into four categories, and the cross-route collection track is the most crowded.
Was this useful? A rating helps me pick the next topic.
Click a star to rate · Only anonymous fingerprint + timestamp stored