Daily AI Highlights · 2026-07-22
20 Papers · Multi-source Aggregation + AI Summaries
- Leading overseas AI vendors announced major updates in a concentrated period: OpenAI launched services for small businesses, DeepMind released new models in the Gemini 3.6 series
- Multiple new achievements emerged in the AI academic field, covering directions including world modeling, diffusion control, multi-objective optimization, and embodied simulation
- China’s AI sector delivered outstanding progress: Xiaohongshu’s large model won a gold medal at IMO, and heavyweight achievements including intelligent computing and agents were released at WAIC
Hugging Face Daily Papers
Masked Visual Actions for Unified World Modeling
HF ★ 0 · Hadi Alzayer, Wenlong Huang, Haonan Chen… · HF Mirror
To address the pain point that action representation and visual pre-training space are hard to align in robotic world modeling, this research proposes a pixel-level control interface called “Masked Visual Actions”, which represents actions as partially visible trajectories of any entity in a video. A single model fine-tuned with only 15 hours of real and simulated masked samples can handle both forward dynamics prediction and inverse kinematics solving, and can effectively assist policy evaluation and planning decision-making in downstream manipulation tasks.
Appearance Pointers — Multimodal Region Control of Diffusion Transformers
HF ★ 0 · Rahul Sajnani, Yulia Gryaditskaya, Radomír Měch… · HF Mirror
To solve the problem of lack of precise regional control in controllable image generation of Diffusion Transformers (DiT), this paper proposes an appearance pointer mechanism: it generates compact representations through a region correspondence network, optimizes via spatial aggregation, aligns text/image inputs with user-specified masks, and does not require retraining the base DiT. It is the first modality-agnostic local multimodal control interface for DiT, with the performance of a single model matching or even outperforming modality-specific SOTA models, enabling precise region-controllable generation.
H^2SD: Hybrid Hindsight Self-Distillation
HF ★ 0 · Qiye Cai, Yichuan Ma, Linyang Li… · HF Mirror
To address the issues of sparse supervision in verifiable reward reinforcement learning, and unstable optimization or lack of clear error correction direction in existing self-distillation solutions for large model reasoning tasks such as mathematical reasoning and code generation, this paper proposes the Hybrid Hindsight Self-Distillation framework H²SD: for successful trajectories, it adjusts the update magnitude with teacher signals; for failed trajectories, it aligns with the teacher distribution by referring to key reasoning prompts. Experiments show that its performance on multiple reasoning benchmarks is better than existing baselines, with stable optimization and excellent generation efficiency.
ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning
HF ★ 0 · Ting Huang, Zhenyu Zhang, Wenyuan Huang… · HF Mirror
To address the problems that existing multimodal large models are semantically oriented, making it difficult to integrate consistent spatial evidence and resulting in poor reasoning stability during long-sequence cross-view video spatial reasoning, this paper proposes the ConsiSpace geometric consistency perception framework: it builds geometric consistency memory to efficiently store task-related spatial evidence, and designs unified consistency self-supervised reinforcement learning as a post-training signal. In three types of spatial reasoning benchmark tests, the average score is 12.6 points higher than the optimal baseline.
arXiv cs.LG
Reinforcement Learning-Guided NSGA-II Enhanced with Gray Relational Coefficient for Multi-Objective Optimization: Application to NASDAQ Portfolio Optimization
Zhiyuan Wang, Qinxu Ding, Ding Ding…
This paper targets the multi-objective portfolio optimization problem, and proposes the improved NSGA-II algorithm RL-NSGA-II-GRC, which integrates reinforcement learning online adaptive parameter tuning and grey relational coefficient selection operator. Benchmark tests show that its convergence is 4.4%~5.8% higher than the original NSGA-II. When applied to NASDAQ portfolio optimization, it can obtain a high-density effective frontier, with a maximum annualized Sharpe ratio of 1.92, and can output optimal solutions adapted to different risk preferences.
DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth
Zihan Xu, Puzhen Wu, Lawrence Chun Man Lau…
To address the pain point of difficult selection among numerous OCR and multimodal large model document parsing tools in scenarios with scarce annotations, this research proposes the annotation-free evaluation framework DocOCR-Eval, which adopts a three-stage correction ranking strategy. It can approximate the tool ranking results of manual annotation without ground truth, and aggregating multimodal large models can further improve the matching degree. Experiments verify that it can reliably complete OCR tool selection in annotation-deficient scenarios, providing guidance for system implementation.
Fully-sensorized smart-eyewear platform for on-device Machine Learning
Andrea Giudici, Christian Veronesi, Pietro Bartoli…
This paper proposes the ARGO smart eyewear platform for on-device machine learning, adopting co-design of software, hardware and AI, equipped with the STM32N6 chip integrated with NPU. It proposes a head-parallel attention mechanism to optimize YOLOv11 for hardware adaptation, processes data locally without the cloud to ensure privacy. It achieves a mAP50-95 of 24 for urban obstacle recognition, only occupies 2.483MB of memory, and has a 113-minute battery life with a 200mAh battery at 10FPS, verifying the feasibility of high-performance privacy-friendly wearable AI devices.
OpenAI
Introducing the ChatGPT for small business program
OpenAI
OpenAI officially launched the ChatGPT for small business program, relying on the capabilities of ChatGPT Work to provide AI skill empowerment services for small and medium-sized entrepreneurs, helping them master the implementation methods of AI tools, realize the automation of daily operation, transaction processing and other links, greatly lower the threshold of AI use for small and micro enterprises in digital transformation, and finally achieve the goals of improving efficiency, increasing revenue and driving long-term business growth.
OpenAI and Hugging Face partner to address security incident during model evaluation
OpenAI
Recently, OpenAI and Hugging Face jointly handled a security incident that occurred during the AI model evaluation phase, and simultaneously disclosed the initial investigation conclusions of the incident. This incident reveals that attackers already have high-level targeted cyber attack capabilities targeting the AI R&D process. The incident details and attack and defense experience summarized by the two parties can provide targeted security optimization references for AI system defenders across the industry, and fill the security gaps in the entire model lifecycle.
Anthropic News
Inviting hard questions
Anthropic
This is a public interaction initiative in the AI field. The initiators openly solicit difficult AI-related questions of public concern from the whole society, and make a clear commitment: when responding to these collected questions one by one in the future, they will fully disclose the entire process of relevant research, demonstration and problem-solving, so as to improve the transparency of AI R&D and respond to public concerns about the development of AI technology.
Redeploying Claude Fable 5
Anthropic
AI vendor Anthropic announced that it will restart the deployment of the Claude Fable 5 large model starting July 1 after the official lifting of export controls. This deployment is accompanied by two security upgrades: an updated full-link cybersecurity protection system, and a new industry-grade anti-jailbreak framework, which can effectively reduce the risk of malicious cracking and illegal abuse of the model, balancing compliance requirements and usage security.
Google DeepMind
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Google DeepMind
Google launched three new models in the Gemini series this time: the general-purpose lightweight flagship Gemini 3.6 Flash, the 3.5 Flash-Lite focused on ultra-low power consumption and low latency suitable for edge deployment, and the 3.5 Flash Cyber customized for cybersecurity vertical scenarios. The three new products complete the layout of the Gemini Flash product line, covering different needs such as high-performance general reasoning, edge implementation, and industry-specific scenarios, adapting to diverse computing power conditions.
Introducing Gemini 3.5 Flash Cyber
Google DeepMind
Google has newly launched the lightweight cybersecurity dedicated model Gemini 3.5 Flash Cyber, which is a vertical scenario customized version of the Gemini 3.5 Flash series. It mainly serves cybersecurity operation and maintenance needs, can automatically identify various system vulnerabilities, and supports generating suitable patch solutions for remediation, which can significantly reduce the workload of cybersecurity practitioners in vulnerability management.
Hugging Face Blog
The State of Simulation for Physical AI: An Overview
Hugging Face
This review “The State of Simulation for Physical AI” sorts out the development context of physical AI simulation technology, compares the adaptation capabilities of existing simulation tools for multi-physics scenarios such as rigid bodies, fluids, and flexible bodies, analyzes the two core bottlenecks of large reality-virtual domain gap and high computational cost, summarizes cutting-edge progress such as differential simulation and reality-virtual closed-loop calibration, and points out that a high-fidelity low-cost general simulation framework is the core path to support the implementation of embodied intelligence and physical robots in the future.
Grabette: an open system to record robot-manipulation data
Hugging Face
This paper introduces Grabette, an open-source system for recording robot manipulation data. The system supports low-cost construction, is compatible with various devices such as robotic arms, RGBD cameras, and force-tactile sensors, realizes synchronous collection and automatic annotation of multimodal manipulation data, adapts to grasping and dexterous manipulation scenarios, and the generated datasets can be directly used for training robot manipulation policies, solving the pain points of existing similar systems being closed-source, poor adaptability, and high cost.
The Gradient
After Orthogonality: Virtue-Ethical Agency and AI Alignment
The Gradient
This paper conducts AI alignment research from the perspective of virtue ethics, refutes the presupposition that “rational agents must be anchored to end goals”, and points out that human rational actions are not directed at fixed goals, but adapt to a practical network including elements such as action norms and evaluation standards. The paper proposes that to achieve smooth collaboration and alignment between AI and humans, the AI decision-making logic needs to match the “type signature” of human practice-based actions, and this path can cover both ethical alignment and core security requirements.
Lil’Log
Harness Engineering for Self-Improvement
Lilian Weng
The concept of Recursive Self-Improvement (RSI) was first proposed by I.J. Good in 1965, referring to superintelligent machines that can surpass all human intellectual activities and can independently design better systems to achieve self-iteration. In 2008, Eliezer Yudkowsky clarified that its core is the feedback loop where AI uses existing intelligence to optimize its own cognitive mechanism. Currently, such feedback for AI can be manifested as directly rewriting its own weights, or optimizing its own training process.
量子位 (QbitAI)
Cowar Technology debuts at WAIC 2026, unveils the industry’s first dual-layer agent world model
量子位 (QbitAI)
During WAIC 2026, Cowar Technology released the industry’s first full-stack self-developed COOWAM dual-layer agent world model. On-site demonstration showed that the robotic arm driven by this model can autonomously complete multi-step complex drink preparation tasks in dynamic environments, and its first heavy-duty quadruped robot dog X0 also made its offline debut. Based on this model, Cowar is promoting urban embodied intelligence to shift from single-unit operation to multi-agent collaboration, accelerating its implementation in urban scenarios such as public services and logistics distribution.
Xiaohongshu’s large model wins full-score gold medal at IMO, its solution to the third problem is praised as elegant by champion contestants
量子位 (QbitAI)
Xiaohongshu’s self-developed large model dots-note-3.0 won a full-score gold medal at the 2026 International Mathematical Olympiad (IMO), making it the first large model in China and the second in the world to receive official IMO gold medal certification, outperforming Google’s Gemini which previously only answered 5 out of 6 questions correctly. The model solves problems entirely in natural language, and the unconventional solutions provided for some questions are concise, elegant and hit the essence, winning high recognition from multiple CMO gold medalists.
WAIC heavyweight achievement | Shanghai INESA Intelligent Computing leads the establishment of “Intelligent Computing System Architecture Alliance” and releases “Super Node System Architecture Specification”
量子位 (QbitAI)
At WAIC 2026, under the guidance of the Shanghai Municipal Commission of Economy and Informatization, Shanghai INESA Intelligent Computing led 23 upstream and downstream entities in the intelligent computing industry chain to officially establish the Intelligent Computing System Architecture Alliance and release the “Super Node System Architecture Specification”. The alliance is operated by the INESA-ZTE Joint Laboratory, which will aggregate full-stack ecological resources, solve pain points such as imbalance between supply and demand of computing power and software and hardware adaptation barriers, and promote the construction of an independent and controllable intelligent computing base.
Was this useful? A rating helps me pick the next topic.
Click a star to rate · Only anonymous fingerprint + timestamp stored