AI Daily Highlights · 2026-08-08
13 papers · multi-source aggregation + AI summaries
arXiv cs.LG
MS-MLB: An Open Machine Learning Benchmark for Blood-Based MS Classification
Adam Simson, Ankush Dutta, Quang Bui
This paper introduces MS-MLB, an open machine learning benchmark for multiple sclerosis (MS) classification based on whole blood RNA expression data. Built on the public dataset GSE17048, it designs a classification task to distinguish MS patients from healthy populations, adopts a leak-proof unified evaluation pipeline, and covers multi-dimensional validation metrics. In empirical tests, the gradient boosting algorithm performs the best with an AUC of 0.989. As the first reproducible open benchmark in this scenario, it is only applicable for scientific research comparison and has not completed clinical validation.
When Do Corrective Features Help? An Agent for Corrective Feature Discovery on Black-Box Forecasters
Fangxin Wang, Ziyi Zhang, Diyi Zhuang…
To address the issues that frozen pretrained black-box time series forecasters often have structured repeated errors and high repair costs via fine-tuning, this paper proposes the CRAFTER agent: it keeps the forecasting backbone frozen, mines interpretable corrective features from prediction residuals, combines two types of candidates from raw input combination search and LLM generation, and uses a lightweight post-corrector to adjust predictions after empirical verification. Multiple tests show its effect far exceeds existing feature engineering systems, reducing the error of the weakest backbone by up to 27% with stable gains.
PPDL: LLM-Based Flows as Probabilistic Programs
Louis Mandel, Guillaume Baudart, Mandana Vaziri…
To solve the pain point that LLM application outputs lack confidence scores and the superposition of uncertainties during multi-step calls makes results untrustworthy, this paper proposes PPDL, a probabilistic programming language for LLM workflows. Developers only need to write core process logic, and can realize full-link uncertainty quantification and transmission without additional coding, and flexibly test different inference scaling strategies. Experiments have verified the capability of the solution, which has been deployed to build a dedicated proof agent for the Rocq theorem prover.
OpenAI
Responding to the next frontier of critical cyber capabilities
OpenAI
This report released by OpenAI addresses the frontier development needs of next-generation critical cyber capabilities and focuses on related risk response. Its core content includes two parts: first, it publishes the preliminary cybersecurity performance evaluation results of its model Astra; second, it discloses the specific measures it is currently taking to strengthen security protection mechanisms and improve the security management and control system, providing vendor practice references for the secure and compliant implementation of AI cyber capabilities.
How HSP GRUPPE builds AI capabilities for tax advisory
OpenAI
This article focuses on the topic of AI capability building in the tax consulting field, and introduces the practice plan of the organization HSP GRUPPE: by deploying ChatGPT Enterprise adapted to the entire process of tax consulting business, it has effectively improved business processing efficiency and professional service quality, and the released human resources can be allocated to high-value-added customized customer services, providing a reference for the digital and intelligent transformation of the same industry.
Anthropic News
Introducing Claude Opus 5
Anthropic
The newly released Claude Opus 5 is a stepwise iteration product of Anthropic’s high-end Opus LLM product line. The core of this upgrade focuses on capability improvement in three core scenarios: first, it greatly strengthens the underlying support capability for long-running agents; second, it significantly improves coding task processing performance; third, it optimizes work efficiency in various professional scenarios, which can better support high-complexity production-level requirements.
Inviting hard questions
Anthropic
This is a public interaction project launched by the AI research team, with two core arrangements: first, it openly collects the most challenging difficult questions about AI from the whole society; second, it publicly promises that in the whole process of answering all subsequent questions, it will fully disclose research work details and ensure research transparency, so as to promote AI research to directly address public concerns.
Google DeepMind
WeatherNext: AI model achieves breakthrough in forecasting cyclones
Google DeepMind
The newly launched AI weather forecasting model WeatherNext has achieved a key breakthrough in the field of cyclone forecasting. Trained on multi-source historical meteorological observations and reanalysis data, it adopts a spatio-temporal attention architecture to fit the law of atmospheric evolution, and its computing power demand is far lower than that of traditional numerical forecasting. Actual measurements show that it can accurately predict cyclone generation, path and intensity 1-2 weeks in advance, with an accuracy rate of more than 25% higher than traditional methods, which can leave a sufficient response window for disaster early warning.
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
Google DeepMind
The newly released Gemini Robotics ER 2 is a dedicated capability support system for robotics scenarios, achieving three core technological leaps: greatly upgraded video understanding capability, which can support robots to complete environmental perception and inference; built-in task orchestration tools, which can decompose and implement real-world scenario tasks; support for multi-robot collaborative operation. This system can effectively empower robots to complete complex real-world tasks, providing key technical support for the implementation of robotics applications.
Hugging Face Blog
TutorMoments: Do AI tutors know when to help and when to hold back?
Hugging Face
Currently only the paper title is provided, and the specific English text of the abstract has not been pasted, so the translation and refinement work cannot be completed. Please supplement the full original text of the paper abstract, and I will refine the key points of about 120 words as required, clearly highlighting its core research methods and final conclusions.
Baseten on Hugging Face Inference Providers 🔥
Hugging Face
Note: Since the abstract text is not provided, the following is summarized based on the public information of this official announcement: Baseten has officially joined the Hugging Face inference provider ecosystem. Users can directly call all types of open-source LLMs and multi-modal models hosted by it on the HF platform. Compared with similar general deployment solutions, the inference latency is reduced by up to 40%, deployment costs are cut by 30%, and custom scaling is also supported, which greatly lowers the threshold for AI developers to deploy models.
The Gradient
After Orthogonality: Virtue-Ethical Agency and AI Alignment
The Gradient
This AI alignment research is based on the perspective of virtue-ethical agency, refutes the default assumption that “rational agents need to take fixed ultimate goals as action guidance”, points out that the rationality of human behavior comes from adapting to the practice network including action rules and evaluation systems, and proposes that AI alignment requires the decision-making logic of agents to be isomorphic with this kind of human practice logic. This path can not only align with ethical requirements such as human well-being, but also guarantee core security attributes.
Lil’Log
Harness Engineering for Self-Improvement
Lilian Weng
The concept of Recursive Self-Improvement (RSI) was first proposed by I.J. Good in 1965, referring to a superintelligent system that can surpass all human intellectual activities and design better machines by itself to achieve iterative upgrades. In 2008, Yudkowsky clarified that its core is the feedback loop where AI optimizes its own cognitive mechanism relying on existing intelligence. Currently, this mechanism in the AI field can be reflected as the model directly rewriting its own weights, or the model optimizing its own training process.
Was this useful? A rating helps me pick the next topic.
Click a star to rate · Only anonymous fingerprint + timestamp stored