The introduction is organized in a unique didactic manner developed by the authors, starting from more simple concepts such as linear programming and single-point methods, and advancing from these to more difficult concepts such as optimality conditions for nonlinear optimization and set-oriented solution algorithms.
The goal of the OpenThoughts project is to create open-source datasets for training reasoning models and to create the first model trained on public reasoning data to match DeepSeek-R1-Distill-Qwen-7B on standard reasoning benchmarks such as AIME and LiveCodeBench.
For production use cases, a new distillation engine is introduced that converts TabPFN-2.5 into a compact MLP or tree ensemble, preserving most of its accuracy while delivering orders-of-magnitude lower latency and plug-and-play deployment.
The experimental results demonstrate the efficacy of RG in both evaluating and reinforcement learning of reasoning models, and its key innovation is the ability to generate virtually infinite training data with adjustable complexity, unlike most previous reasoning datasets, which are typically fixed.
AlphaProof is presented, an AlphaZero-inspired2 agent that learns to find formal proofs through RL by training on millions of auto-formalized problems, and substantially improves state-of-the-art results on historical mathematics competition problems.
A framework to evaluate sycophantic behavior in ChatGPT-4o, Claude-Sonnet, and Gemini-1.5-Pro across AMPS (mathematics) and MedQuad (medical advice) datasets is introduced.
This study studies layer pruning via parameter-efficient finetuning methods, specifically quantization and Low Rank Adapters (QLoRA), such that each of the experiments can be performed on a single 40GB A100 GPU.
Nemotron-Cascade 2 is introduced, an open 30B MoE model with 3B activated parameters that delivers best-in-class reasoning and strong agentic capabilities that aims to achieve Gold Medal-level performance in the 2025 International Mathematical Olympiad, the International Olympiad in Informatics (IOI), and the ICPC World Finals.
FGN is presented, a simple, scalable and flexible modeling approach which significantly outperforms the current state-of-the-art models and produces state-of-the-art ensemble forecasts as measured by a range of deterministic and probabilistic metrics.
A new LMO-based method called $\sf Gluon$ is proposed, capturing prior theoretically analyzed methods as special cases, and a new refined generalized smoothness model is introduced that captures the layer-wise geometry of neural networks, matches the layer-wise practical implementation of $\sf Muon$ and $\sf Scion$, and leads to convergence guarantees with strong practical predictive power.
This work identifies systematic issues that have resulted in a distorted playing field in Chatbot Arena and offers actionable recommendations to reform the Chatbot Arena's evaluation framework and promote fairer, more transparent benchmarking for the field.
This study conducts a systematic literature review (SLR) to investigate the applications and trends of AI in mathematics education by examining articles published in reputable journals indexed in Web of Science and Scopus.
It is shown that accurate unconstrained models can be applied with confidence, especially since simple inference-time modifications can be used to recover observables that are consistent with the relevant physical symmetries when compared to physically constrained models.
Meta Flow Maps are introduced, a framework extending consistency models and flow maps into the stochastic regime that helps solve bottlenecks in both paradigms: enabling inference-time steering without inner rollouts, and facilitating unbiased, off-policy fine-tuning to general rewards.
An extensive systematic and bibliometric literature review on hybrid methods involving optimization and machine learning techniques for clustering and classification aims to identify the potential of methods and algorithms to overcome the difficulties of one or both methodologies when combined.
PutnamBench is presented, a new multi-language benchmark for evaluating the ability of neural theorem-provers to solve competition mathematics problems, which requires significant problem-solving ability and proficiency in a broad range of topics taught in undergraduate mathematics courses.
The scaling laws of CMs under ECT are investigated, showing that they obey the classic power law scaling, hinting at their ability to improve efficiency and performance at larger scales.
SciCode demonstrates both contemporary LMs' progress towards becoming helpful scientific assistants and sheds light on the development and evaluation of scientific AI in the future.
This work presents soft inductive biases as a key unifying principle in explaining these phenomena: rather than restricting the hypothesis space to avoid overfitting, embrace a flexible hypothesis space, with a soft preference for simpler solutions that are consistent with the data.
The sudden cessation of PEPFAR funding likely results in tens of thousands of HIV deaths and new infections, and should compel the United States government to rapidly and fully re-instate one of the most successful health programs in history.
It is argued that a simple shift to training value functions with categorical cross-entropy can yield substantial improvements in the scalability of deep RL at little-to-no cost.
The successful application of SciML to the simulation of the human cardiac function, a field of significant socioeconomic importance that poses numerous challenges on both the mathematical and computational fronts.
This paper theoretically characterize the behavior of AUROC and AUPRC in the presence of model mistakes, establishing clearly that AUPRC is not generally superior in cases of class imbalance and shows that AUPRC can be a harmful metric.
Analyzing reasoning chain length across o1-mini and o3-mini variants on the Omni-MATH benchmark finds that o3-mini (m) achieves superior accuracy without requiring longer reasoning chains than o1-mini, and highlights that while o3-mini (h) achieves a marginal accuracy gain over o3-mini (m), it does so by allocating substantially more reasoning tokens across all problems, even the ones that o3-mini (m) can already solve.
This review delves into the GWO-related research conducted between 2019 and 2022, encompassing over 200 research articles and explores the growth of GWO in terms of publications, citations, and the domains that leverage its potential.
This tutorial provides an overview of new-field NGAT technology, which shifts from conventional far-field channel models to new near-field channel models, and discusses recent advances in semantic-aware NGAT technologies, which can utilize new metrics for advanced transceiver designs.
A systematic, fair and comparable benchmarking framework for quantum optimization methods by presenting ten model-independent problem classes that are challenging for classical methods, which enables fair, reproducible benchmarks of quantum heuristics for ten difficult combinatorial optimization classes with baseline results to track progress towards quantum advantage.
It is shown that if data are replaced, the test error increases with the number of model-fitting iterations, but if data instead accumulate, the test error has a finite upper bound independent of the number of iterations, meaning model collapse no longer occurs.
This review explores the critical role of hyperparameter tuning in ML, detailing its importance, applications, and various optimization techniques, and various tuning methods, including grid search, random search, Bayesian optimization, and meta-learning.
It is proved that LeJEPA (alignment plus Gaussian regularization) linearly recovers the world's latent variables from nonlinear observations, a property known as linear identifiability, in a broad class of worlds where latents evolve under stationary, additive-noise transitions, and the Gaussian is the unique latent distribution for which this guarantee holds.
The first large-scale evaluation of AI-aided formal proof search's ability to solve open problems demonstrates the power of AI-aided formal proof search and sheds light on the agent designs that enable it.
The results show that agents exceed human SOTA in four tasks but fail to match it in sixteen others, indicating that AIRS-Bench is far from saturated and offers substantial room for improvement.
This letter addresses, for the first time, the uplink performance optimization of multi-user pinching-antenna (PA) systems, recently developed for next-generation wireless networks, and proposes an effective approach that separately optimizes the positions of the PAs and the resource allocation.
LLM-SR is introduced, a novel approach that leverages the extensive scientific knowledge and robust code generation capabilities of Large Language Models to discover scientific equations from data that significantly outperform state-of-the-art symbolic regression baselines, particularly in out-of-domain test settings.
The R package sdmTMB is introduced, which extends the flexible interface familiar to users of lme4, glmmTMB, and mgcv to include spatial and spatiotemporal latent GMRFs using an SPDE-(stochastic partial differential equation) based approach and is hoped to help open this useful class of models to a wider field of geostatistical analysts.
It is argued that without regularization, grokking tasks push models to the edge of numerical stability, introducing floating point errors in the Softmax function, which the paper refers to as Softmax Collapse (SC), and that mitigating SC enables grokking without regularization.
Random forests learn local probability peaks that often yield near perfect training AUCs without strongly affecting AUCs on test data, which goes against the common recommendation to use fully grown trees in random forest models.
A review of the current literature on generative AI in mathematics education, focusing on four areas: generative AI for mathematics problem‐solving, generative AI for mathematics tutoring and feedback, generative AI to adapt mathematical tasks, and generative AI to assist mathematics teachers in planning.
An analysis of the dynamical mechanism underlying memorization is presented, highlighting the need for regularization to avoid reproducing the analytically tractable minimizer; and laying the foundations for a principled understanding of how to regularize.
The results show that the Silhouette coefficient and the Davies-Bouldin index are more informative and reliable than the other analyzed rates, when assessing convex-shaped and non-nested clusters in the Euclidean space.
This work demonstrates physics-based machine learning as a promising technique for efficient and reliable fatigue life prediction in engineering structures, with possible integration into digital twin models for real-time assessment.
Article Galaxy Pages is a free service from Research Solutions, a company that offers access to content in collaboration with publishing partners, online repositories and discovery services.