search_query=cat:astro-ph.*+AND+lastUpdatedDate:[202608282000+TO+202609032000]&start=0&max_results=5000

New astro-ph.* submissions cross listed on cs.LG, cs.AI, physics.data-an, stat.* staritng 202608282000 and ending 202609032000

Feed last updated: 2026-09-03T08:21:42Z

Accelerating Chemical Kinetics for Exoplanet Atmospheres using Neural Networks

Authors: Isaac Malsky, Xi Zhang, Tiffany Kataria, Matthew Graham, Ziyu Huang, Boris Bonev, Shang-Min Tsai, Elspeth K. H. Lee
Comments: 15 pages, 7 figures. Accepted for publication in ApJ
Primary Category: astro-ph.EP
All Categories: astro-ph.EP, cs.LG

Observations increasingly reveal the coupled radiative, chemical, and dynamical processes that shape exoplanet atmospheres. Interpreting these atmospheres requires models that can capture this complexity. However, multidimensional models remain fundamentally limited by computational cost, and answering key questions requires simulating the governing physical mechanisms at speeds classical methods cannot achieve. As a result, models often rely on simplifying approximations, such as equilibrium chemistry, even when those assumptions miss important effects. There is a pressing need for fast and accurate chemical kinetics solvers to model planetary atmospheres. Here we present a machine learning local-box chemical kinetics solver for exoplanet atmospheres using a residual flow-map architecture. We demonstrate that this surrogate model is several orders of magnitude faster than a classical solver, achieving microsecond-scale inference while retaining percent-level accuracy. The surrogate model covers a parameter space that spans $T=300$-$3000$ K, $P=10^{-6}$-$10^{4}$ bar, $Δt=10^{-3}$-$10^{8}$ s, and compositions ranging from $10^{-2}$ to $10^{3}$ times solar in both C/O ratio and metallicity. Our model outperforms several commonly used machine learning architectures and performs robustly under the extreme stiffness characteristic of atmospheric chemistry. The machine learning framework presented here is a flexible and efficient approach to emulating state-to-state flow-map problems that commonly arise in numerical simulations.


Cosmic variance and ergodicity in finite systems with correlations

Authors: Dipayan Mukherjee, Syksy Rasanen
Comments: 18+3 pages, 16 figures
Primary Category: astro-ph.CO
All Categories: astro-ph.CO, gr-qc, physics.data-an

We consider the difference between ensemble and volume average in cosmology. It is known that for sufficiently weak long-range correlations the root mean square of the difference, which we call ergodicity bias, decays like $R^{-3/2}$ in the limit of large volume $R^3$. We calculate the condition this imposes on the power spectrum of a Gaussian random field. We quantify the bias for finite $R$, and show that the $R\to\infty$ limit is of little relevance for cosmological observations when the measured scales and correlations extend to the size of the observable universe. We consider curvature, density, and velocity perturbations. On large scales the bias is important in all three cases. For the density perturbations, which are observationally the most relevant, the relative bias first exceeds 100% at the separation $r=177$ Mpc, and is larger than 100% for all $r>560$ Mpc. It should be taken into account when comparing ensemble and volume averages for large-scale structure. The bias is also large for the cosmic microwave background temperature perturbations on large angular scales, but this is not relevant for observations, as their analysis does not involve volume averaging.


Generative Diffusion Surrogates with Analytical Variance Schedule

Authors: Patrick Reichherzer, Gianluca Gregori, David N. Hosking, Subir Sarkar
Comments: Accepted for publication in Nature Communications
Primary Category: cs.LG
All Categories: cs.LG, astro-ph.IM, physics.plasm-ph

Stochastic transport describes physical systems in which an initially structured distribution spreads under unresolved forcing, scattering, or heterogeneous media. Useful surrogates for such systems should be probabilistic, time-resolved, and able to represent non-Gaussian distributional structure. Generative diffusion models, which corrupt data with Gaussian noise and learn a reverse flow back to structured states, have these properties. Their noise schedules, however, are usually chosen heuristically: image and audio generation---the canonical use cases---provide no physical clock. In transport, by contrast, the variance, or mean-square displacement, is often known from macroscopic theory or empirical scaling even when the full distribution is not. Here we prescribe the forward noising rate as the time derivative of this variance, turning generative time into a calibrated transport clock. The variance path is enforced by construction, while the learned score field represents how non-Gaussian structure inherited from entrance data is smoothed along that path, requiring no intermediate-time physical transport data. For ballistic-to-diffusive transport in turbulent plasmas, the surrogate matches test-particle distributions, reproduces the laboratory-measured variance scale, and tracks the simulated kurtosis evolution without schedule tuning, enabling calibrated emulation and likelihood-based inference.


Efficient Context-Limited Telescope Bibliography Classification for the WASP-2025 Shared Task Using SciBERT

Authors: Madhusudhana Naidu
Comments: 3 pages, 2 tables. 1st place system description for the TRACS shared task at WASP 2025 (Third Workshop for Artificial Intelligence for Scientific Publications), co-located with IJCNLP-AACL 2025. Published version: https://aclanthology.org/2025.wasp-main.21/ . Code: https://github.com/E0NIA/TRACS-WASP-2025-1st-Place
Primary Category: cs.LG
All Categories: cs.LG, astro-ph.IM

The creation of telescope bibliographies is a crucial part of assessing the scientific impact of observatories and ensuring reproducibility in astronomy. This task involves identifying, categorizing, and linking scientific publications that reference or use specific telescopes. However, this process remains largely manual and resource intensive. In this work, we present an efficient SciBERT-based approach for automatic classification of scientific papers into four categories - science, instrumentation, mention, and not telescope. Despite strict context-length constraints (maximum 512 tokens) and limited compute resources, our approach achieved a macro F1 score of 0.89, ranking at the top of the WASP-2025 leaderboard. We analyze the effect of truncation and show that even with half the samples exceeding the token limit, SciBERT's domain alignment enables robust classification. We discuss trade-offs between truncation, chunking, and long-context models, providing insights into the efficiency frontier for scientific text curation.


A generalized likelihood model for segmented muon counters

Authors: Joaquín de Jesús, Juan Manuel Figueira, Federico Sanchez, Darko Veberic
Comments: 16 pages, 10 figures
Primary Category: astro-ph.IM
All Categories: astro-ph.IM, physics.data-an

Measurements of the muonic component of extensive air showers constrain cosmic-ray mass composition and hadronic interactions at energies beyond those accessible at accelerators. Arrays of segmented detectors with binary readout are widely used for this purpose: they sample the muon density at different distances from the shower core to reconstruct the muon lateral distribution function (LDF). Each detector response is summarized by the number of activated segments, $k$, whose probability distribution provides the likelihood relating the observation to the expected muon content. Signal pile-up, detector inefficiency, corner-clipping muons, and background signals shape this distribution, and neglecting them can bias the reconstruction. Existing analytical models include pile-up but otherwise assume an ideal detector response. In this work, we develop a unified statistical framework that incorporates inefficiency, corner clipping, and background through a small set of physically interpretable parameters. We derive exact expressions for the detector response and the likelihood required for muon-LDF reconstruction, together with a simple binomial approximation that preserves the main statistical properties of the exact distribution. Dedicated Monte Carlo simulations are used to assess the impact of the assumptions underlying the analytical treatment and show that it is negligible over the parameter range considered. They also show that the exact and approximate likelihoods yield similar performance in terms of estimator bias and confidence-interval coverage. Although motivated by the Underground Muon Detector of the Pierre Auger Observatory, the framework applies more broadly to segmented particle detectors with binary readout in which particle content is inferred from the number of activated segments.


Do Tabular Foundation Models Know Physics? Contamination, Units, and the Deterministic Limit

Authors: Wassim Tenachi, Yashar Hezaveh, Laurence Perreault Levasseur, Pierre-Luc Bacon
Comments: 5 pages (4 figures, 1 table). Submitted to the Representations for the Physical Sciences Workshop @ NeurIPS 2026
Primary Category: cs.LG
All Categories: cs.LG, astro-ph.IM

Tabular foundation models (TFMs) learn to fill in tables the way language models fill in text, and tables are arguably the format in which most physical measurement arrives. Did they learn any physics in the process? They are Bayesian by construction, so the question is what their prior contains. We probe it directly, evaluating four of them (TabPFN-3, TabICLv2, TabDPT and Real-TabPFN-2.5) against six baselines on datasets sampled from 316 physical equations, in and out of domain. TFMs dominate, out of the box and after tuning. But we show that their prior can represent neither a noiseless mechanism nor physical units, which is why they interpolate physics without yet being able to act as physical models.