RLJ Volume 6 (2025) Published as part of RLJ 2025: Volume 6, DOI: 10.5281/zenodo.21537949 .
Reinforcement Learning for Finite Space Mean-Field Type Game Kai Shao, Jiacheng Shen, Mathieu Lauriere ; 6:1−39, 2025. [abs ][pdf ][bib ]Understanding Behavioral Metric Learning: A Large-Scale Study on Distracting Reinforcement Learning Environments Ziyan Luo, Tianwei Ni, Pierre-Luc Bacon, Doina Precup, Xujie Si ; 6:40−96, 2025. [abs ][pdf ][bib ]Which Experiences Are Influential for RL Agents? Efficiently Estimating The Influence of Experiences Takuya Hiraoka, Takashi Onishi, Guanquan Wang, Yoshimasa Tsuruoka ; 6:97−138, 2025. [abs ][pdf ][bib ][supp ]Online Intrinsic Rewards for Decision Making Agents from Large Language Model Feedback Qinqing Zheng, Mikael Henaff, Amy Zhang, Aditya Grover, Brandon Amos ; 6:139−164, 2025. [abs ][pdf ][bib ]A Finite-Time Analysis of Distributed Q-Learning Han-Dong Lim, Donghwan Lee ; 6:165−200, 2025. [abs ][pdf ][bib ]Finite-Time Analysis of Minimax Q-Learning Narim Jeong, Donghwan Lee ; 6:201−230, 2025. [abs ][pdf ][bib ]Collaboration Promotes Group Resilience in Multi-Agent RL Ilai Shraga, Guy Azran, Matthias Gerstgrasser, Ofir Abu, Jeffrey Rosenschein, Sarah Keren ; 6:231−243, 2025. [abs ][pdf ][bib ]Bayesian Meta-Reinforcement Learning with Laplace Variational Recurrent Networks Joery A. de Vries, Jinke He, Mathijs de Weerdt, Matthijs T. J. Spaan ; 6:244−275, 2025. [abs ][pdf ][bib ]Foundation Model Self-Play: Open-Ended Strategy Innovation via Foundation Models Aaron Dharna, Cong Lu, Jeff Clune ; 6:276−342, 2025. [abs ][pdf ][bib ]Action Mapping for Reinforcement Learning in Continuous Environments with Constraints Mirco Theile, Lukas Dirnberger, Raphael Trumpp, Marco Caccamo, Alberto Sangiovanni-Vincentelli ; 6:343−363, 2025. [abs ][pdf ][bib ]Chargax: A JAX Accelerated EV Charging Simulator Koen Ponse, Jan Felix Kleuker, Thomas M. Moerland, Aske Plaat ; 6:364−383, 2025. [abs ][pdf ][bib ]Effect of a slowdown correlated to the current state of the environment on an asynchronous learning architecture Idriss Abdallah, Laurent CIARLETTA, Patrick HENAFF, Jonathan Champagne, Matthieu BONAVENT ; 6:384−398, 2025. [abs ][pdf ][bib ]Cascade - A sequential ensemble method for continuous control tasks Robin Schmöcker, Alexander Dockhorn ; 6:399−411, 2025. [abs ][pdf ][bib ][supp ]Average-Reward Soft Actor-Critic Jacob Adamczyk, Volodymyr Makarenko, Stas Tiomkin, Rahul V Kulkarni ; 6:412−430, 2025. [abs ][pdf ][bib ]Burning RED: Unlocking Subtask-Driven Reinforcement Learning and Risk-Awareness in Average-Reward Markov Decision Processes Juan Sebastian Rojas, Chi-Guhn Lee ; 6:431−477, 2025. [abs ][pdf ][bib ]Your Learned Constraint is Secretly a Backward Reachable Tube Mohamad Qadri, Gokul Swamy, Jonathan Francis, Michael Kaess, Andrea Bajcsy ; 6:478−492, 2025. [abs ][pdf ][bib ]Improved Regret Bound for Safe Reinforcement Learning via Tighter Cost Pessimism and Reward Optimism Kihyun Yu, Duksang Lee, William Overman, Dabeen Lee ; 6:493−546, 2025. [abs ][pdf ][bib ][supp ]Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Jiaqi Chen, Ji Shi, Cansu Sancaktar, Jonas Frey, Georg Martius ; 6:547−584, 2025. [abs ][pdf ][bib ]Uncertainty Prioritized Experience Replay Rodrigo Antonio Carrasco-Davis, Sebastian Lee, Claudia Clopath, Will Dabney ; 6:585−623, 2025. [abs ][pdf ][bib ]RL$^3$: Boosting Meta Reinforcement Learning via RL inside RL$^2$ Abhinav Bhatia, Samer B. Nashed, Shlomo Zilberstein ; 6:624−649, 2025. [abs ][pdf ][bib ]Pareto Optimal Learning from Preferences with Hidden Context Ryan Bahlous-Boldi, Li Ding, Lee Spector, Scott Niekum ; 6:650−670, 2025. [abs ][pdf ][bib ]WOFOSTGym: A Crop Simulator for Learning Annual and Perennial Crop Management Strategies William Solow, Sandhya Saisubramanian, Alan Fern ; 6:671−689, 2025. [abs ][pdf ][bib ][supp ]When and Why Hyperbolic Discounting Matters for Reinforcement Learning Interventions Ian M. Moore, Eura Nofshin, Siddharth Swaroop, Susan Murphy, Finale Doshi-Velez, Weiwei Pan ; 6:690−712, 2025. [abs ][pdf ][bib ]Reinforcement Learning from Human Feedback with High-Confidence Safety Guarantees Yaswanth Chittepu, Blossom Metevier, Will Schwarzer, Austin Hoag, Scott Niekum, Philip S. Thomas ; 6:713−736, 2025. [abs ][pdf ][bib ]AVID: Adapting Video Diffusion Models to World Models Marc Rigter, Tarun Gupta, Agrin Hilmkil, Chao Ma ; 6:737−764, 2025. [abs ][pdf ][bib ]Non-Stationary Latent Auto-Regressive Bandits Anna L. Trella, Walter H. Dempsey, Asim Gazi, Ziping Xu, Finale Doshi-Velez, Susan Murphy ; 6:765−789, 2025. [abs ][pdf ][bib ][supp ]Hierarchical Multi-agent Reinforcement Learning for Cyber Network Defense Aditya Vikram Singh, Ethan Rathbun, Emma Graham, Lisa Oakley, Simona Boboila, Peter Chin, Alina Oprea ; 6:790−810, 2025. [abs ][pdf ][bib ][supp ]The Confusing Instance Principle for Online Linear Quadratic Control Waris Radji, Odalric-Ambrym Maillard ; 6:811−828, 2025. [abs ][pdf ][bib ][supp ]Drive Fast, Learn Faster: On-Board RL for High Performance Autonomous Racing Benedict Hildisch, Edoardo Ghignone, Nicolas Baumann, Cheng Hu, Andrea Carron, Michele Magno ; 6:829−861, 2025. [abs ][pdf ][bib ][supp ]Towards Large Language Models that Benefit for All: Benchmarking Group Fairness in Reward Models Kefan Song, Jin Yao, Runnan Jiang, Rohan Chandra, Shangtong Zhang ; 6:862−873, 2025. [abs ][pdf ][bib ]Pure Exploration for Constrained Best Mixed Arm Identification with a Fixed Budget Dengwang Tang, Rahul Jain, Ashutosh Nayyar, Pierluigi Nuzzo ; 6:874−893, 2025. [abs ][pdf ][bib ][supp ]Quantitative Resilience Modeling for Autonomous Cyber Defense Xavier Cadet, Simona Boboila, Edward Koh, Peter Chin, Alina Oprea ; 6:894−908, 2025. [abs ][pdf ][bib ][supp ]Efficient Information Sharing for Training Decentralized Multi-Agent World Models Xiaoling Zeng, Qi Zhang ; 6:909−922, 2025. [abs ][pdf ][bib ][supp ]Recursive Reward Aggregation Yuting Tang, Yivan Zhang, Johannes Ackermann, Yu-Jie Zhang, Soichiro Nishimori, Masashi Sugiyama ; 6:923−975, 2025. [abs ][pdf ][bib ]A Finite-Sample Analysis of an Actor-Critic Algorithm for Mean-Variance Optimization in a Discounted MDP Tejaram Sangadi, Prashanth L. A., Krishna Jagannathan ; 6:976−1024, 2025. [abs ][pdf ][bib ]Impoola: The Power of Average Pooling for Image-based Deep Reinforcement Learning Raphael Trumpp, Ansgar Schäfftlein, Mirco Theile, Marco Caccamo ; 6:1025−1047, 2025. [abs ][pdf ][bib ]Fast Adaptation with Behavioral Foundation Models Harshit Sikchi, Andrea Tirinzoni, Ahmed Touati, Yingchen Xu, Anssi Kanervisto, Scott Niekum, Amy Zhang, Alessandro Lazaric, Matteo Pirotta ; 6:1048−1074, 2025. [abs ][pdf ][bib ]Multi-Task Reinforcement Learning Enables Parameter Scaling Reginald McLean, Evangelos Chatzaroulas, J K Terry, Isaac Woungang, Nariman Farsad, Pablo Samuel Castro ; 6:1075−1093, 2025. [abs ][pdf ][bib ]Eau De $Q$-Network: Adaptive Distillation of Neural Networks in Deep Reinforcement Learning Théo Vincent, Tim Faust, Yogesh Tripathi, Jan Peters, Carlo D'Eramo ; 6:1094−1119, 2025. [abs ][pdf ][bib ][supp ]Disentangling Recognition and Decision Regrets in Image-Based Reinforcement Learning Alihan Hüyük, Arndt Ryo Koblitz, Atefeh Mohajeri Moghaddam, Matthew Andrews ; 6:1120−1139, 2025. [abs ][pdf ][bib ]Learning to Explore in Diverse Reward Settings via Temporal-Difference-Error Maximization Sebastian Griesbach, Carlo D'Eramo ; 6:1140−1157, 2025. [abs ][pdf ][bib ]Nonparametric Policy Improvement in Continuous Action Spaces via Expert Demonstrations Agustin Castellano, Sohrab Rezaei, Jared Markowitz, Enrique Mallada ; 6:1158−1179, 2025. [abs ][pdf ][bib ]DisDP: Robust Imitation Learning via Disentangled Diffusion Policies Pankhuri Vanjani, Paul Mattes, Xiaogang Jia, Vedant Dave, Rudolf Lioutikov ; 6:1180−1199, 2025. [abs ][pdf ][bib ]Mitigating Goal Misgeneralization via Minimax Regret Karim Abdel Sadek, Matthew Farrugia-Roberts, Usman Anwar, Hannah Erlebach, Christian Schroeder de Witt, David Krueger, Michael D Dennis ; 6:1200−1246, 2025. [abs ][pdf ][bib ]Long-Horizon Planning with Predictable Skills Nico Gürtler, Georg Martius ; 6:1247−1272, 2025. [abs ][pdf ][bib ]HANQ: Hypergradients, Asymmetry, and Normalization for Fast and Stable Deep $Q$-Learning Braham Snyder, Chen-Yu Wei ; 6:1273−1292, 2025. [abs ][pdf ][bib ]Benchmarking Massively Parallelized Multi-Task Reinforcement Learning for Robotics Tasks Viraj Joshi, Zifan Xu, Bo Liu, Peter Stone, Amy Zhang ; 6:1293−1317, 2025. [abs ][pdf ][bib ][supp ]Optimal discounting for offline input-driven MDP Randy Lefebvre, Audrey Durand ; 6:1318−1333, 2025. [abs ][pdf ][bib ]Make the Pertinent Salient: Task-Relevant Reconstruction for Visual Control with Distractions Kyungmin Kim, JB Lanier, Roy Fox ; 6:1334−1364, 2025. [abs ][pdf ][bib ]Reinforcement Learning for Human-AI Collaboration via Probabilistic Intent Inference Yuxin Lin, Seyede Fatemeh Ghoreishi, Tian Lan, Mahdi Imani ; 6:1365−1377, 2025. [abs ][pdf ][bib ]PufferLib 2.0: Reinforcement Learning at 1M steps/s Joseph Suarez ; 6:1378−1388, 2025. [abs ][pdf ][bib ]Uncovering RL Integration in SSL Loss: Objective-Specific Implications for Data-Efficient RL Ömer Veysel Çağatan, Baris Akgun ; 6:1389−1411, 2025. [abs ][pdf ][bib ]Benchmarking Partial Observability in Reinforcement Learning with a Suite of Memory-Improvable Domains Ruo Yu Tao, Kaicheng Guo, Cameron Allen, George Konidaris ; 6:1412−1439, 2025. [abs ][pdf ][bib ]Rectifying Regression in Reinforcement Learning Alex Ayoub, David Szepesvari, Alireza Bakhtiari, Csaba Szepesvari, Dale Schuurmans ; 6:1440−1454, 2025. [abs ][pdf ][bib ]High-Confidence Policy Improvement from Human Feedback Hon Tik Tse, Philip S. Thomas, Scott Niekum ; 6:1455−1479, 2025. [abs ][pdf ][bib ]Adaptive Reward Sharing to Enhance Learning in the Context of Multiagent Teams Kyle Tilbury, David Radke ; 6:1480−1495, 2025. [abs ][pdf ][bib ]MixUCB: Enhancing Safe Exploration in Contextual Bandits with Human Oversight Jinyan Su, Rohan Banerjee, Jiankai Sun, Wen Sun, Sarah Dean ; 6:1496−1520, 2025. [abs ][pdf ][bib ]Efficient Morphology-Aware Policy Transfer to New Embodiments Michael Przystupa, Hongyao Tang, Glen Berseth, Mariano Phielipp, Santiago Miret, Martin Jägersand, Matthew E. Taylor ; 6:1521−1539, 2025. [abs ][pdf ][bib ]Understanding Learned Representations and Action Collapse in Visual Reinforcement Learning Xi Chen, Zhihui Zhu, Andrew Perrault ; 6:1540−1557, 2025. [abs ][pdf ][bib ]Mitigating Suboptimality of Deterministic Policy Gradients in Complex Q-functions Ayush Jain, Norio Kosaka, Xinhu Li, Kyung-Min Kim, Erdem Biyik, Joseph J Lim ; 6:1558−1599, 2025. [abs ][pdf ][bib ][supp ]Leveraging priors on distribution functions for multi-arm bandits Sumit Vashishtha, Odalric-Ambrym Maillard ; 6:1600−1623, 2025. [abs ][pdf ][bib ][supp ]ProtoCRL: Prototype-based Network for Continual Reinforcement Learning Michela Proietti, Peter R. Wurman, Peter Stone, Roberto Capobianco ; 6:1624−1646, 2025. [abs ][pdf ][bib ]Finer Behavioral Foundation Models via Auto-Regressive Features and Advantage Weighting Edoardo Cetin, Ahmed Touati, Yann Ollivier ; 6:1647−1680, 2025. [abs ][pdf ][bib ]Pretraining Decision Transformers with Reward Prediction for In-Context Multi-task Structured Bandit Learning Subhojyoti Mukherjee, Josiah P. Hanna, Qiaomin Xie, Robert D Nowak ; 6:1681−1723, 2025. [abs ][pdf ][bib ]Multi-task Representation Learning for Fixed Budget Pure-Exploration in Linear and Bilinear Bandits Subhojyoti Mukherjee, Qiaomin Xie, Robert D Nowak ; 6:1724−1772, 2025. [abs ][pdf ][bib ]Offline Reinforcement Learning with Domain-Unlabeled Data Soichiro Nishimori, Xin-Qiang Cai, Johannes Ackermann, Masashi Sugiyama ; 6:1773−1793, 2025. [abs ][pdf ][bib ][supp ]Multi-Agent Reinforcement Learning for Inverse Design in Photonic Integrated Circuits Yannik Mahlau, Maximilian Schier, Christoph Reinders, Frederik Schubert, Marco Bügling, Bodo Rosenhahn ; 6:1794−1815, 2025. [abs ][pdf ][bib ][supp ]Syllabus: Portable Curricula for Reinforcement Learning Agents Ryan Sullivan, Ryan Pégoud, Ameen Ur Rehman, Xinchen Yang, Junyun Huang, Aayush Verma, Nistha Mitra, John P Dickerson ; 6:1816−1855, 2025. [abs ][pdf ][bib ][supp ]Exploration-Free Reinforcement Learning with Linear Function Approximation Luca Civitavecchia, Matteo Papini ; 6:1856−1879, 2025. [abs ][pdf ][bib ]SPEQ: Offline Stabilization Phases for Efficient Q-Learning in High Update-To-Data Ratio Reinforcement Learning Carlo Romeo, Girolamo Macaluso, Alessandro Sestini, Andrew D. Bagdanov ; 6:1880−1893, 2025. [abs ][pdf ][bib ]Value Bonuses using Ensemble Errors for Exploration in Reinforcement Learning Abdul Wahab, Raksha Kumaraswamy, Martha White ; 6:1894−1915, 2025. [abs ][pdf ][bib ]Gaussian Process Q-Learning for Finite-Horizon Markov Decision Processes Maximilian Bloor, Tom Savage, Calvin Tsay, Antonio Del rio chanona, Max Mowbray ; 6:1916−1930, 2025. [abs ][pdf ][bib ]On the Effect of Regularization in Policy Mirror Descent Jan Felix Kleuker, Aske Plaat, Thomas M. Moerland ; 6:1931−1950, 2025. [abs ][pdf ][bib ]Concept-Based Off-Policy Evaluation Ritam Majumdar, Jack Teversham, Sonali Parbhoo ; 6:1951−1989, 2025. [abs ][pdf ][bib ]Investigating the Utility of Mirror Descent in Off-policy Actor-Critic Samuel Neumann, Jiamin He, Adam White, Martha White ; 6:1990−2022, 2025. [abs ][pdf ][bib ]Hybrid Classical/RL Local Planner for Ground Robot Navigation Vishnu Dutt Sharma, Jeongran Lee, Matthew Andrews, Ilija Hadžić ; 6:2023−2037, 2025. [abs ][pdf ][bib ][supp ]How Should We Meta-Learn Reinforcement Learning Algorithms? Alexander David Goldie, Zilin Wang, Jaron Cohen, Jakob Nicolaus Foerster, Shimon Whiteson ; 6:2038−2081, 2025. [abs ][pdf ][bib ]Seldonian Reinforcement Learning for Ad Hoc Teamwork Edoardo Zorzi, Alberto Castellini, Leonidas Bakopoulos, Georgios Chalkiadakis, Alessandro Farinelli ; 6:2082−2100, 2025. [abs ][pdf ][bib ][supp ]Offline Reinforcement Learning with Wasserstein Regularization via Optimal Transport Maps Motoki Omura, Yusuke Mukuta, Kazuki Ota, Takayuki Osa, Tatsuya Harada ; 6:2101−2114, 2025. [abs ][pdf ][bib ][supp ]Intrinsically Motivated Discovery of Temporally Abstract Graph-based Models of the World Akhil Bagaria, Anita De Mello Koch, Rafael Rodriguez-Sanchez, Sam Lobel, George Konidaris ; 6:2115−2134, 2025. [abs ][pdf ][bib ][supp ]An Optimisation Framework for Unsupervised Environment Design Nathan Monette, Alistair Letcher, Michael Beukman, Matthew Thomas Jackson, Alexander Rutherford, Alexander David Goldie, Jakob Nicolaus Foerster ; 6:2135−2158, 2025. [abs ][pdf ][bib ]Epistemically-guided forward-backward exploration Núria Armengol Urpí, Marin Vlastelica, Georg Martius, Stelian Coros ; 6:2159−2173, 2025. [abs ][pdf ][bib ][supp ]Rethinking the Foundations for Continual Reinforcement Learning Esraa Elelimy, David Szepesvari, Martha White, Michael Bowling ; 6:2174−2194, 2025. [abs ][pdf ][bib ]Modelling human exploration with light-weight meta reinforcement learning algorithms Thomas D. Ferguson, Alona Fyshe, Adam White ; 6:2195−2209, 2025. [abs ][pdf ][bib ]Zero-Shot Reinforcement Learning Under Partial Observability Scott Jeen, Tom Bewley, Jonathan Cullen ; 6:2210−2233, 2025. [abs ][pdf ][bib ]Building Sequential Resource Allocation Mechanisms without Payments Sihan Zeng, Sujay Bhatt, Alec Koppel, Sumitra Ganesh ; 6:2234−2255, 2025. [abs ][pdf ][bib ][supp ]From Explainability to Interpretability: Interpretable Reinforcement Learning Via Model Explanations Peilang Li, Umer Siddique, Yongcan Cao ; 6:2256−2270, 2025. [abs ][pdf ][bib ]Joint-Local Grounded Action Transformation for Sim-to-Real Transfer in Multi-Agent Traffic Control Justin Turnau, Longchao Da, Khoa Vo, Ferdous Al Rafi, Shreyas Bachiraju, Tiejin Chen, Hua Wei ; 6:2271−2290, 2025. [abs ][pdf ][bib ]Sampling from Energy-based Policies using Diffusion Vineet Jain, Tara Akhound-Sadegh, Siamak Ravanbakhsh ; 6:2291−2307, 2025. [abs ][pdf ][bib ]Multiple-Frequencies Population-Based Training Waël Doulazmi, Auguste Lehuger, Marin Toromanoff, Valentin Charraut, Thibault Buhet, Fabien Moutarde ; 6:2308−2326, 2025. [abs ][pdf ][bib ][supp ]TransAM: Transformer-Based Agent Modeling for Multi-Agent Systems via Local Trajectory Encoding Conor Wallace, Umer Siddique, Yongcan Cao ; 6:2327−2341, 2025. [abs ][pdf ][bib ]Towards Improving Reward Design in RL: A Reward Alignment Metric for RL Practitioners Calarina Muslimani, Kerrick Johnstonbaugh, Suyog Chandramouli, Serena Booth, W. Bradley Knox, Matthew E. Taylor ; 6:2342−2367, 2025. [abs ][pdf ][bib ]Optimistic critics can empower small actors Olya Mastikhina, Dhruv Sreenivas, Pablo Samuel Castro ; 6:2368−2387, 2025. [abs ][pdf ][bib ]PAC Apprenticeship Learning with Bayesian Active Inverse Reinforcement Learning Ondrej Bajgar, Dewi Sid William Gould, Jonathon Liu, Alessandro Abate, Konstantinos Gatsis, Michael A Osborne ; 6:2388−2414, 2025. [abs ][pdf ][bib ]AVG-DICE: Stationary Distribution Correction by Regression Fengdi Che, Bryan Chan, Chen Ma, A. Rupam Mahmood ; 6:2415−2426, 2025. [abs ][pdf ][bib ][supp ]V-Max: A RL Framework for Autonomous Driving Valentin Charraut, Waël Doulazmi, Thomas Tournaire, Thibault Buhet ; 6:2427−2451, 2025. [abs ][pdf ][bib ][supp ]Offline Action-Free Learning of Ex-BMDPs by Comparing Diverse Datasets Alexander Levine, Peter Stone, Amy Zhang ; 6:2452−2484, 2025. [abs ][pdf ][bib ]One Goal, Many Challenges: Robust Preference Optimization Amid Content-Aware, Multi-Source Noise Amirabbas Afzali, Amirhossein Afsharrad, Seyed Shahabeddin Mousavi, Sanjay Lall ; 6:2485−2507, 2025. [abs ][pdf ][bib ][supp ]A Timer-Based Hybrid Supervisor for Robust, Chatter-Free Policy Switching Jan de Priester, Ricardo Sanfelice ; 6:2508−2529, 2025. [abs ][pdf ][bib ][supp ]Deep Reinforcement Learning with Gradient Eligibility Traces Esraa Elelimy, Brett Daley, Andrew Patterson, Marlos C. Machado, Adam White, Martha White ; 6:2530−2550, 2025. [abs ][pdf ][bib ]On Slowly-varying Non-stationary Bandits Ramakrishnan K, Aditya Gopalan ; 6:2551−2584, 2025. [abs ][pdf ][bib ]Focused Skill Discovery: Learning to Control Specific State Variables while Minimizing Side Effects Jonathan Colaço Carr, Qinyi Sun, Cameron Allen ; 6:2585−2599, 2025. [abs ][pdf ][bib ]Goals vs. Rewards: A Preliminary Comparative Study of Objective Specification Mechanisms Septia Rani, Serena Booth, Sarath Sreedharan ; 6:2600−2618, 2025. [abs ][pdf ][bib ]An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Brett Daley, Prabhat Nagarajan, Martha White, Marlos C. Machado ; 6:2619−2636, 2025. [abs ][pdf ][bib ]PEnGUiN: Partially Equivariant Graph NeUral Networks for Sample Efficient MARL Joshua McClellan, Greyson Brothers, Furong Huang, Pratap Tokekar ; 6:2637−2651, 2025. [abs ][pdf ][bib ]Shaping Laser Pulses with Reinforcement Learning Francesco Capuano, Davorin Peceli, Gabriele Tiboni ; 6:2652−2666, 2025. [abs ][pdf ][bib ]Reinforcement Learning with Adaptive Temporal Discounting Sahaj Singh Maini, Zoran Tiganj ; 6:2667−2684, 2025. [abs ][pdf ][bib ]Human-Level Competitive Pokémon via Scalable Offline Reinforcement Learning with Transformers Jake Grigsby, Yuqi Xie, Justin Sasek, Steven Zheng, Yuke Zhu ; 6:2685−2719, 2025. [abs ][pdf ][bib ]Adaptive Submodular Policy Optimization Branislav Kveton, Anup Rao, Viet Dac Lai, Nikos Vlassis, David Arbour ; 6:2720−2736, 2025. [abs ][pdf ][bib ][supp ]Learning Fair Pareto-Optimal Policies in Multi-Objective Reinforcement Learning Umer Siddique, Peilang Li, Yongcan Cao ; 6:2737−2761, 2025. [abs ][pdf ][bib ][supp ]Representation Learning and Skill Discovery with Empowerment Andrew Levy, Alessandro G Allievi, George Konidaris ; 6:2762−2787, 2025. [abs ][pdf ][bib ][supp ]Empirical Bound Information-Directed Sampling for Norm-Agnostic Bandits Piotr M. Suder, Eric Laber ; 6:2788−2819, 2025. [abs ][pdf ][bib ]Thompson Sampling for Constrained Bandits Rohan Deb, Mohammad Ghavamzadeh, Arindam Banerjee ; 6:2820−2843, 2025. [abs ][pdf ][bib ]AI in a vat: Fundamental limits of efficient world modelling for agent sandboxing and interpretability Fernando Rosas, Alexander Boyd, Manuel Baltieri ; 6:2844−2881, 2025. [abs ][pdf ][bib ]Achieving Limited Adaptivity for Multinomial Logistic Bandits Sukruta Prakash Midigeshi, Tanmay Goyal, Gaurav Sinha ; 6:2882−2896, 2025. [abs ][pdf ][bib ][supp ]