PaperScope
LIVE · 2026-09-03 05:40 UTC

Reinforcement Learning and Rule-Based Peer-to-Peer Pricing in Residential PV-BES Communities

Pablo Benalcazar, Maciej Kalka, Wilian Guamán, Jacek Kamiński

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.01680 v1
Category
Submitted
2026-09-01

Abstract

This paper compares rule-based and learning-based pricing mechanisms for peer-to-peer (P2P) electricity trading in residential photovoltaic communities. The rule-based benchmarks comprise bill-sharing as an ex post allocation mechanism, the mid-market rate, and supply-demand-ratio pricing. The reinforcement-learning (RL) formulation is implemented through a Deep Q-Network and evaluated under multiplier-based and learnable SDR-shaped pricing, with a fixed-parameter SDR variant as a non-learning control. Performance is assessed through community savings together with complementary financial and operational indicators. In the base PV-only configuration, the rule-based benchmarks outperform the best RL policy. With battery energy storage, evaluated for the RL policies only, community savings under the best RL policy increase from EUR 734.23 to EUR 978.52. Across the learning-based modes and in both configurations, SDR-shaped pricing outperforms the multiplier-based parameterization considered. The results indicate that rule-based pricing remains highly competitive wherever the two families are compared directly, and that storage substantially improves the learning-based outcomes under this accounting, while the distribution of benefits remains heterogeneous across households.

Comment: 22 pages, 2 figures, submitted to ARTIIS 2026 (Conference on Advanced Research in Technologies, Information, Innovation and Sustainability) https://www.artiis.org/special-sessions/iwet-2026

arXiv abs page · PDF · same-day batch