Outline

Ingegneria Sismica

Ingegneria Sismica

Application of Reinforcement Learning in Optimizing New Media Content Recommendation Systems

Author(s): Ruofei Gu1
1School of Digital Media and Design Arts, Beijing University of Posts and Telecommunications, Beijing, 102206, China
Gu, Ruofei. “Application of Reinforcement Learning in Optimizing New Media Content Recommendation Systems.” Ingegneria Sismica Volume 43 Issue 3: 1-20, doi:10.65102/is20261253.

Abstract

To address the conflict between immediate click-through rate and long-term retention in new media recommendation, this paper proposes DRL-MOREC, a deep reinforcement learning framework. A hybrid state encoder fuses users’ short-term behavior sequences with long-term interest graphs to capture dual temporal scales of interest. A two-stage reward function allocates session-level 7-day retention prediction signals to each step via a discount factor, alleviating delayed reward sparsity. Conservative Q-learning and inverse propensity score weighting are introduced to mitigate distribution shift and popularity bias, respectively. Offline experiments on a short-video platform dataset show that the proposed method achieves a 7-day retention rate 2.0, 2.7, and 1.3 percentage points higher than DeepFM, DDPG-TD3, and SAC-Rec, respectively, while improving catalog coverage (ECC) by 0.09. Online A/B testing demonstrates an 11.3% lift in daily active user retention over DeepFM. Ablation studies reveal that the delayed reward contributes 4.0 percentage points to the retention improvement. These results validate the effectiveness of reinforcement learning in optimizing long-term user value under dynamic recommendation scenarios.

Keywords
Deep Reinforcement Learning; Recommender System; New Media Content; Multi-Objective Optimization; User Retention

Related Articles

Zhihao Jiang1,2, Limi Chen1,2, Jing Yang1
1Hainan Vocational University of Science and Technology, Haikou 571126, China
2Institute for Mathematical Research, Universiti Putra Malaysia, Serdang 43400, Malaysia
Limi Chen1,2, Zhihao Jiang1,2, Jing Yang1
1Hainan Vocational University of Science and Technology, Haikou 571126, China
2Institute for Mathematical Research, Universiti Putra Malaysia, Serdang 43400, Malaysia
Hui Yuan1, Minjie Chai2, Siqing Xu1, Jinsong Li1, Jinwan Zheng1
1Electric Power Research Institute, State Grid Shanxi Electric Power Co., Ltd., Taiyuan, 030001, Shanxi, China
2Jincheng Power Supply Branch, State Grid Shanxi Electric Power Co., Ltd., Jincheng, 048000, Shanxi, China
Yanhan Zhu1,2
1China Academy of Cultural Heritage, Chaoyang District, 100029, Beijing, China
2Beijing University of Civil Engineering and Architecture, Xicheng District, 100044, Beijing, China
Ken Wang1, Jinhan Shu2, Kan Yuan1
1School of Digital Media, Shenzhen Polytechnic University, Shenzhen 518055, Guangdong, China
2Postdoctoral Mobile Station of Journalism and communication, Fudan University, Shanghai 200433, Shanghai, China