RALI-OLST-ILFC Weekly Talks

On Wednesdays at 11:30 a.m. (Montreal time), we hold a one-hour talk on a language processing or linguistics topic. It is typically offered in hybrid mode (in person/videoconference). Once a month, the talk is organized by the French research group Linguistique Informatique, Formelle et de Terrain.

Sentiment-Augmented Deep Reinforcement Learning for Active Trading - An Alpha Reward Approach

Andrei Neagu (andrei (point) neagu <at> concordia (point) ca)

Concordia University

Wednesday 4 November 2026 at 11:30 AM

Salle 3195, Pavillon André-Aisenstadt — En présentiel, avec diffusion simultanée sur Zoom


We present our system for Task 3 of the CLEF 2026 FinMMEval Lab, which requires participants to submit a daily trading decision (long, flat, or short) for Bitcoin (BTC) and Tesla (TSLA) based on news articles and historical market pricing data. We frame the problem as a discrete-action Markov Decision Process and train four Deep Reinforcement Learning (DRL) algorithms: Policy Gradient (PG), Proximal Policy Optimization (PPO), Deep Q-learning (DQL), and Deep Deterministic Policy Gradient (DDPG), using a rich feature set that combines technical indicators (EMA, RSI, MACD, BollingerB, volume change), cyclical date encodings, and daily sentiment scores derived from news articles using LLaMA 3.2 1B. To reduce overfitting and align the training objective with the goal of outperforming a buy-and-hold baseline, we introduce an alpha reward that replaces the raw log-return with excess return over the market, and we randomize episode start days during training. Hyperparameters are optimized over 180 trials per algorithm-asset pair using Ray Tune, with model selection based on validation Sharpe ratio (SR) and early stopping. Evaluation on the CLEF Task 3 test set demonstrates that DDPG yields the best performance across both assets. Note, however, that DQL was selected a priori for the live competition endpoint based on its highest validation Sharpe ratio; the endpoint model was deliberately selected blind to the test period to avoid selection bias. For TSLA, DDPG and DQL achieved cumulative returns of 54.96% and 52.62% respectively, substantially beating the 16.45% buy-and-hold baseline. On the more challenging BTC test set, DDPG mitigated severe market losses, achieving a positive return of 1.58% against a baseline decline of -34.27%. Results reveal a pronounced validation-to-test generalization gap, pointing to the difficulty of adapting policies trained in a bull market validation period to bear market test conditions.



Join with Zoom at this url.


Follow this link to subscribe to future RALI-OLST announcements.
http://rali.iro.umontreal.ca/rali/?q=fr/node/1631

See all the weekly talks for the year:

1991 1992 1993 1994 1995 1997 1998 1999 2000 2001 2002 2003 2004 2005 2006 2007 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2026