Article navigation
Purpose

This study proposes a deep reinforcement learning (DRL) framework designed to optimize agricultural investment portfolios under uncertainty. It specifically addresses the challenges faced by agribusiness investors and policymakers in emerging markets – namely, commodity price volatility, transaction costs and dynamic market disruptions.

Design/methodology/approach

We implement three state-of-the-art DRL algorithms – twin delayed deep deterministic policy gradient (TD3), deep deterministic policy gradient (DDPG), and soft actor-critic (SAC) – combined with recurrent neural networks (LSTM and GRU). These models dynamically allocate capital across key agricultural commodities (maize, rice, coffee and cotton) and benchmark assets (S&P 500, gold). The empirical analysis spans the period 2014–2024 using daily data, and stress-tests portfolio performance under varying levels of risk aversion and transaction cost regimes.

Findings

DRL-based strategies significantly outperform traditional portfolio methods (minimum variance, risk parity and maximum sharpe) in terms of cumulative return, sharpe ratio, and robustness to transaction costs. SAC delivers the highest returns and adaptability in volatile, high-cost environments, while TD3 provides enhanced stability in low-cost scenarios. The ablation study also highlights the impact of systematic risk variation, showing that shifts in macroeconomic uncertainty and market-wide shocks directly affect the DRL models' sensitivity to risk constraints and portfolio allocation. Furthermore, the ablation study confirms that removing transaction costs or risk aversion constraints results in increased volatility and unrealistic rebalancing, underscoring the value of incorporating real-world frictions into the reward structure.

Research limitations/implications

This study focuses on a subset of agricultural commodities and benchmark indices, assuming access to high-frequency data and computational resources. Future research could extend the framework to include climate shocks, ESG factors or multi-agent learning setups. Application in data-constrained contexts – such as smallholder or cooperative-level investment settings – also presents a promising avenue.

Originality/value

This study is among the first to incorporate both investor risk preferences and transaction cost penalties directly into the DRL reward function within the context of agricultural portfolio optimization. It bridges algorithmic finance and agribusiness strategy, offering scalable, real-world solutions for portfolio allocation in volatile and underdeveloped markets.

Licensed re-use rights only
You do not currently have access to this content.
Don't already have an account? Register

Purchased this content as a guest? Enter your email address to restore access.

Pay-Per-View Access
$39.00
Rental

or Create an Account

Close subscription notice
Close access options