SGD3QN: Using Stochastic Games and Deep Reinforcement Learning to Defend IoT Systems
Published:
Malware propagation across Internet of Things (IoT) devices can lead to data leakage, financial loss, service disruption, and the compromise of wider cyber-physical environments. Defending these systems is difficult because malware and defensive actions continually alter the security state of the network.
In this work, we investigated how stochastic game theory and deep reinforcement learning can be combined to learn active malware-defence strategies.
This blog post discusses our paper, SGD3QN: Joint Stochastic Games and Dueling Double Deep Q-Networks for Defending Malware Propagation in Edge Intelligence-Enabled Internet of Things https://doi.org/10.1109/TIFS.2024.3420233, published in IEEE Transactions on Information Forensics and Security.
What problem did we investigate?
IoT malware defence involves a sequence of connected decisions. A defender may need to decide when to intervene, which devices to prioritise, and how to respond as malware spreads across the network.
Each decision changes the environment. A successful defensive action may reduce the number of infected devices, while an unsuccessful action may allow the infection to spread further.
This makes malware propagation a dynamic decision problem rather than a conventional one-time classification task.
Why use stochastic games?
We used a stochastic game to represent the cyber conflict between IoT system nodes and defensive edge devices.
At each stage of the game, the participants select strategies and receive rewards based on the current system state and their chosen actions. The game then transitions to another state according to defined probabilities and the strategies selected by the participants.
This process continues as the participants adapt their decisions. The theoretical objective is to reach a stable strategy represented by a Nash equilibrium.
The stochastic game therefore provides a formal way to describe the changing relationship between malware propagation and defensive intervention.
What is SGD3QN?
Following the game-theoretic analysis, we developed SGD3QN, which stands for Stochastic Games and Dueling Double Deep Q-Networks.
The Dueling Double Deep Q-Network functions as an end-to-end decision-making mechanism. It receives information about the malware propagation environment and learns from successful and unsuccessful defensive experiences.
The Double DQN component helps reduce the overestimation of action values. The dueling architecture separately learns the value of the current state and the relative advantage of the available actions.
Combining these mechanisms helps the defender learn which actions are most appropriate under different malware propagation conditions.
How was the method evaluated?
We performed simulation experiments to study the behaviour of the proposed method. In particular, we examined how batch size and replay memory size affected the learning process and the selection of malware-defence strategies.
Replay memory allows the algorithm to store previous experiences and reuse them during training. Its size can affect the diversity and relevance of the experiences from which the model learns.
Batch size determines how many stored experiences are used during an individual learning update. Selecting these parameters appropriately is important for stable and effective reinforcement learning.
What did we find?
The experimental results demonstrated the ability of SGD3QN to learn effective strategies for mitigating IoT malware propagation.
The study also showed that learning performance depends on more than the underlying algorithm. Configuration choices such as replay memory size and batch size can influence convergence and the quality of the selected defence strategy.
The results support the wider use of reinforcement learning for adaptive IoT security, while also emphasising the importance of systematic parameter evaluation.
Perspective
SGD3QN connects formal security modelling with practical automated decision-making. The stochastic game describes how malware and defenders interact, while the Dueling Double DQN learns how to respond to the resulting security states.
This is an important distinction from machine learning approaches that only classify network activity as malicious or benign. The purpose of SGD3QN is to support the next defensive decision after the system observes a changing malware environment.
Such decision-oriented AI could contribute to future autonomous cyber-defence systems, particularly in large IoT deployments where exclusively manual incident response may not scale.
The paper was co-authored with Yizhou Shen, Carlton Shepherd, Shigen Shen, and Shui Yu.
Research collaboration and consultancy
Our research explores adaptive and intelligent cybersecurity for IoT, edge computing, and cyber-physical systems. We are interested in applying game theory and reinforcement learning to practical problems involving malware, attack response, resource allocation, and autonomous cyber defence.
We welcome enquiries from:
- organisations developing connected and embedded products;
- IoT and edge-computing platform providers;
- critical-infrastructure and industrial operators;
- cybersecurity vendors;
- government and public-sector organisations;
- research laboratories and universities; and
- organisations recruiting expertise in IoT security, AI security, and cyber-physical systems.
Potential areas for research collaboration and consultancy include:
- IoT malware propagation modelling;
- automated malware containment;
- reinforcement learning for cyber defence;
- game-theoretic attack and defence modelling;
- evaluation of AI-based security controls;
- security of edge intelligence systems;
- autonomous incident-response strategies;
- cyber-physical systems security;
- security testbed design and experimental evaluation; and
- independent technical review of IoT security solutions.
For consultancy, collaborative research, advisory activities, invited talks, or related opportunities, contact Dr Mujeeb Ahmed, Senior Lecturer in Computing at Newcastle University:
Email: mujeeb.ahmed@newcastle.ac.uk
Reference
https://doi.org/10.1109/TIFS.2024.3420233. IEEE Transactions on Information Forensics and Security 19 (2024), 6978–6990. https://doi.org/10.1109/TIFS.2024.3420233
