Which Deep Q-Learning Method Best Suppresses Malware in Edge-Enabled IoT Networks?

2 minute read

Published:

Internet of Things (IoT) environments may contain large numbers of interconnected and resource-constrained devices. Once malware compromises part of such an environment, it can propagate between devices and cause data leakage, service disruption, and wider system compromise.

Static defence rules may struggle to respond effectively because the number of infected devices and the behaviour of the malware can change over time. This motivates the use of adaptive decision-making methods that can learn how to suppress malware in changing conditions.

This blog post discusses our paper, Comparative DQN-Improved Algorithms for Stochastic Games-Based Automated Edge Intelligence-Enabled IoT Malware Spread-Suppression Strategies https://doi.org/10.1109/JIOT.2024.3381281, published in the IEEE Internet of Things Journal.

What problem did we investigate?

Malware propagation and defence can be viewed as a continuing interaction. Malware attempts to compromise additional IoT devices, while defensive edge nodes must decide how to allocate their limited resources to suppress the infection.

The system does not remain in a fixed state. Each successful infection or defensive action changes the environment and affects the decisions available at the next stage.

Deep reinforcement learning can help automate these decisions. However, different extensions of the Deep Q-Network, or DQN, address different weaknesses in the original algorithm. It is therefore important to determine how these improved methods perform when applied to IoT malware suppression.

What did we propose?

We represented the confrontation between IoT malware and defensive edge nodes as a stochastic game. In this game, the participants select their strategies, receive rewards, and move between different system states according to transition probabilities.

The stochastic game provides the theoretical model for understanding how malware propagation and defensive intervention influence one another over time.

We then applied and compared three improved DQN algorithms:

  1. DDQMS, which applies Double DQN to malware spread suppression.
  2. D2QMS, which applies Dueling DQN to malware spread suppression.
  3. D3QMS, which combines Dueling and Double DQN for malware spread suppression.

These algorithms learn from their interactions with the simulated malware environment and attempt to identify effective suppression decisions.

Why compare these DQN variants?

A standard DQN can overestimate the expected value of possible actions. Double DQN helps reduce this overestimation by separating action selection from action evaluation.

A Dueling DQN separately estimates the value of the current state and the advantage of each available action. This can improve learning when several actions have similar effects.

D3QMS combines these ideas. The comparison therefore helps explain whether the individual or combined improvements are more suitable for automated malware defence in edge intelligence-enabled IoT systems.

What did we find?

Our experiments compared the learning behaviour and performance of the three algorithms. We also investigated how different parameter settings influenced the selection of malware suppression strategies.

The results showed that improved DQN methods can support adaptive decision-making for IoT malware defence. They also demonstrated that the choice of algorithm and learning parameters can materially affect the quality