FARANE-Q: Fast Parallel and Pipeline Q-Learning Accelerator for Configurable Reinforcement Learning SoC

This paper proposes a FAst paRAllel and pipeliNE Q-learning accelerator (FARANE-Q) for a configurable Reinforcement Learning (RL) algorithm implemented in a System on Chip (SoC). The proposed work offers flexibility, configurability, and scalability while maintaining computation speed and accuracy t...

Full description

Saved in:

Bibliographic Details
Published in:	IEEE access Vol. 11; pp. 144 - 161
Main Authors:	Sutisna, Nana, Ilmy, Andi M. Riyadhus, Syafalni, Infall, Mulyawan, Rahmat, Adiono, Trio
Format:	Journal Article
Language:	English
Published:	Piscataway IEEE 2023 The Institute of Electrical and Electronics Engineers, Inc. (IEEE)
Subjects:	Accuracy Algorithms Computation Computer architecture Energy efficiency Field programmable gate arrays Flexibility FPGA Heuristic algorithms HW accelerator Machine learning Microprocessors Navigation Pipelines Pipelining (computers) Predictive control Predictive maintenance Q-learning reinforcement learning Robot control SoC Software System on chip Throughput
Online Access:	Get full text
Tags:	Add Tag No Tags, Be the first to tag this record!

Description
Summary:	This paper proposes a FAst paRAllel and pipeliNE Q-learning accelerator (FARANE-Q) for a configurable Reinforcement Learning (RL) algorithm implemented in a System on Chip (SoC). The proposed work offers flexibility, configurability, and scalability while maintaining computation speed and accuracy to overcome the challenges of a dynamic environment and increasing complexity. The proposed method includes a Hardware/Software (HW/SW) design methodology for the SoC architecture to achieve flexibility. We also propose joint optimizations on the algorithm, architecture, and implementation to obtain optimum (high efficiency) performance, specifically in energy and area efficiency. Furthermore, we implemented the proposed design in a real-time Zynq Ultra96-V2 FPGA platform to evaluate the functionality with an actual use case of smart navigation. Experimental results confirm that the proposed accelerator FARANE-Q outperforms state-of-the-art works by achieving a throughput of up to 148.55 MSps. It corresponds to the energy efficiency of 1747.64 MSps/W per agent for 32-bit and 2424.33 MSps/W per agent for 16-bit FARANE-Q. Moreover, the proposed 16-bit FARANE-Q outperforms other related works by an improvement of at least <inline-formula> <tex-math notation="LaTeX">1.23\times </tex-math></inline-formula> in energy efficiency. The designed system also maintains an error accuracy of less than 0.4% with optimized bit precision for more than eight fraction bits. The proposed FARANE-Q also offers a speed up of processing time up to <inline-formula> <tex-math notation="LaTeX">1795\times </tex-math></inline-formula> compared to embedded SW computation executed on ARM Zynq processor and <inline-formula> <tex-math notation="LaTeX">280\times </tex-math></inline-formula> of computation of full software executed on i7 processor. Hence, the proposed work has the potential to be used for smart navigation, robotic control, and predictive maintenance.
ISSN:	2169-3536 2169-3536
DOI:	10.1109/ACCESS.2022.3232853