Skip to main content
Have a personal or library account? Click to login
Radix-4 carry-save adder based accumulator for high-performance MAC units in factored systolic array accelerators Cover

Radix-4 carry-save adder based accumulator for high-performance MAC units in factored systolic array accelerators

Open Access
|Aug 2026

References

  1. K. Inayat, I. Ullah, and J. Chung, “Factored Systolic Arrays Based on Radix-8 Multiplication for Machine Learning Acceleration,” IEEE Trans. Very Large Scale Integr. (VLSI) Syst., vol. 32, no. 7, pp. 1205–1215, 2024.
  2. S. N. Kumar, S. V. Kumar, J. Abishek, and R. Sakthivel, “Design of High-Speed Machine Learning Accelerators using Hybrid Accumulator in Systolic Arrays,” in Proc. 6th Int. Conf. Electronics Sustainable Commun. Syst. (ICESC), pp. 500–507, 2025.
  3. F. B. Muslim, K. Inayat, M. Z. Siddiqi, S. Khan, T. Mahmood, and I. ul Islam, “SAPER-AI accelerator: asystolic array-based power-efficient reconfigurable AI accelerator,” Front. Inf. Technol. Electron. Eng., vol. 26, no. 9, pp. 1624–1636, 2025.
  4. V. Sze, Y.-H. Chen, T.-J. Yang, and J. S. Emer, “Efficient processing of deep neural networks: A tutorial and survey,” Proc. IEEE, vol. 105, no. 12, pp. 2295–2329, 2017.
  5. H. T. Kung, “Why systolic architectures?” Computer, vol. 15, no. 1, pp. 37–46, 1982.
  6. N. P. Jouppi et al., “In-datacenter performance analysis of a tensor processing unit,” in Proc. ACM/IEEE 44th Annu. Int. Symp. Comput. Archit. (ISCA), pp. 1–12, 2017.
  7. J. Abdelmaksoud, C. Sestito, S. Wang and T. Prodromakis, “ADiP: Adaptive-Precision Systolic Array for Matrix Multiplication Acceleration,” in IEEE Open Journal of the Solid-State Circuits Society, pp. 1–1, 2026.
  8. D. K. J. Rajanendiran, C.G. Babu and K. Priyadharsini, “A certain examination on heterogeneous systolic array (HSA) design for deep learning accelerations with low power computations”, Sustainable Computing: Informatics and Systems, 44, 101042, 2017.
  9. T. Suguna, B. K. Suresh, and A. M. Malim, “Design and performance analysis of multiply and accumulate (MAC) unit for machine learning acceleration,” J. Comput. Sci., vol. 19, no. 3, pp. 19–33, 2026.
  10. D. N. Devi, G. Ajay Kumar, B. G. Gowda, and M. Rao, “Integrated MAC-based systolic arrays: Design and performance evaluation,” in Proc. Great Lakes Symp. VLSI, pp. 292–295, 2024.
  11. Del Barrio and R. Hermida, “A slack-based approach to efficiently deploy radix-8 Booth multipliers,” in Proc. Design, Autom. Test Eur. Conf. (DATE), pp. 1153–1158, 2017.
  12. N. H. E. Weste and D. M. Harris, CMOS VLSI Design: A Circuits and Systems Perspective, 4th ed. Boston, MA, USA: Addison-Wesley, 2010.
  13. Q. Song, Y. Dai, H. Lu and G. Jin, “High-throughput systolic array-based accelerator for hybrid transformer -CNN networks,” J. King Saud Univ. Comput. Inf. Sci., vol. 36, no. 8, p. 102194, 2024.
  14. M. S. Kumar, D. A. Kumar and P. Samundiswary, “Design and performance analysis of Multiply-Accumulate (MAC) unit,” in Proc. Int. Conf. Circuits, Power Comput. Technol. (ICCPCT), pp. 1084–1089,2014.
  15. S. Zhang, J. Gu, S. Yin, L. Liu, and S. Wei, “A multiple-precision multiply and accumulation design with multiply-add merged strategy for AI accelerating,” in Proc. 26th Asia South Pacific Design Autom. Conf. (ASPDAC), pp. 229–234, 2021.
  16. S. Agasthiya, S. Balambigai, and R. Sarankumar, “An efficient Radix-8 factored systolic array for machine learning acceleration,” in Proc. 5th Int. Conf. Trends Mater. Sci. Inventive Mater. (ICTMIM), pp. 1437–1444, 2025
  17. S. Zhang, J. Gu, S. Yin, L. Liu, and S. Wei, “A multiple-precision multiply and accumulation design with multiply-add merged strategy for AI accelerating,” in Proc. 26th Asia South Pacific Design Autom. Conf. (ASPDAC), pp. 229–234,2021.
  18. P. Gurjar, R. Solanki, P. Kansliwal, and M. Vucha, “VLSI implementation of adders for high speed ALU,” in Proc. Annu. IEEE India Conf., pp. 1–6, 2011.
  19. R. A. Javali, R. J. Nayak, A. M. Mhetar, and M. C. Lakkannavar, “Design of high speed carry save adder using carry lookahead adder,” in Proc. Int. Conf. Circuits, Commun., Control Comput., pp. 33–36, 2014.
  20. M. D. Ercegovac and T. Lang, Digital Arithmetic. San Francisco, CA, USA: Morgan Kaufmann, 2003.
  21. M. Hassan, M. S. Ansari, and W. A. Khan, “Performance analysis of carry select adder based MAC unit for DSP applications,” Microelectron. J., vol. 90, pp. 1–8, 2019.
  22. D. N. Devi, G. A. Kumar, B. G. Gowda, and M. Rao, “Performance-aware design of approximate integrated MAC factored systolic array accelerators,” in Proc. 25th Int. Symp. Qual. Electron. Des. (ISQED), pp. 1–8, 2024.
  23. P. Jaswal, L. H. Krishna, and B. Srinivasu, “Energy efficient exact and approximate systolic array architecture for matrix multiplication,” in Proc. 39th Int. Conf. VLSI Des. & 25th Int. Conf. Embedded Syst. (VLSID), pp. 524–529, 2026.
  24. P. M. Mohan and S. S. Kumar, “Design and analysis of high speed low power radix-4 carry save adder,” in Proc. Int. Conf. Commun. Signal Process. (ICCSP), pp. 1016–1020, 2016.
DOI: https://doi.org/10.2478/jee-2026-0037 | Journal eISSN: 1339-309X (formerly 1335-3632) | Journal ISSN: 1335-3632
Language: English
Page range: 384 - 391
Submitted on: Apr 13, 2026
Published on: Aug 27, 2026
Published by: Slovak University of Technology in Bratislava
In partnership with: Paradigm Publishing Services

© 2026 Komathy Vanitha Krishnan, Angeline Felicia Moses, Sowmya Pravin Sishu, Prediksha Dorai, Rajesh Kana Santharaj, published by Slovak University of Technology in Bratislava
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.