A Comparative Analysis of Spark GraphX and GraphFrames for Healthcare Data Analytics

Authors

  • Soz Raouf Hama Information Technology Department, Technical College of Informatics, Sulaimani Polytechnic University, Sulaymaniyah, Iraq
  • Alaa Khalil Jumaa Database Technology Department, Technical College of Informatics, Sulaimani Polytechnic University, Sulaymaniyah, Iraq

DOI:

https://doi.org/10.21928/uhdjst.v10n2y2026.pp68-79

Keywords:

In-Degree, Out-Degree, Total-Degree, Graph Analytics, GraphFrames, GraphX, Healthcare Data Analytics, PageRank, PySpark

Abstract

Healthcare data analytics is essential for identifying disease patterns, understanding patient relationships, and supporting data-driven decision-making in modern healthcare systems. As healthcare datasets continue to grow in size and complexity, scalable graph-processing frameworks have become increasingly important for analyzing interconnected patient data. This paper compares two Apache Spark graph-processing frameworks, GraphX and GraphFrames, using the Centers for Disease Control and Prevention Diabetes Health Indicators dataset. A patient similarity graph was constructed by representing 70,692 patients as vertices and connecting highly similar patients through cosine–similarity relationships, resulting in 16,009,101 graph edges. The two frameworks were evaluated using the same graph structure and execution environment with respect to graph construction time, execution performance, memory consumption, application programming interface usability, and scalability. In addition to the framework comparison, graph-based features, including in-degree, out-degree, total degree, and PageRank, were extracted to examine the structure of the patient similarity network. The experimental results showed that GraphX completed graph construction, PageRank computation, and degree calculations considerably faster than GraphFrames. However, GraphFrames offered a higher-level programming interface, simpler integration with Spark SQL, and easier implementation of graph analytics workflows. The extracted graph measures also highlighted highly connected and structurally important patients within the network, demonstrating the usefulness of graph-based feature extraction for healthcare data analysis. GraphX completed graph construction approximately 33 times faster (92.5 s vs. 3037 s), PageRank computation approximately 780 times faster (9.1 s vs. 7132 s), and degree computation over 10,000 times faster (1.0 s vs. 10202 s) than GraphFrames, whereas GraphFrames used substantially less memory; these results indicate that GraphX is more suitable for performance-oriented graph workloads, whereas GraphFrames is advantageous when development flexibility and DataFrame integration are primary considerations.

References

M. Newman. “Measures and metrics”. In: Networks. 2nd ed., Ch. 7. Oxford University Press, Oxford, UK, 2018.

W. Raghupathi and V. Raghupathi. “Big data analytics in healthcare: Promise and potential”. Health Information Science and Systems, vol. 2, p. 3, 2014.

M. Zaharia, M. Chowdhury, T. Das, A. Dave, J. Ma, M. McCauly, M. J. Franklin, S. Shenker and I. Stoica. “Resilient distributed datasets: A Fault-Tolerant Abstraction for in-Memory Cluster Computing”. 9th USENIX Symposium on Networked Systems Design and Implementation, San Jose, CA, 2012.

R. S. Xin, Joseph E. Gonzalez, M. J. Franklin and I. Stoica. “GraphX: A Resilient Distributed Graph System on Spark”. Association for Computing Machinery, New York, 2013.

J. E. Gonzalez, R. S. Xin, A. Dave and D. Crankshaw. “GraphX: Graph Processing in a Distributed Dataflow Framework”. USENIX Association, Berkeley 2014.

A. Dave, A. Jindal, E. L. Li, R. Xin, J. Gonzalez and M. Zaharia. “GraphFrames: An integrated API for mixing graph and relational queries”. In: ACM International Conference Proceeding Series, Association for Computing Machinery, 2016.

R. K. Mishra, S. R. Raman. “GraphFrames”. In: PySpark SQL Recipes. Apress, Berkeley, CA, 2019.

K. Ammar and M. T. Özsu. “Experimental analysis of distributed graph systems”. Proceedings of the VLDB Endowment, vol. 11, no. 10, pp. 1151-1164, 2018.

S. Mary Arul, G. Senthil, S. Jayasudha, A. Alkhayyat, K. Azam and R. Elangovan. “Graph Theory and Algorithms for Network Analysis”. E3S Web of Conferences, vol. 399, p. 08002, 2023.

Y. Low, J. Gonzalez, A. Kyrola, D. Bickson, C. Guestrin and J. M. Hellerstein. “GraphLab: A New Framework for Parallel Machine Learning”. In: Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence (UAI), 2010, pp. 340-349.

J. E. Gonzalez, Y. Low, H. Gu, D. Bickson, and C. Guestrin.“PowerGraph: Distributed Graph-Parallel Computation on Natural Graphs”. In: Proceedings of the 10th USENIX Conference on Operating Systems Design and Implementation (OSDI’12), 2012, pp. 17-30.

M. Dayarathna and T. Suzumura. “Benchmarking Graph Data Management and Processing Systems: A Survey”. [arXiv Preprint], 2021.

M. Bin Hannan Siam, M. R. Khan, M. F. Elahe, M. S. Arman and S. Akter. “Effects of similarity networks in graph-based multi-omics classification”. PLoS One, vol. 21, no. 3, p. e0344754, 2026.

M. Zaharia, R. S. and Wendell, P. Das, T. Armbrust, M. Dave, A. Meng, X. Rosen, J. Venkataraman, S. Franklin, M. J. Ghodsi, A. Gonzalez, J. Shenker, S. S. Ion. “Apache spark: A unified engine for big data processing”. Communications of the ACM, vol. 59, no. 11, pp. 56-65, 2016.

H. Khayrolla Omar, A. K. Jumaa. “Distributed big data analysis using spark parallel data processing”. Bulletin of Electrical Engineering and Informatics, Vol 11, no 3, pp. 1505-1515, 2022.

S. Sahu, A. Mhedhbi, S. Salihoglu, J. Lin and M. T. Özsu. “The ubiquity of large graphs and surprising challenges of graph processing”. Proceedings of the VLDB Endowment, vol. 11, no. 4, pp. 420-431, 2017.

G. Malewicz, M. H. Austern, A. J. C. Bik, J. C. Dehnert, I. Horn, N. Leiser, and G. Czajkowski. “Pregel: A system for large-scale graph processing”. In: A. Elmagarmid and D. Agrawal, Eds. Proceedings of the 2010 ACM SIGMOD International Conference on Management of Data. New York, USA: ACM, 2010, pp. 135– 146.

M. Armbrust, R. S. Xin, C. Lian, Y. Huai, D. Liu, J. K. Bradley, X. Meng, T. Kaftan, M. J. Franklin, A. Ghodsi and M. Zaharia. “Spark SQL: Relational data processing in Spark”. In: Proceedings of the ACM SIGMOD International Conference on Management of Data, Association for Computing Machinery, pp. 1383-1394, 2015.

C. D. Manning, P. Raghavan and H. Schütze. “An Introduction to Information Retrieval”. Cambridge University Press, Cambridge, 2009.

S. Bonner, J. Brennan, G. Theodoropoulos, I. Kureshi and A. S. McGough. “GFP-X: A parallel approach to massive graph comparison using spark”. 2016 IEEE International Conference on Big Data (Big Data), Washington, DC, USA, 2016, pp. 3298-3307.

W. M. S. Mohammed and A. K. Jumaa. “An efficient approach to extract and store big semantic web data using hadoop and apache spark GraphX”. Advances in Distributed Computing and Artificial Intelligence Journal, vol. 13, e31506, 2024.

E. S. Apostol, A. C. Cojocaru and C. O. Truică. “Large-Scale Graphs Community Detection using Spark GraphFrames”. In: Proceedings 2024 23rd Proceedings of an International Symposium. Parallel and Distributed Computing (ISPDC), 2024, pp. 1-5.

L. Vyakaranam, R. Raman, H. Shah, N. Mishra, R. Meenakshi and S. Murugan. “Simplifying Graph-Parallel Computation in Apache Spark with GraphX”. In: 2024 8th International Conference on Electronics, Communication and Aerospace Technology (ICECA), IEEE, 2024, pp. 763-769.

P. R. Parvathy, J. Lenin and S. Mishra. “Implementing apache spark GraphX in big data using breadth-first search”. Innovations in Intelligent Systems and Advanced Engineering, 1(1), 19-28,2025.

L. Theodorakopoulos, A. Karras, A. Theodoropoulou and G. Kampiotis. “Benchmarking big data systems: Performance and Decision-making implications in emerging technologies”. Technologies, vol. 12, no. 11, p. 217, 2024.

Z. Lian, X. Lin, and L. Yin. “Optimizing apache spark for healthcare big data management”. Transactions on Computer Science and Intelligent Systems Research, vol. 7, pp. 113-118, 2024.

A. Teboul. “Diabetes Health Indicators Dataset (BRFSS 2015)”. Kaggle, San Francisco, 2021. Available from: https://www.kaggle. com/datasets/alexteboul/diabetes-health-indicators-dataset [Last accessed on 2026 May 25].

Published

2026-08-19

How to Cite

Hama, S. R., & Jumaa, A. K. (2026). A Comparative Analysis of Spark GraphX and GraphFrames for Healthcare Data Analytics. UHD Journal of Science and Technology, 10(2), 68–79. https://doi.org/10.21928/uhdjst.v10n2y2026.pp68-79

Issue

Section

Articles