Data Quality Assurance Methods for Integrating Heterogeneous Sources in High-Load Analytical Systems

Authors

  • Anton Aleksandrovych

Keywords:

anomaly detection; data integration; data quality; deduplication; entity resolution; ETL; heterogeneous sources; schema validation

Abstract

Modern analytical platforms consolidate information from many independent sources, and the value of every downstream report depends on whether the consolidated data can be trusted. Integration across systems with their own storage rules, update cycles, and identifiers produces incompleteness, duplication, inconsistent reference catalogs, divergent value formats, and broken relationships between objects. This study systematizes methods for addressing data quality during the integration of heterogeneous sources into high-load analytical systems, drawing on a practice-based view of a multi-layer cloud data platform that processes multi-million-record volumes in a regulated domain. The analysis organizes quality control into a layered model in which schema validation, value standardization, deduplication, entity resolution across systems, and automated anomaly detection operate as connected stages that feed a single source of truth. Each method is mapped to the class of quality problem it resolves, the pipeline stage at which it acts, and the effect it has on consistency and reproducibility. The work shows that a layered, reproducible control process improves cross-source consistency, reduces the reliance on manual verification, and shifts analysts' roles from collection and reconciliation to interpretation and decision-making. The systematization offers data engineers a transferable reference for designing quality assurance in large-scale integration scenarios where regulatory monitoring and risk assessment require dependable, auditable data.

References

[1] W. Z. Alma'aitah, A. Quraan, F. N. AL-Aswadi, R. S. Alkhawaldeh, M. Alazab, and A. Awajan, “Integration approaches for heterogeneous big data: A survey,” Cybernetics and Information Technologies, vol. 24, no. 1, pp. 3–20, Mar. 2024, doi: 10.2478/cait-2024-0001.

[2] G. Fusco and L. Aversano, “An approach for semantic integration of heterogeneous data sources,” PeerJ Computer Science, vol. 6, Mar. 2020, Art. no. e254, doi: 10.7717/peerj-cs.254.

[3] V. Wenz, A. Kesper, and G. Taentzer, “Clustering heterogeneous data values for data quality analysis,” ACM Journal of Data and Information Quality, vol. 15, no. 3, pp. 1–33, 2023, doi: 10.1145/3603710.

[4] C. Koutras et al., “Valentine: Evaluating matching techniques for dataset discovery,” in Proc. 2021 IEEE 37th International Conference on Data Engineering (ICDE), Chania, Greece, 2021, pp. 468–479, doi: 10.1109/ICDE51399.2021.00047.

[5] L. Dinesh and K. G. Devi, “An efficient hybrid optimization of ETL process in data warehouse of cloud architecture,” Journal of Cloud Computing, vol. 13, 2024, Art. no. 12, doi: 10.1186/s13677-023-00571-y.

[6] G. Papadakis, D. Skoutas, E. Thanos, and T. Palpanas, “Blocking and filtering techniques for entity resolution: A survey,” ACM Computing Surveys, vol. 53, no. 2, pp. 1–42, 2020, doi: 10.1145/3377455.

[7] S. Thirumuruganathan et al., “Deep learning for blocking in entity matching: A design space exploration,” Proceedings of the VLDB Endowment, vol. 14, no. 11, pp. 2459–2472, 2021, doi: 10.14778/3476249.3476294.

[8] Y. Zhang et al., “Schema matching using pre-trained language models,” in Proc. 2023 IEEE 39th International Conference on Data Engineering (ICDE), Anaheim, CA, USA, 2023, pp. 1558–1571, doi: 10.1109/ICDE55515.2023.00123.

[9] Q. Chen et al., “Adaptive deep learning for entity resolution by risk analysis,” Knowledge-Based Systems, vol. 260, Jan. 2023, Art. no. 110118, doi: 10.1016/j.knosys.2022.110118.

[10] Y. Nafa et al., “Active deep learning on entity resolution by risk sampling,” Knowledge-Based Systems, vol. 236, 2022, Art. no. 107729, doi: 10.1016/j.knosys.2021.107729.

Downloads

Published

2026-09-09

Issue

Section

Articles

How to Cite

Aleksandrovych, A. . (2026). Data Quality Assurance Methods for Integrating Heterogeneous Sources in High-Load Analytical Systems. International Journal of Computer (IJC), 57(1), 635-647. https://www.ijcjournal.org/InternationalJournalOfComputer/article/view/2572