DOI QR코드

DOI QR Code

An Efficient Method for Imputing Missing Values in Incomplete Process Data from High-Cost Data Acquisition Environments

고비용 공정데이터 획득 환경에서 불완비 공정데이터의 효율적인 결측치 대체 방법

  • Received : 2025.11.25
  • Accepted : 2025.12.11
  • Published : 2025.12.31

Abstract

This study addresses the challenge of imputing missing values in incomplete process data collected from high-cost data acquisition environments. Such missingness arises due to insufficient completeness, accuracy, and consistency, which significantly affect the quality of critical-to-quality (CTQ) attributes in manufacturing processes. We systematically evaluate three state-of-the-art imputation methods-Multiple Imputation by Chained Equations (MICE), the machine learning-based missForest algorithm, and a deep learning-based one-dimensional convolutional neural network (1D-CNN)-using real-world industrial data. Our analysis aims to identify the most effective imputation technique for handling complex and noisy process datasets typical in manufacturing settings. The results highlight the strengths and limitations of each method, providing practical guidance for selecting appropriate imputation approaches to improve the reliability of quality prediction and decision-making in industrial applications.

Keywords

Acknowledgement

This work was supported by the Starting growth Technological R&D Program (TIPS Program, (No. RS-2025- 25440393)) funded by the Ministry of SMEs and Startups(MSS, Korea) in 2025.

References

  1. Alshammari, M.S. and Alharbi, M.A., Forecasting the required quantity of cement manufacturing materials using time series and Q-network techniques, Ecological Chemistry and Engineering, 2025, Vol. 32, No. 2, pp. 323-336. https://doi.org/10.2478/eces-2025-0021.
  2. Azur, M.J., Stuart, E.A., Frangakis, C., and Leaf, P.J., Multiple imputation by chained equations: what is it and how does it work?, International Journal of Methods in Psychiatric Research, 2011, Vol. 20, No.1, pp.40–49. https://doi.org/10.1002/mpr.329
  3. Bianchi, A.D., MissForest Algorithm for Income and Living Conditions Survey Imputation, Swiss Statistics Series, 2022.
  4. Cho, B.S., Ahn, J.C., and Park, D.C., Rheological evaluation of blast furnace slag cement pastes over setting time, Journal of the Korea Institute of Building Construction, 2016, Vol. 16, No. 6, pp. 505–512. https://doi.org/10.5345/JKIBC.2016.16.6.505
  5. Fichman, M. and Cummings, J.N., Multiple Imputation for Missing Data in Social Research, Carnegie Mellon University, 1999.
  6. Graham, J.W., Missing Data: Analysis and Design. Statistics for Social and Behavioral Sciences[series], New York, USA : Springer, 2012.
  7. Hudes, E. and Neilands, T., Reconstruction of Missing Daily Streamflow Data Using MissForest, Journal of Water and Marine Research, 2022, Vol. 15, No. 2, pp. 49-55. https://doi.org/10.61186/jwmr.15.2.49
  8. Hwang, B.I., Kang, S.P., and Kim, S.J., Study on the strength development factors of alkali-activated slag binder, J. Korea Inst. Resour. Recycl., 2018, Vol. 27, No. 3, pp. 35-42. https://doi.org/10.7844/KIRR.2018.27.3.35
  9. Joel, L.O., Doorsamy, W., and Paul, B.S., Imputation Techniques Performance on Healthcare Data. arXiv:2403.14687v1, 2024.
  10. Kim, B.S., Choi, S.M., and Kim, J.M., Fundamental properties of mortar using magnetically separated basic oxygen furnace (BOF) slag powder as binder, Journal of the Korean Recycled Construction Resources Institute, 2023, Vol. 11, No. 3, pp. 168-176. https://doi.org/10.9721/JKRCRI.2023.11.3.168
  11. Kim, J.M., Choi, S.M., and Kim, J.H., Evaluation of applicability of ladle furnace slag (LFS) produced from various manufacturing processes as construction materials, Journal of the Korea Concrete Institute, 2012, Vol. 24, No. 6, pp. 695-703. https://doi.org/10.7854/JKCI.2012.24.6.695
  12. Kim, J.M., Kwak, E.G., Choi, S.M., Kim, J.H., Lee, W.Y., and Oh, S.Y., Properties of mortar according to gradation change of electric arc furnace oxidizing slag fine aggregate made by rapidly cooled method, Journal of the Korea Recycled Construction Resources Institute, 2011, Vol. 6, No. 4, pp. 112-118. https://doi.org/10.7854/JKRCRI.2011.6.4.112
  13. Kim, T.W. and Kang, C.H., The influence of Al₂O₃ on the properties of alkali-activated slag cement, Journal of the Korea Concrete Institute, 2016, Vol. 28, No. 2, pp. 243-251. https://doi.org/10.7844/JKCI.2016.28.2.243
  14. LeCun, Y., Bengio, Y., and Hinton, G., Deep learning, Nature, 2015, Vol. 521, No. 7553, pp. 436-444. https://doi.org/10.1038/nature14539
  15. Lee, S.R., Comparison of Algorithms for the Missing data Imputation Methods [Master's thesis]. [Seoul, Korea]: Hankook University of Foreign Studies, 2019.
  16. Little, R.J.A. and Rubin, D.B., Incomplete data. Methods and Applications of Statistical in the Life and Health Science [Book Chapter], John Wiley & Sons, Inc, 2014, pp. 441-449.
  17. Lowke, D., Gehlen, C., Plank, J., Pott, U., and Seidel, A., Concrete 4.0—Sustainable concrete construction with digital quality control, CE/Papers, 2023, Vol. 6, No. 5, pp. 976-982. https://doi.org/10.1002/cepa.2965
  18. Luo, Y., Szolovits, P., Dighe, A.S., and Baron, J.M., 3D‑MICE: Imputation for Longitudinal Clinical Data. Journal of the American Medical Informatics Association, 2018, Vol.25, No.6, pp. 645-653. https://doi.org/10.1093/jamia/ocx133
  19. Mengesha, K.F. and Mehari, Y., Advances in statistical quality control chart techniques and their limitations to cement industry, Cogent Engineering, 2022, Vol. 9, No. 1, Article 2088463. https://doi.org/10.1080/23311916.2022.2088463
  20. Mishra, R., Wang, S., Tao, Y., and Monteiro, P.J.M., Industrial-scale prediction of cement clinker phases using machine learning, arXiv preprint, 2024, https://arxiv.org/abs/2412.11981
  21. Park, S.S., Kang, H.Y., and Han, K.S., Development of fly ash/slag cement using alkali activation (I) Compressive strength and acid resistance, Journal of Korean Society of Environmental Engineers, 2007, Vol. 29, No. 7, pp. 801-809.
  22. Rubin, D.B., Multiple imputation for nonresponse in surveys, 3rd ed., New York, USA : John Wiley & Sons. 2019.
  23. Samad, M.D., Abrar, S., and Diawara, N., Missing Value Imputation with Clustering and Deep Learning, Knowledge-Based Systems, 2022, No. 249, 108968.
  24. Schafer, J.L., Analysis of Incomplete Multivariate Data, New Yor, USA: Champman & Hall/CRC, 1997. https://doi.org/10.1201/9780367803025
  25. Shadbahr, T., Roberts, M., Stanczuk, J., Gilbey, J., Teare, P., et al., The impact of imputation quality on machine learning classifiers for datasets with missing values. Communications Medicine, 2023, Vol. 3, No. 139. https://doi.org/10.1038/s43856-023-00356-z
  26. Stekhoven, D.J. and Bühlmann, P., MissForest: nonparametric missing value imputation for mixed-type data. Bioinformatics, 2012, Vol. 28, No. 1, pp. 112-118. https://doi.org/10.1093/bioinformatics/btr597
  27. Sun, Y., Li, J., Xu, Y., Zhang, T., and Wang, X., Deep learning versus conventional methods for missing data imputation: A review and comparative study, Expert Systems with Applications, Vol.227, 2023, https://doi.org/10.1016/j.eswa.2023.120201.
  28. van Buuren, S. and Groothuis-Oudshoorn, K., MICE: Multivariate Imputation by Chained Equations in R, Journal of Statistical Software, 2011. Vol. 45, No.3, pp.1-67. https://doi.org/10.18637/jss.v045.i03
  29. van Buuren, S., Flexible Imputation of Missing Data, 2nd ed, Boca Raton, FL, USA : Chapman & Hall/CRC Press, 2018.
  30. Wang, Z., Akande, O., Poulos, J., and Li, F., Are deep learning models superior for missing data imputation in surveys? Evidence from an Empirical Comparison, arXiv: 2103.09316, 2021.
  31. White, T.K., Reiter, J.P., and Petrin, A., Plant‑Level Productivity and Missing Data Imputation in U.S. Census Manufacturing Data, NBER Working Paper 17816, 2012.
  32. Yoon, J., Jordon, J., and Schaar, M., Gain: Missing data imputation using generative adversarial nets, In International Conference on Machine Learning, 2018, pp. 5689-5698. PMLR.