Screening Vital Few Variables and Development of Logistic Regression Model on a Large Data Set

Lim, Yong-B.;Cho, J.;Um, Kyung-A;Lee, Sun-Ah;

Journal of Korean Society for Quality Management (품질경영학회지)

Volume 34 Issue 2
/
Pages.129-135
/
2006
/
1229-1889(pISSN)
/
2287-9005(eISSN)

Korean Society for Quality Management (한국품질경영학회)

Screening Vital Few Variables and Development of Logistic Regression Model on a Large Data Set

대용량 자료에서 핵심적인 소수의 변수들의 선별과 로지스틱 회귀 모형의 전개

Lim, Yong-B. (Department of Statistics, Ewha Womans) ;
Cho, J. (Department of Statistics, Ewha Womans) ;
Um, Kyung-A (Department of Statistics, Ewha Womans) ;
Lee, Sun-Ah (Department of Statistics, Ewha Womans)

임용빈 (이화여대 통계학과) ;
조재연 (이화여대 통계학과) ;
엄경아 (이화여대 통계학과) ;
이선아 (이화여대 통계학과)

Published : 2006.06.30

PDF KSCI

Download PDF

⟨ Previous Next ⟩

Abstract

In the advance of computer technology, it is possible to keep all the related informations for monitoring equipments in control and huge amount of real time manufacturing data in a data base. Thus, the statistical analysis of large data sets with hundreds of thousands observations and hundred of independent variables whose some of values are missing at many observations is needed even though it is a formidable computational task. A tree structured approach to classification is capable of screening important independent variables and their interactions. In a Six Sigma project handling large amount of manufacturing data, one of the goals is to screen vital few variables among trivial many variables. In this paper we have reviewed and summarized CART, C4.5 and CHAID algorithms and proposed a simple method of screening vital few variables by selecting common variables screened by all the three algorithms. Also how to develop a logistics regression model on a large data set is discussed and illustrated through a large finance data set collected by a credit bureau for th purpose of predicting the bankruptcy of the company.

Keywords

References

강현철 등(1999), '데이터마이닝, 방법론 및 활용', 자유아카데미
임용빈, 오만숙(2002), '분류와 회귀나무 분석에 관한 소고', '품질경영학회지', 30권, 1호, pp. 152-161
허명회, 이용구(2003), '데이터마이닝 모델링과 사례', SPSS 아카데미
Abt, M., Lim, Y. B., Sacks, J., Xie, M., and Young, S.(2001), 'A sequential approach for identifying lead compounds in large chemical databases', Statistical Science, Vol. 16, No. 2, pp. 154-168 https://doi.org/10.1214/ss/1009213288
Breiman, L., Friedman, J. H., Olshen, R. A., and Stone, C. J.(1984), Classification and regression trees, Chapman and Hall, Belmont, CA, Wadsworth
Kass, G.(1980), 'An exploratory technique for investigating large quantities of categorical data', Applied Statistics, Vol. 29, pp, 119-127 https://doi.org/10.2307/2986296
Quinlan, J. R.(1993), C4.5 Programs for machine learning, San Mateo: Morgan Kaufmann

Journal of Korean Society for Quality Management (품질경영학회지)

Screening Vital Few Variables and Development of Logistic Regression Model on a Large Data Set

대용량 자료에서 핵심적인 소수의 변수들의 선별과 로지스틱 회귀 모형의 전개

Abstract

Keywords

References

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)