• Title/Summary/Keyword: 주성분 분석(PCA)

Search Result 655, Processing Time 0.018 seconds

Feature Selection with Non-linear PCA in Text Categorization (대용량 문서분류에서의 비선형 주성분 분석을 이용한 특징 추출)

  • 신형주;장병탁;김영택
    • Proceedings of the Korean Information Science Society Conference
    • /
    • 1999.10b
    • /
    • pp.146-148
    • /
    • 1999
  • 문서분류의 문제점 중의 하나는 사용하는 데이터의 차원이 매우 크다는 것이다. 그러므로 문서에서 필요한 단어만을 자동적으로 추출하여 문서데이터의 차원을 축소하는 작업이 문서분류에서는 필수적이다. DF(Document Frequency)는 문서의 차원축소의 대표적인 통계적 방법 중 하나인데, 본 논문에서는 문서의 차원축소에 DF와 주성분 분석(PCA)을 비교하여 주성분 분석이 문서의 차원축소에 적합함을 실험적으로 보인다. 그리고 비선형 주성분 분석(nonlinear PCA) 방법 중 locally linear PCA와 kenel PCA를 적용하여 비선형 주성분 분석을 이용하여 문서의 차원을 줄이는 것이 선형 주성분 분석을 이용하는 것 보다 문서분류에 더 적합함을 실험적으로 보인다.

  • PDF

Principal component analysis in the frequency domain: a review and their application to climate data (주파수공간에서의 주성분분석: 리뷰와 기상자료에의 적용)

  • Jo, You-Jung;Oh, Hee-Seok;Lim, Yaeji
    • The Korean Journal of Applied Statistics
    • /
    • v.30 no.3
    • /
    • pp.441-451
    • /
    • 2017
  • In this paper, we review principal component analysis (PCA) procedures in the frequency domain and apply them to analyze sea surface temperature data. The classical PCA defined in the time domain is a popular dimension reduction technique. Extending the conventional PCA to the frequency domain makes it possible to define PCA in the frequency domain, which is useful for dimension reduction as well as a feature extraction of multiple time series. We focus on two PCA methods in the frequency domain, Hilbert PCA (HPCA) and frequency domain PCA (FDPCA). We review these two PCAs in order for potential readers to easily understand insights as well as perform a numerical study for comparison with conventional PCA. Furthermore, we apply PCA methods in the frequency domain to sea surface temperature data on the tropical Pacific Ocean. Results from numerical experiments demonstrate that PCA in the frequency domain is effective for the analysis of time series data.

A Way of Securing the Access By Using PCA (주성분분석(PCA)을 이용한 출입인원관리에 대한 보안성 확보 방안)

  • Kim, Min-Su;Lee, Dong-Hwi
    • Convergence Security Journal
    • /
    • v.12 no.3
    • /
    • pp.3-10
    • /
    • 2012
  • This study aimed at making a way of securing the access by using PCA. We got our result through using Box-Plot and PCA with the access data of the area of security level A~E at K(IPS)center. In order to perform PCA, We confirmed the extracted value of commonality has no problem in performing PCA because VIF is below 2.902. Based on this result, We classified people into Green-list, Blue-list, Red-list, and Black-list in a standard of security level with 1.453, as the eigen value of 1 main element, 1.283, as eigen value of 2 main elementm, 1.142, as the eigen value of 3 main element.

Principal Component Analysis with Coefficient of Variation Matrix (변동계수행렬을 이용한 주성분분석)

  • Kim, Ji-Hyun
    • The Korean Journal of Applied Statistics
    • /
    • v.28 no.3
    • /
    • pp.385-392
    • /
    • 2015
  • Principal component analysis (PCA), a dimension-reduction technique, is usually implemented after the variables are standardized when the measurement unit of variables are different. To standardize a variable we divide it by its standard deviation. But there is another way to transform a variable to be independent of its measurement unit. It is to divide it by its mean rather than standard deviation. Implementing PCA on standardized variables is equivalent to implementing PCA with a correlation matrix of original variables. Similarly, implementing PCA on the transformed variables divided by their means is equivalent to implementing PCA with a matrix related to the coefficients of variation of the original variables. We explain why we need to implement PCA on the variables transformed by their means.

A Non-linear Variant of Improved Robust Fuzzy PCA (잡음 민감성이 향상된 주성분 분석 기법의 비선형 변형)

  • Heo, Gyeong-Yong;Seo, Jin-Seok;Lee, Im-Geun
    • Journal of the Korea Society of Computer and Information
    • /
    • v.16 no.4
    • /
    • pp.15-22
    • /
    • 2011
  • Principal component analysis (PCA) is a well-known method for dimensionality reduction and feature extraction while maintaining most of the variation in data. Although PCA has been applied in many areas successfully, it is sensitive to outliers and only valid for Gaussian distributions. Several variants of PCA have been proposed to resolve noise sensitivity and, among the variants, improved robust fuzzy PCA (RF-PCA2) demonstrated promising results. RF-PCA, however, is still a linear algorithm that cannot accommodate non-Gaussian distributions. In this paper, a non-linear algorithm that combines RF-PCA2 and kernel PCA (K-PCA), called improved robust kernel fuzzy PCA (RKF-PCA2), is introduced. The kernel methods make it to accommodate non-Gaussian distributions. RKF-PCA2 inherits noise robustness from RF-PCA2 and non-linearity from K-PCA. RKF-PCA2 outperforms previous methods in handling non-Gaussian distributions in a noise robust way. Experimental results also support this.

Utilizing UPCA and SPCA in Unsupervised Classification Using Landsat TM data

  • Lee, Byung-Gul;Kang, In-Joon
    • Proceedings of the Korean Society of Surveying, Geodesy, Photogrammetry, and Cartography Conference
    • /
    • 2003.04a
    • /
    • pp.167-170
    • /
    • 2003
  • 본 연구는 무감독영상해석(Unsupervised Classification)에서 주성분 분석법(Principal Component Analysis)의 응용성을 연구하기 위하여, 주성분 분석법을 K-means, ISODATA 두가지 무감독분류법에 적용하였다. 적용대상지역은 제주도이다. 본 연구에서 주성분 분석 방법중에서 비정규형 주성분 분석방법 (Unstandardized PCA)과 정규형 주성분 분석방법(Standardized PCA) 두가지 경우로 나누어서 각각 연구하였다. 이를 위하여 제주도의 Landsat TM영상과 국토연구원에서 조사한 제주도 식생분류 조사자료와 현장조사 자료 그리고 1/25,000 수치지도를 이용하였다. 그리고 분석된 자료의 정확도를 평가하기 위하여 오차행렬(Error Matrix)을 도입하여 계산하였다. 우선 비정규형 주성분 분석법으로 구한 주성분 영상과 Landsat TM 원래 영상을 오차행렬을 이용하여 제주도의 식생 분류에 각각 적용하였다. 그 결과, K-means 무감독분류법에서는 Landsat TM 자료를 직접 이용한 경우에는 바다와 육상의 분류가 잘 되지 않았으며, 또한 전반적인 영상분류결과가 관측치와 많은 차이를 보였다. 그러나, 주성분 분석법으로 계산된 주성분 영상으로 K-means방법으로 분류 한 결과는 관측치와 잘 일치를 하였다. ISODATA의 경우, Landsat TM 원래영상을 계산하면, K-means으로 분류한 결과보다는 좋은 값을 나타냈으나, 주성분 분석법으로 구한 영상의 계산결과와 비교하면, 주성분 영상으로 구한 분류결과의 정확도가 약 15%정도 높게 나타났다. 정규형 주성분 분석법의 경우를 보면 K-means에서는 Landsat TM원래 자료보다 우수한 결과를 보여주었으나, 비정규형 주성분 분석법으로 계산된 결과보다는 정확도가 다소 떨어지는 단점이 있었고, ISODATA의 경우도 Landsat TM원래 자료보다 약 7%정도의 높은 정확도를 보였으나, 비정규형 영상보다는 약8%정도 낮은 정확도를 보였다. 본 연구에서 주성분 분석법으로 계산된 결과에서 주목되는 것은, 주성분 분석법으로 구한 주성분 영상은 분류방법(K-means, ISODATA, artificial neural networks)에 따라 분류된 결과값이 비슷하게 나타난 반면, Landsat TM원래 자료는 분류방법에 따라 결과값이 많은 차이를 보여 주었다. 그리고 주성분 분석 방법 중에서도 비정규형 주성분 분석법(Unstandardized PCA)이 정규형 주성분 분석법(Standardized PCA)보다 영상분석에서 더 좋은 결과를 보여주는 것으로 나타났다.

  • PDF

An Improved Robust Fuzzy Principal Component Analysis (잡음 민감성이 개선된 퍼지 주성분 분석)

  • Heo, Gyeong-Yong;Woo, Young-Woon;Kim, Seong-Hoon
    • Journal of the Korea Institute of Information and Communication Engineering
    • /
    • v.14 no.5
    • /
    • pp.1093-1102
    • /
    • 2010
  • Principal component analysis (PCA) is a well-known method for dimension reduction while maintaining most of the variation in data. Although PCA has been applied to many areas successfully, it is sensitive to outliers. Several variants of PCA have been proposed to resolve the problem and, among the variants, robust fuzzy PCA (RF-PCA) demonstrated promising results. RF-PCA uses fuzzy memberships to reduce the noise sensitivity. However, there are also problems in RF-PCA and the convergence property is one of them. RF-PCA uses two different objective functions to update memberships and principal components, which is the main reason of the lack of convergence property. The difference between two functions also slows the convergence and deteriorates the solutions of RF-PCA. In this paper, a variant of RF-PCA, called RF-PCA2, is proposed. RF-PCA2 uses an integrated objective function both for memberships and principal components. By using alternating optimization, RF-PCA2 is guaranteed to converge on a local optimum. Furthermore, RF-PCA2 converges faster than RF-PCA and the solutions found are more similar to the desired solutions than those of RF-PCA. Experimental results also support this.

Efficient Primary-Ambient Decomposition Algorithm for Audio Upmix (오디오 업믹스를 위한 효율적인 주성분-주변성분 분리 알고리즘)

  • Baek, Yong-Hyun;Jeon, Se-Woon;Lee, Seok-Pil;Park, Young-Cheol
    • Journal of Broadcast Engineering
    • /
    • v.17 no.6
    • /
    • pp.924-932
    • /
    • 2012
  • Decomposition of a stereo signal into the primary and ambient components is a key step to the stereo upmix and it is often based on the principal component analysis (PCA). However, major shortcoming of the PCA-based method is that accuracy of the decomposed components is dependent on both the primary-to-ambient power ratio (PAR) and the panning angle. Previously, a modified PCA was suggested to solve the PAR-dependent problem. However, its performance is still dependent on the panning angle of the primary signal. In this paper, we proposed a new PCA-based primary-ambient decomposition algorithm whose performance is not affected by the PAR as well as the panning angle. The proposed algorithm finds scale factors based on a criterion that is set to preserve the powers of the mixed components, so that the original primary and ambient powers are correctly retrieved. Simulation results are presented to show the effectiveness of the proposed algorithm.

A study on the design of fault diagnostic system based on PCA (PCA-기반 고장 진단 시스템 설계에 관한 연구)

  • Kim, Sung-Ho;Lee, Young-Sam;Han, Yoon-Jong
    • Journal of the Korean Institute of Intelligent Systems
    • /
    • v.13 no.5
    • /
    • pp.600-605
    • /
    • 2003
  • PCA(Principle Component Analysis) has emerged as a useful tool for process monitoring and fault diagnosis. The general approach requires the user to identify the root cause by interpreting the residual or principle components. This could be tedious and often impossible for a large process. In this paper, PCA scheme is combined with the FCM-based fault diagnostic algorithm to enhance the diagnostic results. The implementation of the FCM-based fault diagnostic system by using PCA is done and its application is illustrated on the two-tank system.

Numerical taxonomy of Rhus sensu lato (Anacardiaceae) in Korea (한국산 광의의 붉나무속(Rhus L. sensu lato)의 수리분류학적 연구)

  • Tho, Jae-Hwa;Kim, Joo-Hwan
    • Korean Journal of Plant Taxonomy
    • /
    • v.34 no.3
    • /
    • pp.205-220
    • /
    • 2004
  • Numerical analysis based on the 67 morphological characters from 28 populations of 6 species of Korean Rhus sensu lato (Anacardiaceae) was performed for the taxonomic delimitation. Based on the results of PCA with 47 quantitative characters, the sum of contributions for the total variance of three major principal components was 77,9% (PCl 35.2%, PC2 22.5% and PC3 20.2%). The sum of contributions for the total variance of three major principal components were 90,7% (PCl 37.7%, PC2 33.0% and PC3 20.0%) based on the results of PCA with 20 qualitative The characters. Two dimensional plotting from PCA results recognized six distinct species. UPGMA phenogram based on simple matching coefficient method recognized clear taxonomic delimitations among six taxa. On the cluster analysis, qualitative characters were more useful for grouping the species treated. Numerical analysis was very valuable to delimit the Korean taxa of Rhus s.l.