• 제목/요약/키워드: Feature Selection Methods

검색결과 318건 처리시간 0.028초

양자 유전알고리즘을 이용한 특징 선택 및 성능 분석 (Feature Selection and Performance Analysis using Quantum-inspired Genetic Algorithm)

  • 허기수;정현태;박아론;백성준
    • 스마트미디어저널
    • /
    • 제1권1호
    • /
    • pp.36-41
    • /
    • 2012
  • 특징 선택은 패턴 인식의 성능을 향상시키기 위해 부분집합을 구성하는 중요한 문제다. 특징 선택에는 순차 탐색 알고리즘으로부터 확률 기반의 유전 알고리즘까지 다양한 접근 방법이 적용 되었다. 본 연구에서는 특징 선택을 위해 양자 비트, 상태의 중첩 등 양자 컴퓨터 개념을 기반으로 하는 양자 기반 유전 알고리즘(QGA: Quantum-inspired Genetic Algorithm)을 적용하였다. QGA 성능은 전통적인 유전 알고리즘(CGA: Conventional Genetic Algorithm)을 적용한 특징 선택 방법과 분류율 및 평균 특징 개수의 비교를 통해 이루어졌으며, UCI 데이터를 이용한 실험 결과 QGA를 적용한 특징 선택 방법이 CGA를 적용한 경우에 비해 전반적으로 좋은 성능을 보임을 확인 할 수 있었다.

  • PDF

Set Covering 기반의 대용량 오믹스데이터 특징변수 추출기법 (Set Covering-based Feature Selection of Large-scale Omics Data)

  • 마정우;안기동;김광수;류홍서
    • 한국경영과학회지
    • /
    • 제39권4호
    • /
    • pp.75-84
    • /
    • 2014
  • In this paper, we dealt with feature selection problem of large-scale and high-dimensional biological data such as omics data. For this problem, most of the previous approaches used simple score function to reduce the number of original variables and selected features from the small number of remained variables. In the case of methods that do not rely on filtering techniques, they do not consider the interactions between the variables, or generate approximate solutions to the simplified problem. Unlike them, by combining set covering and clustering techniques, we developed a new method that could deal with total number of variables and consider the combinatorial effects of variables for selecting good features. To demonstrate the efficacy and effectiveness of the method, we downloaded gene expression datasets from TCGA (The Cancer Genome Atlas) and compared our method with other algorithms including WEKA embeded feature selection algorithms. In the experimental results, we showed that our method could select high quality features for constructing more accurate classifiers than other feature selection algorithms.

Deep Learning Method for Identification and Selection of Relevant Features

  • Vejendla Lakshman
    • International Journal of Computer Science & Network Security
    • /
    • 제24권5호
    • /
    • pp.212-216
    • /
    • 2024
  • Feature Selection have turned into the main point of investigations particularly in bioinformatics where there are numerous applications. Deep learning technique is a useful asset to choose features, anyway not all calculations are on an equivalent balance with regards to selection of relevant features. To be sure, numerous techniques have been proposed to select multiple features using deep learning techniques. Because of the deep learning, neural systems have profited a gigantic top recovery in the previous couple of years. Anyway neural systems are blackbox models and not many endeavors have been made so as to examine the fundamental procedure. In this proposed work a new calculations so as to do feature selection with deep learning systems is introduced. To evaluate our outcomes, we create relapse and grouping issues which enable us to think about every calculation on various fronts: exhibitions, calculation time and limitations. The outcomes acquired are truly encouraging since we figure out how to accomplish our objective by outperforming irregular backwoods exhibitions for each situation. The results prove that the proposed method exhibits better performance than the traditional methods.

FAFS: A Fuzzy Association Feature Selection Method for Network Malicious Traffic Detection

  • Feng, Yongxin;Kang, Yingyun;Zhang, Hao;Zhang, Wenbo
    • KSII Transactions on Internet and Information Systems (TIIS)
    • /
    • 제14권1호
    • /
    • pp.240-259
    • /
    • 2020
  • Analyzing network traffic is the basis of dealing with network security issues. Most of the network security systems depend on the feature selection of network traffic data and the detection ability of malicious traffic in network can be improved by the correct method of feature selection. An FAFS method, which is short for Fuzzy Association Feature Selection method, is proposed in this paper for network malicious traffic detection. Association rules, which can reflect the relationship among different characteristic attributes of network traffic data, are mined by association analysis. The membership value of association rules are obtained by the calculation of fuzzy reasoning. The data features with the highest correlation intensity in network data sets are calculated by comparing the membership values in association rules. The dimension of data features are reduced and the detection ability of malicious traffic detection algorithm in network is improved by FAFS method. To verify the effect of malicious traffic feature selection by FAFS method, FAFS method is used to select data features of different dataset in this paper. Then, K-Nearest Neighbor algorithm, C4.5 Decision Tree algorithm and Naïve Bayes algorithm are used to test on the dataset above. Moreover, FAFS method is also compared with classical feature selection methods. The analysis of experimental results show that the precision and recall rate of malicious traffic detection in the network can be significantly improved by FAFS method, which provides a valuable reference for the establishment of network security system.

특징 추출과 분석 기법에 기반한 단백질 상호작용 데이터 신뢰도 향상 시스템 (Protein-Protein Interaction Reliability Enhancement System based on Feature Selection and Classification Technique)

  • 이민수;박승수;이상호;용환승;강성희
    • 정보처리학회논문지B
    • /
    • 제13B권7호
    • /
    • pp.679-688
    • /
    • 2006
  • 대용량 실험으로부터 산출된 단백질 상호작용 데이터는 위양성(false positive) 데이터의 비율이 높다는 단점을 가지고 있다. 본 논문에서는 오류가 섞여있는 단백질 상호작용 데이터를 입력으로 받아 각 단백질 상호작용의 신뢰도를 검증하는 시스템을 제안하고 구현하였다. 제안 시스템은 단백질 상호작용 데이터에 상호작용의 근거로서 사용될 수 있는 다양한 생물학적 특징들에 관한 데이터를 통합하고 특징 선택 방법을 사용하여 통합된 속성들 중 위양성 여부를 판별하는데 가장 적합한 특징들을 선택한 후 데이터 마이닝 분류 알고리즘을 적용하여 대용량 실험으로부터 산출된 단백질 상호작용 데이터의 신뢰도를 평가한다. 특징 선택의 결과와 분류 기법의 성능은 데이터 특성에 매우 의존하므로, 제안시스템에 가장 적합한 속성 부분집합과 가장 좋은 성능을 내는 분류 알고리즘을 찾기 위해 다양한 특징 선택 방법과 데이터 마이닝 분류 알고리즘들을 적용하고 그 성능을 다각적으로 비교분석 하였다. 실험 결과, 특징 선택 방법과 분류 알고리즘을 결합시킨 제안 시스템은 오류 데이터가 섞여있는 단백질 상호작용 데이터에서 실제로 상호작용하는 단백질 쌍을 골라내는 작업에 있어 기존 연구들에 비해 매우 뛰어난 성능을 보여줬다. 또한 본 연구를 통해 단백질 상호작용 데이터의 신뢰도를 검증함에 있어서 다양한 특징 선택 방법들과 분류 알고리즘들이 성능에 미치는 영향에 관해서도 정리할 수 있었다.

단백체 스펙트럼 데이터의 분류를 위한 랜덤 포리스트 기반 특성 선택 알고리즘 (Feature Selection for Classification of Mass Spectrometric Proteomic Data Using Random Forest)

  • 온승엽;지승도;한미영
    • 한국시뮬레이션학회논문지
    • /
    • 제22권4호
    • /
    • pp.139-147
    • /
    • 2013
  • 본 논문에서는 질량 분석 방법에 의하여 산출된 단백체 데이터(mass spectrometric proteomic data)의 분류 분석(classification analysis)을 위한 새로운 특성 선택(feature selection) 방법을 제안한다. 이 방법은 i)높은 상관관계를 가지는 중복된 특성을 효과적으로 제거하는 전처리 단계와 ii)토너먼트(tournament) 전략을 사용하여 최적 특성 부분집합(optimal feature subset)을 탐색해 내는 단계로 구성되어 있다. 제안되는 방법을 실제 암진단에 사용되는 공개된 혈액 단백체 데이터에 적용하였으며 널리 사용되는 타 방법과 비교할 때 우수한 성능과 균형된 특이도와 민감도를 달성함을 실증하였다.

Language- Independent Sentence Boundary Detection with Automatic Feature Selection

  • Lee, Do-Gil
    • Journal of the Korean Data and Information Science Society
    • /
    • 제19권4호
    • /
    • pp.1297-1304
    • /
    • 2008
  • This paper proposes a machine learning approach for language-independent sentence boundary detection. The proposed method requires no heuristic rules and language-specific features, such as part-of-speech information, a list of abbreviations or proper names. With only the language-independent features, we perform experiments on not only an inflectional language but also an agglutinative language, having fairly different characteristics (in this paper, English and Korean, respectively). In addition, we obtain good performances in both languages. We have also experimented with the methods under a wide range of experimental conditions, especially for the selection of useful features.

  • PDF

Elastic net 기반 특징 선택을 적용한 fNIRS 기반 뇌-컴퓨터 인터페이스 데이터셋 분류 정확도 평가 (Assessment of Classification Accuracy of fNIRS-Based Brain-computer Interface Dataset Employing Elastic Net-Based Feature Selection)

  • 신재영
    • 대한의용생체공학회:의공학회지
    • /
    • 제42권6호
    • /
    • pp.268-276
    • /
    • 2021
  • Functional near-infrared spectroscopy-based brain-computer interface (fNIRS-based BCI) has been receiving much attention. However, we are practically constrained to obtain a lot of fNIRS data by inherent hemodynamic delay. For this reason, when employing machine learning techniques, a problem due to the high-dimensional feature vector may be encountered, such as deteriorated classification accuracy. In this study, we employ an elastic net-based feature selection which is one of the embedded methods and demonstrate the utility of which by analyzing the results. Using the fNIRS dataset obtained from 18 participants for classifying brain activation induced by mental arithmetic and idle state, we calculated classification accuracies after performing feature selection while changing the parameter α (weight of lasso vs. ridge regularization). Grand averages of classification accuracy are 80.0 ± 9.4%, 79.3 ± 9.6%, 79.0 ± 9.2%, 79.7 ± 10.1%, 77.6 ± 10.3%, 79.2 ± 8.9%, and 80.0 ± 7.8% for the various values of α = 0.001, 0.005, 0.01, 0.05, 0.1, 0.2, and 0.5, respectively, and are not statistically different from the grand average of classification accuracy estimated with all features (80.1 ± 9.5%). As a result, no difference in classification accuracy is revealed for all considered parameter α values. Especially for α = 0.5, we are able to achieve the statistically same level of classification accuracy with even 16.4% features of the total features. Since elastic net-based feature selection can be easily applied to other cases without complicated initialization and parameter fine-tuning, we can be looking forward to seeing that the elastic-based feature selection can be actively applied to fNIRS data.

회전기계 결함신호 진단을 위한 신호처리 기술 개발 (Signal Processing Technology for Rotating Machinery Fault Signal Diagnosis)

  • 최병근;안병현;김용휘;이종명;이정훈
    • 한국소음진동공학회:학술대회논문집
    • /
    • 한국소음진동공학회 2013년도 추계학술대회 논문집
    • /
    • pp.331-337
    • /
    • 2013
  • Acoustic Emission technique is widely applied to develop the early fault detection system, and the problem about a signal processing method for AE signal is mainly focused on. In the signal processing method, envelope analysis is a useful method to evaluate the bearing problems and Wavelet transform is a powerful method to detect faults occurred on rotating machinery. However, exact method for AE signal is not developed yet. Therefore, in this paper two methods which are Hilbert transform and DET for feature extraction. In addition, we evaluate the classification performance with varying the parameter from 2 to 15 for feature selection DET, 0.01 to 1.0 for the RBF kernel function of SVR, and the proposed algorithm achieved 94% classification accuracy with the parameter of the RBF 0.08, 12 feature selection.

  • PDF

특징 선택과 융합 방법을 이용한 음성 감정 인식 (Speech Emotion Recognition using Feature Selection and Fusion Method)

  • 김원구
    • 전기학회논문지
    • /
    • 제66권8호
    • /
    • pp.1265-1271
    • /
    • 2017
  • In this paper, the speech parameter fusion method is studied to improve the performance of the conventional emotion recognition system. For this purpose, the combination of the parameters that show the best performance by combining the cepstrum parameters and the various pitch parameters used in the conventional emotion recognition system are selected. Various pitch parameters were generated using numerical and statistical methods using pitch of speech. Performance evaluation was performed on the emotion recognition system using Gaussian mixture model(GMM) to select the pitch parameters that showed the best performance in combination with cepstrum parameters. As a parameter selection method, sequential feature selection method was used. In the experiment to distinguish the four emotions of normal, joy, sadness and angry, fifteen of the total 56 pitch parameters were selected and showed the best recognition performance when fused with cepstrum and delta cepstrum coefficients. This is a 48.9% reduction in the error of emotion recognition system using only pitch parameters.