• 제목/요약/키워드: feature ranking

검색결과 48건 처리시간 0.019초

텍스트 분류를 위한 자질 순위화 기법에 관한 연구 (An Experimental Study on Feature Ranking Schemes for Text Classification)

  • 김판준
    • 정보관리학회지
    • /
    • 제40권1호
    • /
    • pp.1-21
    • /
    • 2023
  • 본 연구는 텍스트 분류를 위한 효율적인 자질선정 방법으로 자질 순위화 기법의 성능을 구체적으로 검토하였다. 지금까지 자질 순위화 기법은 주로 문헌빈도에 기초한 경우가 대부분이며, 상대적으로 용어빈도를 사용한 경우는 많지 않았다. 따라서 텍스트 분류를 위한 자질선정 방법으로 용어빈도와 문헌빈도를 개별적으로 적용한 단일 순위화 기법들의 성능을 살펴본 다음, 양자를 함께 사용하는 조합 순위화 기법의 성능을 검토하였다. 구체적으로 두 개의 실험 문헌집단(Reuters-21578, 20NG)과 5개 분류기(SVM, NB, ROC, TRA, RNN)를 사용하는 환경에서 분류 실험을 진행하였고, 결과의 신뢰성 확보를 위해 5-fold cross validation과 t-test를 적용하였다. 결과적으로, 단일 순위화 기법으로는 문헌빈도 기반의 단일 순위화 기법(chi)이 전반적으로 좋은 성능을 보였다. 또한, 최고 성능의 단일 순위화 기법과 조합 순위화 기법 간에는 유의한 성능 차이가 없는 것으로 나타났다. 따라서 충분한 학습문헌을 확보할 수 있는 환경에서는 텍스트 분류의 자질선정 방법으로 문헌빈도 기반의 단일 순위화 기법(chi)을 사용하는 것이 보다 효율적이라 할 수 있다.

Computer Aided Diagnosis System based on Performance Evaluation Agent Model

  • Rhee, Hyun-Sook
    • 한국컴퓨터정보학회논문지
    • /
    • 제21권1호
    • /
    • pp.9-16
    • /
    • 2016
  • In this paper, we present a performance evaluation agent based on fuzzy cluster analysis and validity measures. The proposed agent is consists of three modules, fuzzy cluster analyzer, performance evaluation measures, and feature ranking algorithm for feature selection step in CAD system. Feature selection is an important step commonly used to create more accurate system to help human experts. Through this agent, we get the feature ranking on the dataset of mass and calcification lesions extracted from the public real world mammogram database DDSM. Also we design a CAD system incorporating the agent and apply five different feature combinations to the system. Experimental results proposed approach has higher classification accuracy and shows the feasibility as a diagnosis supporting tool.

밝기순위 특징을 이용한 적외선 정지영상 내 물체검출기법 (Object Detection in a Still FLIR Image using Intensity Ranking Feature)

  • 박재희;최학훈;김성대
    • 대한전자공학회논문지SP
    • /
    • 제42권2호
    • /
    • pp.37-48
    • /
    • 2005
  • 본 논문에서는 적외선 영상에서 밝기변화를 예측하기 어려운 일정한 크기의 관심 물체를 검출하기 위하여, 밝기순위 특징과 이론 이용한 물체식별기법을 제안한다. 제안하는 밝기순위 특징은 밝기값의 분포가 균일하도록 영상을 정규화하여 나타낸 것으로, 적외선 영상과 같이 검출대상 물체의 밝기분포를 쉽게 예측하기 어려운 경우에 적합한 특징이다. 제안하는 식별기법은 주어진 후보영역이 검출대상 물체의 학습영상들에 대해 밝기순위가 부합하는 정도를 수치화하여 각각의 후보영역을 물체와 비물체로 식별한다 제안하는 기법을 통하여 별도의 후보영역 선정과정 없이도 일정한 크기의 관심 물체에 대해 화소단위의 검출결과를 획득할 수 있다. 실험에서는 적외선 자동차 영상을 이용하여 밝기순위특징이 적외선 영상 내 물체식별에 적합함을 보이고, 잡음 및 물체의 크기변화, 기울어짐이 존재하는 상황에서의 검출결과를 보인다.

거리순위를 이용한 얼굴검출 (Face Detection using Distance Ranking)

  • 박재희;김성대
    • 대한전자공학회:학술대회논문집
    • /
    • 대한전자공학회 2005년도 추계종합학술대회
    • /
    • pp.363-366
    • /
    • 2005
  • In this paper, for detecting human faces under variations of lighting condition and facial expression, distance ranking feature and detection algorithm based on the feature are proposed. Distance ranking is the intensity ranking of a distance transformed image. Based on statistically consistent edge information, distance ranking is robust to lighting condition change. The proposed detection algorithm is a matching algorithm based on FFT and a solution of discretization problem in the sliding window methods. In experiments, face detection results in the situation of varying lighting condition, complex background, facial expression change and partial occlusion of face are shown

  • PDF

특징 순위 방법을 이용한 혈소판 라만 스펙트럼에서 퇴행성 뇌신경질환과 혈관성 인지증 분류 (Feature Ranking for Detection of Neuro-degeneration and Vascular Dementia in micro-Raman spectra of Platelet)

  • 박아론;백성준
    • 전자공학회논문지CI
    • /
    • 제48권4호
    • /
    • pp.21-26
    • /
    • 2011
  • 특징 순위 방법은 데이터에 대한 정보와 관련된 특징을 구별하는데 유용하게 사용된다. 본 논문에서는 혈소판으로부터 측정된 라만 스펙트럼에서 퇴행성 뇌신경질환과 혈관성 인지증의 분류에 특징 순위를 이용하는 방법을 제안하였다. 퇴행성 뇌신경 질환인 알츠하이머병(Alzheimer's disease)과 파킨슨병(Parkinson's disease) 그리고 혈관성 인지증(vascular dementia)을 유도한 실험용 쥐의 혈소판에서 측정한 스펙트럼은 가우시안 모델을 이용한 커브 피팅으로 노이즈를 제거하고 로컬 최저점에 선형 보간법(linear interpolation)으로 배경 잡음을 제거한다. 전처리 과정을 수행한 스펙트럼에서 분류정확도와 계산복잡도를 개선하기 위해 특징 순위 방법을 이용하여 주요 특징을 선택하였다. 선택된 특징들은 PCA(principal component analysis) 방법으로 변환하여 주성분의 수를 변화시키며 MAP(maximum a posteriori)으로 분류하고 전체 특징을 사용한 경우의 분류 결과와 비교하였다. 실험 결과에서 제안한 방법을 적용한 모든 실험에서 분류 시스템의 계산복잡도를 현저하게 감소시키고 분류정확도는 부분적으로 증가하였다. 특히 파킨슨병과 정상을 분류하는 실험에서 제안한 방법이 전체 특징을 사용한 경우보다 모든 주성분의 수에서 분류정확도가 높았으며 평균 1.7 %의 성능이 향상되었다. 이 결과에서 분류정확도와 계산복잡도의 개선을 고려하면 제안한 방법이 혈소판 라만 스펙트럼에서 퇴행성 뇌신경질환과 혈관성 인지증의 분류 시스템에 효율적으로 사용될 수 있음을 확인하였다.

Relevancy contemplation in medical data analytics and ranking of feature selection algorithms

  • P. Antony Seba;J. V. Bibal Benifa
    • ETRI Journal
    • /
    • 제45권3호
    • /
    • pp.448-461
    • /
    • 2023
  • This article performs a detailed data scrutiny on a chronic kidney disease (CKD) dataset to select efficient instances and relevant features. Data relevancy is investigated using feature extraction, hybrid outlier detection, and handling of missing values. Data instances that do not influence the target are removed using data envelopment analysis to enable reduction of rows. Column reduction is achieved by ranking the attributes through feature selection methodologies, namely, extra-trees classifier, recursive feature elimination, chi-squared test, analysis of variance, and mutual information. These methodologies are ranked via Technique for Order of Preference by Similarity to Ideal Solution (TOPSIS) using weight optimization to identify the optimal features for model building from the CKD dataset to facilitate better prediction while diagnosing the severity of the disease. An efficient hybrid ensemble and novel similarity-based classifiers are built using the pruned dataset, and the results are thereafter compared with random forest, AdaBoost, naive Bayes, k-nearest neighbors, and support vector machines. The hybrid ensemble classifier yields a better prediction accuracy of 98.31% for the features selected by extra tree classifier (ETC), which is ranked as the best by TOPSIS.

영어 lC 자음군에 관한 역사적 조명과 음운적 고찰 (A phonological study and historical view on IC clusters in English)

  • 오관영
    • 영어어문교육
    • /
    • 제16권4호
    • /
    • pp.201-222
    • /
    • 2010
  • The purpose of this study is to investigate /l/-deletion in lC clusters which are composed of a lateral followed by consonants at syllable-final position in English. For this, I have analyzed /l/-deletion in words depending on conditions and theoretical analyses such as Sonority Sequencing Generalization, Cluster Simplification, Complex sounds and merger, and Feature Geometry, but they didn't offer a very satisfactory explanation to the phenomenon. Therefore, I adopted a historical approach in order to determine the cause and origin of /l/-deletion in lC clusters, and then as a phonological analysis tool, I relied on the constraints and their ranking in Optimal Theory framework for explaining /l/-deletion in the clusters more consistently. As a result, I can explain the phenomenon more explicitly than from the above mentioned analyses.

  • PDF

An approach for improving the performance of the Content-Based Image Retrieval (CBIR)

  • Jeong, Inseong
    • 한국측량학회지
    • /
    • 제30권6_2호
    • /
    • pp.665-672
    • /
    • 2012
  • Amid rapidly increasing imagery inputs and their volume in a remote sensing imagery database, Content-Based Image Retrieval (CBIR) is an effective tool to search for an image feature or image content of interest a user wants to retrieve. It seeks to capture salient features from a 'query' image, and then to locate other instances of image region having similar features elsewhere in the image database. For a CBIR approach that uses texture as a primary feature primitive, designing a texture descriptor to better represent image contents is a key to improve CBIR results. For this purpose, an extended feature vector combining the Gabor filter and co-occurrence histogram method is suggested and evaluated for quantitywise and qualitywise retrieval performance criterion. For the better CBIR performance, assessing similarity between high dimensional feature vectors is also a challenging issue. Therefore a number of distance metrics (i.e. L1 and L2 norm) is tried to measure closeness between two feature vectors, and its impact on retrieval result is analyzed. In this paper, experimental results are presented with several CBIR samples. The current results show that 1) the overall retrieval quantity and quality is improved by combining two types of feature vectors, 2) some feature is better retrieved by a specific feature vector, and 3) retrieval result quality (i.e. ranking of retrieved image tiles) is sensitive to an adopted similarity metric when the extended feature vector is employed.

Improved Feature Selection Techniques for Image Retrieval based on Metaheuristic Optimization

  • Johari, Punit Kumar;Gupta, Rajendra Kumar
    • International Journal of Computer Science & Network Security
    • /
    • 제21권1호
    • /
    • pp.40-48
    • /
    • 2021
  • Content-Based Image Retrieval (CBIR) system plays a vital role to retrieve the relevant images as per the user perception from the huge database is a challenging task. Images are represented is to employ a combination of low-level features as per their visual content to form a feature vector. To reduce the search time of a large database while retrieving images, a novel image retrieval technique based on feature dimensionality reduction is being proposed with the exploit of metaheuristic optimization techniques based on Genetic Algorithm (GA), Extended Binary Cuckoo Search (EBCS) and Whale Optimization Algorithm (WOA). Each image in the database is indexed using a feature vector comprising of fuzzified based color histogram descriptor for color and Median binary pattern were derived in the color space from HSI for texture feature variants respectively. Finally, results are being compared in terms of Precision, Recall, F-measure, Accuracy, and error rate with benchmark classification algorithms (Linear discriminant analysis, CatBoost, Extra Trees, Random Forest, Naive Bayes, light gradient boosting, Extreme gradient boosting, k-NN, and Ridge) to validate the efficiency of the proposed approach. Finally, a ranking of the techniques using TOPSIS has been considered choosing the best feature selection technique based on different model parameters.

Image Retrieval Based on the Weighted and Regional Integration of CNN Features

  • Liao, Kaiyang;Fan, Bing;Zheng, Yuanlin;Lin, Guangfeng;Cao, Congjun
    • KSII Transactions on Internet and Information Systems (TIIS)
    • /
    • 제16권3호
    • /
    • pp.894-907
    • /
    • 2022
  • The features extracted by convolutional neural networks are more descriptive of images than traditional features, and their convolutional layers are more suitable for retrieving images than are fully connected layers. The convolutional layer features will consume considerable time and memory if used directly to match an image. Therefore, this paper proposes a feature weighting and region integration method for convolutional layer features to form global feature vectors and subsequently use them for image matching. First, the 3D feature of the last convolutional layer is extracted, and the convolutional feature is subsequently weighted again to highlight the edge information and position information of the image. Next, we integrate several regional eigenvectors that are processed by sliding windows into a global eigenvector. Finally, the initial ranking of the retrieval is obtained by measuring the similarity of the query image and the test image using the cosine distance, and the final mean Average Precision (mAP) is obtained by using the extended query method for rearrangement. We conduct experiments using the Oxford5k and Paris6k datasets and their extended datasets, Paris106k and Oxford105k. These experimental results indicate that the global feature extracted by the new method can better describe an image.