• Title/Summary/Keyword: Selective sampling

Search Result 83, Processing Time 0.028 seconds

A Study on Incremental Learning Model for Naive Bayes Text Classifier (Naive Bayes 문서 분류기를 위한 점진적 학습 모델 연구)

  • 김제욱;김한준;이상구
    • Proceedings of the Korea Database Society Conference
    • /
    • 2001.06a
    • /
    • pp.331-341
    • /
    • 2001
  • 본 논문에서는 Naive Bayes 문서 분류기를 위한 새로운 학습모델을 제안한다. 이 모델에서는 라벨이 없는 문서들의 집합으로부터 선택한 적은 수의 학습 문서들을 이용하여 문서 분류기를 재학습한다. 본 논문에서는 이러한 학습 방법을 따를 경우 작은 비용으로도 문서 분류기의 정확도가 크게 향상될 수 있다는 사실을 보인다. 이와 같이, 알고리즘을 통해 라벨이 없는 문서들의 집합으로부터 정보량이 큰 문서를 선택한 후, 전문가가 이 문서에 라벨을 부여하는 방식으로 학습문서를 결정하는 것을 selective sampling이라 한다. 본 논문에서는 이러한 selective sampling 문제를 Naive Bayes 문서 분류기에 적용한다. 제안한 학습 방법에서는 라벨이 없는 문서들의 집합으로부터 재학습 문서를 선택하는 기준 측정치로서 평균절대편차(Mean Absolute Deviation), 엔트로피 측정치를 사용한다. 실험을 통해서 제안한 학습 방법이 기존의 방법인 신뢰도(Confidence measure)를 이용한 학습 방법보다 Naive Bayes 문서 분류기의 성능을 더 많이 향상시킨다는 사실을 보인다.

  • PDF

New and Improved Time-selective Self-triggering Water Sampler: AUTTLE

  • Jin, Jae-Youll;Hwang, Kuen-Choon;Park, Jin-Soon;Eo, Young-Sang;Kim, Seong-Eun;Yum, Ki-Dai;Oh, Jae-Kyung
    • Ocean and Polar Research
    • /
    • v.22 no.2
    • /
    • pp.57-67
    • /
    • 2000
  • Time-selective self-triggering water sampler, AUTTLE developed by Jin et al. (1999) has been improved in order to prevent pre-deposition of suspended sediments (SS) before sampling. By using two solenoids, the improved sampler is able to be moored or deployed with inclination. Its position is changed to horizontal position by activating the first solenoid, and then the endcaps of the sampling bottle are closed by the second solenoid that is driven three times to minimize possible failure of sampling. An external control unit for setting sampling time has been also constructed. Additionally, the electric circuit housing of the sampler has been modified to be detached from the sampling bottle when operating manually. Its performance has been confirmed through flume tests and a field experiment. It will serve as a valuable tool in the various fields of oceanography and environmental engineering, especially where seawater sampling synchronized at several sites and/or the information in storm period is important.

  • PDF

Accelerating the EM Algorithm through Selective Sampling for Naive Bayes Text Classifier (나이브베이즈 문서분류시스템을 위한 선택적샘플링 기반 EM 가속 알고리즘)

  • Chang Jae-Young;Kim Han-Joon
    • The KIPS Transactions:PartD
    • /
    • v.13D no.3 s.106
    • /
    • pp.369-376
    • /
    • 2006
  • This paper presents a new method of significantly improving conventional Bayesian statistical text classifier by incorporating accelerated EM(Expectation Maximization) algorithm. EM algorithm experiences a slow convergence and performance degrade in its iterative process, especially when real online-textual documents do not follow EM's assumptions. In this study, we propose a new accelerated EM algorithm with uncertainty-based selective sampling, which is simple yet has a fast convergence speed and allow to estimate a more accurate classification model on Naive Bayesian text classifier. Experiments using the popular Reuters-21578 document collection showed that the proposed algorithm effectively improves classification accuracy.

Low-Sampling Rate UWB Channel Characterization and Synchronization

  • Maravic, Irena;Kusuma, Julius;Vetterli, Martin
    • Journal of Communications and Networks
    • /
    • v.5 no.4
    • /
    • pp.319-327
    • /
    • 2003
  • We consider the problem of low-sampling rate high-resolution channel estimation and timing for digital ultrawideband (UWB) receivers. We extend some of our recent results in sampling of certain classes of parametric non-bandlimited signals and develop a frequency domain method for channel estimation and synchronization in ultra-wideband systems, which uses sub-Nyquist uniform sampling and well-studied computational procedures. In particular, the proposed method can be used for identification of more realistic channel models, where different propagation paths undergo different frequency-selective fading. Moreover, we show that it is possible to obtain high-resolution estimates of all relevant channel parameters by sampling a received signal below the traditional Nyquist rate. Our approach leads to faster acquisition compared to current digital solutions, allows for slower A/D converters, and potentially reduces power consumption of digital UWB receivers significantly.

Robust Localization Algorithm for Mobile Robots in a Dynamic Environment with an Incomplete Map (동적 환경에서 불완전한 지도를 이용한 이동로봇의 강인한 위치인식 알고리즘의 개발)

  • Lee, Jung-Suk;Chung, Wan Kyun;Nam, Sang Yep
    • IEMEK Journal of Embedded Systems and Applications
    • /
    • v.3 no.2
    • /
    • pp.109-118
    • /
    • 2008
  • We present a robust localization algorithm using particle filter for mobile robots in a dynamic environment. It is difficult to describe moving obstacles like people or other robots on the map and the environment is changed after mapping. A mobile robot cannot estimate its pose robustly with this incomplete map because sensor observations are corrupted by un-modeled obstacles. The proposed algorithms provide robustness in such a dynamic environment by suppressing the effect of corrupted sensor observations with a selective update or a sampling from non-corrupted window. A selective update method makes some particles keep track of the robot, not affected by the corrupted observation. In a sampling from non-corrupted window method, particles are always sampled from several particle sets which use only non-corrupted observation. The robustness of proposed algorithm is validated with experiments and simulations.

  • PDF

Performance Analysis of a Fractionally Spaced Equalizer using Selective Normalized CMA (선택적 NCMA 방법을 이용한 분할 블라인드 적응 등화기의 성능 분석)

  • Hong, Ji-Hun;Jang, Tae-Jeong
    • Journal of Industrial Technology
    • /
    • v.21 no.B
    • /
    • pp.99-105
    • /
    • 2001
  • In this paper, the selective normalized constant modulus algorithm(SNCMA) is applied to a fractionally spaced equalizer. The fractionally spaced equalizer is insensitive to the sampling timing because it processes received signals with the sampling rate larger than the symbol rate. The SNCMA improves the convergence rate by using the large step size for the most outer covering symbol belonging to the trust-level. This blind equalizer exhibits a fast start-up convergence rate as well as a reduced steady-state residual error compared to the fractionally spaced blind equalizer and the T-spaced blind equalizer using conventional blind algorithms.

  • PDF

A Study on Incremental Learning Model for Naive Bayes Text Classifier (Naive Bayes 문서 분류기를 위한 점진적 학습 모델 연구)

  • 김제욱;김한준;이상구
    • The Journal of Information Technology and Database
    • /
    • v.8 no.1
    • /
    • pp.95-104
    • /
    • 2001
  • In the text classification domain, labeling the training documents is an expensive process because it requires human expertise and is a tedious, time-consuming task. Therefore, it is important to reduce the manual labeling of training documents while improving the text classifier. Selective sampling, a form of active learning, reduces the number of training documents that needs to be labeled by examining the unlabeled documents and selecting the most informative ones for manual labeling. We apply this methodology to Naive Bayes, a text classifier renowned as a successful method in text classification. One of the most important issues in selective sampling is to determine the criterion when selecting the training documents from the large pool of unlabeled documents. In this paper, we propose two measures that would determine this criterion : the Mean Absolute Deviation (MAD) and the entropy measure. The experimental results, using Renters 21578 corpus, show that this proposed learning method improves Naive Bayes text classifier more than the existing ones.

  • PDF

A Study on Selective Sampling using SOM (SOM을 적용한 선택적 샘플링에 관한 연구)

  • Kim, Man-Sun;Yang, Hyung-Jeong;Kim, Jeong-Sik;Kim, Sun-Hee
    • Proceedings of the Korea Information Processing Society Conference
    • /
    • 2007.11a
    • /
    • pp.38-41
    • /
    • 2007
  • 데이타 마이닝을 위하여 수집된 대용량의 데이타를 여과 없이 기계학습에 적용하는 것은 많은 시간과 비용이 요구될 뿐만 아니라 저장 공간면에서도 비효율적이다. 선별적 샘플링은 이러한 상황에서 매우 효율적으로 적용할 수 있도록 원본 데이타의 특성을 가능한 반영하여 새로운 훈련 데이타를 생성하는 방법이다. 본 연구에서는 신경망의 하나인 SOM을 적용한 선별적 샘플링을 수행하는데 있어서 여러 가지 선택 문제를 효과적으로 해결하기 위한 실험을 수행한다. 실험 결과로는 두 가지 결과를 얻었다. 1) 충분한 맵 사이즈를 선택해야 학습 데이타의 함축적인 특성을 잘 반영한다, 2) 선택적 샘플링을 위한 유닛선택 방법에서는 의미없는 유닛을 제거함으로서 분류 성능향상을 얻을 수 있다.

  • PDF

Selective collecting device utilizing the ecological characteristics of Ephemera orientalis (Ephemeroptera: Ephemeridae) (동양하루살이(하루살이목: 하루살이과) 성충의 생태적 특성을 활용한 선택적 포집 장치)

  • Jin Seok Byeon;Seong Uk Son;Jang Ho Lee;Min Kyung Kim;Rong Jin Jung;Dong Sik Ryu;Dong Gun Kim
    • Korean Journal of Environmental Biology
    • /
    • v.41 no.3
    • /
    • pp.247-255
    • /
    • 2023
  • The occurrence of sudden strike pest events in urban areas is increasing as global warming intensifies, consequently, re causing harmful impacts. Studies on these incidents are fewer in number and insufficient compared to research on other nuisances such as mosquitoes and flies. Therefore, we conducted a study on the development of a selective collection method, using a filter layer to establish a monitoring system for Ephemera orientalis (Ephemeroptera: Ephemeridae), a species frequently identified as a sudden strike pest. Three sampling points were selected along the Hangang River in Namyangju, where E. orientalis outbreaks occur. Prototypes, consisting of four layers and with a light source attached to attract insects, were installed at each sampling point. Sampling was performed every 30 minutes between 19:00 and 22:30 in the month of June. The filter interval of each layer was adjusted so that the collected mayflies were distributed into specific layers. To evaluate the collection efficiency in line with the materials and the filter intervals, the optimal collection efficiency was investigated by combining two types of layer materials (stainless and acrylic) and filter intervals (1-5 mm). The optimal conditions were as follows: The selective collection efficiency was found to be highest at 96.5% when the interval of the selective target filter was 2.0 mm and there was one upper filter.

A Study on the Total, Particle Size-Selective Mass Concentration of Airborne Manganese, and Blood Manganese Concentration of Welders in a Shipbuilding Yard (조선업 용접작업자의 공기 중 총 망간 및 입경별 망간 농도와 혈중 망간농도에 관한 연구)

  • Park, Jong Su;Kim, Pan Gyi;Jeong, Jee Yeon
    • Journal of Korean Society of Occupational and Environmental Hygiene
    • /
    • v.25 no.4
    • /
    • pp.472-481
    • /
    • 2015
  • Objectives: Welding is a major task in shipbuilding yards that generates welding fumes. A significant amount of welding in shipbuilding yards is done on steel. Inevitably, manganese is present in the base metals being joined and the filler wire being used and, consequently, in the fumes to which workers are exposed. The objective of this work was to characterize manganese exposure associated with work area, total and particle size-selective mass concentration, and compare the mass concentrations obtained using a three-piece cassette sampler, size-selective impactor sampler and blood manganese concentrations. Materials: All samples were collected from the main work areas at one shipbuilding yard. We used a three piece cassette sampler and the eight stage cascade impactor sampler for the airborne manganese mass concentration of total and all size fractions, respectively. In addition, we used the results of health examination of workers sampled for airborne manganese. Results: The oder of high concentration of airborne manganese in shipbuilding processes was as follows; block assembly, block erection, outfitting installation, steel cutting, and outfitting preparation. The percentages of samples that exceeded the OES of the ministry of employment and labor by the cassette sampling method was 12.5%, however 59.1% of sampled workers by the impactor sampling method exceeded the TLV of the ACGIH. Conclusions: Even though the manganese concentrations in blood of workers exposed to higher airborne manganese concentration were higher than among those exposed to lower concentrations, there was no difference in blood manganese concentrations among work duration. The data analyzed here by characterizing size-selective mass concentrations indicates that the inhaled manganese of welders in shipbuilding yards could be mostly manganese-containing respirable particle sizes.