• Title/Summary/Keyword: cluster method

Search Result 2,498, Processing Time 0.029 seconds

Design and implementation of data mining tool using PHP and WEKA (피에이치피와 웨카를 이용한 데이터마이닝 도구의 설계 및 구현)

  • You, Young-Jae;Park, Hee-Chang
    • Journal of the Korean Data and Information Science Society
    • /
    • v.20 no.2
    • /
    • pp.425-433
    • /
    • 2009
  • Data mining is the method to find useful information for large amounts of data in database. It is used to find hidden knowledge by massive data, unexpectedly pattern, relation to new rule. We need a data mining tool to explore a lot of information. There are many data mining tools or solutions; E-Miner, Clementine, WEKA, and R. Almost of them are were focused on diversity and general purpose, and they are not useful for laymen. In this paper we design and implement a web-based data mining tool using PHP and WEKA. This system is easy to interpret results and so general users are able to handle. We implement Apriori algorithm of association rule, K-means algorithm of cluster analysis, and J48 algorithm of decision tree.

  • PDF

More effective application of importance-performance analysis in the case of cyber lecture (중요도-실행도 분석의 효율적 활용에 대한 연구 - 온라인 수능강의에 대한 사례 연구)

  • Pak, Ro-Jin
    • Journal of the Korean Data and Information Science Society
    • /
    • v.20 no.2
    • /
    • pp.329-338
    • /
    • 2009
  • The importance performance analysis is a simple and condensed analytic method for decision making based on the level of performance or satisfaction. Many researches already have witnessed usefulness of the importance performance analysis, but it also has some drawbacks from the statistical points of view. In this article, some additional techniques dealing the importance performance analysis are introduced and it is shown that these techniques would turn out to be very informative. The importance performance analysis uses the arithmetic average as the main statistic, but by the use of the median, the frequency and the cluster analysis it is shown that the importance performance analysis can be carried out with more crucial information. In addtion to that, it is demonstrated that the combination of the analytic hierarchy process and importance performance analysis could enable more reliable decision making.

  • PDF

A synchronous/asynchronous hybrid parallel method for some eigenvalue problems on distributed systems

  • 박필성
    • Proceedings of the Korean Society of Computational and Applied Mathematics Conference
    • /
    • 2003.09a
    • /
    • pp.11-11
    • /
    • 2003
  • 오늘날 단일 슈퍼컴퓨터로는 처리가 불가능한 거대한 문제들의 해법이 시도되고 있는데, 이들은 지리적으로 분산된 슈퍼컴퓨터, 데이터베이스, 과학장비 및 디스플레이 장치 등을 초고속 통신망으로 연결한 GRID 환경에서 효과적으로 실행시킬 수 있다. GRID는 1990년대 중반 과학 및 공학용 분산 컴퓨팅의 연구 과정에서 등장한 것으로, 점차 응용분야가 넓어지고 있다. 그러나 GRID 같은 분산 환경은 기존의 단일 병렬 시스템과는 많은 점에서 다르며 이전의 기술들을 그대로 적용하기에는 무리가 있다. 기존 병렬 시스템에서는 주로 동기 알고리즘(synchronous algorithm)이 사용되는데, 직렬 연산과 같은 결과를 얻기 위해 동기화(synchronization)가 필요하며, 부하 균형이 필수적이다. 그러나 부하 균형은 이질 클러스터(heterogeneous cluster)처럼 프로세서들의 성능이 서로 다르거나, 지리적으로 분산된 계산자원을 사용하는 GRID 환경에서는 이기종의 문제뿐 아니라 네트워크를 통한 메시지의 전송 지연 등으로 유휴시간이 길어질 수밖에 없다. 이처럼 동기화의 필요성에 의한 연산의 지연을 해결하는 하나의 방안으로 비동기 반복법(asynchronous iteration)이 나왔으며, 지금도 활발히 연구되고 있다. 이는 알고리즘의 동기점을 가능한 한 제거함으로써 빠른 프로세서의 유휴 시간을 줄이는 것이 목적이다. 즉 비동기 알고리즘에서는, 각 프로세서는 다른 프로세서로부터 갱신된 데이터가 올 때까지 기다리지 않고 계속 다음 작업을 수행해 나간다. 따라서 동시에 갱신된 데이터를 교환한 후 다음 단계로 진행하는 동기 알고리즘에 비해, 미처 갱신되지 않은 데이터를 사용하는 경우가 많으므로 전체적으로는 연산량 대비의 수렴 속도는 느릴 수 있다 그러나 각 프로세서는 거의 유휴 시간이 없이 연산을 수행하므로 wall clock time은 동기 알고리즘보다 적게 걸리며, 때로는 50%까지 빠른 결과도 보고되고 있다 그러나 현재까지의 연구는 모두 어떤 수렴조건을 만족하는 선형 시스템의 해법에 국한되어 있으며 비교적 구현하기 쉬운 공유 메모리 시스템에서의 연구만 보고되어 있다. 본 연구에서는 행렬의 주요 고유쌍을 구하는 데 있어 비동기 반복법의 적용 가능성을 타진하기 위해 우선 이론적으로 단순한 멱승법을 사용하여 실험하였고 그 결과 순수한 비동기 반복법은 수렴하기 어렵다는 결론을 얻었다 그리하여 동기 알고리즘에 비동기적 요소를 추가한 혼합 병렬 알고리즘을 제안하고, MPI(Message Passing Interface)를 사용하여 수원대학교의 Hydra cluster에서 구현하였다. 그 결과 특정 노드의 성능이 다른 것에 비해 현저하게 떨어질 때 전체적인 알고리즘의 수렴 속도가 떨어지는 것을 상당히 완화할 수 있음이 밝혀졌다.

  • PDF

Upper Body Type Classification of Elementary School Boys Using 3D Data (3차원 데이터를 활용한 학령기 남아의 상반신 체형 분류)

  • Kim, Hyun Wook;Nam, Yun Ja
    • Fashion & Textile Research Journal
    • /
    • v.21 no.6
    • /
    • pp.789-799
    • /
    • 2019
  • This study classified and analyzed the upper body types of 7-13 years old elementary school boys, using 3D data from the 6th Size Korea. The results of this study are as follows. Seven factors were extracted from the factorial analysis as an independent factor for a cluster analysis. The cluster analysis generated four body types. Type 1 has large ratio of front and back depth as well as circumference, with a front protrusion. In Type 2, the vertical value of upper torso is longer than average; in addition, its flatness is the largest and produces a thin body type. Type 3 has a smaller flatness in the bust, waist, abdomen and hip than other types, while also having the largest BMI. Type 4 is characterized by a greater shoulder angle than other types and its other factors are close to average. As a result of the logistic regression analysis, the prediction model used eight variables to generate and its accuracy is 88.679%. The classification of upper body types from this study can be used as basic data to improve patternmaking for each body type. The generated prediction model is also expected to be used as a method to help classify upper body types using the eight variables.

DNA fingerprinting analysis for soybean (Glycine max) varieties in Korea using a core set of microsatellite marker (핵심 Microsatellite 마커를 이용한 한국 콩 품종에 대한 Fingerprinting 분석)

  • Kwon, Yong-Sham
    • Journal of Plant Biotechnology
    • /
    • v.43 no.4
    • /
    • pp.457-465
    • /
    • 2016
  • Microsatellites are one of the most suitable markers for identification of variety, as they have the capability to discriminate between narrow genetic variations. The polymorphism level between 120 microsatellite primer pairs and 148 soybean varieties was investigated through the fluorescence based automatic detection system. A set of 16 primer pairs showed highly reproducible polymorphism in these varieties. A total of 204 alleles were detected using the 16 microsatellite markers. The number of alleles per locus ranged from 6 to 28, with an average of 12.75 alleles per locus. The average polymorphism information content (PIC) was 0.86, ranging from 0.75 to 0.95. The unweighted pair group method using the arithmetic averages (UPGMA) cluster analysis for 148 varieties were divided into five distinctive groups, reflecting the varietal types and pedigree information. All the varieties were perfectly discriminated by marker genotypes. These markers may be useful to complement a morphological assessment of candidate varieties in the DUS (distinctness, uniformity and stability) test, intervening of seed disputes relating to variety authentication, and testing of genetic purity in soybean varieties.

Redescription and Multivariate Analysis of Genus Phintella (Araneae, Salticidae) from Korea (한국산 Phintella속(거미목, 깡충거미과)의 재기재와 다변량분석)

  • Bo-Keun Seo
    • Animal Systematics, Evolution and Diversity
    • /
    • v.11 no.2
    • /
    • pp.183-197
    • /
    • 1995
  • Description and identifications of 6 species belonging to genus Phintella from Korea are in insufficient and inaccurate situation. In the present paper, redescriptions illustrations and identification key are provided for 7 species of genus Phintella including P. popovi newly recorded in Korean spider fauna, and Ocius munitus described by Wesolowska (1981s) was synonymized to P.cavaleriei. For the author's identiication and pairing to be valid multivariate analysis was performed with 13 RVCs below STD 0.05 to 134 individuals. The result of discriminant analysis carried out with 13 RVCs of 134 individuals was not satisfactory, but cluster analysis performed with mean ratio values of 14 OTUs to 13 RVCs showed the same result with author's pairing except P.abnormis , which has larger dissimilarity than the pairs of the others. So pairing of 7 species was possible as a whole because one species only failed in pairing , even though this is imperful result. This method to be helpful to pairing test and identification if it were to improve.

  • PDF

Evaluation of Water Quality Characteristics and Grade Classification of Yeongsan River Tributaries (영산강 수계 지류.지천의 수질 특성 평가 및 등급화 방안)

  • Jung, Soojung;Kim, Kapsoon;Seo, Dongju;Kim, Junghyun;Lim, Byungjin
    • Journal of Korean Society on Water Environment
    • /
    • v.29 no.4
    • /
    • pp.504-513
    • /
    • 2013
  • Water quality trends for major tributaries (66 sites) in the Yeongsan River basin of Korea were examined for 12 parameters based on water quality data collected every month over a period of 12 months. The complex data matrix was treated with multivariate analysis such as PCA, FA and CA. PCA/FA identified four factors, which are responsible for the structure explaining 78.2% of the total variance. The first factor accounting 27.3% of the total variance was correlated with BOD, TN, TP, and TOC, and weighting values were allowed to these parameters for grade classification. CA rendered a dendrogram, where monitoring sites were grouped into 5 clusters. Cluster 2 corresponds to high pollution from domestic wastewater, wastewater treatment and run-off from livestock farms. For grade classification of tributaries, scores to 10 indexes were calculated considering the weighting values to 3 parameters as BOD, TN and TP which were categorized as the first factor after FA. The highest-polluted group included 10 tributaries such as Gwangjucheon, Jangsucheon, Daejeoncheon, Gamjungcheon, Yeongsancheon. The results indicate that grade classification method suggested in this study is useful in reliable classification of tributaries in the study area.

Regional Division According to the Annual Change of Sunshine Duration in Korea (일조시간의 연변화에 따른 한국의 지역구분)

  • 문영수
    • Journal of Environmental Science International
    • /
    • v.5 no.3
    • /
    • pp.253-263
    • /
    • 1996
  • This study is an attempt to classify climatic regions of Korea based on the data of sunshine duration and to clarify the characteristics of sunshine for each divided regions. The data used in this study are the mean values of monthly and ten-daily sunshine duration, sunshine percentage, solar radiation and proud amount obtained from 63 weather stations of the Korea Meteorological Administration during the period of 1974~ 1993. The characteristics of annual change of sunshine percentage, annual duration of sunshine, percentage of sunshine, annual radiation, amount of cloud, days of sunshine percentage above 80% and-days of sunless are investigated by the mean values of -the stations belong to divided regions. The ward method of hierarchical cluster analysis is adopted to the analysis of data for the regional division. The results obtained in this study are summarized as follows. (1) The sunshine regions of Korea can be divided into six regions of the central west, central east, south west, souls east, Ullung-do and Cheju-do. These are strongly affected by the dirtribution of inclined slopes taking account of the topographic characteristics of Korea. (2) Annual distribution shows the sunshine duration of 1777~ 2287 hours, sunshine percentage of 40~53%, solar radiation of 3469~4637 MJ/$m^2$, cloud amount of 5.0~6.1, days of sunshine perrentage above 80% of 53~116days and sunless days of 46~71days. (3) The types of annual change of sunshine percentages is classified with four types of minimum in July and maximum in October, minimum in July and maximum in December, high in May and October and low in July and January, high in May and November and low in June and January. (4) The long-term trend of sunshine duration decrease in peninsula area but increase in island area and the Tong-term inclination of cloud amount is almost zero. The author believe this tendency is related to a pollutional turbidity than a cloud amount in inland area.

  • PDF

Context-aware Connectivity Analysis Method using Context Data Prediction Model in Delay Tolerant Networks (Delay Tolerant Networks에서 속성정보 예측 모델을 이용한 상황인식 연결성 분석 기법)

  • Jeong, Rae-Jin;Oh, Young-Jun;Lee, Kang-Whan
    • Journal of the Korea Institute of Information and Communication Engineering
    • /
    • v.19 no.4
    • /
    • pp.1009-1016
    • /
    • 2015
  • In this paper, we propose EPCM(Efficient Prediction-based Context-awareness Matrix) algorithm analyzing connectivity by predicting cluster's context data such as velocity and direction. In the existing DTN, unrestricted relay node selection causes an increase of delay and packet loss. The overhead is occurred by limited storage and capability. Therefore, we propose the EPCM algorithm analyzing predicted context data using context matrix and adaptive revision weight, and selecting relay node by considering connectivity between cluster and base station. The proposed algorithm saves context data to the context matrix and analyzes context according to variation and predicts context data after revision from adaptive revision weight. From the simulation results, the EPCM algorithm provides the high packet delivery ratio by selecting relay node according to predicted context data matrix.

Chacterization and Preparation of SiO2-TiO2-AgO thin Films by the Chemical Solution Process (용액법에 의한 SiO2-TiO2-AgO계 박막의 제조 및 특성에 관한 연구)

  • Kim, Sangmoon;Shim, Moon-Sik;Lim, Yongmu;Hwang, Kyuseog
    • Journal of Korean Ophthalmic Optics Society
    • /
    • v.3 no.1
    • /
    • pp.217-222
    • /
    • 1998
  • Coating films of $SiO_2-TiO_2-AgO$ have been prepared on soda-lime-silica slide glasses and single crystal silicon wafer by the sol-gel method using a spin-coating technique. Commercially available tetraethyl orthosilicate, titanium trichloride, and silver-nitrates were used as starting materials. The heat treatment temperature of this coating films was $500^{\circ}C$ properly, obtained from TG-DTA result. The films with thickness of 310 nm were prepared by 5 times coating. In the case of l0 mol% AgO, the film showed a crack-free and smooth surface, but the higher Ago content exhibited the more pin hole and the segregated cluster of AgO. The IR absorbance of the films decreased in the range of 400 nm to 700 nm with the increase of annealing temperature. And the reflectance of the coating films decreased and the color was changed light yellow to white yellow with the increase of Ago content.

  • PDF