• Title/Summary/Keyword: Data interpretation, statistical

Search Result 173, Processing Time 0.023 seconds

Autoregressive Cholesky Factor Modeling for Marginalized Random Effects Models

  • Lee, Keunbaik;Sung, Sunah
    • Communications for Statistical Applications and Methods
    • /
    • v.21 no.2
    • /
    • pp.169-181
    • /
    • 2014
  • Marginalized random effects models (MREM) are commonly used to analyze longitudinal categorical data when the population-averaged effects is of interest. In these models, random effects are used to explain both subject and time variations. The estimation of the random effects covariance matrix is not simple in MREM because of the high dimension and the positive definiteness. A relatively simple structure for the correlation is assumed such as a homogeneous AR(1) structure; however, it is too strong of an assumption. In consequence, the estimates of the fixed effects can be biased. To avoid this problem, we introduce one approach to explain a heterogenous random effects covariance matrix using a modified Cholesky decomposition. The approach results in parameters that can be easily modeled without concern that the resulting estimator will not be positive definite. The interpretation of the parameters is sensible. We analyze metabolic syndrome data from a Korean Genomic Epidemiology Study using this method.

Standardizing Unstructured Big Data and Visual Interpretation using MapReduce and Correspondence Analysis (맵리듀스와 대응분석을 활용한 비정형 빅 데이터의 정형화와 시각적 해석)

  • Choi, Joseph;Choi, Yong-Seok
    • The Korean Journal of Applied Statistics
    • /
    • v.27 no.2
    • /
    • pp.169-183
    • /
    • 2014
  • Massive and various types of data recorded everywhere are called big data. Therefore, it is important to analyze big data and to nd valuable information. Besides, to standardize unstructured big data is important for the application of statistical methods. In this paper, we will show how to standardize unstructured big data using MapReduce which is a distribution processing system. We also apply simple correspondence analysis and multiple correspondence analysis to nd the relationship and characteristic of direct relationship words for Samsung Electronics and The Korea Economic Daily newspaper as well as Apple Inc.

Application of Plackett-Burman model in welding experiments : effects of welding parameters on bead shape in Cu-Ni PULSE MIG process (PLACKETT-BURMAN MODEL을 이용한 Cu-Ni합금의 PULSE MIG 용접 변수해석)

  • 문영훈;이기학;허성도
    • Journal of Welding and Joining
    • /
    • v.5 no.4
    • /
    • pp.47-53
    • /
    • 1987
  • The purpose of this study is to reexamine our test method in the light ofstatistical methods for data interpretation. Our trial to apply Plackett-Burman statistical model in multifactorial welding experiments shows that is saves much time and cost and extracts very accurate results. In this study, the parametric effects of bead shape on pulse MIG process in Cu-Ni alloy are investigated for verifying our trial.

  • PDF

A Study on the Conformity of KS Standards according to Agreement on WTO/TBT (WTO/TBT 협정에 따른 KS 규격의 부합화에 관한 연구)

  • 김진규
    • Journal of Korean Society of Industrial and Systems Engineering
    • /
    • v.26 no.4
    • /
    • pp.12-22
    • /
    • 2003
  • The purposes of this study are to investigate the conformity of Korean Standards(KS) according to agreement on WTO/TBT, and to propose systematic frameworks of preparation, adoption, and application for KS in our enterprises. Significant changes in this establishment, revision, and abrogation include the following divisions; i) statistics-vocabulary and symbols, ii) Shewhart control chart, iii) statistical interpretation of data, iv) sampling procedures for inspection by attributes, v) sequential sampling plans for inspection.

Data Mining Research on Maehwado Painting Poetry in the Early Joseon Dynasty

  • Haeyoung Park;Younghoon An
    • Journal of Information Processing Systems
    • /
    • v.19 no.4
    • /
    • pp.474-482
    • /
    • 2023
  • Data mining is a technique for extracting valuable information from vast amounts of data by analyzing statistical and mathematical operations, rules, and relationships. In this study, we employed data mining technology to analyze the data concerning the painting poetry of Maehwado (plum blossom paintings) from the early Joseon Dynasty. The data was extracted from the Hanguk Munjip Chonggan (Korean Literary Collections in Classical Chinese) in the Hanguk Gojeon Jonghap database (Korea Classics DB). Using computer information processing techniques, we carried out web scraping and classification of the painting poetry from the Hanguk Munjip Chonggan. Subsequently, we narrowed down our focus to the painting poetry specifically related to Maehwado in the early Joseon Dynasty. Based on this, refined dataset, we conducted an in-depth analysis and interpretation of the text data at the syllable corpus level. As a result, we found a direct correlation between the corpus statistics for each syllable in Maehwado painting poetry and the symbolic meaning of plum blossoms.

Assessment of Water Quality using Multivariate Statistical Techniques: A Case Study of the Nakdong River Basin, Korea

  • Park, Seongmook;Kazama, Futaba;Lee, Shunhwa
    • Environmental Engineering Research
    • /
    • v.19 no.3
    • /
    • pp.197-203
    • /
    • 2014
  • This study estimated spatial and seasonal variation of water quality to understand characteristics of Nakdong river basin, Korea. All together 11 parameters (discharge, water temperature, dissolved oxygen, 5-day biochemical oxygen demand, chemical oxygen demand, pH, suspended solids, electrical conductivity, total nitrogen, total phosphorus, and total organic carbon) at 22 different sites for the period of 2003-2011 were analyzed using multivariate statistical techniques (cluster analysis, principal component analysis and factor analysis). Hierarchical cluster analysis grouped whole river basin into three zones, i.e., relatively less polluted (LP), medium polluted (MP) and highly polluted (HP) based on similarity of water quality characteristics. The results of factor analysis/principal component analysis explained up to 83.0%, 81.7% and 82.7% of total variance in water quality data of LP, MP, and HP zones, respectively. The rotated components of PCA obtained from factor analysis indicate that the parameters responsible for water quality variations were mainly related to discharge and total pollution loads (non-point pollution source) in LP, MP and HP areas; organic and nutrient pollution in LP and HP zones; and temperature, DO and TN in LP zone. This study demonstrates the usefulness of multivariate statistical techniques for analysis and interpretation of multi-parameter, multi-location and multi-year data sets.

Statistical Literacy of Fifth and Sixth Graders in Elementary School about the Beginning Inference from a Pictograph Task ('그림그래프에서 추론하기' 과제에서 나타나는 초등학교 5, 6학년 학생들의 통계적 소양)

  • Moon, Eunhye;Lee, Kwangho
    • Education of Primary School Mathematics
    • /
    • v.22 no.3
    • /
    • pp.149-166
    • /
    • 2019
  • The purpose of this study is to analyze the statistical literacy in elementary school students when they beginning inference. Picto-graphs provide statistical information and often data-related arguments they certainly qualify as objects for interpretation, for critical evaluation, and for discussion or communication of the conclusions presented. For research, the inference from pictograph task was designed and statistical literacy standards for evaluating the student's level was presented based on prior studies. Evaluating student's statistical literacy is meaningful in that it can check their current level. To know the student's current level can help them achieve a higher level of performance. The outcomes of this research indicate that pictograph can provide a basis for rich tasks displaying not only student's counting skills but also their appreciation of variation and uncertainty in prediction. Raising statistical thinking by students is an important goal in statistical education, and the experience of informal statistical reasoning can help with formal statistical reasoning that will be learned later. Therefore, the task about the inference from a pictograph, discussions on statistical learning of elementary school children are expected to present meaningful implications for statistical education.

Multi-block Analysis of Genomic Data Using Generalized Canonical Correlation Analysis

  • Jun, Inyoung;Choi, Wooree;Park, Mira
    • Genomics & Informatics
    • /
    • v.16 no.4
    • /
    • pp.33.1-33.9
    • /
    • 2018
  • Recently, there have been many studies in medicine related to genetic analysis. Many genetic studies have been performed to find genes associated with complex diseases. To find out how genes are related to disease, we need to understand not only the simple relationship of genotypes but also the way they are related to phenotype. Multi-block data, which is a summation form of variable sets, is used for enhancing the analysis of the relationships of different blocks. By identifying relationships through a multi-block data form, we can understand the association between the blocks in comprehending the correlation between them. Several statistical analysis methods have been developed to understand the relationship between multi-block data. In this paper, we will use generalized canonical correlation methodology to analyze multi-block data from the Korean Association Resource project, which has a combination of single nucleotide polymorphism blocks, phenotype blocks, and disease blocks.

Estimation methods and interpretation of competing risk regression models (경쟁 위험 회귀 모형의 이해와 추정 방법)

  • Kim, Mijeong
    • The Korean Journal of Applied Statistics
    • /
    • v.29 no.7
    • /
    • pp.1231-1246
    • /
    • 2016
  • Cause-specific hazard model (Prentice et al., 1978) and subdistribution hazard model (Fine and Gray, 1999) are mostly used for the right censored survival data with competing risks. Some other models for survival data with competing risks have been subsequently introduced; however, those models have not been popularly used because the models cannot provide reliable statistical estimation methods or those are overly difficult to compute. We introduce simple and reliable competing risk regression models which have been recently proposed as well as compare their methodologies. We show how to use SAS and R for the data with competing risks. In addition, we analyze survival data with two competing risks using five different models.

A Study on the Activation Plan of 4-H Club in Korea (농촌 청소년조직(4-H)의 활성화 방안에 관한 연구)

  • Yang, Seung Choon;Choi, Chang Wook
    • Journal of Agricultural Extension & Community Development
    • /
    • v.8 no.1
    • /
    • pp.41-58
    • /
    • 2001
  • The purpose of this study was to develop plans for activating the 4-H clubs in Korea. Data for this study were collected from 125 members in 4-H clubs and 140 extension educators who participated in 4-H activity. Total of 265 responses were analyzed after data screening. The Statistical Package for Social Sciences(SPSS) for Windows for the personal computer were used to analyze the data. Frequency, percentage, ANOVA, and LSD test for post-hoc interpretation were employed to analyze the data with a statistical significance level of .05. Based on the conclusions of this study, following recommendations were offered: To activate rural youth organizations, especially 4-H clubs in Korea the following measures should be included in the plans for activation; 1) To classify membership into student and non-student clubs to focus on the needs of active members; 2) To establish clear objectives for club activities; 3) To enhance field-oriented operation of clubs; 4) To develop various activity programs that members could be fascinated; 5) To promote subject matter specialists in order to support club activities effectively; and 6) To clarify functions and roles of extension service centers and non-governmental organizations in order to support club activites.

  • PDF