Search | Korea Science

Binary classification on compositional data

Joo, Jae Yun;Lee, Seokho
- Communications for Statistical Applications and Methods
- /
- v.28 no.1
- /
- pp.89-97
- /
- 2021
Due to boundedness and sum constraint, compositional data are often transformed by logratio transformation and their transformed data are put into traditional binary classification or discriminant analysis. However, it may be problematic to directly apply traditional multivariate approaches to the transformed data because class distributions are not Gaussian and Bayes decision boundary are not polynomial on the transformed space. In this study, we propose to use flexible classification approaches to transformed data for compositional data classification. Empirical studies using synthetic and real examples demonstrate that flexible approaches outperform traditional multivariate classification or discriminant analysis.
https://doi.org/10.29220/CSAM.2021.28.1.089 인용 PDF KSCI

Statistical Methods for Multivariate Missing Data in Health Survey Research (보건조사연구에서 다변량결측치가 내포된 자료를 효율적으로 분석하기 위한 통계학적 방법)

Kim, Dong-Kee;Park, Eun-Cheol;Sohn, Myong-Sei;Kim, Han-Joong;Park, Hyung-Uk;Ahn, Chae-Hyung;Lim, Jong-Gun;Song, Ki-Jun
- Journal of Preventive Medicine and Public Health
- /
- v.31 no.4 s.63
- /
- pp.875-884
- /
- 1998
Missing observations are common in medical research and health survey research. Several statistical methods to handle the missing data problem have been proposed. The EM algorithm (Expectation-Maximization algorithm) is one of the ways of efficiently handling the missing data problem based on sufficient statistics. In this paper, we developed statistical models and methods for survey data with multivariate missing observations. Especially, we adopted the EM algorithm to handle the multivariate missing observations. We assume that the multivariate observations follow a multivariate normal distribution, where the mean vector and the covariance matrix are primarily of interest. We applied the proposed statistical method to analyze data from a health survey. The data set we used came from a physician survey on Resource-Based Relative Value Scale(RBRVS). In addition to the EM algorithm, we applied the complete case analysis, which uses only completely observed cases, and the available case analysis, which utilizes all available information. The residual and normal probability plots were evaluated to access the assumption of normality. We found that the residual sum of squares from the EM algorithm was smaller than those of the complete-case and the available-case analyses.
PDF

Optimal Designs for Multivariate Nonparametric Kernel Regression with Binary Data

Park, Dong-Ryeon
- Communications for Statistical Applications and Methods
- /
- v.2 no.2
- /
- pp.243-248
- /
- 1995
The problem of optimal design for a nonparametric regression with binary data is considered. The aim of the statistical analysis is the estimation of a quantal response surface in two dimensions. Bias, variance and IMSE of kernel estimates are derived. The optimal design density with respect to asymptotic IMSE is constructed.
PDF

MULTIVARIATE JOINT NORMAL LIKELIHOOD DISTANCE

Kim, Myung-Geun
- Journal of applied mathematics & informatics
- /
- v.27 no.5_6
- /
- pp.1429-1433
- /
- 2009
The likelihood distance for the joint distribution of two multivariate normal distributions with common covariance matrix is explicitly derived. It is useful for identifying outliers which do not follow the joint multivariate normal distribution with common covariance matrix. The likelihood distance derived here is a good ground for the use of a generalized Wilks statistic in influence analysis of two multivariate normal data.
PDF

AUTOMATED ELECTROFACIES DETERMINATION USING MULTIVARIATE STATISTICAL ANALYSIS

Kim Jungwhan;Lim Jong-Se
- 한국석유지질학회:학술대회논문집
- /
- spring
- /
- pp.10-14
- /
- 1998
A systematic methodology is developed for the electrofacies determination from wireline log data using multivariate statistical analysis. To consider corresponding contribution of each log and reduce the computational dimension, multivariate logs are transformed into a single variable through principal components analysis. Resultant principal components logs are segmented using the statistical zonation method to enhance the efficiency and quality of the interpreted results. Hierarchical cluster analysis is then used to group the segments into electrofacies. Optimal number of groups is determined on the basis of the ratio of within-group variance to total variance and core data. This technique is applied to the wells in the Korea Continental Shelf. The results of field application demonstrate that the prediction of lithology based on the electrofacies classification matches well to the core and the cutting data with high reliability This methodology for electrofacies classification can be used to define the reservoir characteristics which are helpful to the reservoir management.
PDF

USING MULTIVARIATE DATA ANALYSIS FOR PROCESS TROUBLE SHOOTING

Winchell, Patricia
- Proceedings of the Korea Technical Association of the Pulp and Paper Industry Conference
- /
- 2006.06b
- /
- pp.191-195
- /
- 2006
Multivariate data analysis tools were used to improve the understanding of the wet end chemistry and white water system of the Papermill at NorskeCanada Crofton Division. Specifically, the analysis was aimed at identifying what variables were contributing to increased retention aid use and wet end instability. Several models were developed using data sets with up to 88 process variables and over 3000 observations. It was found that increased retention aid use was driven primarily by PCC and TMP usage as well as the addition of Alaskan White Spruce to the TMP furnish.
PDF

Multivariate Analysis of Variance for Fuzzy Data

Kang, Man-Ki;Han, Sung-Il
- International Journal of Fuzzy Logic and Intelligent Systems
- /
- v.4 no.1
- /
- pp.97-100
- /
- 2004
We propose some properties of fuzzy multivariate analysis of variance by fuzzy vector operation with agreement index. We deals fuzzy null hypotheses and fuzzy alternative hypothesis and define the agreement index for the grades of the judgements that the hypothesis is rejection or acceptance. Finally, we provide an example to evaluate the judgements.
https://doi.org/10.5391/IJFIS.2004.4.1.097 인용 PDF KSCI

Multivariate Control Chart for Autocorrelated Process (자기상관자료를 갖는 공정을 위한 다변량 관리도)

Nam, Gook-Hyun;Chang, Young-Soon;Bai, Do-Sun
- Journal of Korean Institute of Industrial Engineers
- /
- v.27 no.3
- /
- pp.289-296
- /
- 2001
This paper proposes multivariate control chart for autocorrelated data which are common in chemical and process industries and lead to increase in the number of false alarms when conventional control charts are applied. The effect of autocorrelated data is modeled as a vector autoregressive process, and canonical analysis is used to reduce the dimensionality of the data set and find the canonical variables that explain as much of the data variation as possible. Charting statistics are constructed based on the residual vectors from the canonical variables which are uncorrelated over time, and therefore the control charts for these statistics can attenuate the autocorrelation in the process data. The charting procedures are illustrated with a numerical example and Monte Carlo simulation is conducted to investigate the performances of the proposed control charts.
PDF

Relevance of Multivariate Analysis in Management Research

Ojha, Sateesh Kumar
- Journal of Information Technology Applications and Management
- /
- v.23 no.3
- /
- pp.25-34
- /
- 2016
Often we receive misled conclusion in the research if properly variables are not analyzed. In different functional issues of management it is very essential that all the latent and observed variable are properly understood so management decisions will be relevant and effective. The objective of this paper is to investigate the use of different multivariate tools for analyzing in the management research : applied or basic. The sources of data is primary as well as secondary. The primary includes the observation of different research articles of the proceedings of different conferences. And the secondary includes different publications related to multivariate analysis. The study has revealed the reasons of not using such tools of research. The preliminary finding reveals that most of the researches do not use such analytical tools in a comprehensive manner. Carelessness in design while fixing the design aspect is the main reasons of not using appropriate design.
https://doi.org/10.21219/jitam.2016.23.3.025 인용 PDF KSCI

Residuals Plots for Repeated Measures Data

PARK TAESUNG
- Proceedings of the Korean Statistical Society Conference
- /
- 2000.11a
- /
- pp.187-191
- /
- 2000
In the analysis of repeated measurements, multivariate regression models that account for the correlations among the observations from the same subject are widely used. Like the usual univariate regression models, these multivariate regression models also need some model diagnostic procedures. In this paper, we propose a simple graphical method to detect outliers and to investigate the goodness of model fit in repeated measures data. The graphical method is based on the quantile-quantile(Q-Q) plots of the $X^2$ distribution and the standard normal distribution. We also propose diagnostic measures to detect influential observations. The proposed method is illustrated using two examples.
PDF

Search Result 1,427, Processing Time 0.024 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)