• 제목/요약/키워드: text visualization

검색결과 216건 처리시간 0.024초

천문학 카탈로그 자료의 통합검색 DB 구축 (JAVA BASED WEB APPLICATION FOR THE ASTRONOMICAL CATALOGUES)

  • 성현일;;김봉규;임인성;김상철;안영숙
    • 천문학논총
    • /
    • 제20권1호
    • /
    • pp.85-95
    • /
    • 2005
  • We collected eleven large astronomical catalogues, which include 2MASS, USNO B1.0, GSC 2 catalogues and so on. Most of these catalogues are the frequently used by astronomers for all sorts of applications. But the researches are faced with the problem of accessing these databases because these catalogues contain from tens millions up to thousands of millions of records. So we developed a web application system to manage these large catalogues, the main purpose of the web application is to allow a powerful and efficient querying activity on these catalogues through internet by using a simple web interface. User could retrieve the query result in variety of formats including plain text, HTML, Microsoft Excel format (XLS), and VOTable. Furthermore, user also could display and analyze result graphically by using a powerful interactive visualization tools named VOPlot which was developed by the Virtual Observatory-India (VOI) project.

Topographic non-negative matrix factorization에 기반한 텍스트 문서로부터의 토픽 가시화 (Topographic Non-negative Matrix Factorization for Topic Visualization from Text Documents)

  • 장정호;엄재홍;장병탁
    • 한국정보과학회:학술대회논문집
    • /
    • 한국정보과학회 2006년도 가을 학술발표논문집 Vol.33 No.2 (B)
    • /
    • pp.324-329
    • /
    • 2006
  • Non-negative matrix factorization(NMF) 기법은 음이 아닌 값으로 구성된 데이터를 두 종류의 양의 행렬의 곱의 형식으로 분할하는 데이터 분석기법으로서, 텍스트마이닝, 바이오인포매틱스, 멀티미디어 데이터 분석 등에 활용되었다. 본 연구에서는 기본 NMF 기법에 기반하여 텍스트 문서로부터 토픽을 추출하고 동시에 이를 가시적으로 도시하기 위한 Topographic NMF (TNMF) 기법을 제안한다. TNMF에 의한 토픽 가시화는 데이터를 전체적인 관점에서 보다 직관적으로 파악하는데 도움이 될 수 있다. TNMF는 생성모델 관점에서 볼 때, 2개의 은닉층을 갖는 계층적 모델로 표현할 수 있으며, 상위 은닉층에서 하위 은닉층으로의 연결은 토픽공간상에서 토픽간의 전이확률 또는 이웃함수를 정의한다. TNMF에서의 학습은 전이확률값의 연속적 스케줄링 과정 속에서 반복적 파리미터 갱신 과정을 통해 학습이 이루어지는데, 파라미터 갱신은 기본 NMF 기반 학습 과정으로부터 유사한 형태로 유도될 수 있음을 보인다. 추가적으로 Probabilistic LSA에 기초한 토픽 가시화 기법 및 희소(sparse)한 해(解) 도출을 목적으로 한 non-smooth NMF 기법과의 연관성을 분석, 제시한다. NIPS 학회 논문 데이터에 대한 실험을 통해 제안된 방법론이 문서 내에 내재된 토픽들을 효과적으로 가시화 할 수 있음을 제시한다.

  • PDF

데이터마이닝을 이용한 동의보감에서 경락의 주치특성 분석 (An Analysis of Indications of Meridians in DongUiBoGam Using Data Mining)

  • 채윤병;류연희;정원모
    • Korean Journal of Acupuncture
    • /
    • 제36권4호
    • /
    • pp.292-299
    • /
    • 2019
  • Objectives : DongUiBoGam is one of the representative medical literatures in Korea. We used text mining methods and analyzed the characteristics of the indications of each meridian in the second chapter of DongUiBoGam, WaeHyeong, which addresses external body elements. We also visualized the relationships between the meridians and the disease sites. Methods : Using the term frequency-inverse document frequency (TF-IDF) method, we quantified values regarding the indications of each meridian according to the frequency of the occurrences of 14 meridians and 14 disease sites. The spatial patterns of the indications of each meridian were visualized on a human body template according to the TF-IDF values. Using hierarchical clustering methods, twelve meridians were clustered into four groups based on the TF-IDF distributions of each meridian. Results : TF-IDF values of each meridian showed different constellation patterns at different disease sites. The spatial patterns of the indications of each meridian were similar to the route of the corresponding meridian. Conclusions : The present study identified spatial patterns between meridians and disease sites. These findings suggest that the constellations of the indications of meridians are primarily associated with the lines of the meridian system. We strongly believe that these findings will further the current understanding of indications of acupoints and meridians.

Big Data Smoothing and Outlier Removal for Patent Big Data Analysis

  • Choi, JunHyeog;Jun, Sunghae
    • 한국컴퓨터정보학회논문지
    • /
    • 제21권8호
    • /
    • pp.77-84
    • /
    • 2016
  • In general statistical analysis, we need to make a normal assumption. If this assumption is not satisfied, we cannot expect a good result of statistical data analysis. Most of statistical methods processing the outlier and noise also need to the assumption. But the assumption is not satisfied in big data because of its large volume and heterogeneity. So we propose a methodology based on box-plot and data smoothing for controling outlier and noise in big data analysis. The proposed methodology is not dependent upon the normal assumption. In addition, we select patent documents as target domain of big data because patent big data analysis is a important issue in management of technology. We analyze patent documents using big data learning methods for technology analysis. The collected patent data from patent databases on the world are preprocessed and analyzed by text mining and statistics. But the most researches about patent big data analysis did not consider the outlier and noise problem. This problem decreases the accuracy of prediction and increases the variance of parameter estimation. In this paper, we check the existence of the outlier and noise in patent big data. To know whether the outlier is or not in the patent big data, we use box-plot and smoothing visualization. We use the patent documents related to three dimensional printing technology to illustrate how the proposed methodology can be used for finding the existence of noise in the searched patent big data.

A Novel Interactive Power Electronics Seminar (iPES) Developed at the Swiss Federal Institute of Technology (ETH) Zurich

  • Drofenik, Uwe;Kolar, Johann W.
    • Journal of Power Electronics
    • /
    • 제2권4호
    • /
    • pp.250-257
    • /
    • 2002
  • This paper introduces the Interactive Power Electronics Seminar - iPES - a new software package for teaching of fundamentals of power electronic circuits and systems. iPES is constituted by HTML text with Java applets for interactive animation, circuit design and simulation and visualization of electromagnetic fields and thermal issues in power electronics. It does comprise an easy-to-use self-explaining graphical user interface. The software does need just a standard web-browser, i.e. no installations are required. iPES can be accessed via the World Wide Web or from a CD-ROM in a stand-alone PC by students and professionals. Due to the underlying software technology iPES is very flexible and could be used for on-line learning and could easily be integrated into an e-learning platform. The aim of this paper Is to give an introduction to the iPES-project and to show the different areas covered. The e- learning software is available at no costs at $\underline{www.ipes.ethz.ch}$ in English, German, Japanese, Korean, Chinese and Spanish. The project is still under development and the web page is updated in about 4 weeks intervals.

A New Communication Network Model for Chat Agents in Virtual Space

  • Kim, Jong-Woo;Ji, Seong-Hyun;Kim, Seon-Yeong;Cho, Hwan-Gue
    • KSII Transactions on Internet and Information Systems (TIIS)
    • /
    • 제5권2호
    • /
    • pp.287-312
    • /
    • 2011
  • Internet chat programs and instant messaging services are becoming increasingly popular among Internet users. One of the crucial issues with Internet chat is how to manage the corresponding pairs of questions and answers in a sequence of conversations. Although many novel methodologies have been introduced to cope with this problem, most are poor in managing interruptions, organizing turn-taking, and conveying comprehension. The Internet environment is recently evolving into a 3D environment, but the problems with managing chat dialogues with the standard 2D text-based chat have remained. Therefore, we propose a more realistic communication model for chat agents in 3D virtual space in this paper. First, we propose a new method to measure the capacity of communication between chat agents and a novel visualization method to depict the hierarchical structure of chat dialogues. In addition, we are concerned with communication networks for virtual people (avatars) living in virtual worlds. In this paper we consider a microscopic aspect of a social network in a relatively short period of time. Our experiments show that our model is highly effective in a virtual chat environment, and the communication network based on our model greatly facilitates investigation of a very large and complicated communication network.

신문기사 분석을 통한 이슈 라이프사이클에 관한 연구 - 삼성 갤럭시노트7 사례 - (A Study on the Issue Lifecycle through the Analysis of News Texts - A Case of Samsung Galaxy Note 7 -)

  • 허필희;김양석;이충권
    • 스마트미디어저널
    • /
    • 제7권4호
    • /
    • pp.99-105
    • /
    • 2018
  • 시장에 내놓은 제품이나 서비스가 문제를 일으켜서 해당 기업의 비즈니스와 이미지에 큰 타격을 입히는 사례가 자주 발생하고 있다. 발생한 문제에 적절하게 대응하고 피해를 최소화하는 것은 기업들에게 매우 중요한 일이다. 본 연구는 스마트폰 시장을 주도하고 있는 삼성이 개발하여 출시한 갤럭시노트7의 리콜 사태를 다루었던 뉴스를 수집하고 분석하였다. 이슈 라이프사이클에 기반하여 단계별로 뉴스의 특징을 표현하고 연관규칙을 이용하여 텍스트의 내용을 분석하고 시각화하였다. 본 연구의 결과는 이슈의 변화와 흐름을 이해하고 대응책을 모색해야 하는 기업들의 비즈니스 활동에 도움을 줄 것으로 기대된다.

소셜데이터 감성분석을 통한 사용자의 호감도 분석 (Favorable analysis of users through the social data analysis based on sentimental analysis)

  • 이민규;손효정;성백민;김종배
    • 한국정보통신학회:학술대회논문집
    • /
    • 한국정보통신학회 2014년도 추계학술대회
    • /
    • pp.438-440
    • /
    • 2014
  • 최근 폭발적으로 증가하는 SNS서비스의 상업적으로 이용하려는 움직임이 활발하다. 따라서 본 논문은 실시간 SNS 환경에서 제조기업과 제품의 평판에 관련된 정보를 정확하게 분석 할 수 있는 방안을 제시한다. 크롤링 방식으로 수집된 SNS의 텍스트 데이터들에 대한 형태소 분석을 수행하여 단어 간 연관성을 파악한다. 또, 문장에서 추출된 형태소는 구축된 감성사전을 통해 통계적으로 분석하여 이를 시각화 하여 보여준다. 이때, 추출된 단어가 감성사전에 존재하지 않을 경우 이를 자동으로 추가하는 기법을 제안한다.

  • PDF

현대 소비자의 공간소비행동에 관한 연구 -소셜미디어 데이터 분석을 중심으로- (A Study on Space Consumption Behavior of Contemporary Consumers -Focusing on Analysis of Social Media Big Data-)

  • 안서영;고애란
    • 한국의류학회지
    • /
    • 제44권5호
    • /
    • pp.1019-1035
    • /
    • 2020
  • This study examines the millennial generation, who express themselves and share information on social media after experiencing constantly changing 'hot places' (places of interest) in contemporary cities, with the goal of analyzing space consumption behaviors. Data were collected via an Instagram crawler application developed with Python 3.4 administered to 19,262 posts using the term 'hot places' from November 1 and December 15, 2019. Issues were derived from a text mining technique using Textom 2.0; in addition, semantic network analysis using Ucinet6 and the NetDraw program were also conducted. The results are as follows. First, a frequency analysis of keywords for hot places indicated words frequently found in nouns were related to food, local names, SNS and timing. Words related to positive emotions felt in experience, and words related to behavior in hot places appeared in predicate. Based on importance, communication is the most important keyword and influenced all issues. Second, the results of visualization of semantic network analysis revealed four categories in the scope of the definition of "hot place": (1) culinary exploration, (2) atmosphere of cafés, (3) happy daily life of 'me' expressed in images, (4) emotional photos.

Analysis of Infertility Keywords in the Largest Domestic Mom Cafe Bulletin Board in Korea Using Text Mining

  • Sangmin Lee
    • 인터넷정보학회논문지
    • /
    • 제24권4호
    • /
    • pp.137-144
    • /
    • 2023
  • The purpose of this study is to examine consumers' perceptions of domestic infertility support policies based on infertility-related keywords and the trends of their changes. To this end, Momsholic, a mom cafe which has the most active infertility-related bulletin boards on Naver, was selected as the analysis target, and 'infertility' was selected as a keyword for data search. The data was collected for three months. In addition, network analysis and visualization were performed using R for data collection and analysis, and cross-validation was attempted using the NetDraw function of 'textom 1.0' and the UCINET6 program. As a result of the analysis, the main keywords were cost, artificial insemination, in vitro fertilization, freezing, harvest, ovulation, and how much. Next, looking at the central value of the degree of connection, it was found that the degree of connection between the words cost, cost, how much, problem, public health center, and artificial insemination was high. According to the results of this study, women who visit mom cafes due to infertility in Korea are more interested in the cost. It is believed to be closely related to infertility treatment as well as in vitro fertilization and egg freezing. Therefore, by examining keywords related toinfertility, it has academic significance in that it is possible to identify major factors that end users are interested in. Furthermore, it is possible to redefine the guidelines for domestic infertility support policies by presenting infertility support policies that reflect the factors of interest of end consumers.