• Title/Summary/Keyword: R 텍스트 마이닝

Search Result 89, Processing Time 0.021 seconds

Visualizing the Results of Opinion Mining from Social Media Contents: Case Study of a Noodle Company (소셜미디어 콘텐츠의 오피니언 마이닝결과 시각화: N라면 사례 분석 연구)

  • Kim, Yoosin;Kwon, Do Young;Jeong, Seung Ryul
    • Journal of Intelligence and Information Systems
    • /
    • v.20 no.4
    • /
    • pp.89-105
    • /
    • 2014
  • After emergence of Internet, social media with highly interactive Web 2.0 applications has provided very user friendly means for consumers and companies to communicate with each other. Users have routinely published contents involving their opinions and interests in social media such as blogs, forums, chatting rooms, and discussion boards, and the contents are released real-time in the Internet. For that reason, many researchers and marketers regard social media contents as the source of information for business analytics to develop business insights, and many studies have reported results on mining business intelligence from Social media content. In particular, opinion mining and sentiment analysis, as a technique to extract, classify, understand, and assess the opinions implicit in text contents, are frequently applied into social media content analysis because it emphasizes determining sentiment polarity and extracting authors' opinions. A number of frameworks, methods, techniques and tools have been presented by these researchers. However, we have found some weaknesses from their methods which are often technically complicated and are not sufficiently user-friendly for helping business decisions and planning. In this study, we attempted to formulate a more comprehensive and practical approach to conduct opinion mining with visual deliverables. First, we described the entire cycle of practical opinion mining using Social media content from the initial data gathering stage to the final presentation session. Our proposed approach to opinion mining consists of four phases: collecting, qualifying, analyzing, and visualizing. In the first phase, analysts have to choose target social media. Each target media requires different ways for analysts to gain access. There are open-API, searching tools, DB2DB interface, purchasing contents, and so son. Second phase is pre-processing to generate useful materials for meaningful analysis. If we do not remove garbage data, results of social media analysis will not provide meaningful and useful business insights. To clean social media data, natural language processing techniques should be applied. The next step is the opinion mining phase where the cleansed social media content set is to be analyzed. The qualified data set includes not only user-generated contents but also content identification information such as creation date, author name, user id, content id, hit counts, review or reply, favorite, etc. Depending on the purpose of the analysis, researchers or data analysts can select a suitable mining tool. Topic extraction and buzz analysis are usually related to market trends analysis, while sentiment analysis is utilized to conduct reputation analysis. There are also various applications, such as stock prediction, product recommendation, sales forecasting, and so on. The last phase is visualization and presentation of analysis results. The major focus and purpose of this phase are to explain results of analysis and help users to comprehend its meaning. Therefore, to the extent possible, deliverables from this phase should be made simple, clear and easy to understand, rather than complex and flashy. To illustrate our approach, we conducted a case study on a leading Korean instant noodle company. We targeted the leading company, NS Food, with 66.5% of market share; the firm has kept No. 1 position in the Korean "Ramen" business for several decades. We collected a total of 11,869 pieces of contents including blogs, forum contents and news articles. After collecting social media content data, we generated instant noodle business specific language resources for data manipulation and analysis using natural language processing. In addition, we tried to classify contents in more detail categories such as marketing features, environment, reputation, etc. In those phase, we used free ware software programs such as TM, KoNLP, ggplot2 and plyr packages in R project. As the result, we presented several useful visualization outputs like domain specific lexicons, volume and sentiment graphs, topic word cloud, heat maps, valence tree map, and other visualized images to provide vivid, full-colored examples using open library software packages of the R project. Business actors can quickly detect areas by a swift glance that are weak, strong, positive, negative, quiet or loud. Heat map is able to explain movement of sentiment or volume in categories and time matrix which shows density of color on time periods. Valence tree map, one of the most comprehensive and holistic visualization models, should be very helpful for analysts and decision makers to quickly understand the "big picture" business situation with a hierarchical structure since tree-map can present buzz volume and sentiment with a visualized result in a certain period. This case study offers real-world business insights from market sensing which would demonstrate to practical-minded business users how they can use these types of results for timely decision making in response to on-going changes in the market. We believe our approach can provide practical and reliable guide to opinion mining with visualized results that are immediately useful, not just in food industry but in other industries as well.

Analysis of Research Trends in SIAM Journal on Applied Mathematics Using Topic Modeling (토픽모델링을 활용한 SIAM Journal on Applied Mathematics의 연구 동향 분석)

  • Kim, Sung-Yeun
    • Journal of the Korea Academia-Industrial cooperation Society
    • /
    • v.21 no.7
    • /
    • pp.607-615
    • /
    • 2020
  • The purpose of this study was to analyze the research status and trends related to the industrial mathematics based on text mining techniques with a sample of 4910 papers collected in the SIAM Journal on Applied Mathematics from 1970 to 2019. The R program was used to collect titles, abstracts, and key words from the papers and to analyze topic modeling techniques based on LDA algorithm. As a result of the coherence score on the collected papers, 20 topics were determined optimally using the Gibbs sampling methods. The main results were as follows. First, studies on industrial mathematics were conducted in a variety of mathematics fields, including computational mathematics, geometry, mathematical modeling, topology, discrete mathematics, probability and statistics, with a focus on analysis and algebra. Second, 5 hot topics (mathematical biology, nonlinear partial differential equation, discrete mathematics, statistics, topology) and 1 cold topic (probability theory) were found based on time series regression analysis. Third, among the fields that were not reflected in the 2015 revised mathematics curriculum, numeral system, matrix, vector in space, and complex numbers were extracted as the contents to be covered in the high school mathematical curriculum. Finally, this study suggested strategies to activate industrial mathematics in Korea, described the study limitations, and proposed directions for future research.

Text-Mining Analysis on the Interaction between the American Consumers Aged over 60 and Companion Pets Robots: Focused on Amazon Reviews for Joy For All Companion Pets (텍스트 마이닝을 활용한 미국 노년 소비자와 애완용 로봇 간 상호작용에 대한 분석: Joy For All Companion Pets에 대한 아마존 리뷰를 중심으로)

  • Chung, Yea-Eun;Lee, Yu Lim;Chung, Jae-Eun
    • Journal of Digital Convergence
    • /
    • v.19 no.10
    • /
    • pp.469-489
    • /
    • 2021
  • This study explores consumers' responses to socially assistive robotics by using text-mining method focusing on Companion Pets from Hasbro as it gives emotional support. We conducted text frequency analysis, LDA analysis using R programming. The key findings are 1)the most frequently used words the mimicry of living pets and the appearance of companion pets, 2)the five topics were derived from the LDA analysis and classified keywords in each topic split between positive and negative, 3)user, product, environment affect the interaction between consumer and companion pets, 4)consumers who have difficulty in cognition and physical conditions use companion pets to replace living pets. This study provides an understanding of consumer responses in companion pets and gives practical implications that may improve the efficacy of usage for consumers and understand the companion robot, which provides emotional support in COVID-19.

An Analysis of Changes in Social Issues Related to Patient Safety Using Topic Modeling and Word Co-occurrence Analysis (토픽 모델링과 동시출현 단어 분석을 활용한 환자안전 관련 사회적 이슈의 변화)

  • Kim, Nari;Lee, Nam-Ju
    • The Journal of the Korea Contents Association
    • /
    • v.21 no.1
    • /
    • pp.92-104
    • /
    • 2021
  • This study aims to analyze online news articles to identify social issues related to patient safety and compare the changes in these issues before and after the implementation of the Patient Safety Act. This study performed text mining through the R program, wherein 7,600 online news articles were collected from January 1, 2010, to March 5, 2020, and examined using keyword analysis, topic modeling, and word co-occurrence network analysis. A total of 2,609 keywords were categorized into 8 topics: "medical practice", "medical personnel", "infection and facilities", "comprehensive nursing service", "medicine and medical supplies", "system development and establishment for improvement", "Patient Safety Act" and "healthcare accreditation". The study revealed that keywords such as "patient safety awareness", "infection control" and "healthcare accreditation" appeared before the implementation of the Patient Safety Act. Meanwhile, keywords such as "patient safety culture". and "administration and injection" appeared after the act's implementation with improved ranking of importance pertaining to nursing-related terminology. Interest in patient safety has increased in the medical community as well as among the public. In particular, nursing plays an important role in improving patient safety. Therefore, the recognition of patient safety as a core competency of nursing and the persistent education of the public are vital and inevitable.

Analysis on Research Trends in Sport Facilities: Focusing on SCOPUS DB (스포츠시설에 관한 연구 동향 분석: SCOPUS DB를 중심으로)

  • Kim, Il-Gwang;Park, Seong-Taek;Park, Su-Sun;Kim, Mi-Suk;Park, Jong-Chul;Jiang, Jialei
    • Journal of Industrial Convergence
    • /
    • v.19 no.6
    • /
    • pp.11-19
    • /
    • 2021
  • The purpose of this study is to explore trends in research at home and abroad related to "Sport Facilities", and seek the direction of further research. 1,801 abstracts of papers including "Sport Facilities" were collected from the SCOPUS DB from 2016 to 2020. Topic modeling techniques based on Latent Dirichlet Allocation (LDA) algorithm implemented in R language, TD-IDF techniques, and word cluds using Tagxedo was conducted to analyze the data. As a result, 8 topics were optimally determined, and "sports", "facilities", "health", "physical", "data", and "using" were derived as the main keywords for topics. This results indicated that studies on physical activity, health and using facilities regarding sports facilities at home and abroad have been actively carried out in recent years. This indicates that papers in SCOPUS DB are paying attention to the instrumental value of sport facilities, such as health promotion and improving the quality of life. Therefore, various studies that help participants who use sport facilities for a healthy life should be continuously conducted in the future.

Study on the Analysis of National Paralympics by Utilizing Social Big Data Text Mining (소셜 빅데이터 텍스트 마이닝을 활용한 전국장애인체육대회 분석 연구)

  • Kim, Dae kyung;Lee, Hyun Su
    • 한국체육학회지인문사회과학편
    • /
    • v.55 no.6
    • /
    • pp.801-810
    • /
    • 2016
  • The purpose of the study was to conduct a text mining examining keywords related to the National Paralympics and provide the fundamental information that would be used to change perception of people without disabilities toward disabilities and to promote the social participation of people with and without disabilities in the National Paralympics. Social big data regarding the National Paralympics were retrieved from news articles and blog postings identified by search engines, Naver, Daum, and Google. The data were then analysed using R-3.3.1 Version Program. The analysing techniques were cloud analysis, correlation analysis and social network analysis. The results were as follows. First, news were mainly related to game results, sports events, team participation and host avenue of the 33rd ~ 36th National Paralympics. Second, search results about the 33rd ~ 36th National Paralympics between Naver, Daum, and Google were similar to one another. Thirds, the keywrods, National Paralympics, sports for the disabled, and sports, demonstrated a high close centrality. Further, degree centrality and betweenness centrality were associated in the keywords such as sports for all, participation, research, development, sports-disabled, research-disabled, sports for all-participation, disabled-participation, sports for all-disabled, and host-paralympics.

The Characteristics and Improvement Directions of Regional Climate Change Adaptation Policies in accordance with Damage Cases (지자체 기후변화 적응 대책 특성 및 개선 방향)

  • Ahn, Yoonjung;Kang, Youngeun;Park, Chang Sug;Kim, Ho Gul
    • Journal of Environmental Impact Assessment
    • /
    • v.25 no.4
    • /
    • pp.296-306
    • /
    • 2016
  • There is a growing interest in establishing a regional climate change adaptation policy as the climate change impact in the region and local scale increases. This study focused on the analysis of 32 regions on its characteristics of local climate change adaptation plans. First, statistic program R was used for conducting cluster analysis based on the frequency and budgets of adaptation plan. Further, we analyzed damage frequency from newspapers regarding climate change impacts in eight categories which were caused by extreme weather events on 2,565 cases for 24 years. Lastly, the characteristics of climate change adaptation plan was compared with damage frequency patterns for evaluating the adequacy of climate change adaptation plan on each cluster. Four different clusters were created by cluster analysis. Most clusters clearly have their own characteristics on certain sectors. There was a high frequency of damage in 'disaster' and 'health' sectors. Climate change adaptation plan and budget also invested a lot on those sectors. However, when comparing the relative rate among regional governments, there was a difference between types of damage and climate change adaptation plan. We assumed that the difference could come from that each region established their adaptation plans based on not only the frequency of damage, but vulnerability assessment, and expert opinions as well. The result of study could contribute to policy making of climate change adaptation plan.

6G Technology Competitiveness and Network Analysis: Focusing on GaN Integrated Circuit Patent Data (6G의 기술경쟁력 및 네트워크 분석: GaN 집적회로 특허 데이터 중심)

  • Woo-Seok Choi;Jin-Yong Kim;Jung-Hwan Lee;Sang-Hyun Choi
    • Journal of Industrial Convergence
    • /
    • v.21 no.3
    • /
    • pp.1-15
    • /
    • 2023
  • Expectations for wireless communication technology are rising as a base technology that promotes innovation in various industries in line with the paradigm of digital transformation in the 21st century beyond the stage of being used only for communication service itself. In this study, in order to compare 6G technological competitiveness between Korea and leading countries, technological competitiveness was confirmed through PFS, CPP, and network analysis based on GaN Integrated Circuit patent data. Korea's 6G technological competitiveness was 0.62 in PFS and 3.93 in CPP, which were 32.8% and 19.9%, respectively, compared to leading countries. In addition, as a result of network analysis, the collaboration rate in the 6G field was 7.2%, and the collaboration ecosystem was very insufficient in most countries. In contrast, it was confirmed that Korea, unlike leading countries, has established a small-scale collaboration ecosystem linked by industry and academia. Thus, it is necessary to establish a strategy for 6G communication technology at the national level so that communication technology can be advanced based on a relatively well-established collaborative ecosystem.

Text Mining of Successful Casebook of Agricultural Settlement in Graduates of Korea National College of Agriculture and Fisheries - Frequency Analysis and Word Cloud of Key Words - (한국농수산대학 졸업생 영농정착 성공 사례집의 Text Mining - 주요단어의 빈도 분석 및 word cloud -)

  • Joo, J.S.;Kim, J.S.;Park, S.Y.;Song, C.Y.
    • Journal of Practical Agriculture & Fisheries Research
    • /
    • v.20 no.2
    • /
    • pp.57-72
    • /
    • 2018
  • In order to extract meaningful information from the excellent farming settlement cases of young farmers published by KNCAF, we studied the key words with text mining and created a word cloud for visualization. First, in the text mining results for the entire sample, the words 'CEO', 'corporate executive', 'think', 'self', 'start', 'mind', and 'effort' are the words with high frequency among the top 50 core words. Their ability to think, judge and push ahead with themselves is a result of showing that they have ability of to be managers or managers. And it is a expression of how they manages to achieve their dream without giving up their dream. The high frequency of words such as "father" and "parent" is due to the high ratio of parents' cooperation and succession. Also 'KNCAF', 'university', 'graduation' and 'study' are the results of their high educational awareness, and 'organic farming' and 'eco-friendly' are the result of the interest in eco-friendly agriculture. In addition, words related to the 6th industry such as 'sales' and 'experience' represent their efforts to revitalize farming and fishing villages. Meanwhile, 'internet', 'blog', 'online', 'SNS', 'ICT', 'composite' and 'smart' were not included in the top 50. However, the fact that these words were extracted without omission shows that young farmers are increasingly interested in the scientificization and high-tech of agriculture and fisheries Next, as a result of grouping the top 50 key words by crop, the words 'facilities' in livestock, vegetables and aquatic crops, the words 'equipment' and 'machine' in food crops were extracted as main words. 'Eco-friendly' and 'organic' appeared in vegetable crops and food crops, and 'organic' appeared in fruit crops. The 'worm' of eco-friendly farming method appeared in the food crops, and the 'certification', which means excellent agricultural and marine products, appeared only in the fishery crops. 'Production', which is related to '6th industry', appeared in all crops, 'processing' and 'distribution' appeared in the fruit crops, and 'experience' appeared in the vegetable crops, food crops and fruit crops. To visualize the extracted words by text mining, we created a word cloud with the entire samples and each crop sample. As a result, we were able to judge the meaning of excellent practices, which are unstructured text, by character size.