• Title/Summary/Keyword: support vector regression machine

Search Result 386, Processing Time 0.028 seconds

Movie Popularity Classification Based on Support Vector Machine Combined with Social Network Analysis

  • Dorjmaa, Tserendulam;Shin, Taeksoo
    • Journal of Information Technology Services
    • /
    • v.16 no.3
    • /
    • pp.167-183
    • /
    • 2017
  • The rapid growth of information technology and mobile service platforms, i.e., internet, google, and facebook, etc. has led the abundance of data. Due to this environment, the world is now facing a revolution in the process that data is searched, collected, stored, and shared. Abundance of data gives us several opportunities to knowledge discovery and data mining techniques. In recent years, data mining methods as a solution to discovery and extraction of available knowledge in database has been more popular in e-commerce service fields such as, in particular, movie recommendation. However, most of the classification approaches for predicting the movie popularity have used only several types of information of the movie such as actor, director, rating score, language and countries etc. In this study, we propose a classification-based support vector machine (SVM) model for predicting the movie popularity based on movie's genre data and social network data. Social network analysis (SNA) is used for improving the classification accuracy. This study builds the movies' network (one mode network) based on initial data which is a two mode network as user-to-movie network. For the proposed method we computed degree centrality, betweenness centrality, closeness centrality, and eigenvector centrality as centrality measures in movie's network. Those four centrality values and movies' genre data were used to classify the movie popularity in this study. The logistic regression, neural network, $na{\ddot{i}}ve$ Bayes classifier, and decision tree as benchmarking models for movie popularity classification were also used for comparison with the performance of our proposed model. To assess the classifier's performance accuracy this study used MovieLens data as an open database. Our empirical results indicate that our proposed model with movie's genre and centrality data has by approximately 0% higher accuracy than other classification models with only movie's genre data. The implications of our results show that our proposed model can be used for improving movie popularity classification accuracy.

Landslide susceptibility assessment using feature selection-based machine learning models

  • Liu, Lei-Lei;Yang, Can;Wang, Xiao-Mi
    • Geomechanics and Engineering
    • /
    • v.25 no.1
    • /
    • pp.1-16
    • /
    • 2021
  • Machine learning models have been widely used for landslide susceptibility assessment (LSA) in recent years. The large number of inputs or conditioning factors for these models, however, can reduce the computation efficiency and increase the difficulty in collecting data. Feature selection is a good tool to address this problem by selecting the most important features among all factors to reduce the size of the input variables. However, two important questions need to be solved: (1) how do feature selection methods affect the performance of machine learning models? and (2) which feature selection method is the most suitable for a given machine learning model? This paper aims to address these two questions by comparing the predictive performance of 13 feature selection-based machine learning (FS-ML) models and 5 ordinary machine learning models on LSA. First, five commonly used machine learning models (i.e., logistic regression, support vector machine, artificial neural network, Gaussian process and random forest) and six typical feature selection methods in the literature are adopted to constitute the proposed models. Then, fifteen conditioning factors are chosen as input variables and 1,017 landslides are used as recorded data. Next, feature selection methods are used to obtain the importance of the conditioning factors to create feature subsets, based on which 13 FS-ML models are constructed. For each of the machine learning models, a best optimized FS-ML model is selected according to the area under curve value. Finally, five optimal FS-ML models are obtained and applied to the LSA of the studied area. The predictive abilities of the FS-ML models on LSA are verified and compared through the receive operating characteristic curve and statistical indicators such as sensitivity, specificity and accuracy. The results showed that different feature selection methods have different effects on the performance of LSA machine learning models. FS-ML models generally outperform the ordinary machine learning models. The best FS-ML model is the recursive feature elimination (RFE) optimized RF, and RFE is an optimal method for feature selection.

VKOSPI Forecasting and Option Trading Application Using SVM (SVM을 이용한 VKOSPI 일 중 변화 예측과 실제 옵션 매매에의 적용)

  • Ra, Yun Seon;Choi, Heung Sik;Kim, Sun Woong
    • Journal of Intelligence and Information Systems
    • /
    • v.22 no.4
    • /
    • pp.177-192
    • /
    • 2016
  • Machine learning is a field of artificial intelligence. It refers to an area of computer science related to providing machines the ability to perform their own data analysis, decision making and forecasting. For example, one of the representative machine learning models is artificial neural network, which is a statistical learning algorithm inspired by the neural network structure of biology. In addition, there are other machine learning models such as decision tree model, naive bayes model and SVM(support vector machine) model. Among the machine learning models, we use SVM model in this study because it is mainly used for classification and regression analysis that fits well to our study. The core principle of SVM is to find a reasonable hyperplane that distinguishes different group in the data space. Given information about the data in any two groups, the SVM model judges to which group the new data belongs based on the hyperplane obtained from the given data set. Thus, the more the amount of meaningful data, the better the machine learning ability. In recent years, many financial experts have focused on machine learning, seeing the possibility of combining with machine learning and the financial field where vast amounts of financial data exist. Machine learning techniques have been proved to be powerful in describing the non-stationary and chaotic stock price dynamics. A lot of researches have been successfully conducted on forecasting of stock prices using machine learning algorithms. Recently, financial companies have begun to provide Robo-Advisor service, a compound word of Robot and Advisor, which can perform various financial tasks through advanced algorithms using rapidly changing huge amount of data. Robo-Adviser's main task is to advise the investors about the investor's personal investment propensity and to provide the service to manage the portfolio automatically. In this study, we propose a method of forecasting the Korean volatility index, VKOSPI, using the SVM model, which is one of the machine learning methods, and applying it to real option trading to increase the trading performance. VKOSPI is a measure of the future volatility of the KOSPI 200 index based on KOSPI 200 index option prices. VKOSPI is similar to the VIX index, which is based on S&P 500 option price in the United States. The Korea Exchange(KRX) calculates and announce the real-time VKOSPI index. VKOSPI is the same as the usual volatility and affects the option prices. The direction of VKOSPI and option prices show positive relation regardless of the option type (call and put options with various striking prices). If the volatility increases, all of the call and put option premium increases because the probability of the option's exercise possibility increases. The investor can know the rising value of the option price with respect to the volatility rising value in real time through Vega, a Black-Scholes's measurement index of an option's sensitivity to changes in the volatility. Therefore, accurate forecasting of VKOSPI movements is one of the important factors that can generate profit in option trading. In this study, we verified through real option data that the accurate forecast of VKOSPI is able to make a big profit in real option trading. To the best of our knowledge, there have been no studies on the idea of predicting the direction of VKOSPI based on machine learning and introducing the idea of applying it to actual option trading. In this study predicted daily VKOSPI changes through SVM model and then made intraday option strangle position, which gives profit as option prices reduce, only when VKOSPI is expected to decline during daytime. We analyzed the results and tested whether it is applicable to real option trading based on SVM's prediction. The results showed the prediction accuracy of VKOSPI was 57.83% on average, and the number of position entry times was 43.2 times, which is less than half of the benchmark (100 times). A small number of trading is an indicator of trading efficiency. In addition, the experiment proved that the trading performance was significantly higher than the benchmark.

Strain demand prediction of buried steel pipeline at strike-slip fault crossings: A surrogate model approach

  • Xie, Junyao;Zhang, Lu;Zheng, Qian;Liu, Xiaoben;Dubljevic, Stevan;Zhang, Hong
    • Earthquakes and Structures
    • /
    • v.20 no.1
    • /
    • pp.109-122
    • /
    • 2021
  • Significant progress in the oil and gas industry advances the application of pipeline into an intelligent era, which poses rigorous requirements on pipeline safety, reliability, and maintainability, especially when crossing seismic zones. In general, strike-slip faults are prone to induce large deformation leading to local buckling and global rupture eventually. To evaluate the performance and safety of pipelines in this situation, numerical simulations are proved to be a relatively accurate and reliable technique based on the built-in physical models and advanced grid technology. However, the computational cost is prohibitive, so one has to wait for a long time to attain a calculation result for complex large-scale pipelines. In this manuscript, an efficient and accurate surrogate model based on machine learning is proposed for strain demand prediction of buried X80 pipelines subjected to strike-slip faults. Specifically, the support vector regression model serves as a surrogate model to learn the high-dimensional nonlinear relationship which maps multiple input variables, including pipe geometries, internal pressures, and strike-slip displacements, to output variables (namely tensile strains and compressive strains). The effectiveness and efficiency of the proposed method are validated by numerical studies considering different effects caused by structural sizes, internal pressure, and strike-slip movements.

Understanding the Sentiment on Gig Economy: Good or Bad?

  • NORAZMI, Fatin Aimi Naemah;MAZLAN, Nur Syazwani;SAID, Rusmawati;OK RAHMAT, Rahmita Wirza
    • The Journal of Asian Finance, Economics and Business
    • /
    • v.9 no.10
    • /
    • pp.189-200
    • /
    • 2022
  • The gig economy offers many advantages, such as flexibility, variety, independence, and lower cost. However, there are also safety concerns, lack of regulations, uncertainty, and unsatisfactory services, causing people to voice their opinion on social media. This paper aims to explore the sentiments of consumers concerning gig economy services (Grab, Foodpanda and Airbnb) through the analysis of social media. First, Vader Lexicon was used to classify the comments into positive, negative, and neutral sentiments. Then, the comments were further classified into three machine learning algorithms: Support Vector Machine, Light Gradient Boosted Machine, and Logistic Regression. Results suggested that gig economy services in Malaysia received more positive sentiments (52%) than negative sentiments (19%) and neutral sentiments (29%). Based on the three algorithms used in this research, LGBM has been the best model with the highest accuracy of 85%, while SVM has 84% and LR 82%. The results of this study proved the power of text mining and sentiment analysis in extracting business value and providing insight to businesses. Additionally, it aids gig managers and service providers in understanding clients' sentiments about their goods and services and making necessary adjustments to optimize satisfaction.

Quality Estimation of Net Packaged Onions during Storage Periods using Machine Learning Techniques

  • Nandita Irsaulul, Nurhisna;Sang-Yeon, Kim;Seongmin, Park;Suk-Ju, Hong;Eungchan, Kim;Chang-Hyup, Lee;Sungjay, Kim;Jiwon, Ryu;Seungwoo, Roh;Daeyoung, Kim;Ghiseok, Kim
    • KOREAN JOURNAL OF PACKAGING SCIENCE & TECHNOLOGY
    • /
    • v.28 no.3
    • /
    • pp.237-244
    • /
    • 2022
  • Onions are a significant crop in Korea, and cultivation is increasing every year along with high demand. Onions are planted in the fall and mainly harvested in June, the rainy season, therefore, physiological changes in onion bulbs during long-term storage might have happened. Onions are stored in cold room and at adequate relative humidity to avoid quality loss. In this study, bio-yield stress and weight loss were measured as the quality parameters of net packaged onions during 10 weeks of storage, and the storage environmental conditions are monitored using sensor networks systems. Quality estimation of net packaged onion during storage was performed using the storage environmental condition data through machine learning approaches. Among the suggested estimation models, support vector regression method showed the best accuracy for the quality estimation of net packaged onions.

Comparative Analysis of Machine Learning Models for Crop's yield Prediction

  • Babar, Zaheer Ud Din;UlAmin, Riaz;Sarwar, Muhammad Nabeel;Jabeen, Sidra;Abdullah, Muhammad
    • International Journal of Computer Science & Network Security
    • /
    • v.22 no.5
    • /
    • pp.330-334
    • /
    • 2022
  • In light of the decreasing crop production and shortage of food across the world, one of the crucial criteria of agriculture nowadays is selecting the right crop for the right piece of land at the right time. First problem is that How Farmers can predict the right crop for cultivation because famers have no knowledge about prediction of crop. Second problem is that which algorithm is best that provide the maximum accuracy for crop prediction. Therefore, in this research Author proposed a method that would help to select the most suitable crop(s) for a specific land based on the analysis of the affecting parameters (Temperature, Humidity, Soil Moisture) using machine learning. In this work, the author implemented Random Forest Classifier, Support Vector Machine, k-Nearest Neighbor, and Decision Tree for crop selection. The author trained these algorithms with the training dataset and later these algorithms were tested with the test dataset. The author compared the performances of all the tested methods to arrive at the best outcome. In this way best algorithm from the mention above is selected for crop prediction.

Diagnosis Atherosclerosis Model Using Radiomics Approach in Carotid Vessel MRI (경동맥 혈관 MRI에서 라디오믹스를 이용한 동맥경화증 진단 모델)

  • Kim, Jong-hun;Park, Hyunjin
    • Proceedings of the Korean Institute of Information and Commucation Sciences Conference
    • /
    • 2022.10a
    • /
    • pp.289-290
    • /
    • 2022
  • Arteriosclerosis is a disease in which the carotid vessel wall becomes thick, and it is important to monitor the thickness of the vessel wall for diagnosis. In this study, we propose a model for extracting 324 radiomics features from carotid MRI images and diagnosing arteriosclerosis using machine learning techniques. We learned a total of four classification models: logistic regression, support vector machine, random forest, and XGBoost through radiomics features. XGBoost model, which showed the highest performance in 5-fold cross-validation, shows the results of accuracy 0.9023, sensitivity 0.9517, specificity 0.8035, AUC 0.8776.

  • PDF

Machine learning-based techniques to facilitate the production of stone nano powder-reinforced manufactured-sand concrete

  • Zanyu Huang;Qiuyue Han;Adil Hussein Mohammed;Arsalan Mahmoodzadeh;Nejib Ghazouani;Shtwai Alsubai;Abed Alanazi;Abdullah Alqahtani
    • Advances in nano research
    • /
    • v.15 no.6
    • /
    • pp.533-539
    • /
    • 2023
  • This study aims to examine four machine learning (ML)-based models for their potential to estimate the splitting tensile strength (STS) of manufactured sand concrete (MSC). The ML models were trained and tested based on 310 experimental data points. Stone nanopowder content (SNPC), curing age (CA), and water-to-cement (W/C) ratio were also studied for their impacts on the STS of MSC. According to the results, the support vector regression (SVR) model had the highest correlation with experimental data. Still, all of the optimized ML models showed promise in estimating the STS of MSC. Both ML and laboratory results showed that MSC with 10% SNPC improved the STS of MSC.

Quality Prediction Model for Manufacturing Process of Free-Machining 303-series Stainless Steel Small Rolling Wire Rods (쾌삭 303계 스테인리스강 소형 압연 선재 제조 공정의 생산품질 예측 모형)

  • Seo, Seokjun;Kim, Heungseob
    • Journal of Korean Society of Industrial and Systems Engineering
    • /
    • v.44 no.4
    • /
    • pp.12-22
    • /
    • 2021
  • This article suggests the machine learning model, i.e., classifier, for predicting the production quality of free-machining 303-series stainless steel(STS303) small rolling wire rods according to the operating condition of the manufacturing process. For the development of the classifier, manufacturing data for 37 operating variables were collected from the manufacturing execution system(MES) of Company S, and the 12 types of derived variables were generated based on literature review and interviews with field experts. This research was performed with data preprocessing, exploratory data analysis, feature selection, machine learning modeling, and the evaluation of alternative models. In the preprocessing stage, missing values and outliers are removed, and oversampling using SMOTE(Synthetic oversampling technique) to resolve data imbalance. Features are selected by variable importance of LASSO(Least absolute shrinkage and selection operator) regression, extreme gradient boosting(XGBoost), and random forest models. Finally, logistic regression, support vector machine(SVM), random forest, and XGBoost are developed as a classifier to predict the adequate or defective products with new operating conditions. The optimal hyper-parameters for each model are investigated by the grid search and random search methods based on k-fold cross-validation. As a result of the experiment, XGBoost showed relatively high predictive performance compared to other models with an accuracy of 0.9929, specificity of 0.9372, F1-score of 0.9963, and logarithmic loss of 0.0209. The classifier developed in this study is expected to improve productivity by enabling effective management of the manufacturing process for the STS303 small rolling wire rods.