• Title/Summary/Keyword: language resources

Search Result 428, Processing Time 0.022 seconds

A Study of Fine Tuning Pre-Trained Korean BERT for Question Answering Performance Development (사전 학습된 한국어 BERT의 전이학습을 통한 한국어 기계독해 성능개선에 관한 연구)

  • Lee, Chi Hoon;Lee, Yeon Ji;Lee, Dong Hee
    • Journal of Information Technology Services
    • /
    • v.19 no.5
    • /
    • pp.83-91
    • /
    • 2020
  • Language Models such as BERT has been an important factor of deep learning-based natural language processing. Pre-training the transformer-based language models would be computationally expensive since they are consist of deep and broad architecture and layers using an attention mechanism and also require huge amount of data to train. Hence, it became mandatory to do fine-tuning large pre-trained language models which are trained by Google or some companies can afford the resources and cost. There are various techniques for fine tuning the language models and this paper examines three techniques, which are data augmentation, tuning the hyper paramters and partly re-constructing the neural networks. For data augmentation, we use no-answer augmentation and back-translation method. Also, some useful combinations of hyper parameters are observed by conducting a number of experiments. Finally, we have GRU, LSTM networks to boost our model performance with adding those networks to BERT pre-trained model. We do fine-tuning the pre-trained korean-based language model through the methods mentioned above and push the F1 score from baseline up to 89.66. Moreover, some failure attempts give us important lessons and tell us the further direction in a good way.

Exploring Variables of Korean Language Education for Preschooler With Multicultural Family Background (다문화가정 취학 전 유아 한국어교육 지원을 위한 기초 연구)

  • Kim, Min Hwa;Shin, Hye Eun
    • Korean Journal of Child Studies
    • /
    • v.29 no.2
    • /
    • pp.155-176
    • /
    • 2008
  • This study explored variables related to Korean language education for preschool children with multicultural family backgrounds. Participants were 21 Korean language teachers and 14 women who immigrated from China, Japan, Mongolia, Philippines, and Vietnam to marry Korean men. They were mothers of children 2 to 7 years of age and had lived in Korea an average of five years. Mean age of mothers was 37(range of 30 to 43). Half had college and none had less then middle school education. They were interviewed with a series of semi-structured questionnaires. The children were reported to have a low level of vocabulary and articulation because their mothers could not provide fruitful oral language experiences. Supporting systems including family literacy were discussed.

  • PDF

Tester Structure Expression Language and Its Application to the Environment for VLSI Tester Program Development

  • Sato, Masayuki;Wakamatsu, Hiroki;Arai, Masayuki;Ichino, Kenichi;Iwasaki, Kazuhiko;Asakawa, Takeshi
    • Journal of Information Processing Systems
    • /
    • v.4 no.4
    • /
    • pp.121-132
    • /
    • 2008
  • VLSI chips have been tested using various automatic test equipment (ATE). Although each ATE has a similar structure, the language for ATE is proprietary and it is not easy to convert a test program for use among different ATE vendors. To address this difficulty we propose a tester structure expression language, a tester language with a novel format. The developed language is called the general tester language (GTL). Developing an interpreter for each tester, the GTL program can be directly applied to the ATE without conversion. It is also possible to select a cost-effective ATE from the test program, because the program expresses the required ATE resources, such as pin counts, measurement accuracy, and memory capacity. We describe the prototype environment for the GTL and the tester selection tool. The software size of the prototype is approximately 27,800 steps and 15 manmonths were required. Using the tester selection tool, the number of man-hours required in order to select an ATE could be reduced to 1/10. A GTL program was successfully executed on actual ATE.

Burmese Sentiment Analysis Based on Transfer Learning

  • Mao, Cunli;Man, Zhibo;Yu, Zhengtao;Wu, Xia;Liang, Haoyuan
    • Journal of Information Processing Systems
    • /
    • v.18 no.4
    • /
    • pp.535-548
    • /
    • 2022
  • Using a rich resource language to classify sentiments in a language with few resources is a popular subject of research in natural language processing. Burmese is a low-resource language. In light of the scarcity of labeled training data for sentiment classification in Burmese, in this study, we propose a method of transfer learning for sentiment analysis of a language that uses the feature transfer technique on sentiments in English. This method generates a cross-language word-embedding representation of Burmese vocabulary to map Burmese text to the semantic space of English text. A model to classify sentiments in English is then pre-trained using a convolutional neural network and an attention mechanism, where the network shares the model for sentiment analysis of English. The parameters of the network layer are used to learn the cross-language features of the sentiments, which are then transferred to the model to classify sentiments in Burmese. Finally, the model was tuned using the labeled Burmese data. The results of the experiments show that the proposed method can significantly improve the classification of sentiments in Burmese compared to a model trained using only a Burmese corpus.

Learning a Foreign Language Using Information Technologies for Comfortable Implementation of the Professional Position of a Future Specialist in a Foreign Language Environment

  • Postolenko, Iryna;Biletska, Iryna;Kmit', Olena;Paltseva, Valentyna;Mykhailenko, Olena;Yatsyna, Svitlana;Kuchai, Tetiana
    • International Journal of Computer Science & Network Security
    • /
    • v.22 no.11
    • /
    • pp.63-70
    • /
    • 2022
  • At the present stage, the main directions of the professional position of a specialist in the implementation of English-language Education are to improve and spread the practice of learning languages throughout a person's life by involving information, communication and digital technologies in the educational process. Computerization of the educational process in Higher Education Institutions is considered as one of the first and most promising areas for improving the quality of education in Higher Education Institutions. The necessity of ensuring timely training and retraining of specialists of various profiles (in particular teachers) on the effective use of domestic and foreign electronic resources with the help of modern information technologies for the implementation of the professional position of a future specialist in a foreign-language environment is noted. The main goal of teaching a foreign language (the formation of students' communicative competence, which means mastering the language as a means of intercultural communication) is defined. The types of speech activity that cover the content of teaching a foreign language are highlighted. The main types of assessment in a foreign language are shown - current (non-classroom), thematic, semester, annual assessment and final state certification. The task of the teacher is drawn, which is to create conditions for practical language acquisition for each student, to choose such teaching methods by means of information technologies that would allow each student to show their activity, their creativity; to activate the cognitive activity of the student in the process of learning a foreign language.

The Context and Reality of Memes as Information Resources: Focused on Analysis of Research Trends in South Korea (정보자원으로서 '밈'의 맥락과 실재 - 국내 연구동향 분석을 중심으로 -)

  • Soram Hong
    • Journal of the Korean BIBLIA Society for library and Information Science
    • /
    • v.34 no.3
    • /
    • pp.227-253
    • /
    • 2023
  • The study is a preliminary study to conceptualize memes as information resources for literacy education in information environment changed with digital revolution. The study is to explain the context and reality of memes in order to promote the utilization of memes as information resources. The research questions are as follows: First, what topics are 'memes' studied with? Second, what things are captured and studied as 'memes'? The study conducted frequency and co-occurrence network analysis on 145 domestic studies and contents analysis on 73 domestic studies. The results are as follows: First, memes were mainly studied in the fields of 'humanities', 'social sciences', 'interdiciplinary studies', and 'arts and kinesiology'. Studies based on Dawkins' concept of memes (around 2012), studies on introducing the concept of memes to explain the spread of Korean Wave content (around 2015), and independent studies of memes as a major research topic in cultural sociology (around 2019) were performed. Second, memes are linguistic. Language memes (L-memes) are 102 (37%), language-visual memes (LV-memes) are 23 (8%), language-visual-musical memes (LVM-memes) are 21 (8%). Keyword 'language meme' ranked high in frequency, degree centrality and betweenness centrality of co-occurrence network. In other words, memes are expanding as a unique information phenomenon of cultural sociology based on linguistic characteristics. It is necessary to conceptualize meme literacy in terms of information literacy.

Enhancing LoRA Fine-tuning Performance Using Curriculum Learning

  • Daegeon Kim;Namgyu Kim
    • Journal of the Korea Society of Computer and Information
    • /
    • v.29 no.3
    • /
    • pp.43-54
    • /
    • 2024
  • Recently, there has been a lot of research on utilizing Language Models, and Large Language Models have achieved innovative results in various tasks. However, the practical application faces limitations due to the constrained resources and costs required to utilize Large Language Models. Consequently, there has been recent attention towards methods to effectively utilize models within given resources. Curriculum Learning, a methodology that categorizes training data according to difficulty and learns sequentially, has been attracting attention, but it has the limitation that the method of measuring difficulty is complex or not universal. Therefore, in this study, we propose a methodology based on data heterogeneity-based Curriculum Learning that measures the difficulty of data using reliable prior information and facilitates easy utilization across various tasks. To evaluate the performance of the proposed methodology, experiments were conducted using 5,000 specialized documents in the field of information communication technology and 4,917 documents in the field of healthcare. The results confirm that the proposed methodology outperforms traditional fine-tuning in terms of classification accuracy in both LoRA fine-tuning and full fine-tuning.

Network-based Language Teaching and Learning - The Internet and Classroom -

  • Hong, Sung-Ryong
    • Journal of Digital Contents Society
    • /
    • v.7 no.3
    • /
    • pp.175-182
    • /
    • 2006
  • The Internet is now of the fastest growing areas of telecommunications and of Computer Assisted Language Learning. It is rapidly becoming more integrated into society and accessible to people form around the world. A number of educators believe there is potential for language teaching and learning opportunities through the Internet, and have already developed uses and resources for this purpose. The range of what is available is growing continually. The purpose of this study is to research CMC via the Internet and other long-distance networks, to investigate the analyse best and worst things about studying English on the internet and to suggest some findings from the comparison between internet and classroom learning by means of questionnaire.

  • PDF

A Semi-supervised Learning of HMM to Build a POS Tagger for a Low Resourced Language

  • Pattnaik, Sagarika;Nayak, Ajit Kumar;Patnaik, Srikanta
    • Journal of information and communication convergence engineering
    • /
    • v.18 no.4
    • /
    • pp.207-215
    • /
    • 2020
  • Part of speech (POS) tagging is an indispensable part of major NLP models. Its progress can be perceived on number of languages around the globe especially with respect to European languages. But considering Indian Languages, it has not got a major breakthrough due lack of supporting tools and resources. Particularly for Odia language it has not marked its dominancy yet. With a motive to make the language Odia fit into different NLP operations, this paper makes an attempt to develop a POS tagger for the said language on a HMM (Hidden Markov Model) platform. The tagger judiciously considers bigram HMM with dynamic Viterbi algorithm to give an output annotated text with maximum accuracy. The model is experimented on a corpus belonging to tourism domain accounting to a size of approximately 0.2 million tokens. With the proportion of training and testing as 3:1, the proposed model exhibits satisfactory result irrespective of limited training size.

A Case Study of KSL Learner-Learner Dialogue as a Cognitive Activity in Speaking Tasks (말하기 과제 수행에서 인지적 활동으로서의 학습자 대화 사례 연구)

  • Son, Hyejin
    • Journal of Korean language education
    • /
    • v.29 no.2
    • /
    • pp.73-100
    • /
    • 2018
  • The purpose of this study is to investigate learner-learner dialogue during speaking tasks. In the Korean language classroom, conversation between learners is an important activity as speaking practice. However, learner dialogue is also a tool to enable learners to collaboratively conduct various cognitive activities in the classroom. In previous research, it was unfolded that through learner-learner dialogue, learners can solve second-language related problems and set a goal to carry out tasks. Therefore, this study analyzed learner-learner dialogue to investigate what kinds of cognitive activities are activated during the role-play task. As a result, the learners collaboratively generated and monitored language and content for role play. Also, in order to accomplish tasks more successfully, learners shared the same understanding about the goal of the task, and tried to manage the task procedure. Through learner-learner dialogue, learners can participate in cognitive activities such as content, language construction, and task management voluntarily without the help from teachers. This means that learner-learner dialogue can be an activity to support language learning tasks. Also, it can make learners actively involved in learning and by sharing resources with each other. It is also important that learners can experience language use that participates in real-world communication activities, such as learning in the classroom and collaborating with peer learners. This study is an exploratory study for a basic understanding of learner's conversation as a cognitive activity, and the scope of the study is limited to clarifying contents of learner-learner dialogue as a cognitive activity in speaking tasks. Based on the findings of this study, future research should be conducted on the function of learner-learner dialogue as a cognitive activity in Korean language learning and its role in the classroom of Korean language education.