• 제목/요약/키워드: corpus study

검색결과 683건 처리시간 0.025초

An Attempt to Measure the Familiarity of Specialized Japanese in the Nursing Care Field

  • Haihong Huang;Hiroyuki Muto;Toshiyuki Kanamaru
    • 아시아태평양코퍼스연구
    • /
    • 제4권2호
    • /
    • pp.57-74
    • /
    • 2023
  • Having a firm grasp of technical terms is essential for learners of Japanese for Specific Purposes (JSP). This research aims to analyze Japanese nursing care vocabulary based on objective corpus-based frequency and subjectively rated word familiarity. For this purpose, we constructed a text corpus centered on the National Examination for Certified Care Workers to extract nursing care keywords. The Log-Likelihood Ratio (LLR) was used as the statistical criterion for keyword identification, giving a list of 300 keywords as target words for a further word recognition survey. The survey involved 115 participants of whom 51 were certified care workers (CW group) and 64 were individuals from the general public (GP group). These participants rated the familiarity of the target keywords through crowdsourcing. Given the limited sample size, Bayesian linear mixed models were utilized to determine word familiarity rates. Our study conducted a comparative analysis of word familiarity between the CW group and the GP group, revealing key terms that are crucial for professionals but potentially unfamiliar to the general public. By focusing on these terms, instructors can bridge the knowledge gap more efficiently.

사전과 말뭉치를 이용한 한국어 단어 중의성 해소 (Korean Word Sense Disambiguation using Dictionary and Corpus)

  • 정한조;박병화
    • 지능정보연구
    • /
    • 제21권1호
    • /
    • pp.1-13
    • /
    • 2015
  • 빅데이터 및 오피니언 마이닝 분야가 대두됨에 따라 정보 검색/추출, 특히 비정형 데이터에서의 정보 검색/추출 기술의 중요성이 나날이 부각되어지고 있다. 또한 정보 검색 분야에서는 이용자의 의도에 맞는 결과를 제공할 수 있는 검색엔진의 성능향상을 위한 다양한 연구들이 진행되고 있다. 이러한 정보 검색/추출 분야에서 자연어처리 기술은 비정형 데이터 분석/처리 분야에서 중요한 기술이고, 자연어처리에 있어서 하나의 단어가 여러개의 모호한 의미를 가질 수 있는 단어 중의성 문제는 자연어처리의 성능을 향상시키기 위해 우선적으로 해결해야하는 문제점들의 하나이다. 본 연구는 단어 중의성 해소 방법에 사용될 수 있는 말뭉치를 많은 시간과 노력이 요구되는 수동적인 방법이 아닌, 사전들의 예제를 활용하여 자동적으로 생성할 수 있는 방법을 소개한다. 즉, 기존의 수동적인 방법으로 의미 태깅된 세종말뭉치에 표준국어대사전의 예제를 자동적으로 태깅하여 결합한 말뭉치를 사용한 단어 중의성 해소 방법을 소개한다. 표준국어대사전에서 단어 중의성 해소의 주요 대상인 전체 명사 (265,655개) 중에 중의성 해소의 대상이 되는 중의어 (29,868개)의 각 센스 (93,522개)와 연관된 속담, 용례 문장 (56,914개)들을 결합 말뭉치에 추가하였다. 품사 및 센스가 같이 태깅된 세종말뭉치의 약 79만개의 문장과 표준국어대사전의 약 5.7만개의 문장을 각각 또는 병합하여 교차검증을 사용하여 실험을 진행하였다. 실험 결과는 결합 말뭉치를 사용하였을 때 정확도와 재현율에 있어서 향상된 결과가 발견되었다. 본 연구의 결과는 인터넷 검색엔진 등의 검색결과의 성능향상과 오피니언 마이닝, 텍스트 마이닝과 관련한 자연어 분석/처리에 있어서 문장의 내용을 보다 명확히 파악하는데 도움을 줄 수 있을 것으로 기대되어진다.

신문 기사의 언어 사용 양상: 코퍼스언어학적 접근 (Aspects of Language Use in Newspaper Articles: A Corpus Linguistic Perspective)

  • 송경화;강범모
    • 인지과학
    • /
    • 제17권4호
    • /
    • pp.255-269
    • /
    • 2006
  • 본 연구는 신문 기사에 대한 실증적 언어 분석을 목적으로 한다. <21세기 세종계획>에 의해 구축된 대용량의 신문 기사 말뭉치를 형태, 어절, 절, 문장 등의 단위로 계량화하여 분석하였다. 신문 기사를 표제, 전문, 본문의 세 구성 성분으로 나누고 표제의 표시성과 압축성의 실현 양상, 전문과 표제의 연관성, 본문의 문장 구조와 일반명사 구성 비율 등을 살펴보았다. 이 연구를 통하여 기존의 비계량적 연구 방법들과 차별화 된 실증적 연구로서 신문 이론을 검증하고, 신문 기사의 새로운 언어 현상을 발견할 수 있었다. 신문 기사와 같은 텍스트는 인간의 인지적 언어 처리의 결과이며 동시에 인지적 언어 형성에 영향을 미칠 것이다.

  • PDF

Reduction and Frequency Analyses of Vowels and Consonants in the Buckeye Speech Corpus

  • Yang, Byung-Gon
    • 말소리와 음성과학
    • /
    • 제4권3호
    • /
    • pp.75-83
    • /
    • 2012
  • The aims of this study were three. First, to examine the degree of deviation from dictionary prescribed symbols and actual speech made by American English speakers. Second, to measure the frequency of vowel and consonant production of American English speakers. And third, to investigate gender differences in the segmental sounds in a speech corpus. The Buckeye Speech Corpus was recorded by forty American male and female subjects for one hour per subject. The vowels and consonants in both the phonemic and phonetic transcriptions were extracted from the original files of the corpus and their frequencies were obtained using codes of a free software R. Results were as follows: Firstly, the American English speakers produced a reduced number of vowels and consonants in daily conversation. The reduction rate from the dictionary transcriptions to the actual transcriptions was around 38.2%. Secondly, the American English speakers used more front high and back low vowels while three-fourths of the consonants accounted for stops, fricatives, and nasals. This indicates that the segmental inventory has nonlinear frequency distribution in the speech corpus. Thirdly, the two gender groups produced vowels and consonants similarly even though there were a few noticeable differences in their speech. From these results we propose that English teachers consider pronunciation education reflecting the actual speech sounds and that linguists find a way to establish unmarked segmentals from speech corpora.

Endothelial Cells Isolated from the Bovine Corpus Luteum Synthesize Prostaglandin $F_{2{\alpha}}$ Receptor

  • Gwon, Sun-Yeong;Rhee, Ki-Jong;Lee, Seunghyung
    • 대한의생명과학회지
    • /
    • 제19권3호
    • /
    • pp.261-265
    • /
    • 2013
  • The corpus luteum is a transient endocrine gland essential for regulation of the ovarian cycle as well as for establishing and maintaining pregnancy. Prostaglandin $F_{2{\alpha}}$ (PGF) initiates functional and structural regression of the corpus luteum and therefore is an important regulator of the estrous cycle. It is a matter of debate whether the endothelial cells of the bovine corpus luteum express PGFR, the cognate receptor for PGF. Therefore, the aim of this study was to assess the expression of PGFR in bovine endothelial cells. Endothelial cells were isolated from the bovine corpus luteum of the mid-luteal stage using magnetic beads and cultured in vitro. We demonstrate that this isolation procedure generates a pure culture of endothelial cells as confirmed by synthesis of Factor VIII and lack of expression of $3{\beta}$-hydroxysteroid dehydrogenase. By RT-PCR, Western blot and immunofluorescence analyses, we further show that the cultured endothelial cells produced PGFR. This model system can be utilized to provide an experimental system to investigate the role of PGF on endothelial cells during the reproductive cycle.

The Ratios of CEFR-J Vocabulary Usage Compared with GSL and AWL in Elementary EFL Classrooms and Suggestions of Vocabulary Items to be Taught

  • Ohashi, Yukiko;Katagiri, Noriaki
    • 아시아태평양코퍼스연구
    • /
    • 제1권1호
    • /
    • pp.61-94
    • /
    • 2020
  • The present study examined vocabulary usage in elementary English classrooms in Japan using elementary school corpus. The authors used three wordlists to benchmark the lexical items for four classes in the corpus: the CEFR-J, the General Service List (GSL), and Academic Word List (AWL). The percentage of vocabulary usage belonging to the Level A1 in the CEFR-J was below 15% (Class A: 12.1%, Class B: 12.6%, Class C: 8.9%, and Class D: 13.6%) with no statistical difference between levels. The mean ratio of Level A2 vocabulary items was below 10%, and all classes showed less than 1% of vocabulary usage for the Levels B1 and B2. Over 70% of all vocabulary items in the corpus belonged to the most frequent 1,000-word band (level 1) of the GSL, while the next most frequent word band (level 2 of the GSL and AWL) accounted for less than 10%. The results suggest that elementary school English teachers should use more vocabulary items in the CEFR-J Level A1. The findings demonstrate that elementary school teachers are less likely to expose their pupils to grammatically well-structured sentences with an abundance of lexical items since the teachers repeatedly use the same lexemes in each class.

Modal Auxiliary Verbs in Japanese EFL Learners' Conversation: A Corpus-based Study

  • Nakayama, Shusaku
    • 아시아태평양코퍼스연구
    • /
    • 제2권1호
    • /
    • pp.23-34
    • /
    • 2021
  • This research examines Japanese non-native speakers' (JNNS) modal auxiliary verb use from two different perspectives: frequency of use and preferences for modalities. Additionally, error analysis is carried out to identify errors in modal use common among JNNSs. Their modal use is compared to that of English native speakers within a spoken dialogue corpus which is part of the International Corpus Network of Asian Learners' English. Research findings show at a statistically significant level that when compared to native speakers, JNNSs underuse past forms of modals and infrequently convey epistemic modality, indicating the possibility that JNNSs fail to express their opinions or thoughts indirectly when needed or to convey politeness appropriately. Error analysis identifies the following three types of common errors: (1) the use of incorrect tenses of modal verb phrases, (2) the use of inflected verb forms after modals, and (3) the non-use of main verbs after modals. The first type of error is largely because JNNSs do not master how to express past meanings of modals. The second and third types of errors seem to be due to first language transfer into second language acquisition and JNNSs' overgeneralization of the subject-verb agreement rules to modals respectively.

Effects of Corpus Use on Error Identification in L2 Writing

  • Yoshiho Satake
    • 아시아태평양코퍼스연구
    • /
    • 제4권1호
    • /
    • pp.61-71
    • /
    • 2023
  • This study examines the effects of data-driven learning (DDL)-an approach employing corpora for inductive language pattern learning-on error identification in second language (L2) writing. The data consists of error identification instances from fifty-five participants, compared across different reference materials: the Corpus of Contemporary American English (COCA), dictionaries, and no use of reference materials. There are three significant findings. First, the use of COCA effectively identified collocational and form-related errors due to inductive inference drawn from multiple example sentences. Secondly, dictionaries were beneficial for identifying lexical errors, where providing meaning information was helpful. Finally, the participants often employed a strategic approach, identifying many simple errors without reference materials. However, while maximizing error identification, this strategy also led to mislabeling correct expressions as errors. The author has concluded that the strategic selection of reference materials can significantly enhance the effectiveness of error identification in L2 writing. The use of a corpus offers advantages such as easy access to target phrases and frequency information-features especially useful given that most errors were collocational and form-related. The findings suggest that teachers should guide learners to effectively use appropriate reference materials to identify errors based on error types.

Sentence-Chain Based Seq2seq Model for Corpus Expansion

  • Chung, Euisok;Park, Jeon Gue
    • ETRI Journal
    • /
    • 제39권4호
    • /
    • pp.455-466
    • /
    • 2017
  • This study focuses on a method for sequential data augmentation in order to alleviate data sparseness problems. Specifically, we present corpus expansion techniques for enhancing the coverage of a language model. Recent recurrent neural network studies show that a seq2seq model can be applied for addressing language generation issues; it has the ability to generate new sentences from given input sentences. We present a method of corpus expansion using a sentence-chain based seq2seq model. For training the seq2seq model, sentence chains are used as triples. The first two sentences in a triple are used for the encoder of the seq2seq model, while the last sentence becomes a target sequence for the decoder. Using only internal resources, evaluation results show an improvement of approximately 7.6% relative perplexity over a baseline language model of Korean text. Additionally, from a comparison with a previous study, the sentence chain approach reduces the size of the training data by 38.4% while generating 1.4-times the number of n-grams with superior performance for English text.

벅아이 코퍼스의 모음 길이 연구 (A Study on the Vowel Duration of the Buckeye Corpus)

  • 정혜정;윤규철
    • 말소리와 음성과학
    • /
    • 제7권4호
    • /
    • pp.103-110
    • /
    • 2015
  • The purpose of this study is to assess the vowel property by examining the vowel duration of the American English vowles found in the Buckeye corpus[6]. The vowel durations were analyzed in terms of various linguistic factors including the number of syllables of the word containing the vowel, the location of the vowel in a word, types of stress, function versus content word, the word frequency in the corpus and the speech rate calculated from the three consecutive words. The findings from this work agreed mostly with those from earlier studies, but with some exceptions. The relationship between the speech rate and the vowel duration proved non-linear.