• Title/Summary/Keyword: 텍스트 개념

Search Result 379, Processing Time 0.028 seconds

텍스트 마이닝의 개념과 응용

  • Jo, Tae-Ho
    • Journal of Scientific & Technological Knowledge Infrastructure
    • /
    • s.5
    • /
    • pp.76-85
    • /
    • 2001
  • 정보검색시스템은 물론 텍스트 데이터를 대상으로하는 지식관리 시스템, 문서관리시스템, 그리고 전자도서관등에서 텍스트 마이닝에 대한 기술에 대한 수요가 증가하고 있는 추세이다. 이 글에서는 텍스트 마이닝의 개념을 소개하고, 텍스트 마이닝의 주요기능, 그리고, 응용사례등을 기술할것이다. 텍스트 마이닝은 텍스트 데이터를 대상으로 하여 그들간의 암묵적인 정보를 추출하는 과정으로 정의할 수 있다. 데이터마이닝과 텍스트 마이닝의 차이는 대상이 텍스트 데이터와 수치 데이터하는 점에서 구분되고 텍스트 마이닝은 데이터 마이닝과 달리 이를 구조화시키는 과정이 필요하다. 텍스트마이닝에 있어서 구조화하는 과정에서 가장 보편적으로 사용되는것은 문서색인이다.

  • PDF

Building Concept Networks using a Wikipedia-based 3-dimensional Text Representation Model (위키피디아 기반의 3차원 텍스트 표현모델을 이용한 개념망 구축 기법)

  • Hong, Ki-Joo;Kim, Han-Joon;Lee, Seung-Yeon
    • KIISE Transactions on Computing Practices
    • /
    • v.21 no.9
    • /
    • pp.596-603
    • /
    • 2015
  • A concept network is an essential knowledge base for semantic search engines, personalized search systems, recommendation systems, and text mining. Recently, studies of extending concept representation using external ontology have been frequently conducted. We thus propose a new way of building 3-dimensional text model-based concept networks using the world knowledge-level Wikipedia ontology. In fact, it is desirable that 'concepts' derived from text documents are defined according to the theoretical framework of formal concept analysis, since relationships among concepts generally change over time. In this paper, concept networks hidden in a given document collection are extracted more reasonably by representing a concept as a term-by-document matrix.

A Semantic Text Model with Wikipedia-based Concept Space (위키피디어 기반 개념 공간을 가지는 시멘틱 텍스트 모델)

  • Kim, Han-Joon;Chang, Jae-Young
    • The Journal of Society for e-Business Studies
    • /
    • v.19 no.3
    • /
    • pp.107-123
    • /
    • 2014
  • Current text mining techniques suffer from the problem that the conventional text representation models cannot express the semantic or conceptual information for the textual documents written with natural languages. The conventional text models represent the textual documents as bag of words, which include vector space model, Boolean model, statistical model, and tensor space model. These models express documents only with the term literals for indexing and the frequency-based weights for their corresponding terms; that is, they ignore semantical information, sequential order information, and structural information of terms. Most of the text mining techniques have been developed assuming that the given documents are represented as 'bag-of-words' based text models. However, currently, confronting the big data era, a new paradigm of text representation model is required which can analyse huge amounts of textual documents more precisely. Our text model regards the 'concept' as an independent space equated with the 'term' and 'document' spaces used in the vector space model, and it expresses the relatedness among the three spaces. To develop the concept space, we use Wikipedia data, each of which defines a single concept. Consequently, a document collection is represented as a 3-order tensor with semantic information, and then the proposed model is called text cuboid model in our paper. Through experiments using the popular 20NewsGroup document corpus, we prove the superiority of the proposed text model in terms of document clustering and concept clustering.

Empirical Analysis on the Holy Bible Texts' Cliche for English-Korean Interpretation and Translation (영·한 통번역을 위한 성경 텍스트 클리셰(cliche)의 실증적 분석)

  • You, Seon-Young
    • The Journal of the Korea Contents Association
    • /
    • v.17 no.10
    • /
    • pp.54-64
    • /
    • 2017
  • The purpose of this study was to analyze the cliche for English-Korean interpretation and translation with special reference to the cliche based on the Holy Bible texts. Cliches are figurative or literal expressions and are overused expressions in various different cultures. In addition, cliches are languages, a tool of communication in an appealing way. Therefore, cliches are must be clearly distinguished from the term of idioms that are figurative phrases with an implied meaning; the phrase is not to be taken literally. Also, cliches are the single most important factor that characterizes socioculturally. Through this empirical analysis on cliches we see that this study has conceptualized the meaning of cliche. Based on this result, I expect that anyone who researches English-Korean interpretation and translation field should be concerned about cliches. I hope this study will be a guide to the right uses of cliches in English language fields.

Concept based Image Retrieval Using Similarity Measurement Between Concepts (개념간 유사성 측정을 이용한 개념 기반 이미지 검색)

  • 조미영;최춘호;신주현;김판구
    • Proceedings of the Korean Information Science Society Conference
    • /
    • 2003.04c
    • /
    • pp.253-255
    • /
    • 2003
  • 기존의 개념 기반 이미지 검색에서는 이미지의 의미적 내용 인식을 위해 일반적으로 어휘적 정보나 텍스트 정보를 이용했다. 이러한 텍스트 정보 기반 이미지 검색은 전통적인 검색 방법인 키워드 검색 기술을 그대로 사용하여 쉽게 구현할 수 있으나 텍스트의 개념적 매칭이 아닌 스트링 매칭이므로 주석처리된 단어와 정확한 매칭이 없다면 찾을 수가 없었다. 이에 본 논문에서는 ontology의 일종인 WordNet을 이용하여 깊이 정보량 링크 타입, 밀도 등을 고려한 개념간 유사성 측정으로 패턴 매칭의 문제를 해결하고자 했다. 또한 키워드로 주석처리 되어 있는 Microsofts Design Gallery Live의 이미지를 이용하여 개념간 유사성 측정법을 실질적으로 개념 기반 이미지 검색에 적용해 보았다.

  • PDF

Applying Method WordNet for Concept based Image Retrieval system (개념 기반 이미지 검색 시스템을 위한 WordNet 적용 방안)

  • 조미영;최준호;김판구
    • Proceedings of the Korean Information Science Society Conference
    • /
    • 2002.10d
    • /
    • pp.487-489
    • /
    • 2002
  • 기존의 키워드 기반 이미지 검색에서는 의미적 내용 인식을 위해 일반적으로 어휘적 정보나 텍스트 정보를 인간이 주석 형태로 달아주었다. 그러나 이런 텍스트 정보 기반 이미지 검색은 개념적 매칭이 아닌 스트링 매칭이므로 주석을 달아놓은 단어와 정확한 매칭이 없다면 찾을 수가 없다. 이러한 문제를 해결하기 위해 본 논문에서는 개념 기반 이미지 검색 시스템을 위한 WordNet의 적용 방안에 대해 연구했다. WordNet은 단언형이 아닌 단어의 의미 즉 synset이 구성 요소라는 특징을 이용해 각각의 이미지에 텍스트 정보 대신 적합한 개념의 Synset번호를 저장한다. 그리고 검색시 개념간의 유사성 측정을 이용해 검색어와 개념적으로 유사한 모든 이미지를 검색하도록 한다.

  • PDF

The Forming Mechanism of Brain Text and Brain Concept in the Theory of Ethical Literary Criticism (뇌텍스트(Brain Text) 및 뇌개념(Brain Concept)의 형성원리와 문학윤리학비평)

  • Nie, Zhenzhao;Yoon, Seokmin
    • Journal of Popular Narrative
    • /
    • v.25 no.1
    • /
    • pp.193-215
    • /
    • 2019
  • According to ethical literary criticism, every type of literature has its text. The original definition of oral literature refers to the literature disseminated orally. Before the dissemination, the text of oral literature is stored in the human brain, which is termed as "brain text". Brain text is the textual form used before the formation of writing symbols and its application to a recording of information, and it still exists after the creation of writing symbols. Other types of texts are written text and electronic text. Brain text consists of brain concepts, which, according to different sources, can be divided into objective concepts and abstractive concepts. Brain concepts are tools for thinking while thought comes from thinking with understanding and an application of brain concepts. Brain text is the carrier of thought. The termination of the synthesis of brain concepts signifies the completion of thinking, which produces thoughts to form brain text. Brain text determines thinking and behavioral patterns that not only communicate and spread information, but also decide our ideas, thoughts, judgments, choices, actions and emotions. Brain text is also a deciding factor for our lifestyle and moral behaviors. The nature of a person's brain text determines his thoughts and actions, and most importantly determines who he is.

Effects of Collaborative Argumentation and Self-Explanation on Text Comprehension in a Concept Mapping Context (텍스트이해를 위한 개념도사용의 효과적 활용전략:협력적 논쟁과 자기설명의 상호작용 효과)

  • Kim, Jong Baeg
    • (The) Korean Journal of Educational Psychology
    • /
    • v.22 no.2
    • /
    • pp.461-478
    • /
    • 2008
  • This study attempted to test whether or not students' collaborative argumentation and explanation activity while using concept mapping did improve understanding on texts. Total of 52 college students participated in this study. They were randomly assigned to one of four experimental conditions. The experiment lasted for two or three weeks and students were tested on comprehension level of a text material that they have studied over the period. As a result, with two independent factors of explanation and collaboration, there was a significant interaction effect without main effects. That is, individual did better when they did have to explain what they were doing. However, this is not the case when students collaborate. Students in the paired condition, they did better when they do not have to explain what they were doing with concept maps. This study showed efficiency with using computerized software does not always guarantee higher understanding on text materials. Instructional contexts and variables, collaboration and explanation, needs to be considered. Collaborating with others and explaining their own learning processes should be carefully designed when they are combined with concept mapping contexts. How to minimize learning obstacles from discussing ideas with others are a critical issue for future research.

Exploring Teaching Method for Productive Knowledge of Scientific Concept Words through Science Textbook Quantitative Analysis (과학교과서 텍스트의 계량적 분석을 이용한 과학 개념어의 생산적 지식 교육 방안 탐색)

  • Yun, Eunjeong
    • Journal of The Korean Association For Science Education
    • /
    • v.40 no.1
    • /
    • pp.41-50
    • /
    • 2020
  • Looking at the understanding of scientific concepts from a linguistic perspective, it is very important for students to develop a deep and sophisticated understanding of words used in scientific concept as well as the ability to use them correctly. This study intends to provide the basis for productive knowledge education of scientific words by noting that the foundation of productive knowledge teaching on scientific words is not well established, and by exploring ways to teach the relationship among words that constitute scientific concept in a productive and effective manner. To this end, we extracted the relationship among the words that make up the scientific concept from the text of science textbook by using quantitative text analysis methods, second, qualitatively examined the meaning of the word relationship extracted as a result of each method, and third, we proposed a writing activity method to help improve the productive knowledge of scientific concept words. We analyzed the text of the "Force and motion" unit on first grade science textbook by using four methods of quantitative linguistic analysis: word cluster, co-occurrence, text network analysis, and word-embedding. As results, this study suggests four writing activities, completing sentence activity by using the result of word cluster analysis, filling the blanks activity by using the result of co-occurrence analysis, material-oriented writing activities by using the result of text network analysis, and finally we made a list of important words by using the result of word embedding.

An Overview of Hypertext and Its Applications (하이퍼텍스트의 개념과 응용에 관한 고찰)

  • 정영미
    • Journal of the Korean Society for information Management
    • /
    • v.6 no.2
    • /
    • pp.3-20
    • /
    • 1989
  • Hypertext system is a new type of electronic information system which offers users great freedom in writing and reading electronic documents. Hypertext means non-linear or non-sequential text, which consists of a collection of nodes connected by links. Nodes may contain segments of text, video images, and sound. In this paper, the concept and characteristics of hypertext are reviewed, components of hypertext are explored in detail, and Guide is illustrated with application examples.

  • PDF