• Title/Summary/Keyword: 식별데이터

Search Result 1,602, Processing Time 0.026 seconds

De-identifying Unstructured Medical Text and Attribute-based Utility Measurement (의료 비정형 텍스트 비식별화 및 속성기반 유용도 측정 기법)

  • Ro, Gun;Chun, Jonghoon
    • The Journal of Society for e-Business Studies
    • /
    • v.24 no.1
    • /
    • pp.121-137
    • /
    • 2019
  • De-identification is a method by which the remaining information can not be referred to a specific individual by removing the personal information from the data set. As a result, de-identification can lower the exposure risk of personal information that may occur in the process of collecting, processing, storing and distributing information. Although there have been many studies in de-identification algorithms, protection models, and etc., most of them are limited to structured data, and there are relatively few considerations on de-identification of unstructured data. Especially, in the medical field where the unstructured text is frequently used, many people simply remove all personally identifiable information in order to lower the exposure risk of personal information, while admitting the fact that the data utility is lowered accordingly. This study proposes a new method to perform de-identification by applying the k-anonymity protection model targeting unstructured text in the medical field in which de-identification is mandatory because privacy protection issues are more critical in comparison to other fields. Also, the goal of this study is to propose a new utility metric so that people can comprehend de-identified data set utility intuitively. Therefore, if the result of this research is applied to various industrial fields where unstructured text is used, we expect that we can increase the utility of the unstructured text which contains personal information.

A Study on the Identification Metadata of National R&D Information (국가 R&D정보 식별메타데이터에 관한 연구)

  • Kwon, Lee-nam;Kim, Jae-soo
    • Proceedings of the Korea Contents Association Conference
    • /
    • 2007.11a
    • /
    • pp.55-59
    • /
    • 2007
  • A need for sharing the R&D information at a national level regardless of the ministry in charge of it is being aroused. However, this information is being managed by different institutes under different ministries without a standard identifier which makes it difficult to manage the information such as to look up duplicate projects. A URN-based identifier is required for the integrated management and utilization of the national R&D information. In this paper, the concept of the identification metadata for applying the identification system and the status of the information management by each institute will be the subject for discussion. Based on this, the elements required for issuing a unique identification number and identifier to the national R&D projects will be suggested in this paper.

  • PDF

A Novel Way of Context-Oriented Data Stream Segmentation using Exon-Intron Theory (Exon-Intron이론을 활용한 상황중심 데이터 스트림 분할 방안)

  • Lee, Seung-Hun;Suh, Dong-Hyok
    • The Journal of the Korea institute of electronic communication sciences
    • /
    • v.16 no.5
    • /
    • pp.799-806
    • /
    • 2021
  • In the IoT environment, event data from sensors is continuously reported over time. Event data obtained in this trend is accumulated indefinitely, so a method for efficient analysis and management of data is required. In this study, a data stream segmentation method was proposed to support the effective selection and utilization of event data from sensors that are continuously reported and received. An identifier for identifying the point at which to start the analysis process was selected. By introducing the role of these identifiers, it is possible to clarify what is being analyzed and to reduce data throughput. The identifier for stream segmentation proposed in this study is a semantic-oriented data stream segmentation method based on the event occurrence of each stream. The existence of identifiers in stream processing can be said to be useful in terms of providing efficiency and reducing its costs in a large-volume continuous data inflow environment.

Design of Clinical Big data-based Verifiability Identification Process through Characterization of Medical Device (의료기기의 특성 분석을 통한 임상 빅데이터 기반 검증 가능성 식별 프로세스 설계)

  • Choi, Yoo-Rim;Park, Ye-Seul;Lee, Jung-Won
    • Proceedings of the Korea Information Processing Society Conference
    • /
    • 2017.11a
    • /
    • pp.753-756
    • /
    • 2017
  • 의료기기는 사람의 생명과 직접적으로 연관되어 있기 때문에 다른 분야의 기기보다 안전성에 대한 검증이 필수적이다. 의료 분야에서는 안전성 검증을 위해 기기의 허가 심사 조건으로서 소수의 피험자를 대상으로 수행되는 임상 시험이 존재한다. 그러나 임상 시험의 경우 의료기기를 직접 사람에게 적용하여 검증을 진행하기 때문에, 인체에 미칠 위해성을 고려하여 전임상 시험을 수행하고 있다. 하지만 전임상 시험은 동물이나 가상의 물체를 대상으로 수행하여 실제 사람에 대한 적용이 아니기 때문에, 임상 시험에 비해 검증에 대한 효력을 갖지 못한다. 따라서 본 연구에서는 피험자의 안전을 보장할 수 있고, 임상 빅데이터에 축적된 실제 환자의 사례를 활용한 신뢰성 있는 검증 방안을 제안하고자 한다. 그러나 현재 식품의약품안전처에서 제공되고 있는 의료기기 품목군은 개발하고자 하는 의료기기의 임상 빅데이터 기반 검증 가능성을 식별하기 어렵다. 그러므로 본 논문에서는 의료기기에 대한 다양한 특성 분석을 통해 임상 빅데이터 기반 검증 가능성을 식별하기 위한 프로세스를 제안한다. 제안하는 프로세스에서는 의료기기의 검증에 요구되는 데이터의 식별을 통해 임상 빅데이터를 이용한 테스트 데이터 수집 및 이를 활용한 신뢰성 높은 검증을 가능케 한다.

Re-defining Named Entity Type for Personal Information De-identification and A Generation method of Training Data (개인정보 비식별화를 위한 개체명 유형 재정의와 학습데이터 생성 방법)

  • Choi, Jae-hoon;Cho, Sang-hyun;Kim, Min-ho;Kwon, Hyuk-chul
    • Proceedings of the Korean Institute of Information and Commucation Sciences Conference
    • /
    • 2022.05a
    • /
    • pp.206-208
    • /
    • 2022
  • As the big data industry has recently developed significantly, interest in privacy violations caused by personal information leakage has increased. There have been attempts to automate this through named entity recognition in natural language processing. In this paper, named entity recognition data is constructed semi-automatically by identifying sentences with de-identification information from de-identification information in Korean Wikipedia. This can reduce the cost of learning about information that is not subject to de-identification compared to using general named entity recognition data. In addition, it has the advantage of minimizing additional systems based on rules and statistics to classify de-identification information in the output. The named entity recognition data proposed in this paper is classified into twelve categories. There are included de-identification information, such as medical records and family relationships. In the experiment using the generated dataset, KoELECTRA showed performance of 0.87796 and RoBERTa of 0.88.

  • PDF

Wrapper-based Approach for Protein Identification in PPI Network (PPI 네트워크에서의 래퍼 기반 단백질 식별)

  • Lee Yong-Ho;Choi Jae-Hun;Lim Myung-Eun;Park Su-Jun
    • Proceedings of the Korean Information Science Society Conference
    • /
    • 2006.06a
    • /
    • pp.7-9
    • /
    • 2006
  • 단백질 상호작용 관계들은 고 성능 실험 기법을 이용한 생물학적 실험에 의해서 대규모로 추출되고, 동시에 이들을 구성하는 단백질 데이터 역시 공공 데이터베이스에 빈번하게 갱신되고 있다. 이 갱신으로 인하여 인터넷을 통해 공개되는 공공 데이터베이스와 PPI(Protein-Protein interaction) 네트워크에 포함된 단백질 데이터가 서로 일치하지 않게 된다. 본 논문에서는 PPI 네트워크에 존재하는 단백질을 래퍼(Wrapper)를 이용하여 빈번하게 갱신되는 공공 데이터베이스의 단백질로 식별하고, 이 식별을 통해 PPI 네트워크에 존재하는 데이터들을 항상 최신 데이터로 동기화함으로써 데이터의 실시간성을 제공하고 데이터에 대한 신뢰도를 보장할 수 있도록 하였다.

  • PDF

A Study on the Applicability of ISNI for Authority Control (전거제어를 위한 국제표준이름식별자(ISNI)의 활용가능성에 관한 연구)

  • Lee, Mihwa
    • Journal of the Korean Society for information Management
    • /
    • v.31 no.3
    • /
    • pp.133-151
    • /
    • 2014
  • This study was to investigate the concept of ISNI and to find its availability in authority control, realizing importance of ISNI as the bridge identifier including all the information media content industries. ISNI is needed for global and comprehensive name authority control as the bridge identifier for the identification of public identities of parties involved throughout the information media content industries in the creation, production, management and content distribution chains. First of all, it was to inquire ISNI concept, goal, terms and definitions, structure and syntax, allocation of ISNI, administration of the ISNI system, and metadata. Next, it was to suggest the applicability of ISNI in authority control. First, it should be needed to consider in applying ISNI for cooperative authority control. It is possible to interactively use the authority data created in library and other information industries area by constructing KISNI system. Second, it is possible to construct linked data by linking various identifier through ISNI identifier as bridge identifier. Third, it is needed to develop KORMARC for describing ISNI identifier in KORMARC bibliographic and authority record.

Detecting Spam Data for Securing the Reliability of Text Analysis (텍스트 분석의 신뢰성 확보를 위한 스팸 데이터 식별 방안)

  • Hyun, Yoonjin;Kim, Namgyu
    • The Journal of Korean Institute of Communications and Information Sciences
    • /
    • v.42 no.2
    • /
    • pp.493-504
    • /
    • 2017
  • Recently, tremendous amounts of unstructured text data that is distributed through news, blogs, and social media has gained much attention from many researchers and practitioners as this data contains abundant information about various consumers' opinions. However, as the usefulness of text data is increasing, more and more attempts to gain profits by distorting text data maliciously or nonmaliciously are also increasing. This increase in spam text data not only burdens users who want to obtain useful information with a large amount of inappropriate information, but also damages the reliability of information and information providers. Therefore, efforts must be made to improve the reliability of information and the quality of analysis results by detecting and removing spam data in advance. For this purpose, many studies to detect spam have been actively conducted in areas such as opinion spam detection, spam e-mail detection, and web spam detection. In this study, we introduce core concepts and current research trends of spam detection and propose a methodology to detect the spam tag of a blog as one of the challenging attempts to improve the reliability of blog information.

Study on Public Institution Dataset Identification and Evaluation Process : Focusing on the Case of KR Electronic Procurement System (공공기관 데이터세트 식별과 평가 절차 연구 국가철도공단 전자조달시스템 사례를 중심으로)

  • Hwang, jin hyun;Baek, young mi;Yim, jin hee
    • The Korean Journal of Archival Studies
    • /
    • no.70
    • /
    • pp.41-83
    • /
    • 2021
  • After the revision of the Enforcement Decree of the Public Records Act, the archives created a management standard table for data set records management and performed management and control. Therefore, in this study, the data set record identification procedure and evaluation index were developed for systematic data set record management of archives. By applying this, a management standard table was prepared after identifying the records of 8 datasets in kr's electronic procurement system, and the evaluation was carried out according to the evaluation index, and the retention period, transfer, and collection were determined. It is hoped that this case study will be of practical use to the archives at a time when concrete examples of procedures for the management of dataset records are lacking.

전자 카탈로그 식별코드 표준화 방안

  • 성낙현
    • Proceedings of the CALSEC Conference
    • /
    • 2002.01a
    • /
    • pp.307-312
    • /
    • 2002
  • □가장 일상적으로 사용하는 체계 □상품의 모든 처리 단계에서 데이터의 색인을 구성 - 수요예측 - 주문서 작성 - 송장 -입고확인 -재고관리 -판매 □각 기업은 나름의 식별코드 □각 산업은 나름대로의 식별코드(중략)

  • PDF