• 제목/요약/키워드: Visual Recognition

검색결과 826건 처리시간 0.027초

Multimodal audiovisual speech recognition architecture using a three-feature multi-fusion method for noise-robust systems

  • Sanghun Jeon;Jieun Lee;Dohyeon Yeo;Yong-Ju Lee;SeungJun Kim
    • ETRI Journal
    • /
    • 제46권1호
    • /
    • pp.22-34
    • /
    • 2024
  • Exposure to varied noisy environments impairs the recognition performance of artificial intelligence-based speech recognition technologies. Degraded-performance services can be utilized as limited systems that assure good performance in certain environments, but impair the general quality of speech recognition services. This study introduces an audiovisual speech recognition (AVSR) model robust to various noise settings, mimicking human dialogue recognition elements. The model converts word embeddings and log-Mel spectrograms into feature vectors for audio recognition. A dense spatial-temporal convolutional neural network model extracts features from log-Mel spectrograms, transformed for visual-based recognition. This approach exhibits improved aural and visual recognition capabilities. We assess the signal-to-noise ratio in nine synthesized noise environments, with the proposed model exhibiting lower average error rates. The error rate for the AVSR model using a three-feature multi-fusion method is 1.711%, compared to the general 3.939% rate. This model is applicable in noise-affected environments owing to its enhanced stability and recognition rate.

시각정보처리과정을 이용한 인공시각시스템 (Artificial Vision System using Human Visual Information Processing)

  • 서창진
    • 디지털융복합연구
    • /
    • 제12권11호
    • /
    • pp.349-355
    • /
    • 2014
  • 본 논문은 인간의 생물학적 시각정보처리과정 특성과 웨이블릿을 이용한 인공시각시스템을 제안한다. 인공시각시스템은 인간의 생물학적 시각정보 처리과정을 이용하며 맹인의 인공시각시스템 제작 및 일반적인 인식시스템에 적용할 수 있다. 본 논문에서는 인간의 망막에서 신경절 세포까지 정보 처리과정을 모델링하여 구성하였고 신경절 세포에서 후두엽 초기시각피질까지 전달되는 정보 복원과정을 모델링하여 초기시각피질에 표현되는 영상정보를 구성하였다. 실험과정은 인간의 시각정보 처리과정 특성인 저주파, 고주파 분해를 웨이블릿 변환으로 시각 정보를 구현하였고 이를 이용하여 인식 시스템에 적용하였다. 실험에 사용한 데이터는 AT&T 얼굴데이터베이스를 사용하였다. 그리고 제안하는 인간의 시각정보처리 과정 특성을 이용한 방법이 영상인식 시스템의 정확성을 향상시킬 수 있음을 실험을 통하여 증명하고자 한다.

망막 세포 특성에 의한 영상인식에 관한 연구 (A Study on Image Recognition based on the Characteristics of Retinal Cells)

  • 조재현;김도현;김광백
    • 한국정보통신학회논문지
    • /
    • 제11권11호
    • /
    • pp.2143-2149
    • /
    • 2007
  • 최근 시각 장애인을 위한 인공망막 모델 구현에 관한 연구 중 시피질 자극기 기술은 시각 자극 전달의 중간 단계를 생략하고 직접 뇌세포를 자극하는 것이다. 본 논문에서는 망막에서 시각 피질로 시각정보를 전달할 때 발생하는 시각 피질의 특성, 즉 방향성에 대한 반응 특성을 특징 데이터로 구성하여 인식함으로써 인간 시각 정보 처리와 유사한 영상 추출 및 인식 모델을 제안한다. 제안된 방법은 영상의 특징을 추출 한 후 Delta-bar-delta 기반 오류 역전파 알고리즘을 적용하여 영상의 특징들을 인식한다. 제시된 방법의 성능을 분석하기 위하여 다양한 숫자 패턴들을 대상으로 실험한 결과, 제안된 망막 세포로부터 전달된 정보를 방향성에 대한 민감성을 고려하여 영상의 특성을 추출하여 인식하는 모델이 기존의 영상 추출 및 인식 모델보다 인식률에 있어서는 별 차이가 없지만 다양한 실험에서 확인할 수 있듯이 인간 시각과 같이 인식 성능이 민감하지 않는 것을 알 수 있었다.

안면부 여과식 방진마스크와 안경 동시 착용 시 불편감과 밀착계수 비교 (Effects of Wearing between Respirators and Glasses Simultaneously on Physical and Visual Discomforts and Quantitative Fit Factors)

  • 어원석;최영보;신창섭
    • 한국안전학회지
    • /
    • 제33권2호
    • /
    • pp.52-60
    • /
    • 2018
  • This study compares the differences of the fit factor by the order of wearing preference between Particulate filtering facepiece respirators(PFFR) and glasses when participants wore simultaneously and a survey of physical and visual complaint. Recognition level about fit of respirators was investigated and the educational (before- and after-) effect of the fit factor. When participants wore PFFR and glasses, physical complaints were nose pressure, slipping, nose and ear pressure, ear pressure and rim loosen, the most highly physical complaints were nose pressure. Visual complaints were demister, blurry vision, dizziness, visual field, and lens dirty, the most highly visual complaints were demister. But, there was significant difference in physical complaint such as nose pressure(10.3%), slipping (23.0%), nose and ear pressure(14.3%), and rim loosen(16.2%), visual complaint such as visual field(13.8%) and lens dirty(32.4%). For the recognition of fit of respirators, respirators fitness, leak site, an initial point and an object, faulty factor, recognition level was higher. Fit factor was increased after education of proper wearing of respirator. Change of the fit factor was smaller compared to the normal breathing and after 6 actions in case of after education. Questionnaire consisted of general characteristics and physical/visual complaint, recognition of fit. Complaints were measured after the QNFT with multiple choices. Quantitative fit factor was measured by device and compared the result of (before- and after-) educational effect. Also, we selected to 6 actions (Normal breathing, Deep breathing, Bending over, Turning head side to side, Moving head up and down, Normal breathing) among 8 actions OSHA QNFT (Quantitative Fit testing) protocol to measure the fit factors. The fit factor was higher after the training (p=0.000). Descriptive statistics, paired t-test, and Wilcoxon analysis were performed to describe the result of questionnaire and fit test. (P=0.05) Therefore, it is necessary to investigate the quantitative research such as training program and glasses fitting factor about the wearing of PFFR and glasses simultaneously.

방향 정보 처리에 의한 영상 인식에 관한 연구 (A Study on Image Recognition by Orientation Information)

  • 조재현;김진환;이종희
    • 한국정보통신학회:학술대회논문집
    • /
    • 한국해양정보통신학회 2009년도 추계학술대회
    • /
    • pp.308-309
    • /
    • 2009
  • 인간의 시각정보처리는 망막에서 입력된 영상을 시각피질에 전달될 때 많은 특성을 가지고 있다. 그중에서 방향성에 대한 민감한 반응을 분석하여 영상인식에 미치는 영향을 분석하고자 한다. 수직반응의 가중치와 수평 및 대각선반응의 가중치를 여러 가지로 변동하여 영상의 인식률을 비교함으로써 수직반응 정도에 매우 민감함을 보이며 차후 인간시각모델 구성에 적용하고자 한다.

  • PDF

Robust Video-Based Barcode Recognition via Online Sequential Filtering

  • Kim, Minyoung
    • International Journal of Fuzzy Logic and Intelligent Systems
    • /
    • 제14권1호
    • /
    • pp.8-16
    • /
    • 2014
  • We consider the visual barcode recognition problem in a noisy video data setup. Unlike most existing single-frame recognizers that require considerable user effort to acquire clean, motionless and blur-free barcode signals, we eliminate such extra human efforts by proposing a robust video-based barcode recognition algorithm. We deal with a sequence of noisy blurred barcode image frames by posing it as an online filtering problem. In the proposed dynamic recognition model, at each frame we infer the blur level of the frame as well as the digit class label. In contrast to a frame-by-frame based approach with heuristic majority voting scheme, the class labels and frame-wise noise levels are propagated along the frame sequences in our model, and hence we exploit all cues from noisy frames that are potentially useful for predicting the barcode label in a probabilistically reasonable sense. We also suggest a visual barcode tracking approach that efficiently localizes barcode areas in video frames. The effectiveness of the proposed approaches is demonstrated empirically on both synthetic and real data setup.

3차원 물체의 인식 성능 향상을 위한 감각 융합 시스템 (Sensor Fusion System for Improving the Recognition Performance of 3D Object)

  • Kim, Ji-Kyoung;Oh, Yeong-Jae;Chong, Kab-Sung;Wee, Jae-Woo;Lee, Chong-Ho
    • 대한전기학회:학술대회논문집
    • /
    • 대한전기학회 2004년도 학술대회 논문집 정보 및 제어부문
    • /
    • pp.107-109
    • /
    • 2004
  • In this paper, authors propose the sensor fusion system that can recognize multiple 3D objects from 2D projection images and tactile information. The proposed system focuses on improving recognition performance of 3D object. Unlike the conventional object recognition system that uses image sensor alone, the proposed method uses tactual sensors in addition to visual sensor. Neural network is used to fuse these informations. Tactual signals are obtained from the reaction force by the pressure sensors at the fingertips when unknown objects are grasped by four-fingered robot hand. The experiment evaluates the recognition rate and the number of teaming iterations of various objects. The merits of the proposed systems are not only the high performance of the learning ability but also the reliability of the system with tactual information for recognizing various objects even though visual information has a defect. The experimental results show that the proposed system can improve recognition rate and reduce learning time. These results verify the effectiveness of the proposed sensor fusion system as recognition scheme of 3D object.

  • PDF

3차원 물체의 인식 성능 향상을 위한 감각 융합 신경망 시스템 (Neural Network Approach to Sensor Fusion System for Improving the Recognition Performance of 3D Objects)

  • 동성수;이종호;김지경
    • 대한전기학회논문지:시스템및제어부문D
    • /
    • 제54권3호
    • /
    • pp.156-165
    • /
    • 2005
  • Human being recognizes the physical world by integrating a great variety of sensory inputs, the information acquired by their own action, and their knowledge of the world using hierarchically parallel-distributed mechanism. In this paper, authors propose the sensor fusion system that can recognize multiple 3D objects from 2D projection images and tactile informations. The proposed system focuses on improving recognition performance of 3D objects. Unlike the conventional object recognition system that uses image sensor alone, the proposed method uses tactual sensors in addition to visual sensor. Neural network is used to fuse the two sensory signals. Tactual signals are obtained from the reaction force of the pressure sensors at the fingertips when unknown objects are grasped by four-fingered robot hand. The experiment evaluates the recognition rate and the number of learning iterations of various objects. The merits of the proposed systems are not only the high performance of the learning ability but also the reliability of the system with tactual information for recognizing various objects even though the visual sensory signals get defects. The experimental results show that the proposed system can improve recognition rate and reduce teeming time. These results verify the effectiveness of the proposed sensor fusion system as recognition scheme for 3D objects.

A Salient Based Bag of Visual Word Model (SBBoVW): Improvements toward Difficult Object Recognition and Object Location in Image Retrieval

  • Mansourian, Leila;Abdullah, Muhamad Taufik;Abdullah, Lilli Nurliyana;Azman, Azreen;Mustaffa, Mas Rina
    • KSII Transactions on Internet and Information Systems (TIIS)
    • /
    • 제10권2호
    • /
    • pp.769-786
    • /
    • 2016
  • Object recognition and object location have always drawn much interest. Also, recently various computational models have been designed. One of the big issues in this domain is the lack of an appropriate model for extracting important part of the picture and estimating the object place in the same environments that caused low accuracy. To solve this problem, a new Salient Based Bag of Visual Word (SBBoVW) model for object recognition and object location estimation is presented. Contributions lied in the present study are two-fold. One is to introduce a new approach, which is a Salient Based Bag of Visual Word model (SBBoVW) to recognize difficult objects that have had low accuracy in previous methods. This method integrates SIFT features of the original and salient parts of pictures and fuses them together to generate better codebooks using bag of visual word method. The second contribution is to introduce a new algorithm for finding object place based on the salient map automatically. The performance evaluation on several data sets proves that the new approach outperforms other state-of-the-arts.

모바일 환경에서의 시각 음성인식을 위한 눈 정위 기반 입술 탐지에 대한 연구 (A Study on Lip Detection based on Eye Localization for Visual Speech Recognition in Mobile Environment)

  • 송민규;;김진영;황성택
    • 한국지능시스템학회논문지
    • /
    • 제19권4호
    • /
    • pp.478-484
    • /
    • 2009
  • 음성 인식 기술은 편리한 삶을 추구하는 요즘 추세에 HMI를 위해 매력적인 기술이다. 음성 인식기술에 대한 많은 연구가 진행되고 있으나 여전히 잡음 환경에서의 성능은 취약하다. 이를 해결하기 위해 요즘은 청각 정보 뿐 아니라 시각 정보를 이용하는 시각 음성인식에 대한 연구가 활발히 진행되고 있다. 본 논문에서는 모바일 환경에서의 시각 음성인식을 위한 입술의 탐지 방법을 제안한다. 시각 음성인식을 위해서는 정확한 입술의 탐지가 필요하다. 우리는 입력 영상에서 입술에 비해 보다 찾기 쉬운 눈을 이용하여 눈의 위치를 먼저 탐지한 후 이 정보를 이용하여 대략적인 입술 영상을 구한다. 구해진 입술 영상에 K-means 집단화 알고리듬을 이용하여 영역을 분할하고 분할된 영역들 중 가장 큰 영역을 선택하여 입술의 양 끝점과 중심을 얻는다. 마지막으로, 실험을 통하여 제안된 기법의 성능을 확인하였다.