통합 검색 | Korea Science

Representative Batch Normalization for Scene Text Recognition

Sun, Yajie;Cao, Xiaoling;Sun, Yingying
- KSII Transactions on Internet and Information Systems (TIIS)
- /
- 제16권7호
- /
- pp.2390-2406
- /
- 2022
Scene text recognition has important application value and attracted the interest of plenty of researchers. At present, many methods have achieved good results, but most of the existing approaches attempt to improve the performance of scene text recognition from the image level. They have a good effect on reading regular scene texts. However, there are still many obstacles to recognizing text on low-quality images such as curved, occlusion, and blur. This exacerbates the difficulty of feature extraction because the image quality is uneven. In addition, the results of model testing are highly dependent on training data, so there is still room for improvement in scene text recognition methods. In this work, we present a natural scene text recognizer to improve the recognition performance from the feature level, which contains feature representation and feature enhancement. In terms of feature representation, we propose an efficient feature extractor combined with Representative Batch Normalization and ResNet. It reduces the dependence of the model on training data and improves the feature representation ability of different instances. In terms of feature enhancement, we use a feature enhancement network to expand the receptive field of feature maps, so that feature maps contain rich feature information. Enhanced feature representation capability helps to improve the recognition performance of the model. We conducted experiments on 7 benchmarks, which shows that this method is highly competitive in recognizing both regular and irregular texts. The method achieved top1 recognition accuracy on four benchmarks of IC03, IC13, IC15, and SVTP.
https://doi.org/10.3837/tiis.2022.07.015 인용 PDF KSCI HTML

주파수 영역 심층 신경망 기반 음성 향상을 위한 실수 네트워크와 복소 네트워크 성능 비교 평가 (Performance comparison evaluation of real and complex networks for deep neural network-based speech enhancement in the frequency domain)

황서림;박성욱;박영철
- 한국음향학회지
- /
- 제41권1호
- /
- pp.30-37
- /
- 2022
본 논문은 주파수 영역에서 심층 신경망 기반 음성 향상 모델 학습을 위하여 학습 대상과 네트워크 구조에 따라 두 가지 관점에서 성능을 비교 평가한다. 이때, 학습 대상으로는 스펙트럼 매핑과 Time-Frequency(T-F) 마스킹 기법을 사용하였고 네트워크 구조는 실수 네트워크와 복소 네트워크를 사용하였다. 음성 향상 모델의 성능은 데이터 셋 규모에 따라 Perceptual Evaluation of Speech Quality(PESQ)와 Short-Time Objective Intelligibility(STOI) 두 가지 객관적 평가지표를 통해 평가하였다. 실험 결과, 네트워크의 종류와 데이터 셋 종류에 따라 적정한 훈련 데이터의 크기가 다르다는 것을 확인하였다. 또한, 데이터의 크기와 학습 대상에 따라 복소 네트워크보다 실수 네트워크가 비교적 높은 성능을 보이기 때문에 총 파라미터의 수를 고려한다면 경우에 따라 실수 네트워크를 사용하는 것이 보다 현실적인 해결책일 수 있다는 것을 확인하였다.
https://doi.org/10.7776/ASK.2022.41.1.030 인용 PDF KSCI

Diagnostic Accuracy of Magnetic Resonance Imaging Features and Tumor-to-Nipple Distance for the Nipple-Areolar Complex Involvement of Breast Cancer: A Systematic Review and Meta-Analysis

Jung Hee Byon;Seungyong Hwang;Hyemi Choi;Eun Jung Choi
- Korean Journal of Radiology
- /
- 제24권8호
- /
- pp.739-751
- /
- 2023
Objective: This systematic review and meta-analysis evaluated the accuracy of preoperative breast magnetic resonance imaging (MRI) features and tumor-to-nipple distance (TND) for diagnosing occult nipple-areolar complex (NAC) involvement in breast cancer. Materials and Methods: The MEDLINE, Embase, and Cochrane databases were searched for articles published until March 20, 2022, excluding studies of patients with clinically evident NAC involvement or those treated with neoadjuvant chemotherapy. Study quality was assessed using the Quality Assessment of Diagnostic Accuracy Studies 2 tool. Two reviewers independently evaluated studies that reported the diagnostic performance of MRI imaging features such as continuity to the NAC, unilateral NAC enhancement, non-mass enhancement (NME) type, mass size (> 20 mm), and TND. Summary estimates of the sensitivity and specificity curves and the summary receiver operating characteristic (SROC) curve of the MRI features for NAC involvement were calculated using random-effects models. We also calculated the TND cutoffs required to achieve predetermined specificity values. Results: Fifteen studies (n = 4002 breast lesions) were analyzed. The pooled sensitivity and specificity (with 95% confidence intervals) for NAC involvement diagnosis were 71% (58-81) and 94% (91-96), respectively, for continuity to the NAC; 58% (45-70) and 97% (95-99), respectively, for unilateral NAC enhancement; 55% (46-64) and 83% (75-88), respectively, for NME type; and 88% (68-96) and 58% (40-75), respectively, for mass size (> 20 mm). TND had an area under the SROC curve of 0.799 for NAC involvement. A TND of 11.5 mm achieved a predetermined specificity of 85% with a sensitivity of 64%, and a TND of 12.3 mm yielded a predetermined specificity of 83% with a sensitivity of 65%. Conclusion: Continuity to the NAC and unilateral NAC enhancement may help predict occult NAC involvement in breast cancer. To achieve the desired diagnostic performance with TND, a suitable cutoff value should be considered.
https://doi.org/10.3348/kjr.2022.0846 인용 PDF

음성 향상을 위한 최소값 제어 음성 존재 부정확성의 추적기법 (Minima Controlled Speech Presence Uncertainty Tracking Method for Speech Enhancement)

이우정;장준혁
- 한국음향학회지
- /
- 제28권7호
- /
- pp.668-673
- /
- 2009
본 논문에서는 최소값 제어 음성 존재 부정확성의 추정기법을 이용한 음성 향상 기법을 제안한다. 기존의 음성 존재 부정확성 추정기법에서는 간단한 a posteriori SNR에 근거하여 프레임, 채널마다 다른 a priori음성 부재 확률값을 결정하여 음성 부재 확률 계산에 적용하였다. 본 논문에서 제안된 알고리즘은 기존 음성 존재 부정확성 추적방법과는 달리 최소값 제어방법을 이용하여 주파수성분별 최소값에 근거한 강인한 a priori음성 부재 확률값 추정방법을 통해 음성 부재 확률에 적용하여 음성을 향상시킨다. 제안된 음성 향상 기법은 ITU-T P.862 perceptual evaluation of speech quality (PESQ)를 이용하여 평가하였고 기존의 음성 존재 부정확성 추적방법보다 향상된 결과를 나타내었다.
https://doi.org/10.7776/ASK.2009.28.7.668 인용 PDF KSCI

지식기반 영상개선을 위한 지문영상의 품질분석 (Fingerprint Image Quality Analysis for Knowledge-based Image Enhancement)

윤은경;조성배
- 한국정보과학회논문지:소프트웨어및응용
- /
- 제31권7호
- /
- pp.911-921
- /
- 2004
지문영상으로부터 특징점을 정확하게 추출하는 것은 효과적인 지문인식 시스템의 구축에 매우 중요하다. 하지만 지문영상의 품질에 따라 특징점 추출의 정확도가 달라지기 때문에 지문인식 시스템에서의 영상 전처리 과정은 시스템의 성능에 크게 영향을 미친다. 본 논문에서는 지문영상으로부터 명암값의 평균 및 분산, 블록 방향성 차, 방향성 변화도, 융선과 골의 두께 비율 등의 5가지 특징을 추출하고 계층적 클러스터링 알고리즘으로 클러스터링하여 영상의 품질 특성을 분석한 후 습성(oily), 보통(neutral), 건성(dry)의 특성에 적합하게 영상을 개선하는 지식기반 전처리 방법을 제안한다. NIST DB 4와 인하대학교 데이타를 이용하여 실험한 결과, 클러스터링 기법이 영상의 특성을 제대로 구분함을 확인할 수 있었다. 또한 제안한 방법의 성능 평가를 위해 품질 지수와 블록 방향성 차이를 측정하여 일반적인 전처리 방법보다 지식기반 전처리 방법이 품질 지수와 블록 방향성 차이를 향상시킴을 확인할 수 있었다.
PDF KSCI

Signal Enhancement of a Variable Rate Vocoder with a Hybrid domain SNR Estimator

Park, Hyung Woo
- KSII Transactions on Internet and Information Systems (TIIS)
- /
- 제13권2호
- /
- pp.962-977
- /
- 2019
The human voice is a convenient method of information transfer between different objects such as between men, men and machine, between machines. The development of information and communication technology, the voice has been able to transfer farther than before. The way to communicate, it is to convert the voice to another form, transmit it, and then reconvert it back to sound. In such a communication process, a vocoder is a method of converting and re-converting a voice and sound. The CELP (Code-Excited Linear Prediction) type vocoder, one of the voice codecs, is adapted as a standard codec since it provides high quality sound even though its transmission speed is relatively low. The EVRC (Enhanced Variable Rate CODEC) and QCELP (Qualcomm Code-Excited Linear Prediction), variable bit rate vocoders, are used for mobile phones in 3G environment. For the real-time implementation of a vocoder, the reduction of sound quality is a typical problem. To improve the sound quality, that is important to know the size and shape of noise. In the existing sound quality improvement method, the voice activated is detected or used, or statistical methods are used by the large mount of data. However, there is a disadvantage in that no noise can be detected, when there is a continuous signal or when a change in noise is large.This paper focused on finding a better way to decrease the reduction of sound quality in lower bit transmission environments. Based on simulation results, this study proposed a preprocessor application that estimates the SNR (Signal to Noise Ratio) using the spectral SNR estimation method. The SNR estimation method adopted the IMBE (Improved Multi-Band Excitation) instead of using the SNR, which is a continuous speech signal. Finally, this application improves the quality of the vocoder by enhancing sound quality adaptively.
https://doi.org/10.3837/tiis.2019.02.026 인용 PDF KSCI HTML

인간의 감성에 기초한 승합차량 액슬의 음질 인덱스 개발에 대한 연구 (Development of Sound Quality Index of a SUV' Axle for Evaluation of Enhancement of Sound Quality Based on Human Sensibility)

임종태;이상권
- 한국소음진동공학회논문집
- /
- 제17권4호
- /
- pp.298-309
- /
- 2007
There are various sounds in the car as much as cars have many mechanical parts. These sounds make various psychological development. The international competition in car markets has continuously required the research about the sound quality of a car. The domestic car makers have also invested a lot of money for the research and development of sound quality. Car axle plays an important role in a vehicle and its NVH development is also important. By this time, NVH development of car axle is mainly based on the reduction of sound pressure level (dBA), which cannot gives, the satisfaction to the customers in view of the sound quality of a vehicle. Therefore, in this paper, a sound quality index evaluating the sound quality of axle noise based on human sensibility is developed.
https://doi.org/10.5050/KSNVN.2007.17.4.298 인용 PDF KSCI

Noise Reduction Using the Standard Deviation of the Time-Frequency Bin and Modified Gain Function for Speech Enhancement in Stationary and Nonstationary Noisy Environments

Lee, Soo-Jeong;Kim, Soon-Hyob
- The Journal of the Acoustical Society of Korea
- /
- 제26권3E호
- /
- pp.87-96
- /
- 2007
In this paper we propose a new noise reduction algorithm for stationary and nonstationary noisy environments. Our algorithm classifies the speech and noise signal contributions in time-frequency bins, and is not based on a spectral algorithm or a minimum statistics approach. It relies on calculating the ratio of the standard deviation of the noisy power spectrum in time-frequency bins to its normalized time-frequency average. We show that good quality can be achieved for enhancement speech signal by choosing appropriate values for ${\delta}_t\;and\;{\delta}_f$. The proposed method greatly reduces the noise while providing enhanced speech with lower residual noise and somewhat higher mean opinion score (MOS), background intrusiveness (BAK) and signal distortion (SIG) scores than conventional methods.
PDF KSCI

Color Image Enhancement Using a Retinex Algorithm with Bilateral Filtering for Images with Poor Illumination

Mulyantini, Agustien;Choi, Heung-Kook
- 한국멀티미디어학회논문지
- /
- 제19권2호
- /
- pp.233-239
- /
- 2016
Color enhancement basically deals with color manipulation in digital images. Recently, the technique has become widely used as a result of the increasing use of digital cameras. Retinex-based colorenhancement algorithms are a popular technique. In this paper, retinex with bilateral filtering is proposed to improve the quality of poorly illuminated images. Generally, it consists of three main steps: first, a retinex-based algorithm with color restoration; second, transformation mapping using histogram matching; and finally, smoothing the image using a bilateral filter. The experimental results demonstrate that the proposed method can successfully enhance image contrast while avoiding the halo effect and maintaining the color distribution in the image.
https://doi.org/10.9717/kmms.2016.19.2.233 인용 PDF KSCI KPUBS HTML

Gradient Fusion Method for Night Video Enhancement

Rao, Yunbo;Zhang, Yuhong;Gou, Jianping
- ETRI Journal
- /
- 제35권5호
- /
- pp.923-926
- /
- 2013
To resolve video enhancement problems, a novel method of gradient domain fusion wherein gradient domain frames of the background in daytime video are fused with nighttime video frames is proposed. To verify the superiority of the proposed method, it is compared to conventional techniques. The implemented output of our method is shown to offer enhanced visual quality.
https://doi.org/10.4218/etrij.13.0212.0550 인용 PDF KSCI

검색결과 1,488건 처리시간 0.025초

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

자세히 찾기

이미지 검색 (β)