한국어 명사 출현 특성과 후절어를 이용한 명사추출기

한국어 명사 출현 특성과 후절어를 이용한 명사추출기

ㆍ 저자명: 박용현,황재원,고영중,Park. Yong-Hyun,Hwang. Jae-Won,Ko. Young-Joong
ㆍ 간행물명: 정보과학회논문지. Journal of KIISE. 소프트웨어 및 응용
ㆍ 권/호정보: 2010년|37권 12호|pp.919-927 (9 pages)
ㆍ 발행정보: 한국정보과학회
ㆍ 파일정보: 정기간행물|
PDF텍스트
ㆍ 주제분야: 기타

이 논문은 한국과학기술정보연구원과 논문 연계를 통해 무료로 제공되는 원문입니다.

서지반출

기타언어초록

최근 모바일 기기의 발전으로 인하여, PC뿐만 아니라 모바일 기기에서의 정보검색의 요구가 증가하고 있다. 모바일 기기에서 명사를 추출하기 위하여 기존의 언어분석도구를 사용하게 되면, 상대적으로 적은 메모리를 가지고 있는 모바일 기기에는 부담이 되게 된다. 따라서, 모바일 기기에 적합한 언어분석도구의 필요성이 증가하고 있다. 본 논문에서는 대량의 말뭉치로부터 추출한 영사 출현 특성과 후절어를 이용하여 명사를 추출하는 방법을 제안한다. 제안된 명사 추출기는 형태소 분석기를 사용한 기존 명사 추출기의 명사 사전의 약 4% 용량인 146KB의 명사 사전만을 사용함에도 불구하고, 최종적으로 $F_1$-measure 0.86라는 좋은 성능을 얻었다. 또한, 명사 사전에 대한 의존도가 낮으므로, 미등록 명사 추출에 대한 성능이 높을 것으로 예상된다.

기타언어초록

Since the performance of mobile devices is recently improved, the requirement of information retrieval is increased in the mobile devices as well as PCs. If a mobile device with small memory uses a tradition language analysis tool to extract nouns from korean texts, it will impose a burden of analysing language. As a result, the need for the language analysis tools adequate to the mobile devices is increasing. Therefore, this paper proposes a new method for noun extraction using post-noun morpheme sequences and noun patterns from a large corpus. The proposed noun extractor has only the dictionary capacity of 146KB and its performance shows 0.86 $F_1$-measure; the capacity of noun dictionary corresponds to only the 4% capacity of the existing noun extractor with a POS tagger. In addition, it easily extract nouns for unknown word because its dependence for noun dictionaries is low.

키워드

모바일 기기 명사 추출 명사 출현 특성 미등록어 추출 Mobile device Noun extraction Noun pattern Unknown word extraction

다운URL