- 한국어 형태소 분석을 위한 3단계 확률 모델
- ㆍ 저자명
- 이재성,Lee. Jae-Sung
- ㆍ 간행물명
- 정보과학회논문지. Journal of KIISE. 소프트웨어 및 응용
- ㆍ 권/호정보
- 2011년|38권 5호|pp.257-268 (12 pages)
- ㆍ 발행정보
- 한국정보과학회
- ㆍ 파일정보
- 정기간행물| PDF텍스트
- ㆍ 주제분야
- 기타
확률 모델을 기반으로 만들어진 형태소 분석기는 형태소 품사 부착 말뭉치의 다양한 언어 현상과 태깅 원칙을 바로 학습할 수 있으므로 다양한 분야에 대한 적응력이 높다. 본 논문에서는 한국어 형태소 분석을 위한 3단계 확률 모델을 제안한다. 이 모델은 분석 단계를 형태소 복원, 분리, 태깅의 3단계로 나누어 독립된 모듈로 처리함으로써 기존의 2단계 확률 모델보다 처리 복잡도를 줄였다 또한, 음절 대신 자소 단위의 처리를 하고, 형태소 전이 확률을 이용하여 형태소 분리를 함으로써 다양한 품사 태깅 원칙을 학습할 수 있도록 했다. 모델의 성능 평가는 세종 계획 프로젝트에서 개발한 문어체 및 구어체 형태소 부착 발뭉치에 대해 실험하였고 기존의 방법들과 비교하였다.
A morphological analyzer based on probabilistic model can learn easily various language phenomena and tagging principles used in morpheme-tagged corpus, so that it is very portable to various domains. In this paper, we propose a three-step probabilistic model for Korean morphological analysis which consists of original form restoring step, morpheme segmentation step and morpheme tagging step. The three-step method, which uses modular approach, reduces processing complexity compared with two-step probabilistic model. Processing in Jaso unit rather than syllable unit and using morpheme transition probahility for morpheme segmentation increase portahility for various tagging principles. Experiment on Sejong tagged corpus, both of written text corpus and spoken text corpus, was done to show the performance of the model and compare it with other methods.