기관회원 [로그인]
소속기관에서 받은 아이디, 비밀번호를 입력해 주세요.
개인회원 [로그인]

비회원 구매시 입력하신 핸드폰번호를 입력해 주세요.
본인 인증 후 구매내역을 확인하실 수 있습니다.

회원가입
서지반출
한국어 형태소 분석을 위한 3단계 확률 모델
[STEP1]서지반출 형식 선택
파일형식
@
서지도구
SNS
기타
[STEP2]서지반출 정보 선택
  • 제목
  • URL
돌아가기
확인
취소
  • 한국어 형태소 분석을 위한 3단계 확률 모델
저자명
이재성,Lee. Jae-Sung
간행물명
정보과학회논문지. Journal of KIISE. 소프트웨어 및 응용
권/호정보
2011년|38권 5호|pp.257-268 (12 pages)
발행정보
한국정보과학회
파일정보
정기간행물|
PDF텍스트
주제분야
기타
이 논문은 한국과학기술정보연구원과 논문 연계를 통해 무료로 제공되는 원문입니다.
서지반출

기타언어초록

확률 모델을 기반으로 만들어진 형태소 분석기는 형태소 품사 부착 말뭉치의 다양한 언어 현상과 태깅 원칙을 바로 학습할 수 있으므로 다양한 분야에 대한 적응력이 높다. 본 논문에서는 한국어 형태소 분석을 위한 3단계 확률 모델을 제안한다. 이 모델은 분석 단계를 형태소 복원, 분리, 태깅의 3단계로 나누어 독립된 모듈로 처리함으로써 기존의 2단계 확률 모델보다 처리 복잡도를 줄였다 또한, 음절 대신 자소 단위의 처리를 하고, 형태소 전이 확률을 이용하여 형태소 분리를 함으로써 다양한 품사 태깅 원칙을 학습할 수 있도록 했다. 모델의 성능 평가는 세종 계획 프로젝트에서 개발한 문어체 및 구어체 형태소 부착 발뭉치에 대해 실험하였고 기존의 방법들과 비교하였다.

기타언어초록

A morphological analyzer based on probabilistic model can learn easily various language phenomena and tagging principles used in morpheme-tagged corpus, so that it is very portable to various domains. In this paper, we propose a three-step probabilistic model for Korean morphological analysis which consists of original form restoring step, morpheme segmentation step and morpheme tagging step. The three-step method, which uses modular approach, reduces processing complexity compared with two-step probabilistic model. Processing in Jaso unit rather than syllable unit and using morpheme transition probahility for morpheme segmentation increase portahility for various tagging principles. Experiment on Sejong tagged corpus, both of written text corpus and spoken text corpus, was done to show the performance of the model and compare it with other methods.