범주형 시퀀스 데이터의 K-Nearest Neighbor알고리즘

범주형 시퀀스 데이터의 K-Nearest Neighbor알고리즘
A K-Nearest Neighbor Algorithm for Categorical Sequence Data

ㆍ 저자명: 오승준,Oh. Seung-Joon
ㆍ 간행물명: 韓國컴퓨터情報學會論文誌
ㆍ 권/호정보: 2005년|10권 2호|pp.215-221 (7 pages)
ㆍ 발행정보: 한국컴퓨터정보학회
ㆍ 파일정보: 정기간행물|
PDF텍스트
ㆍ 주제분야: 기타

이 논문은 한국과학기술정보연구원과 논문 연계를 통해 무료로 제공되는 원문입니다.

서지반출

기타언어초록

최근에는 단백질 시퀀스, 소매점 거래 데이터, 웹 로그 등과 같은 상업적이거나 과학적인 데이터의 폭발적인 증가를 볼 수 있다. 이런 데이터들은 순서적인 면을 가지고 있는 시퀀스 데이터들이다. 본 논문에서는 이런 시퀀스 데이터들을 분류하는 문제를 다룬다. 분류 기법 으로는 의사결정 나무나 베이지안 분류기, K-NN방법 등 석러 종류가 있는데, 본 연구에서는 또-U방법을 이용하여 시퀀스들을 분류한다. 또한, 시퀀스들간의 유사도를 구하기 위한 새로운 계산 방법과 효율적인 계산 방법도 제안한다.

기타언어초록

TRecently, there has been enormous growth in the amount of commercial and scientific data, such as protein sequences, retail transactions, and web-logs. Such datasets consist of sequence data that have an inherent sequential nature. In this Paper, we study how to classify these sequence datasets. There are several kinds techniques for data classification such as decision tree induction, Bayesian classification and K-NN etc. In our approach, we use a K-NN algorithm for classifying sequences. In addition, we propose a new similarity measure to compute the similarity between two sequences and an efficient method for measuring similarity

키워드

데이터 마이닝 분류 시퀀스 Data Mining Classification Sequences

다운URL