WASET

%0 Journal Article
%A Liping Jing and Michael K. Ng and Xinhua Yang and Joshua Zhexue Huang
%D 2008
%J International Journal of Computer and Information Engineering
%B World Academy of Science, Engineering and Technology
%I Open Science Index 16, 2008
%T A Text Clustering System based on k-means Type Subspace Clustering and Ontology
%U https://publications.waset.org/pdf/2401
%V 16
%X This paper presents a text clustering system developed based on a k-means type subspace clustering algorithm to cluster large, high dimensional and sparse text data. In this algorithm, a new step is added in the k-means clustering process to automatically calculate the weights of keywords in each cluster so that the important words of a cluster can be identified by the weight values. For understanding and interpretation of clustering results, a few keywords that can best represent the semantic topic are extracted from each cluster. Two methods are used to extract the representative words. The candidate words are first selected according to their weights calculated by our new algorithm. Then, the candidates are fed to the WordNet to identify the set of noun words and consolidate the synonymy and hyponymy words. Experimental results have shown that the clustering algorithm is superior to the other subspace clustering algorithms, such as PROCLUS and HARP and kmeans type algorithm, e.g., Bisecting-KMeans. Furthermore, the word extraction method is effective in selection of the words to represent the topics of the clusters.

%P 1296 - 1308