Extensive Incorporation of K-Nearest Neighbor to Support Vector Machine through Correlation Studies for a Better Classification

Authors

  • Doreen Ying Ying Sim

Abstract

Correlation studies are performed to assess the patterns of data for all the applied datasets before k-Nearest Neighbor (kNN) algorithms are incorporated to the Support Vector Machine to form the algorithms to be developed. Based on the correlation studies on the data of the applied datasets, each of them is categorized into three categories of low, medium or high correlations. Weighted distances of kNN are then computed based on the correlation studies in order to tune and adjust the hinge loss function and width of the Support Vector Machine (SVM) kernel. The proposed formulations of SVM which are derived from the studies on the correlations among the patterns of data from the applied datasets are applied to tune the kernel width and therefore adjust the hinge loss function of SVM. When the adjusted SVM hinge loss function and its optimized kernel width after being computed by the proposed formulations, it is found that for all datasets, having low, moderate or high correlation coefficients, the proposed formulations and algorithms can optimally adjust the SVM Gaussian kernel width so that its hinge loss function can be tuned. After implementing the proposed and developed correlation-based k-Nearest Neighbor Support Vector Machine (ckNNSVM) algorithms to the datasets, it is shown to be more accurate in classification when compared with the classical SVM classification and the kNNSVM algorithms without getting the SVM hinge loss function tuned and/or the Gaussian kernel width adjusted accordingly through extensive correlation studies.

Downloads

Published

2020-02-21

Issue

Section

Articles