In this paper, various Kolmogorov-Arnold Network (KAN), namely Smooth KAN (S-KAN), Wavelength KAN (WAV-KAN), Basic KAN (B-KAN), and Temporal KAN (T-KAN) models, is compared. The comparison is done based upon how they are processing the multimodal data. These models are the variation of KAN model either by incorporating the functional or structural changes to enhance their scalability and efficiency in accessing complex multimodal data. Wavelength KAN used the strength of wavelet-based concepts, which helped to access the low- and high-frequency data components to enhance the model. To reduce the overfitting, Smooth KAN utilizes the functional transformations. To address temporal dependencies and sequential flow of data, KAN is modified for time-sensitive information as T-KAN. The datasets used to compare the performance of these models are image, audio, and time-series data. The performance is compared in terms of accuracy, interpretability, computational efficiency, and generalizability and provide the strengths and limitations of each KAN model. The analysis made in this paper can be utilized in the future to develop KAN-based models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative Study of Multimodel Architecture in Generative AI

  • Prabhjot Kaur,
  • Aditya Chauhan

摘要

In this paper, various Kolmogorov-Arnold Network (KAN), namely Smooth KAN (S-KAN), Wavelength KAN (WAV-KAN), Basic KAN (B-KAN), and Temporal KAN (T-KAN) models, is compared. The comparison is done based upon how they are processing the multimodal data. These models are the variation of KAN model either by incorporating the functional or structural changes to enhance their scalability and efficiency in accessing complex multimodal data. Wavelength KAN used the strength of wavelet-based concepts, which helped to access the low- and high-frequency data components to enhance the model. To reduce the overfitting, Smooth KAN utilizes the functional transformations. To address temporal dependencies and sequential flow of data, KAN is modified for time-sensitive information as T-KAN. The datasets used to compare the performance of these models are image, audio, and time-series data. The performance is compared in terms of accuracy, interpretability, computational efficiency, and generalizability and provide the strengths and limitations of each KAN model. The analysis made in this paper can be utilized in the future to develop KAN-based models.