Captioning Images Effectively: Investigating BLEU Scores in CNN-LSTM Models with Different Training Configurations on Flickr8k Dataset
摘要
This study delves into the pivotal role played by batch size and epochs in determining optimal parameter values for achieving peak accuracy in image captioning, as measured by BLEU scores of 1 and 2. Utilizing the CNN-LSTM model, the nuanced relationship between parameter values and BLEU scores is examined, revealing distinct trends in the graph. Larger batch sizes emerge as crucial in mitigating overfitting and underfitting issues, thereby enhancing consistency in accuracy across epochs. The findings offer a roadmap for practitioners to optimize resource utilization, empowering them to extract maximum value from limited computational resources. Through strategic parameter fine-tuning, this research not only advances the field of machine learning but also fosters efficiency and innovation in various projects. In an era of resource constraints, this study illuminates a pathway toward achieving maximal accuracy while efficiently allocating resources. Through meticulous analysis, actionable insights are provided for navigating the complexities of parameter optimization in machine learning applications.