This article explores the performance of the CLIP model during the training process for text-based person retrieval tasks. It investigates whether the model's performance improves or declines after a certain number of epochs.