MALFM-Captioner: A Multipath Alignment Learning for Image Captioning With Feature Mask
{{output}}
Diffusion-based image captioning models effectively mitigate the token dependency issue inherent in autoregressive methods. However, the noise introduced in diffusion methods weakens sentence information, resulting in insufficient ability of image-text feature... ...