首页 正文

ToFe: Lagged Token Freezing and Reusing for Efficient Vision Transformer Inference

{{output}}
Although vision transformers (ViTs) have shown remarkable success in various vision tasks, their computationally expensive self-attention mechanisms hinder their deployment on resource-constrained edge devices. Token reduction, which discards less important to... ...