Supervised Contrastive Frame Aggregation for Video Representation Learning
arxiv.org·11h
📊Learned Metrics
Preview
Report Post

View PDF HTML (experimental)

Abstract:We propose a supervised contrastive learning framework for video representation learning that leverages temporally global context. We introduce a video to image aggregation strategy that spatially arranges multiple frames from each video into a single input image. This design enables the use of pre trained convolutional neural network backbones such as ResNet50 and avoids the computational overhead of complex video transformer models. We then design a contrastive learning objective that directly compares pairwise projections generated by the model. Positive pairs are defined as projections from videos sharing the same label while all other projections are treated as neg…

Similar Posts

Loading similar posts...