Mapping a variable-length sentence to a fixed-length vector using BERT model
Repository
Repository di m-bain
CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
A Clip-Hitchiker's Guide to Long Video Retrieval [Arxiv 2022]
Video embeddings for retrieval - code for the paper "Use What You Have: Video retrieval using representations from collaborative experts"
Conceptual 12M is a dataset containing (image-URL, caption) pairs collected for vision-and-language pre-training.
Story-Based Retrieval with Contextual Embeddings. Largest freely available movie video dataset. [ACCV'20]
Condensed Movies Challenge 2021
Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval [ICCV'21]
An open-source AI agent that brings the power of Gemini directly into your terminal.
Hydra is a framework for elegantly configuring complex applications
LAVIS - A One-stop Library for Language-Vision Intelligence
Awesome multilingual OCR toolkits based on PaddlePaddle (practical ultra lightweight OCR system, support 80+ languages recognition, provide data annotation and synthesis tools, support training and deployment among server, mobile, embedded and IoT devices)
Automated Audiovisual Behaviour Recognition in Wild Primates
Neural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection, speaker embedding
PyTorch image models, scripts, pretrained weights -- (SE)ResNet/ResNeXT, DPN, EfficientNet, MixNet, MobileNet-V3/V2, MNASNet, Single-Path NAS, FBNet, and more
A pytorch implemented classifier for Multiple-Label classification
Multimodal language model benchmark, featuring challenging examples
Simple Diarization model