数据集 / BAAI/BGE-VL-v1.5-zs

BAAI/BGE-VL-v1.5-zs 已完整同步

MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval

</a>
<a href="https://github.com/VectorSpaceLab/MegaPairs">
    
</a>
<a href="https://huggingface.co/datasets/JUNJIE99/MegaPairs">

</a>
<a href="https://huggingface.co/BAAI/BGE-VL-large">
    
</a>
<a href="https://huggingface.co/BAAI/BGE-VL-MLLM-S1">
    
</a>
<a href="https://huggingface.co/BAAI/BGE-VL-MLLM-S2">
    
</a>

News

2025-4-13 🎉🎉 We have uploaded our MegaPairs dataset to 🤗Hugging Face, which contains over 26 million multimodal retrieval instruction-tuning triplets. To reduce upload time and enhance data accessibility, we resized all images to a resolution of 512 × 512 instead of using their original size. This adjustment has minimal impact on performance, considering that most vision-language models (e.g., CLIP) use even smaller input image sizes. Dataset Card

2025-4-2 🌟🌟 BGE-VL models are also available on WiseModel.

2025-3-6 📰📰 Thank you to SyncedTech (机器之心), QbitAI (量子位), and AI Era (新智元) for reporting on our work!

2025-3-4 🚀🚀 We have released the BGE-VL-MLLM models on Huggingface: BGE-VL-MLLM-S1 and BGE-VL-MLLM-S2. BGE-VL-MLLM-S1 is trained exclusively on our MegaPairs dataset, achieving outstanding performance in composed image retrieval, with an 8.1% improvement on the CIRCO benchmark (mAP@5) over the previous state-of-the-art. BGE-VL-MLLM-S2 builds on BGE-VL-MLLM-S1 with an additional epoch of fine-tuning on the MMEB benchmark training set, delivering enhanced performance across a broader range of multimodal embedding tasks.

2024-12-27 🚀🚀 BGE-VL-CLIP models are released on Huggingface: BGE-VL-base and BGE-VL-large.

2024-12-19 🎉🎉 Release our paper: MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval.

Release Plan

  • Paper
  • BGE-VL-base and BGE-VL-large models
  • BGE-VL-MLLM model
  • MegaPairs Dataset
  • Evaluation code examples
  • Fine-tuning code

Introduction

In this work, we introduce MegaPairs, a novel data synthesis method that leverages open-domain images to create heterogeneous KNN triplets for universal multimodal retrieval. Our MegaPairs dataset contains over 26 million triplets, and we have trained a series of multimodal retrieval models, BGE-VL, including BGE-VL-CLIP (base and large) and BGE-VL-MLLM.

BGE-VL achieve state-of-the-a

29 个文件

浏览文件