MINT-1T-HTML mlfoundations

๐Ÿƒ MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens ๐Ÿƒ MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. ๐Ÿƒ MINT-1T is designed to facilitate research in multimodal pretraining. ๐Ÿƒ MINT-1T is created by a team from the University of Washington inโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML.

็จฎๅˆฅ
dataset
ใƒฉใ‚คใ‚ปใƒณใ‚น
cc-by-4.0
่จ€่ชž
en
ใƒ€ใ‚ฆใƒณใƒญใƒผใƒ‰
186,938
ใ„ใ„ใญ
97
ใ‚ขใ‚ฏใ‚ปใ‚น
public
ใƒ•ใ‚กใ‚คใƒซ
0

ๆŸฅ็œ‹ๅฎŒๆ•ด้กต้ข ยท ๆŸฅ็œ‹ๅŽŸๆ–‡