fineweb-tokenized anisoleai

FineWeb Tokenized > 4 trillion tokens of the pre-tokenized data the 🌐 web has to offer What is it? This is a pre-tokenized version of the HuggingFaceFW/fineweb dataset (currently in-progress, tokenization of the ~15 trillion tokens corpus is ongoing). The data is being pre-processed and tokenized using the AnisoleAI BPE tokenizer (52,022 vocabulary size) and packed into compact uint16 Parquet shards. By distributing the pre-tokenized corpus, we eliminate… See the full description on the dataset page: https://huggingface.co/datasets/anisoleai/fineweb-tokenized.

Type
dataset
License
odc-by
Language
en
Downloads
1,717,378
Likes
31
Access
public
Files
0

README

--- license: odc-by task_categories: - text-generation language: - en pretty_name: FineWeb Tokenized (AnisoleAI) size_categories: - n>1T tags: - tabular - text - pre-training configs: - config_name: default data_files: - split: train path: data_*/*.parquet --- # <img src="https://i.postimg.cc/SRBB1…

查看完整页面 · 查看原文