dclm-baseline-1.0 mlfoundations

DCLM-baseline DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks. Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime. Model Params Tokens Open dataset? CORE MMLU EXTENDED Open weights, closed datasets Llama2 7B 2T ✗ 49.2 45.8 34.1 DeepSeek 7B 2T ✗ 50.7 48.5 35.3 Mistral-0.3 7B ? ✗ 57.0 62.7 45.1 QWEN-2 7B ? ✗ 57.5 71.9 50.5 Llama3 8B 15T ✗… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0.

種別
dataset
ライセンス
cc-by-4.0
ダウンロード
292,863
いいね
304
アクセス
public
ファイル
0

README

--- license: cc-by-4.0 dataset_info: features: - name: bff_contained_ngram_count_before_dedupe dtype: int64 - name: language_id_whole_page_fasttext struct: - name: en dtype: float64 - name: metadata struct: - name: Content-Length dtype: string - name: Content-Type dtype: string - name: WARC-Block-D…

查看完整页面 · 查看原文