wikipedia legacy-datasets

Wikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wikipedia dump (https://dumps.wikimedia.org/) with one split per language. Each example contains the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.).

種別
dataset
ライセンス
cc-by-sa-3.0
言語
aa
ダウンロード
113,740
いいね
663
アクセス
public
ファイル
0

查看完整页面 · 查看原文