PersonaHub proj-persona
Scaling Synthetic Data Creation with 1,000,000,000 Personas This repo releases data introduced in our paper Scaling Synthetic Data Creation with 1,000,000,000 Personas: We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data. To fully exploit this methodology at scale, we introduce PERSONA HUB – a collection of 1 billion diverse personas automatically curated from web… See the full description on the dataset page: https://huggingface.co/datasets/proj-persona/PersonaHub.
- 类型
- dataset
- 许可
- cc-by-nc-sa-4.0
- 语言
- en
- 下载量
- 9,861
- 点赞
- 792
- 访问
- public
- 文件
- 0