HuggingFaceFW / fineweb
从CommonCrawl清洗去重后的英文网页文本数据集
社区实测
大规模网页预训练语料(15 万亿 token),在同等规模数据集中表现领先
- 提供 15 万亿 token 的高质量网页预训练语料,在同类规模中表现领先
- FineWeb-Edu 子集提供 1.3 万亿 token 教育类高质量内容,满足对内容质量要求更高的预训练场景
- 原始数据源自 Common Crawl,需经过大量过滤清洗才能达到可用质量
来源
FineWeb: decanting the web for the finest text data at scaleFineWeb-Edu: How to Make a Very High-Quality Dataset to Pre-train ...lmmx/bbcfw: Exploring the BBC News subset of the FineWeb dataset ...
截至 2026-06-20