WithinUsAI / claude_mythos_distilled_25k
模仿Claude Mythos的合成SFT数据集,涵盖网络安全等多个领域。
关键规格
HF 热度3,892 下载 · 143 收藏观测于 2026-07-07
源仓库更新2026-05-18
仓库许可不等于数据内容权利已全部清理,商用仍需核对数据来源、隐私与版权条款。
采用判断
项目资料与外部反馈 · 4 个来源 · 核验于 2026-07-23
该数据集是 Grok/xAI 生成的 100% 合成数据,旨在模拟 Anthropic 的 Claude Mythos 风格,但并非实际 Mythos 输出,且底层模型 Mythos 在社区中广泛被认为能力被夸大,采用风险较高。[1]
适合用来
- 作为合成 SFT 数据引导开源模型获得安全分析和代码生成能力
采用前注意
- 数据集为 100% 合成数据,由 Grok/xAI 生成来模拟 Claude Mythos 的输出风格,而非真实的 Mythos 或 Anthropic 模型输出,README 明确声明'Not actual outputs from Claude Mythos or any Anthropic model'[1]
- 仅限英文,且领域分布严重偏向网络安全(7,000 条)和高级编码(5,500 条),README 承认'heavy cyber focus because that was Mythos' most publicized strength'
- 数据集中存在大量模板化提示词,相近主题的 prompt 仅做轻微变体重复出现,表明合成生成方式多样性有限
- 底层被模拟对象 Claude Mythos 的能力在社区中受到广泛质疑,相关讨论指出其发现的漏洞大多不可利用且被夸大
- 该数据集的 JSONL 文件在 Hugging Face 上被标记为'Suspicious',存在文件完整性或安全风险
本段综合依据
- WithinUsAI/claude_mythos_distilled_25k · Datasets at Hugging Face ↗huggingface.co
- WithinUsAI/claude_mythos_distilled_25k at main ↗huggingface.co
- r/ClaudeAI on Reddit: Anthropic's Claude Mythos isn't a sentient super-hacker, it's a sales pitch — claims of 'thousands' of severe zero-days rely on just 198 manual reviews ↗reddit.com
- r/Anthropic on Reddit: Mythos is Mostly Hype... (also the bugs it found were mostly unexploitable and exaggerated...) ↗reddit.com