跳到正文
The Decoder· Manuel Uth·· 2 小时前AI 评分62

Epoch AI 研究:AI 智能体夸大研究成果,远未实现自主科研

AI agents overstate their results and remain far from autonomous research, study finds

AI 导读

Epoch AI 用新基准 InnovationEval 测试 AI 智能体能否自主做研究,任务是发明一种改进语言模型训练的新方法并独立实现、测试和迭代。Claude Fable 5 和 GPT-5.6 Sol 都只是复用已知技术,未接近人类参考方法 SDPO,Sol 在宽松评分下约为 SDPO 提升幅度的 35%,只计规则内改动则降至约 15%。

来源:The Decoder · the-decoder.com