文章摘要
该文章介绍了CursorBench评估结果,对多个AI模型在模糊多文件任务上的表现进行排名。Fable 5系列模型表现最佳,其中Fable 5 Max以72.9%的得分位居榜首,而GPT-5.5和Opus 4.7等模型紧随其后。
文章总结
我们针对真实Cursor会话中模糊、多文件的任务对智能体进行评估,分数越高越好。
| 排名 | 模型 | 得分 | 平均每任务成本(美元) | 平均每任务Token数 | 任务数 | |------|------|------|----------------------|------------------|--------| | 1 | Fable 5 Max | 72.9% | $18.02 | 63,842 | 76 | | 2 | Fable 5 Extra High | 72.0% | $13.74 | 48,754 | 63 | | 3 | Fable 5 High | 70.6% | $10.81 | 37,173 | 54 | | 4 | Fable 5 Medium | 69.8% | $8.27 | 28,507 | 47 | | 5 | Opus 4.7 Max | 64.8% | $11.02 | 62,989 | 96 | | 6 | GPT-5.5 Extra High | 64.3% | $4.37 | 17,905 | 46 | | 7 | Fable 5 Low | 64.2% | $5.70 | 18,882 | 36 | | 8 | Opus 4.8 Max | 63.8% | $7.59 | 77,370 | 60 | | 9 | Composer 2.5 | 63.2% | $0.55 | 15,152 | 37 | | 10 | GPT-5.5 High | 62.6% | $3.59 | 13,329 | 40 | | 11 | Opus 4.8 Extra High | 62.1% | $6.14 | 55,622 | 54 | | 12 | Opus 4.7 Extra High | 61.6% | $7.11 | 43,942 | 72 | | 13 | Sonnet 5 Max | 61.2% | $6.87 | 93,485 | 93 | | 14 | Opus 4.7 High | 59.4% | $5.01 | 32,227 | 59 | | 15 | GPT-5.5 Medium | 59.2% | $2.22 | 9,065 | 35 | | 16 | Opus 4.8 High | 58.4% | $4.41 | 36,788 | 45 | | 17 | Sonnet 5 Extra High | 58.4% | $5.23 | 58,228 | 86 | | 18 | Sonnet 5 High | 57.0% | $3.74 | 41,735 | 66 | | 19 | Opus 4.8 Medium | 56.6% | $3.83 | 31,684 | 41 | | 20 | Sonnet 5 Medium | 54.9% | $2.57 | 27,469 | 53 | | 21 | GLM 5.2 Max | 54.6% | $3.11 | 51,312 | 83 | | 22 | Opus 4.8 Low | 54.3% | $2.93 | 22,726 | 36 | | 23 | Opus 4.7 Medium | 52.7% | $2.93 | 19,193 | 41 | | 24 | Kimi K2.7 Code | 52.7% | $1.92 | 32,902 | 70 | | 25 | Composer 2 | 52.2% | $0.56 | 14,163 | 40 | | 26 | GLM 5.2 High | 50.7% | $2.46 | 30,621 | 76 | | 27 | Gemini 3.5 Flash | 49.8% | $1.94 | 35,105 | 79 | | 28 | Sonnet 4.6 Max | 49.0% | $3.09 | 40,280 | 55 | | 29 | GPT-5.5 Low | 48.8% | $1.19 | 4,923 | 24 | | 30 | Sonnet 4.6 High | 48.8% | $3.06 | 37,352 | 57 | | 31 | Opus 4.7 Low | 48.3% | $1.87 | 13,164 | 29 | | 32 | Sonnet 5 Low | 47.7% | $1.46 | 17,028 | 37 | | 33 | Kimi 2.6 | 47.6% | $1.27 | 24,783 | 56 | | 34 | Sonnet 4.6 Medium | 46.0% | $2.64 | 31,360 | 50 | | 35 | Sonnet 4.6 Low | 41.5% | $1.89 | 21,211 | 50 | | 36 | Kimi 2.5 | 31.9% | $0.87 | 9,446 | 30 |
更新日志
CursorBench 3.1
- 新增了专注于代码库理解、错误查找、规划和代码审查的问题。
- 改进了部分编辑任务的评分标准。
CursorBench 3.0
- 初始任务集聚焦于编辑、重构和错误修复问题。
平均每任务成本的计算方式为:将每个模型公布的每百万Token定价(包括输入、缓存读取、缓存写入和输出)应用于其在每个CursorBench 3.1任务中使用的Token数,然后对所有任务取平均值。结果存在一定波动,分数上的微小差异可能不具有统计意义。
评论总结
根据评论内容,主要观点和论据总结如下:
1. 对Cursor Composer 2.5性能的质疑(多数观点) - 多位用户认为Cursor的基准测试存在偏见,其声称Composer 2.5与Opus 4.8、GPT-5.5性能相当的说法不可信。 - 关键引用: - "Cursor's benchmark finds that Cursor's model (Composer 2.5) is basically as good as Opus 4.8 max and GPT-5.5 xhigh, but at a fraction of the price... third party benchmarks have it far behind." (mdasen) - "their claim that Composer 2.5 is anywhere close in performance to GPT 5.5 is absolutely farcical." (BugsJustFindMe) - "No shot 2.5 is beating out 4.8" (bfjvibybd6cuvu6)
2. 对Anthropic模型效率的批评(少数观点) - 用户o10449366质疑Anthropic模型(如Claude)在明确任务中过度消耗token,可能影响利润。 - 关键引用: - "I feel like this benchmark reiterates my disbelief that anyone uses the latest Anthropic models for any productive work... spawning unnecessary subagents." (o10449366) - "if the default is less efficient and more token expensive that directly results in higher 'profit' for Anthropic." (o10449366)
3. 对基准测试本身价值的怀疑(部分观点) - 用户认为基准测试易被操纵,且模型表现因任务而异,实际使用体验更重要。 - 关键引用: - "The independent benchmarks are probably part of training data now... The final test of a model is how good it works FOR YOU." (shadeslayer_) - "The only benchmark you can trust is your actual workload!" (andai)
4. 对成本与性能平衡的关注(部分观点) - 用户关注模型性价比,认为成本与性能的权衡是关键。 - 关键引用: - "I wish all these sites would show pareto frontier graphs of cost/performance." (xyzsparetimexyz) - "The most interesting part is costs. Gpt 5.5 and sonnet 5 cost same amount of money as GLM 5.2 but are more capable models." (maxdo)
5. 对图表设计的批评(少数观点) - 用户指出X轴反向设计不直观。 - 关键引用: - "backwards X axis? is there a reason for that? it looks ridiculous" (verse)
总结: 评论普遍质疑Cursor基准测试的客观性,认为Composer 2.5性能被高估;同时关注模型效率、成本与真实任务表现。多数用户建议以实际工作负载为准,而非依赖单一基准。