basic/claude-3-5-sonnet-20241022

模型描述

The Claude 3.5 Sonnet upgrade delivers significant improvements across benchmarks, particularly in coding and agentic tasks. It achieves 49.0% on SWE-bench Verified (up from 33.4%), outperforming all publicly available models, including specialized coding agents. It also excels in tool use, scoring 69.2% in retail and 46.0% in airline domains on TAU-bench. A major innovation is its computer use beta, enabling Claude to navigate UIs, click, type, and automate workflows—though still experimental. Early adopters like Replit and GitLab report 10% better reasoning and efficiency in multi-step coding tasks. Safety remains a priority, with joint testing by US/UK AI Safety Institutes confirming its adherence to ASL-2 risk standards.

全文结束

推荐模型

basic/gpt-4.1

GPT-4.1 是我们针对复杂任务的旗舰模型。它非常适合跨领域的问题解决。

gpt-4o-image

使用逆向工程在官方应用程序中调用模型并将其转换为 API。

gemini-2.5-flash-preview-04-17

Gemini-2.5-Flash-Preview-04-17 是一个大型语言模型,支持文本、图像、视频和音频输入,具有先进的输出和代码执行能力以及高令牌限制。