最新
2026-09-04
实质变化 20成功 61失败 6共 67 个信源
AI 今日要点
- 发布OpenAIOpenAI 发布旗舰模型 GPT-6 Astra,面向企业 Trusted Access Program 开放 API 访问,定价 $10/MTok 输入、$50/MTok 输出,并新增 Async tool calling、Mid-turn steering 等控制功能。
- 发布Google Vertex AIGoogle Vertex AI 上线 Gemini 3.8 Flash,主打长时编码与自主智能体场景,并推出 $0.75/$3.75 每百万 tokens 的入门定价及 50% Provisioned Throughput 促销信用。
- 动态硅基流动硅基流动新增 zai-org/GLM-5.3、MiniMax-M2.5、Qwen3.5-397B-A17B、Nex-N2-Pro 等模型,同时宣布将于 2026-09-11 下线 Nex-N2-Pro、Qwen3.5-397B-A17B、MiniMax-M2.5 等模型。
- 调价硅基流动硅基流动调整 DeepSeek-V4-Flash 分时段定价及 DeepSeek-V4-Pro 缓存命中输入 tokens 价格,并宣布 /user/info 接口将停止服务。
- 调价Google Vertex AIGoogle Vertex AI 新增 Gemini 3.5 Live Translate 定价(音频输入 $3.50/百万、输出 $21.00/百万),并扩展 Gemini 3.8 Flash 的音频 token 计费说明。
- 动态LMArenaLMArena 排行榜新增 claude-fable-5.1-max(Elo 1765)、hy4-preview(1626)等模型,其中 claude-fable-5.1-max 大幅领先其他新上榜模型。
删除 4 行
- SDK 升级与使用说明
- 知识型视频智能问答解决方案
- 视频高光提取及检索解决方案
- 音视频知识问答核心流程
Computer Use Agent 服务规则
新增 2 行
- Computer Use Agent 服务规则
- 用量明细 CSV 解析最佳实践
zai-org/GLM-5.3
zai-org/GLM-5.3新增 1 行
zai-org/GLM-5.3
删除 2 行
- GLM-Z1-9B-0414
- THUDM/GLM-Z1-9B-0414
MiniMaxAI/MiniMax-M2.5
MiniMaxAI/MiniMax-M2.5变更 1 处(旧 → 新)
- 【模型价格调整】DeepSeek-V4-Flash 模型分时段定价调整【接口服务调整】/user/info 接口将停止服务【模型价格调整】DeepSeek-V4-Pro 缓存命中输入 tokens 价格调整【模型价格调整】Nex-N2-Pro、DeepSeek-V4-Pro、DeepSeek-V3.2、Qwen3.6 →【模型服务调整】Nex-N2-Pro、Qwen3.5-397B-A17B、MiniMax-M2.5 等模型将下线【模型价格调整】DeepSeek-V4-Flash 模型分时段定价调整【接口服务调整】/user/info 接口将停止服务【模型价格调整】DeepSeek-V4-Pro 缓存命中输入 tokens 价格调整【
新增 6 行
MiniMaxAI/MiniMax-M2.5Pro/MiniMaxAI/MiniMax-M2.5Qwen/Qwen3.5-397B-A17Bnex-agi/Nex-N2-Pro- 【模型服务调整】Nex-N2-Pro、Qwen3.5-397B-A17B、MiniMax-M2.5 等模型将下线
- 为进一步优化资源配置,提供更先进、优质的技术服务,平台将于 2026-09-11 对下列模型进行下线处理:
$10 / Input MTok
变更 2 处(旧 → 新)
- If you're not sure where to start, use GPT-5.6 Sol, our flagship model for complex reasoning and coding. Choose GPT-5.6 Terra to balance intelligence and cost, →If you're not sure where to start, use GPT-6 Astra, our flagship model for complex reasoning and coding. Choose GPT-5.6 Terra to balance intelligence and cost,
- Using GPT-5.6→Using GPT-6 Astra
新增 13 行
- $10 / Input MTok
- $50 / Output MTok
- Async tool calling
- Compare model capabilities and specifications.
- GPT-6 Astra
- GPT‑6 Astra is rolling out today for enterprises in our Trusted Access Program, with access through API and our Plus, Pro, Business and Enterprise plans coming in the coming days.
- Mid-turn steering
- Misalignment monitoring
- Our most capable model, built for the hardest
end-to-endwork - Safety classifiers
- Under-18 guidance
gpt-6-astra- lowmediumhighxhighmax
删除 2 行
- Start with GPT-5.6 Sol for complex reasoning and coding, choose GPT-5.6 Terra to balance intelligence and cost, or use GPT-5.6 Luna for cost-sensitive, high-volume workloads.
- Under 18 API Guidance
Fast mode is unavailable for GPT-6 Astra with EU data residency. Use Standard processing for those requests. See Fast mode compatibility.
变更 1 处(旧 → 新)
- Using GPT-5.6→Using GPT-6 Astra
新增 10 行
- Async tool calling
- Fast mode is unavailable for GPT-6 Astra with EU data residency. Use Standard processing for those requests. See Fast mode compatibility.
- GPT‑6 Astra is rolling out today for enterprises in our Trusted Access Program, with access through API and our Plus, Pro, Business and Enterprise plans coming in the coming days.
- Mid-turn steering
- Misalignment monitoring
- Safety classifiers
- Under-18 guidance
- |
gpt-6-astra| $10.00 | $1.00 | $12.50 | $50.00 | $20.00 | $2.00 | $25.00 | $75.00 - |
gpt-6-astra| $20.00 | $2.00 | $25.00 | $100.00 | $40.00 | $4.00 | $50.00 | $150.00 - |
gpt-6-astra| $5.00 | $0.50 | $6.25 | $25.00 | $10.00 | $1.00 | $12.50 | $37.50
删除 1 行
- Under 18 API Guidance
Added new controls for long-running work with GPT-6 Astra in the Responses API:
long-running work with GPT-6 Astra in the Responses API:变更 1 处(旧 → 新)
- Using GPT-5.6→Using GPT-6 Astra
新增 22 行
- Added new controls for
long-runningwork with GPT-6 Astra in the Responses API: - Async tool calling
- Async tool calling: Let the model continue working while your application runs function or custom tools, then return results as they become available.
- Change reasoning effort
mid-conversation: Increase effort for difficult work or reduce it for routinefollow-upswhile preserving the cached prompt prefix. - Connections to api.openai.com can now use IPv6.
- GPT-6 Astra does not support custom temperature or top_p values or log probabilities (logprobs).
- GPT-6 Astra does not support the none reasoning effort level.
- Key changes to consider when migrating:
- Mid-turn steering
- Mid-turn steering: Send additional instructions while a response is in progress over WebSockets, so the model can incorporate corrections or changing requirements.
- Misalignment monitoring
- Misalignment monitoring asynchronously checks for potential issues during agent work in supported Responses API requests. Checks can trigger safety alerts or stop a conversation for review.
- Released GPT-6 Astra, our most capable model, built for the hardest
end-to-endwork. - Safety classifiers
- September, 2026
删除 1 行
- Under 18 API Guidance
展示前 15 条新增 / 8 条删除,完整数据见仓库 data/diff/ 目录
新增 1 行
- | Show all details
删除 1 行
- | Show all detailsShow fewer details
The guides for the Claude Enterprise endpoints of the Admin API (user management and spend limits), the Claude Enterprise Analytics API, and the Compliance API now show the anthropic-version header; send it on every request to these endpoints, as in the rest of the Claude API. See API versions.
anthropic-version header; send it on every request to these endpoints, as in the rest of the Claude API. See API versions.变更 1 处(旧 → 新)
- Text generated by Claude Fable 5.1 and Claude Mythos 5.1 carries Anthropic's text watermark, and supported image and video files that Claude produces through th→Text generated by Claude Fable 5.1 and Claude Mythos 5.1 carries Anthropic's text watermark, and supported image, video, and audio files that Claude produces th
新增 1 行
- The guides for the Claude Enterprise endpoints of the Admin API (user management and spend limits), the Claude Enterprise Analytics API, and the Compliance API now show the
anthropic-versionheader; send it on every requ
5 amazing visuals show how the male fruit fly’s brain map is advancing neuroscience
新增 1 行
- 5 amazing visuals show how the male fruit fly’s brain map is advancing neuroscience
删除 1 行
- AMIE, our research medical AI system, demonstrates real-time clinical video consultation capabilities in a first-of-its-kind study.
Built for long-horizon coding and autonomous agents
long-horizon coding and autonomous agents变更 2 处(旧 → 新)
- Last updated 2026-09-01 UTC.→Last updated 2026-09-02 UTC.
- [[["Easy to understand","easyToUnderstand","thumb-up"],["Solved my problem","solvedMyProblem","thumb-up"],["Other","otherUp","thumb-up"]],[["Hard to understand"→[[["Easy to understand","easyToUnderstand","thumb-up"],["Solved my problem","solvedMyProblem","thumb-up"],["Other","otherUp","thumb-up"]],[["Hard to understand"
新增 5 行
- Built for
long-horizoncoding and autonomous agents - Gemini 3.8 Flash
- Improved response quality
- Our intelligent everyday driver for developers, delivering a step forward in software engineering and feeling distinctly better to build with.
- Our most intelligent workhorse model yet, built for
long-horizoncoding and autonomous agents.
删除 2 行
- Enhanced efficiency and practical reasoning, helping you build and iterate with greater ease
- Greater token efficiency
$2.00 / 1M tokens |
变更 7 处(旧 → 新)
- Effective August 13, 2026 through December 31, 2026, Google Cloud is injecting a monthly billing credit equal to 50% of net eligible Provisioned Throughput spen→Effective August 13, 2026 through December 31, 2026, Google Cloud is injecting a monthly billing credit equal to 50% of net eligible Provisioned Throughput spen
- Gemini 3.7 Flash, Gemini 3.6 Flash, and CodeMender using these models are offered with introductory pricing of $0.75 / $3.75 per 1M tokens input / output throug→Gemini 3.5 Live Translate
- Global | $3.00 | $3.00 |→Global | $0.375 | $0.375 | $0.0375 | $0.0375
- Global | - |→Global | $1.35 | $1.35 | $0.135 | $0.135
- Image output*** | - |→Image output*** | $30.00 |
- Input (text, image)*** | - |→Input (text, image)*** | $0.30 |
- Issued credits apply to usage across On-Demand Gemini 3.6 Flash and 3.7 Flash SKUs, as well as Provisioned Throughput SKUs across all models. Credits are issued→Issued credits apply to usage across On-Demand Gemini 3.8 Flash, Gemini 3.7 Flash and Gemini 3.6 Flash SKUs, as well as Provisioned Throughput SKUs across all m
新增 22 行
- $2.00 / 1M tokens |
- * Billing is based on total input and output audio token consumption, calculated at a rate of 25 tokens per second of audio.
- 50% Promotional Credits on Gemini 3.8 Flash, Gemini 3.7 Flash and Gemini 3.6 Flash Provisioned Throughput Usage
- Audio Input | $3.50 / 1,000,000 count
- Audio Output | $21.00 / 1,000,000 count
- Claude Fable 5.1 |
- Gemini 3.5 Live Translate |
- Gemini 3.8 Flash
- Gemini 3.8 Flash*
- Gemini 3.8 Flash, Gemini 3.7 Flash, Gemini 3.6 Flash, and CodeMender using these models are offered with introductory pricing of $0.75 / $3.75 per 1M tokens input / output through December 31, 2026. Starting January 1, 2
- Global | $1.875 | $1.875 |
- Global | $6.75 | $6.75 |
- Low-latency,
real-timespeech to speech translation model that supports 70+ languages. | - Non-global | $0.4125 | $0.4125 | $0.04125 | $0.04125
- Non-global | $0.825 | $0.825 | $0.0825 | $0.0825
删除 3 行
- $0.000001 / 1 hour |
- $2.50 / 1M tokens |
- 50% Promotional Credits on Gemini 3.6 Flash & Gemini 3.7 Flash Provisioned Throughput Usage
展示前 15 条新增 / 8 条删除,完整数据见仓库 data/diff/ 目录
By Nick Galloway • 3-minute read
新增 2 行
- By Nick Galloway • 3-minute read
- Getting started with Mantis, our
open-sourcebugfinding-and-fixingharness
删除 3 行
- By Sergio Villani • 19-minute read
- Developers & Practitioners
- 10 questions every startup should answer before moving to production with their AI prototype
$4,300Pricing | Secure and Scalable Enterprise AI | Cohere
新增 3 行
- $4,300Pricing | Secure and Scalable Enterprise AI | Cohere
- Retrieval optimization model
- Semantic representation model
删除 3 行
- Pricing | Secure and Scalable Enterprise AI | Cohere
- Search and discovery model
- Semantic search ranking
Automation’s Early Footprint
新增 6 行
- Automation’s Early Footprint
- Enterprise AIFor Business
- How small AI models can make a big impact for enterprises
- Retrieval optimization model
- Semantic representation model
- Where AI Agents Are (and Aren’t) Being Built
删除 6 行
- A high-throughput vision parsing model with the strongest price-performance profile on the market.
- Cohere triples UK footprint with new London office to support R&D growth
- LLM Serving Fairness: No more noisy neighbors
- Product LaunchCompany NewsEnterprise AI
- Search and discovery model
- Semantic search ranking
Imagine Image Quality Retirement on Nov 2
新增 1 行
- Imagine Image Quality Retirement on Nov 2
Imagine Image Quality Retirement on Nov 2
新增 1 行
- Imagine Image Quality Retirement on Nov 2
"base_model:finetune:runwayml/stable-diffusion-v1-5",
runwayml/stable-diffusion-v1-5",新增 3 行
- "base_model:finetune:
runwayml/stable-diffusion-v1-5", - "base_model:
runwayml/stable-diffusion-v1-5", - "license:
creativeml-openrail-m",
删除 3 行
- "arxiv:2309.00071",
- "base_model:Qwen/Qwen3-8B-Base",
- "base_model:finetune:Qwen/Qwen3-8B-Base",
All reasoning and effort variants of each selected release · Weighted average cost (USD) per Artificial Analysis Intelligence Index task
新增 19 行
- All reasoning and effort variants of each selected release · Weighted average cost (USD) per Artificial Analysis Intelligence Index task
- AnthropicMetaOpenAISpaceXAIKimiZ AIGoogleAlibabaDeepSeekMBZUAI Institute of Foundation ModelsMiniMaxThinking MachinesMistralCohere
- AnthropicMetaOpenAISpaceXAIKimiZ AIGoogleDeepSeekNVIDIAOther Models
- Arena Elo: average Elo rating of the model · Higher is better
- Artificial Analysis Finance & Accounting Index
- Benchmarking GPT-6 Astra
- Claude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback)See more
- Finance & AccountingStrategy & OpsLegalHealthcare & MedicalEngineeringEconomics
- GPT-6 Astra (Non-reasoning)
- GPT-6 Astra (high)
- GPT-6 Astra (low)
- GPT-6 Astra (max)
- GPT-6 Astra (medium)
- GPT-6 Astra (xhigh)
- Gemini 3.8 Flash (high)
展示前 15 条新增 / 8 条删除,完整数据见仓库 data/diff/ 目录
Bytedancedreamina-seedance-2.5-720p: 1478±10
seedance-2.5-720p: 1478±10变更 2 处(旧 → 新)
- 4Bytedancedreamina-seedance-2.0-720p1477±8→5Bytedancedreamina-seedance-2.0-720p1477±8
- claude-opus-5-high1661+8/-8→claude-opus-5-high1661+7/-7
新增 17 行
- 4Bytedancedreamina-
seedance-2.5-720p1478±10 claude-fable-5.1-max|claude-fable-5.1-max1504±11claude-fable-5.1-max1765+23/-23claude-fable-51507±5claude-opus-5-high1493±5flux-3-video-202608111449±6gemini-3.8-flash-high|gemini-3.8-flash-high1494gemini-omni-flash1462±6grok-imagine-video-1.5-720p1456±5hy4-preview1626minimax-h31497±6muse-spark-1.11492±5- qwen3.8-
max-09021688
删除 7 行
- 3Bytedancedreamina-seedance-2.5-720p1483±12
- claude-fable-51508±5
- claude-opus-5-high1492±5
- flux-3-video-202608111447±6
- gemini-3.7-flash-high1491
- gemini-omni-flash1463±6
- gpt-5.6-sol-xhigh (codex-harness)1616+7/-7
展示前 15 条新增 / 8 条删除,完整数据见仓库 data/diff/ 目录
◌ 1 个信源仅有常规波动Anthropic
- Anthropic · 博客 — 数据抖动 2 行,已过滤
⚠ 6 个信源抓取失败
- 火山方舟 · 模型列表 — 提取后内容过短 (raw_len=11071)
- 腾讯混元/TokenHub · 模型列表 — 提取后内容过短 (raw_len=29182)
- 腾讯混元/TokenHub · 定价 — 提取后内容过短 (raw_len=29182)
- 腾讯混元/TokenHub · 更新日志 — 提取后内容过短 (raw_len=29182)
- 腾讯混元/TokenHub · 博客 — 提取后内容过短 (raw_len=29182)
- OpenAI · 博客 — 提取后内容过短 (raw_len=9951)