DeepSeek API clarifies that holidays and adjusted working weekends will be billed at off-peak rates all day
开发生态
For DeepSeek’s official API, weekends designated as workdays due to schedule adjustments and Chinese statutory holidays are billed as off-peak hours all day. The platform’s announcement banner also says that reclassifying a weekend as a workday does not change this designation.
The DeepSeek API platform clarified in an announcement banner that its official API bills both weekends designated as workdays under schedule adjustments and Chinese statutory holidays as off-peak hours for the entire day. A weekend remains subject to the off-peak classification even when a schedule adjustment makes it a workday. The notice covers two types of dates: weekends requiring work because of schedule adjustments and Chinese statutory holidays. Both are covered for the full day, and the rule applies to users of DeepSeek’s official API.
After a user demanded a banked reset because OpenAI had not released an expected update, Codex lead Tibo confirmed that previously announced content was still coming on Tuesday but did not explicitly promise a reset.
Codex lead Tibo responded after a user posted a demand for OpenAI to provide a banked reset, saying that previously announced content was still coming on Tuesday without directly confirming the reset itself. The request followed OpenAI’s failure to release an expected update during what the user described as this week. Tibo’s exact words were: “OK fine. But it’s also still coming in Tuesday.” Some users interpreted “OK fine” as agreement to provide a banked reset, while another reading holds that it was not a clear commitment and confirmed only that the previously mentioned content remained scheduled for Tuesday. Because the reply did not include the words “banked reset,” the original wording does not establish whether a reset will be provided.
Step 5 Preview appears on the Artificial Analysis website
模型发布
Step 5 Preview was released in September 2026 and scored 44 on the Artificial Analysis Intelligence Index. It supports text and image input and has a 1 million-token context window.
According to the Artificial Analysis page, Step 5 Preview was released in September 2026 and scored 44 on the Artificial Analysis Intelligence Index v4.3.2, compared with a median of 24 for models in the same class. It generated 160M tokens during the evaluation, versus a class median of 92M, and the full evaluation cost $922.84. Pricing is $1.00 per 1 million input tokens and $2.70 per 1 million output tokens, compared with respective medians of $1.88 and $10.00. Its measured output speed was 100 tokens/s, while the class median was 70 tokens/s.
The model accepts text and image input, produces text output, and has a 1M-token context window. The page presents the reasoning version and says a non-reasoning variant may also exist. All metrics are compared with models in the same class. Speed data comes from the first-party API or, when no first-party API is available, the median across providers. Intelligence Index v4.3.2 consists of AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1.
Qwen has released Qwen3.8-LiveTranslate, which provides real-time translation for conversations in 60 languages; according to its data, average per-character latency fell from 2.8 seconds in the previous generation to 2.3 seconds. The model is now available for online testing, and its API is live on the Qwen AI platform.
Qwen3.8-LiveTranslate has now been officially released as Qwen’s simultaneous interpretation model for real-world conversations, with support for 60 languages. Online testing is already available, and the API is also live on the Qwen AI platform. According to Qwen, the model uses an Interleave architecture that arranges audio and text into a single interleaved stream. It reuses previously received audio and generated translations through caching, reducing average per-character latency from 2.8 seconds in the previous generation to 2.3 seconds. New capabilities include real-time speaker diarization, same-frame output of source and translated text, and disambiguation using long context. In evaluation results published by Qwen, the model outperformed both its predecessor and current mainstream systems in tests covering long, multi-speaker audio and multilingual real-time simultaneous interpretation.
阿里巴巴达摩院 open-sources RADAR, an abdominal CT diagnostic model, with the paper published in Science
模型发布
Alibaba DAMO Academy has open-sourced RADAR, a vision-language model for abdominal CT, with the related paper published in Science. The project says the model was trained on more than 400,000 contrast-enhanced abdominal CT examinations and 15 million anatomy-aware image-text pairs without manual annotation.
Alibaba DAMO Academy has open-sourced RADAR, a general-purpose vision-language model for abdominal CT, on GitHub, with the related paper published in Science and the project code released under the Apache License 2.0. According to the project description, RADAR learns directly from clinical reports and was trained on more than 400,000 contrast-enhanced abdominal CT examinations and 15 million anatomy-aware image-text pairs without manual annotation. The organization states that the model covers both routine and complex clinical tasks in radiology AI and achieved expert-level performance in both categories. The repository provides a conda environment, dependencies, and detailed guides, while the code can also be archived on Zenodo. Some portions are derived from third-party open-source projects governed by their own licenses, whose original texts are retained in THIRD_PARTY_LICENSES.md.
Cua open-sources cua-s1-forms, a compact form-filling model
模型发布
Cua has open-sourced cua-s1-forms for GUI form filling, using one forward pass to evaluate candidate actions for UI elements in parallel. In the official same-task comparison, the model scored 99.7% overall, versus 83.6% for the hosted Jev API.
Cua has open-sourced cua-s1-forms as the GUI form-filling decision layer behind cua-driver. The model does not generate text; it processes one UI element and its typed candidates in a single forward pass and returns a probability for each option. Context is truncated to 224 bytes and each option to 96 bytes, with candidates comprising a fill action for each document entity plus the fixed actions check, click, and skip. All actionable elements in a form are scored independently and in parallel within one batch. A downstream executor selects the argmax, orders the actions as fills, checkboxes, and one submit click, then passes them to cua-driver as set_value or click operations. According to the official results, cua-s1-forms scored 99.7% overall on the same task, compared with 83.6% for the hosted Jev API, jev-latest, with zero fine-tuning. Jev scored 96% on fill, check, and click decisions and 74% on recognizing an already filled field as a no-op. Cua states that the latter convention was included in training for cua-s1-forms but not for hosted Jev. The project repository also provides the full write-up, training code, synthetic data generator, Cua Driver integration, and the snapshot → score → order → execute implementation.
China Telecom has open-sourced Xing4.0-29B-A4B, a 29B-parameter model that activates 4B parameters per token and supports a native context length of 256K, extendable to 512K. According to the company, it was trained entirely with Ascend NPU and MindSpore.
China Telecom Artificial Intelligence Technology Co., Ltd. has open-sourced Xing4.0-29B-A4B. The model has 29B total parameters, activates 4B per token, supports a native context length of 256K, and can be extended to 512K. According to the company, it is the first model at this scale to be trained entirely on the Ascend NPU platform using MindSpore. The project reported Coding Agent scores of 75.0 on SWE-bench Verified, 66.0 on SWE-bench Multilingual, and 57.5 on Terminal-Bench 2.1. Its reported General Agent scores were 76.55 on Claw-Eval, 60.80 on DeepresearchBII, and 64.63 on Tau3-Bench.
The project has submitted PRs adding Xing4.0 support to vLLM, SGLang, and KTransformers, but they remain under review and have not been merged into their respective main branches; until then, installation from the corresponding PR branches is required. Once launched, vLLM and SGLang provide an OpenAI-compatible API at http://localhost:8000/v1. The model supports fine-tuning, weight merging, inference, and deployment through LLaMA-Factory. It has also been adapted for Ascend Atlas 800T A3 clusters, enabling distributed training and evaluation with MindSpore + MindFormers. Through FlagOS, the project has integrated mHC Triton-Ascend fused operators into MindFormers and completed single-node, 16-card training validation.
Android Developers has released Android Bench 2.0 to evaluate how AI handles real-world, multi-day Android engineering tasks, adding tests of common Agent-model combinations and 6 new models.
Android Developers has released Android Bench 2.0, and the updated evaluation methodology and model comparison results are now available to view. The benchmark covers real-world, multi-day engineering tasks including building apps from scratch, developing new features, and migrating cross-platform codebases to Android, with continuous scoring based on completion rate, visual fidelity, regressions, and the average cost of each model on each task. Version 2.0 also begins evaluating combinations of commonly used Agents and their corresponding models, adding Gemini 3.8 Flash, Gemini 3.7 Flash, GPT-6 Astra, Fable 5.1, Kimi K3, and Qwen 3.8 Max.
OpenAI and three other companies face an antitrust lawsuit filed by paying users
行业动态
On September 18, 2026, paying users sued Anthropic, OpenAI, SpaceXAI, and Google, alleging that the four companies conspired to slow improvements to AI products, leaving subscribers with slower updates at the same prices, and asking the court to block the arrangement.
On September 18, 2026, paying users sued Anthropic PBC, OpenAI OPCO LLC, SpaceXAI, and Google LLC in the U.S. District Court for the Northern District of California, alleging that the four companies agreed to restrict the pace of AI development in suspected violation of Section 1 of the Sherman Antitrust Act. The case is Buist v. Anthropic PBC, No. 3:26-cv-10693. According to the complaint, competitors jointly determining how quickly their products improve constitutes a restraint on competition, leaving subscribers to pay the same prices while receiving slower product improvements than they would under normal competitive conditions.
According to the complaint, the dispute began when Anthropic CEO Dario Amodei published an essay titled We Must Pace the Frontier and called for industry-wide coordination. SpaceXAI founder Elon Musk then expressed support, OpenAI CEO Sam Altman also endorsed the proposal, and Google DeepMind co-founder Demis Hassabis described it as the right direction. The plaintiffs characterize the arrangement as an output-restricting cartel and seek class certification, an injunction, and a declaratory judgment that the four companies violated federal antitrust law. Trial Lawyers for Justice represents the plaintiffs, and none of the four defendants immediately responded to requests for comment.
Anthropic may launch a new model to compete with Astra in the enterprise market
前瞻与传闻
Reuters, citing three sources, said Anthropic is considering a new model in response to the enterprise attention received by OpenAI GPT-6 Astra, but has not decided whether or when to release it. Leaker leo also said three Claude models are being tested, a claim that has not been officially confirmed.
On September 19, 2026, Reuters reported, citing three sources, that Anthropic was considering releasing a new model in response to the attention OpenAI GPT-6 Astra had received in the enterprise market, but had not decided whether or when to launch it. According to the report, the company was still evaluating the safety of its next model and weighing model investment against profitability amid rising interest rates, intensifying competition, and an anticipated IPO. Separately, leaker leo said new versions of Fable, Opus, and Sonnet were being tested discreetly across different Claude interfaces and accounts. Anthropic has not confirmed the claim, and it does not mean the three models have been made broadly available.
OpenAI team member Tibo said some new material is planned for release next week, with more to follow over the coming months. GPT-6 Sol and GPT-6 Luna were names proposed by a user and have not been officially confirmed.
OpenAI team member Tibo said the team plans to disclose some new material next week and present the remainder over the coming months, but OpenAI has not confirmed the existence or release timing of GPT-6 Sol or GPT-6 Luna. Sam Altman previously said that content originally scheduled for release, and which he was most looking forward to, had been postponed until next week. Tibo also said he was preparing a keynote with Sam Altman and others; because there is a large amount to present, the team needs to explain how the new material fits together. According to Tibo, the existing material could support about three developer events. He then asked the community what else it wanted, and one user proposed prioritizing GPT-6 Sol and GPT-6 Luna. Tibo replied, “What else do you want,” prompting speculation about the two names. OpenAI DevDay is scheduled for September 29 local time, but the source text does not specify the year.
Multiple third-party testers and platforms say Google is covertly testing Gemini 4 Pro under other model names in Arena on coding and image tasks. Google has not confirmed the model, results, pricing, or release timing.
Multiple third-party testers and platforms previously claimed that Google was covertly testing Gemini 4 Pro under other model names in Arena, but Google has not confirmed the claim or announced when it will be publicly available. The testers and platforms said the model performed well in code generation and image tasks. Separately, an image associated with benchmark testing circulated externally and was identified as having been generated using GPT-Image. Google has also not confirmed Gemini 4 Pro’s specific model designation, internal codename, test results, or pricing.
MiniMax-M3.1 has appeared in multiple parts of MiniMax’s open-source MiniMax Code CLI, with test configurations reaching 1,000,000 context and 128,000 output. The model is not yet built in, and no further official information is available.
The MiniMax-M3.1 name appears in multiple files within MiniMax’s officially open-sourced MiniMax Code CLI, but the model has not yet become an actual built-in model, and the company has disclosed no further information. The relevant code covers historical model field restoration, the managed catalog, model lists, parameter parsing, and queue persistence. Test configurations list three context values—450,000, 512,000, and 1,000,000—and three output values—8,192, 16,000, and 128,000—along with multiple thinking effort levels. A separate screenshot circulating in the community shows an official staff member previewing the model’s upcoming release, but no additional official release details are currently available.
Kimi’s verified institutional account on Zhihu posted digits of pi beginning with “4159,” omitting the initial “3.1,” and later deleted the post. Observers interpreted it as a possible hint at Kimi K3.1, but the post neither named the version nor provided a release date.
Kimi’s verified institutional account on Zhihu prompted speculation that Kimi K3.1 would be released by posting digits of pi without the initial “3.1,” but the post did not confirm the product and has since been deleted. The sequence began with “4159,” and the omitted portion corresponded to the “K3.1” version number, leading it to be interpreted as a possible release hint. The post itself did not directly mention Kimi K3.1 or specify a release date, product capabilities, availability, or usage method. The speculation has not received further confirmation.