DeepSeek-V4-Flash Official API Launches Public Beta
要闻
DeepSeek announced the public beta launch of the DeepSeek-V4-Flash official API. Currently, only the API side has been updated; the App and web models remain unchanged, and the model weights have been released on relevant platforms. Compared to the preview version, the model has the same architecture and parameter scale, only re-trained, with significantly improved Agent capabilities, greatly surpassing V4-Pro-Preview on multiple Agent benchmarks. Additionally, the model natively supports the Responses API and is compatible with Codex. Furthermore, the official stated that DeepSeek-V4-Pro is (original text truncated).
DeepSeek announced that the production API of DeepSeek-V4-Flash is now available for public beta testing. The production model ID is DeepSeek-V4-Flash-0731, with model architecture and size consistent with DeepSeek-V4-Flash-preview, only undergoing retraining. According to official data, the production version significantly outperforms V4-Pro-Preview across multiple Agent benchmark tests. For Code Agent tasks in public benchmark suites, the production version was evaluated using the upcoming DeepSeek Harness minimal mode as the framework. Additionally, the model supports the Responses API format and is specifically adapted for Codex. Official documentation also indicates that the DeepSeek-V4-Pro API is expected to support Codex integration in early August 2026. This update only upgrades the DeepSeek-V4-Flash API interface; the app and web versions of the model remain unchanged, and the model weights have been open-sourced on relevant platforms.
ByteDance Seed Officially Releases Seedance 2.5, Single Video Generation Duration Increased to 30 Seconds
要闻
ByteDance Seed officially released Seedance 2.5. The model continues the previous generation architecture, focusing on breakthroughs in long-form narrative ability, multimodal reference ability, and editing ability, with single video generation duration increased from 15 seconds to 30 seconds. The model is being rolled out on Jimeng AI and Doubao Professional Edition, and API services will be launched on Volcano Ark soon.
ByteDance Seed officially released the next-generation video generation model Seedance 2.5, continuing Seedance 2.0's multimodal audio-video joint generation architecture, with a focus on breakthroughs in long-form narrative capabilities, multimodal reference, and editing capabilities. Seedance 2.5 extends single video generation duration from 15 seconds to 30 seconds, and supports multi-round extension, allowing users to input up to 30 images, 10 videos, and 10 audio clips as reference material. The model supports precise editing of audio-video content via timestamps, and enhances multiple editing capabilities such as green screen editing and perspective editing. Seedance 2.5 is being rolled out on Jimeng AI and Doubao Pro platforms, with enterprise-facing API services launching soon on Volcano Ark.
ChatGPT Desktop App Supports Clicking Pet for Quick Launch of Voice Mode
开发生态
ChatGPT desktop app announced a new virtual pet shortcut feature. Users can quickly launch ChatGPT Voice mode by simply clicking the virtual pet on the desktop.
ChatGPT officially announced that its desktop application has added a new feature for quick actions via virtual pets. Currently, users in the desktop app can simply click on the pet to quickly open ChatGPT Voice. Through this shortcut, users can check work progress and approve or stop related tasks. The official statement says this update improves the desktop user experience.
Google AI Studio Cancels Standalone Mobile App, Integrates into Gemini App
开发生态
Google AI Studio canceled its standalone mobile app, which had received over 800,000 reservations. The company decided to integrate the capabilities into the Gemini app, allowing them to emerge naturally in everyday conversations. At the same time, the company will continue to invest in optimizing the web version to meet users' advanced development needs.
Google AI Studio has canceled its plan to release a standalone mobile app. Previously, the app had received over 800,000 pre-orders on iOS and Android platforms. The official decision is to adopt a new approach by integrating the AI Studio experience directly into the mobile and desktop versions of the Gemini app, allowing the app to naturally surface in everyday conversations. Meanwhile, Google AI Studio will continue to invest in optimizing its web version to serve users with deeper development needs.
GPT-5.4 Series Will Be Removed from Codex on August 31
开发生态
OpenAI announced that GPT-5.4 and GPT-5.4 mini will be disabled for Codex users logged in via ChatGPT on August 31, but they will still be available in the API.
OpenAI announced that GPT-5.4 and GPT-5.4 mini will officially be retired from Codex on August 31, and will no longer be available to users who log in with their ChatGPT accounts. These two models will remain available in the OpenAI API and in Codex sessions verified via API keys. The official recommendation is for developers to replace existing gpt-5.4 with gpt-5.6-terra, and gpt-5.4-mini with gpt-5.6-luna, and update all related default settings and custom tasks before this deadline.
Nous Research Releases First Official Hermes Desktop Plugin Kanban and Companion Plugin SDK
开发生态
Nous Research released its first official Hermes Desktop plugin, Kanban, along with the Desktop Plugin SDK. Developers can load custom plugins without build steps to add extension capabilities to desktop applications.
Nous Research officially announced that the first Hermes Desktop plugin, Kanban, has been launched. The desktop native version was released in response to numerous user requests and includes significant upgrades. The accompanying Desktop Plugin SDK module, @hermes/plugin-sdk, allows developers to add independent pages, sidebar navigation, shortcuts, status bar actions, themes, and backend interfaces to Hermes Desktop native applications.
Qwen Releases Qwen-Audio-3.0-ASR-Flash Speech Recognition Large Model
模型发布
Qwen released the speech recognition large model Qwen-Audio-3.0-ASR-Flash, which has been launched on Alibaba Cloud Bailian platform. The company stated that the model has been upgraded in five aspects including contextual memory, industry-specific word recognition, and speech polishing, supporting 30 languages.
Qwen officially released the speech recognition large model Qwen-Audio-3.0-ASR-Flash, now available for invocation through the Alibaba Cloud Bailian platform. It provides three versions: Flash, Filetrans (offline file transcription), and Streaming (real-time speech recognition). The Flash version supports transcription of audio up to 5 minutes long. The new model has undergone systematic upgrades across five dimensions: contextual consistency, industry-specific word recognition, hotword customization, speech polishing, and multilingual recognition. According to official internal evaluations, the recall rate for professional terms in medical scenarios reached 95.36%, and 93.24% in industrial scenarios. Hotword recall exceeded 99% in most scenarios, and the average semantic error rate across seven languages was 17.09%. The Streaming version has a theoretical character output latency of 300 milliseconds, with a character error rate of 7.8% in Chinese industrial scenarios.
Meituan Releases LongCat-Flash-Lite-Sparse, Adopting New Sparse Attention Framework
模型发布
Meituan open-sourced the LongCat-Flash-Lite-Sparse model. It is an MoE model with 69B total parameters and approximately 3B active parameters, introducing LongCat Sparse Attention, natively supporting up to 1M tokens of context length.
Meituan released the LongCat-Flash-Lite-Sparse model on Hugging Face. This model adopts a non-thinking MoE architecture with a total parameter count of 69B and approximately 3B activated parameters per token. It introduces LongCat Sparse Attention (LSA), natively supporting a context length of up to 1M tokens. Compared to its predecessor dense model, LongCat-Flash-Lite-Sparse maintains inference and general knowledge performance while improving long-context inference efficiency and agentic capabilities. The model weights are open-sourced under the MIT license and can be deployed on a single node using SGLang.
SenseTime Open-Sources SenseNova-U1.5-8B-MoT Preview Version
模型发布
SenseNova open-sourced the SenseNova-U1.5-8B-MoT preview version, supporting native 4K image generation and fine-grained image editing, while also releasing pretraining scripts and configuration files. The company stated that a more complete official version will be released soon.
SenseNova released and open-sourced SenseNova-U1.5-8B-MoT (preview version), a natively unified multimodal model built on the NEO-unify architecture, incorporating a new patch encoding layer and ConvDecoder decoding layer. Compared to the previous U1, this preview version supports native 4K image generation, with significant improvements in texture, materials, lighting, and realism. It enhances Chinese and English text rendering and dense, complex layout organization. In image editing, it shows improvements in instruction following, subject identity preservation, and structural consistency, and supports region-controllable fine-grained editing via masks, bounding boxes, and visual markers. The official announcement stated that a more powerful and refined official version will be released soon.
LG AI Research Releases Open-Source Model K-EXAONE 2.0
模型发布
LG AI Research released the open-source model K-EXAONE 2.0. The model has 750 billion total parameters and 37 billion active parameters. According to official data, it performs outstandingly on benchmarks such as long-context understanding and safety, and supports ten languages.
LG AI Research officially released the K-EXAONE 2.0 model on Hugging Face, the second-generation product developed under Korea's Ministry of Science and ICT's sovereign AI foundation model project. The new model adopts a hybrid attention MoE architecture, with 750 billion total parameters and 37 billion activated parameters, making it more than three times the scale of the first generation. The model supports two speculative decoding methods, multi-token prediction and DSpark, which can accelerate generation by approximately 3 to 5 times, and expands its multilingual coverage from six to ten languages. The model is fully open-sourced under the Apache 2.0 license. According to official benchmark data, its performance in long-context retrieval and safety surpasses several mainstream open-source models currently available.
Huawei Open-Sources openPangu-2.0-Pro Model with 505 Billion Total Parameters
模型发布
Huawei open-sourced the openPangu-2.0-Pro model trained on Ascend NPU. The model has approximately 505 billion total parameters, and currently provides model weights, basic inference code, and technical report. The company noted that it still lags behind top models on complex software engineering tasks.
Huawei has open-sourced openPangu-2.0-Pro, a large-scale Mixture-of-Experts model trained on Ascend NPUs, with the model weights, inference code, and technical report now officially available. The model has approximately 505 billion total parameters and about 18 billion active parameters, supporting a 512K context length. It employs a DSA+SWA independently layered hybrid architecture and the Muon optimizer. The technical report discloses evaluation data for both Thinking and Non-Thinking modes. Officials state that the model performs excellently in general capabilities and logical reasoning, but also acknowledge that it still has a significant gap compared to top-tier models when handling complex real-world software engineering tasks.
MiniMax Launches MiniMax Hub and Introduces Media Plan Subscription
产品应用
MiniMax officially launched MiniMax Hub and introduced a new Media Plan subscription. It mainly covers core creative scenarios such as video generation, image creation, soundtrack production, and speech synthesis, supporting mainstream models like Seedance 2.0, MiniMax H3, and Seedream 5.0 Pro.
MiniMax has officially launched MiniMax Hub and introduced a new Media Plan subscription for audio/video creators, film and TV content production teams, and content marketing teams. According to MiniMax's official platform information, the MiniMax API open platform currently processes over one trillion token requests daily and has served over 100,000 developers. The new Media Plan covers core creative scenarios such as video generation, image creation, music composition, and speech synthesis, offering annual subscription tiers of Starter, Plus, and Pro, supporting mainstream models such as Seedance 2.0, MiniMax H3, and Seedream 5.0 Pro.
Google Earth AI Image Generation Feature Withdrawn One Day After Launch
行业动态
Google rolled back Google Earth's AI image generation feature just one day after launch. After discovering that some users shared seemingly policy-violating generated screenshots, Google decided to temporarily withdraw the feature while strengthening protective measures.
Google previously launched an AI image generation feature in Google Earth that allows users to use its Nano Banana 2 image generation model to create fictional images on satellite maps via prompts and overlay them onto real maps. The feature quickly drew criticism after launch, with commentators pointing out that the tool could be used to create and spread false geographic information. Google subsequently issued an official statement saying that it had observed users sharing screenshots of generated images that potentially violated platform policies, and therefore decided to roll back the feature while implementing stronger guardrails. Google also noted that the generated images would not be visible to other users on the Google Earth main interface and are marked with AI-generated watermarks.
Tesla Pushes Vehicle Software Update, Introduces Doubao Large Model Voice Assistant
行业动态
According to reports, Tesla China pushed a new version of vehicle software update to models such as Model 3, with the core addition of the Doubao large model smart voice assistant.
Tesla China has released vehicle software update version 2026.14.13, now being pushed in batches to Model 3, Model Y, Model S, and Model X models. This is the first time Tesla has introduced a third-party large model service into its in-vehicle infotainment system in China. The Doubao large-model-powered intelligent voice assistant enables natural and fluent conversation and real-time information retrieval, offering multiple voice options and personalized roles such as "Know-It-All," "Music Lover," and "Storyteller." However, it requires an advanced in-vehicle entertainment service subscription. According to the official update notes, the voice assistant does not yet support vehicle control commands, and cannot adjust the air conditioning, set navigation, or perform other in-car operations via voice. According to feedback, this update applies to vehicles with HW 3.0 or earlier hardware, i.e., non-AI4 hardware models.
Mathematician Claims GPT-5.6 Sol Found Counterexample to Maxwell's Conjecture
技术与洞察
Mathematician Philip Arathoon stated that GPT-5.6 Sol found a counterexample to Maxwell's conjecture - a triangular bipyramid of 5 point charges producing 24 equilibrium points, exceeding the conjecture's upper bound.
Philip Arathoon announced on social media that Maxwell's conjecture has been disproven. The counterexample was discovered by AI (GPT-5.6 Sol) and organized and written into a paper by human mathematicians, published on arXiv. The counterexample is a triangular bipyramid structure composed of 5 point charges, producing 24 equilibrium points, exceeding the (n-1)^2 upper bound given by the conjecture. Greg Brockman retweeted this news, stating that GPT-5.6 Sol can solve mathematical conjectures that have stood for over a century. Other users pointed out that Maxwell's original text actually stated (n-1)(n-2) rather than (n-1)^2, and that Maxwell himself never referred to this statement as a "conjecture." The name "conjecture" was only proposed by later researchers in 2007. The finding has not yet been independently verified.
Note: This content is AI-assisted and may contain hallucinations and errors.