Alibaba's Qwen teases today's open-source release of Qwen3.8-Flash-Next
要闻
Alibaba’s Qwen team said Qwen3.8-Flash-Next is scheduled to be open-sourced at 23:00 Beijing time on August 26. The multimodal MoE model is based on the next-generation Qwen4 architecture, and its weights have not yet been released.
Alibaba’s Qwen team expects to open-source Qwen3.8-Flash-Next at 23:00 Beijing time on August 26 and has posted previews on the ModelScope community and Hugging Face. The release is still in the countdown stage, and the model weights are not yet available. According to the team, Qwen3.8-Flash-Next is a multimodal MoE model built on the next-generation Qwen4 architecture. The team said it is disclosing these architectural improvements early to help the community prepare for the full Qwen4 model family that will follow. No further details have been released.
OpenAI publishes real-world benchmark data for the Jalapeño inference chip
要闻
OpenAI published the first measured results for Jalapeño, its first custom inference chip. Across three models in InferenceX tests, peak performance per watt was 1.5 to 1.9 times that of comparison systems, while end-to-end latency was reduced by factors of 1.7 to 3.6.
OpenAI has published the first measured performance data for Jalapeño and plans to deploy its first custom in-house inference chip in its computing infrastructure by the end of this year. The tests used the InferenceX benchmark and covered GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. According to the company, Jalapeño delivers higher throughput, lower latency, and more AI compute per watt with a single architecture; its peak performance per watt reached 1.5 to 1.9 times that of the comparison systems, while end-to-end latency was reduced by factors of 1.7 to 3.6. OpenAI said its second-generation chip is in advanced development and its third-generation chip has begun to take shape. The company will also continue deploying accelerators from NVIDIA and other partners.
Codex resets paid users' usage allowances and restores the five-hour limit for Plus and Business subscriptions
要闻
OpenAI adjusted Codex’s paid usage allowances, resetting them for Plus, Team, and Pro users and restoring the 5-hour usage limit for Plus; the community reported that Team is subject to the same restriction. The desktop app has also started natively displaying configured third-party models.
OpenAI has reset usage allowances for Codex Plus, Team, and Pro users and restored the 5-hour usage limit for Plus subscriptions; the Codex desktop app has also started natively displaying configured third-party models. According to the company, restoring the limit is intended to smooth compute load and prevent users from exhausting their weekly allowances too quickly. The community subsequently reported that Team subscriptions are subject to the same restriction. User testing also found that desktop sessions are no longer isolated by provider, while sessions in the CLI remain isolated by provider.
OpenAI officially launches ChatGPT Business Premium seats at $100 per seat
开发生态
OpenAI has launched Business Premium seats in ChatGPT at $100 per seat. The offering removes the five-hour usage limit and provides five times the usage allowance of a standard seat.
OpenAI announced that ChatGPT Business Premium seats are now available in ChatGPT at a price of $100 per seat. Greg Brockman said the launch followed requests from a large number of users. According to the company, Premium seats remove the five-hour usage limit and provide five times the usage allowance of standard seats. Tibo said the offering is similar to the $100 Pro plan but is designed for teams and small companies.
ChatGPT Work adds the ability to log in to websites on users' behalf
开发生态
ChatGPT Work is rolling out website login on users’ behalf to Plus, Pro, and Business users across the web and mobile apps. According to the official account, ChatGPT does not see usernames or passwords during login.
The official ChatGPT account announced that, starting on the date of the announcement, ChatGPT Work would gradually roll out website login on users’ behalf to Plus, Pro, and Business users on the web and mobile apps. The feature uses its computer and browser to open and log in to websites. According to the official explanation, ChatGPT Work handles the login steps, while ChatGPT does not see the user’s username or password at any point in the process. The official account also stated that tasks users would normally complete by opening a browser and clicking through each step themselves can be handed to ChatGPT Work on a trial basis.
ChatGPT browser extension adds support for Edge, Brave, Opera, and Vivaldi
开发生态
OpenAI’s ChatGPT browser extension now supports Edge, Brave, Opera, and Vivaldi, bringing tab context into the desktop app and Codex while enabling browser-based tasks. Side chat is not yet available on Opera.
OpenAI’s ChatGPT browser extension now supports Microsoft Edge, Brave, Opera, and Vivaldi, and can bring context from the current tab into tasks in the ChatGPT desktop app. Codex can work from the content the user is viewing, while the extension can also drive the browser to handle web actions such as canceling subscriptions and entering data into a CRM. Users can access ChatGPT’s full capabilities through side chat from any tab; the feature is not currently available on Opera, and OpenAI says it will be released soon.
ChatGPT scheduled tasks updated with support for triggers based on changes in apps such as Github
开发生态
ChatGPT has expanded scheduled tasks on the web and mobile versions of ChatGPT Work. Plus and Pro users can trigger responses from changes in Slack, Gmail, and Github, while free users will gradually receive an allowance of up to three tasks.
ChatGPT officially announced that scheduled tasks on the web and mobile versions of ChatGPT Work have been updated. Plus and Pro users can set changes in Slack, Gmail, and Github as triggers for automatic task responses instead of relying only on fixed schedules. The scheduled tasks feature is also rolling out to ChatGPT free users, who can create up to three tasks. Users can now share tasks they regularly use, allowing others to customize and use them for their own needs.
OpenAI is adding WebMCP support to the ChatGPT desktop browser and ChatGPT Sites, enabling them to invoke tools from compatible websites. It is also launching the 10-day WebMCP Challenge with six partners.
OpenAI Developers announced that the built-in browser in the ChatGPT desktop app and ChatGPT Sites are adding WebMCP support. At the same time, OpenAI is launching the 10-day WebMCP Challenge hackathon with Google Chrome, Cloudflare, Shopify, Vercel, Render, and Netlify. WebMCP is an experimental open standard that lets websites expose structured tools directly to an Agent. When users visit a compatible website, ChatGPT or Codex can automatically use the tools exposed by that site to complete tasks.
Cursor revises its billing model and plan allowances
开发生态
Cursor is replacing Auto’s flat pricing with charges based on the model each request is routed to, with most requests set to cost more. It is also raising its proprietary-model plan allowances and permanently increasing the included Grok allowance again after previously doubling it.
Cursor has ended Auto’s flat pricing and increased the available allowances for its proprietary-model plans and Grok models, with the Grok change representing another permanent increase. New Auto charges are calculated according to the model each request is actually routed to; according to the company, most requests will cost more than under the previous flat price. The allowance increase for proprietary-model plans includes Auto. For Grok, Cursor said demand for Grok 4.6 has increased, so it is expanding the included Grok model allowance beyond the previous doubling and making this additional increase permanent.
Command Code officially announced that it has added Alipay as a payment option for its product, confirming support for the payment method. At the same time, multiple community users requested a Chinese-language interface and Chinese documentation.
Command Code officially announced that its product now supports Alipay, adding it as a new payment method. The announcement centers on an update to payment options, with Alipay support presented as the confirmed product change. At the same time, multiple community users raised requests related to Chinese localization, specifically asking for a Chinese-language interface and Chinese documentation. The payment support comes from Command Code’s official announcement, while the Chinese interface and documentation remain requests from community users, placing the two items at different stages in the report.
商汤's SenseNova U1.5 Lite image generation model joins Token Plan
开发生态
SenseTime announced that its SenseNova U1.5 Lite image creation model has joined the SenseNova Token Plan, offering 1,500 free requests every 5 hours during public beta and free generation of unwatermarked high-resolution original images.
SenseTime has officially added its new-generation image creation model, SenseNova U1.5 Lite, to the SenseNova Token Plan during the public beta and opened a free usage quota for the model. The quota is calculated at 1,500 requests every 5 hours; setting the watermark parameter to false in an image generation request provides an unwatermarked high-resolution original image at no charge. According to the company, generation of unwatermarked high-resolution originals will later become a paid feature, but the source does not specify when the public beta will end, the exact date charging will begin, or the price.
OpenCode confirmed through its official account that Grok 4.6 is now live on OpenCode Go and available to subscribers. The usage allowance is calculated in 5-hour periods, with up to 169 requests per period.
OpenCode announced through its official account that Grok 4.6 is now live on OpenCode Go and available to OpenCode Go subscribers. The model is currently available, with access limited to OpenCode Go subscribers. Its usage allowance is calculated in 5-hour units, permitting 169 requests in each 5-hour period. The information disclosed by OpenCode covers the model name, platform, availability, eligible users, and request allowance: Grok 4.6, OpenCode Go, currently available, subscribers, and 169 requests every 5 hours, respectively.
OpenCode says $5 introductory offer was withdrawn due to abuse
开发生态
OpenCode has ended its $5 introductory offer for the first month and changed its current starting price to $10 per month. Official staff member dax said the plan attracted widespread abuse and that OpenCode will seek a new approach less susceptible to abuse.
OpenCode has removed its $5 introductory offer for the first month and changed its current starting price to $10 per month. Community users first noticed the pricing change, after which OpenCode staff member dax confirmed that the first-month offer had been removed from the pricing plan. According to dax, the offer was discontinued because it attracted widespread abuse. OpenCode said it still supports the idea of an introductory offer but needs to find a new approach that is less susceptible to abuse. The currently listed starting price is $10 per month, and the $5 first-month offer is no longer included in the pricing.
Anthropic unifies memory across Claude chats and Cowork
产品应用
Anthropic has unified Memory across Claude chat and Cowork, allowing user context to be shared and carried back between them. It is on by default for Free, Pro, and Max across web, desktop, and mobile, while storage of sensitive topics remains off by default.
Anthropic has merged the Memory used by Claude chat and Cowork into a single system, allowing Cowork to access context accumulated in chat when it runs tasks in the cloud and carry information from those tasks back into chat. According to the company, Claude now writes information to Memory by topic during a conversation instead of generating a summary after the conversation ends, so details such as project deadlines, work preferences, and metric definitions can be used in later conversations and Cowork tasks. Users can pause or reset Memory at any time and view, edit, or delete each short file under Topics in Settings > Memory. After a correction is made in one place, all subsequent conversations use the revised information.
According to Anthropic, personal or sensitive topics are not stored by default, including health, race, ethnicity, religious beliefs, political views, and gender identity. Users can enable “include sensitive topics in memory,” after which Claude displays a notice each time it saves information in these categories. Only information encountered after the setting is enabled is recorded; earlier information is not saved retroactively, and the setting can be disabled at any time. Even when it is enabled, Claude does not store sensitive identification numbers such as SSNs and government ID numbers, criminal history, immigration status, or content that violates the AUP, and it notifies users when Memory cannot be updated. Memory is on by default for Free, Pro, and Max across web, desktop, and mobile. The latest version of the app is required on iOS and Android. For Team and Enterprise, admins control availability for the organization, while individual users must enable Memory themselves.
豆包工作 officially launches, offering a 30-day subscription benefit with the desktop app download
产品应用
Doubao has officially launched Doubao Work, an Agent product for productivity tasks. Users who download the desktop app or update Doubao’s desktop app to the latest version can claim a 30-day subscription benefit, while existing subscribers receive a 30-day extension.
Doubao has officially released the productivity Agent Doubao Work, offering a 30-day subscription benefit from the date of launch to users who download the Doubao Work desktop app or update Doubao’s desktop app to the latest version. Existing Doubao subscribers who claim the benefit will have their subscriptions extended by another 30 days. According to the company, the product can break down tasks based on user goals, invoke tools, and continuously execute complex workflows, including work involving documents, spreadsheets, and PPT files. It can also operate computers and browsers, with the process visible to users and available for them to take over at any time. The product is also deeply integrated with Feishu, and the company says the Agent can use enterprise context to complete tasks more accurately.
The “工作助理” in 千问APP and its PC version now supports linking 阿里云 Token Plan allowances
产品应用
Qwen’s app and PC version now support Alibaba Cloud Token Plan. After linking an API Key, users can run complex Work Assistant tasks—including computer and browser operations and calling a skill to deliver office documents—with consumed Tokens deducted directly from their subscription quota.
Qwen has added Alibaba Cloud Token Plan integration to its app and PC version, allowing users to access their Token Plan subscription quota after linking an API Key; the feature is now available. Once linked, Token consumption generated when Qwen’s Work Assistant handles complex tasks is recorded against and deducted directly from the user’s Token Plan subscription quota. Tasks listed in the source include operating a computer and browser, as well as calling a skill to deliver office documents, all of which can use the quota linked by the user.
ChatGPT has launched Stickers, allowing users to turn photos or ideas into custom stickers with transparent backgrounds through ChatGPT Images and share them on iMessage or WhatsApp.
ChatGPT announced the launch of Stickers, which users can access by selecting “Images” in the sidebar and then clicking “Stickers.” The feature uses ChatGPT Images to turn uploaded photos or written ideas into custom stickers and supports transparent backgrounds for generated results. Finished stickers can be shared on iMessage or WhatsApp. According to the company, the transparent-background capability is not limited to Stickers and can also be applied to other images.
Google Labs launches Play with Putty, a real-time collaborative vibe coding tool
产品应用
Google Labs has announced Play with Putty, an experimental collaborative vibe coding tool, and opened its waitlist. According to the company, it lets multiple people build tools and websites together in real time and is currently limited to applicants aged 18 or older in the United States.
Google Labs has announced the experimental Play with Putty project and opened its waitlist, allowing eligible users to register for participation. Play with Putty is positioned as a collaborative vibe coding tool; according to the company, multiple people can use it to build tools and websites together in real time. The experiment is limited to users aged 18 or older in the United States. Google Labs is also inviting people who join the waitlist to share feedback with the project team.
Perplexity launches Portable Computer, a local Agent app
产品应用
Perplexity has introduced Portable Computer, a local AI Agent initially built for NVIDIA DGX Spark. It is available to Pro and Max users who own the device, can access cloud models with permission, and keeps sensitive documents stored locally.
Perplexity released Portable Computer on August 26, 2026, as a local AI Agent for NVIDIA DGX Spark, with access now open to Pro and Max users who own the device. The application includes PPLX 27B and supports Qwen 3.8 27B. When a task requires stronger reasoning, it can call frontier models in the cloud after receiving user authorization, while sensitive documents remain stored on the device. Perplexity also published a series of test results covering knowledge work, web research, and document parsing, which the company said demonstrate the product’s advantages in privacy, cost, and local execution efficiency. According to the company, support will later expand to additional NVIDIA hardware and models.
Apodex releases the Apodex 1.1 model family and open-sources the 35B mini version
模型发布
Apodex has released the Apodex 1.1 model family and integrated the full model into its online workbench. The 35B mini version has open weights and supports local deployment, while the open-source FrontierAgent offers ReAct and Agent Team modes.
Apodex has released the Apodex 1.1 model family, deployed the full model in its official online workbench, and opened the weights of the 35B Apodex 1.1 mini with support for local deployment. According to the company, the full model delivers frontier Agentic performance on complex tasks including professional work, scientific research, financial analysis, and deep search. The accompanying open-source FrontierAgent is a locally deployable research workbench with ReAct and Agent Team modes, and it can run on macOS and Linux with a single command without Docker being preinstalled. Bare-model API access will be rolled out through major API platforms; the company said work on model weights, the technical report, and developer documentation remains in progress.
IBM releases open-source Granite 4.2 for enterprise Agentic AI
模型发布
IBM has released the open model family Granite 4.2 in 3B, 8B, and 30B parameter sizes for enterprise Agentic AI workflows, with a 512K context window and three reasoning modes.
IBM has released Granite 4.2, an open model family for enterprise Agentic AI workflows, in 3B, 8B, and 30B parameter sizes. According to the official model card, the series supports three modes: full thinking, no thinking, and low effort, with a 512K context window. The language models are licensed under Apache 2.0 and, according to IBM, can be deployed in cloud, on-premises, and edge environments; they are also available on platforms including Hugging Face and Ollama. IBM simultaneously released the 470M-parameter Granite Speech 5.0 Turbo CTC and a non-commercial version of the speech model.
腾讯 open-sources the WeMM-Embedding multimodal embedding model
模型发布
Tencent has open-sourced WeMM-Embedding, developed by the WeChat Vision team, to encode text, images, videos, visual documents, and interleaved multimodal inputs in a unified format. At 256 dimensions, its 2B model retains 98.7% of its full-dimensional image and video performance on MMEB-v2.
Tencent has open-sourced WeMM-Embedding, a family of universal multimodal embedding models developed by the WeChat Vision team. The models support text, images, videos, visual documents, and interleaved multimodal inputs, but do not currently support audio. An embedding is taken from the dedicated <embedding> token position in the final-layer hidden state and then undergoes L2 normalization. The project recommends transformers==5.2.0 for inference and reproducibility, explaining that newer versions may have different preprocessing behavior. SentenceTransformer can load the model directly from a local path or a Hugging Face model id and process text, images, and videos through SentenceTransformer.encode(). The MRL dimension is selected with --dimension; omitting the parameter produces the full-dimensional output.
For a supported dimension d, the project specifies truncating the full embedding and then normalizing it again. According to the project, the 2B model at 256 dimensions retains 98.7% of its full-dimensional image and video performance on MMEB-v2. Table 1 of the technical report covers 78 datasets, using Hit@1 for image and video tasks and NDCG@5 for visual-document tasks. Table 2 covers 190 tasks; V3-All includes 78 MMEB-v2 tasks, 53 text tasks, 47 agent tasks, 11 audio tasks, and MCMR, with unsupported tasks assigned a score of zero. The repository provides MMEB-v3 evaluation code based on the TIGER-AI-Lab/VLM2Vec pipeline with a minimal set of modifications. Tested versions are vLLM 0.27.0 and SGLang 0.5.9. Unless otherwise noted, Tencent-authored code is released under the Apache License 2.0, while third-party components retain their original licenses.
MiniMax releases an index repository for H3 ecosystem integrations
模型发布
MiniMax has released a GitHub integration index for H3 projects. It covers quantization guides for INT8, NVFP4, GGUF, and an 8GB VRAM constraint, along with inference acceleration, Agent skills, and Apple Silicon support.
MiniMax released a GitHub index repository named “Awesome MiniMax H3 Integrations” to collect and track projects built around H3. Its quantization guides cover INT8, NVFP4, and GGUF, including material for an 8GB VRAM constraint; the inference acceleration section lists tools including LoRA acceleration and Sol-Attn, while the developer tooling includes native Agent skills and h3.c with Apple Silicon support. According to the company, the H3 ecosystem is developing rapidly.
Apple equipped Mac Studio with M5 Max or M5 Ultra, offering up to 512GB of unified memory and 4.3 times the peak AI compute performance of M3 Ultra. It goes on sale on September 22, 2026.
Apple announced that the new Mac Studio with M5 Max and the all-new M5 Ultra will go on sale on September 22, 2026. The M5 Max model has an 18-core CPU and an up-to-40-core GPU, with Neural Accelerators built into every GPU core, and supports up to 128GB of unified memory. The M5 Ultra model scales to a 36-core CPU, an up-to-80-core GPU, and up to 512GB of unified memory. According to Apple, M5 Ultra delivers up to 4.3 times the peak AI compute performance of M3 Ultra and 9.8 times that of M1 Ultra, while memory bandwidth reaches 1.2TB/s, 50% higher than the previous generation.
The new Mac Studio adds Wi-Fi 7 and Bluetooth 6 for the first time and includes Thunderbolt 5. Multiple Mac Studio systems can form a cluster using Thunderbolt 5 and RDMA; according to Apple, a four-system Mac Studio cluster delivers up to three times the AI inference performance of a single system. On the software side, the new Core AI framework is designed to build, run, and deploy AI models on Apple silicon, using unified memory, the CPU, GPU, and Neural Engine. It also supports deploying full-scale LLMs locally and integrating custom models into apps. Apple’s open-source MLX framework is used to run, train, and fine-tune models on Mac.