Daily AI Digest

2026-08-25

Source:橘鸦 AI 早报 · 14 items

2026-08-25
2026-08-25 2026-08-24 2026-08-23 2026-08-22 2026-08-21 2026-08-20 2026-08-19 2026-08-18 2026-08-17 2026-08-16 2026-08-15 2026-08-14

MiniMax partners with GMI Cloud to offer 14-day free access to models including M3

开发生态

MiniMax partners with GMI Cloud to offer 14-day free access to models including M3

MiniMax and GMI Cloud announced 14 days of free access from August 24 to September 6, with the year unspecified in the source, covering M3, M2.7, Speech 2.8, and Music 3.0 under anti-abuse controls and rate limits.

MiniMax and GMI Cloud announced that they will provide 14 days of free access to MiniMax models and features through GMI Cloud from August 24 to September 6, with the year unspecified in the source. The offer covers MiniMax M3, M2.7, Speech 2.8, and Music 3.0, and developers must connect using a GMI API key. According to the companies, these services will have no usage-volume cap during the campaign, while anti-abuse controls and rate limits will still apply.

Read original →

阿里云’s 云工开物 offers Qoder CN Professional Edition to university students

开发生态

阿里云’s 云工开物 offers Qoder CN Professional Edition to university students

Alibaba Cloud’s Yun Gong Kai Wu program has begun offering Qoder CN Professional benefits to verified students enrolled at Chinese higher-education institutions, with one RMB 300 no-threshold coupon available per user. The coupon remains valid for 1 year after collection and covers designated public cloud products; the end date will be announced on the official website.

Alibaba Cloud has launched its Yun Gong Kai Wu Qoder CN Professional campaign for university students, allowing verified students enrolled at Chinese higher-education institutions to claim one RMB 300 no-threshold coupon. The campaign has already started, and its end date will be announced on the official website. According to the campaign rules, eligible users include students in junior college, bachelor’s, master’s, doctoral, and part-time postgraduate programs, while multiple linked accounts belonging to the same user are managed as associated accounts. The coupon is valid for 1 year from the collection date and can be used for selected public cloud products listed in the campaign’s coupon section. It may be used repeatedly until the balance reaches 0 when each order does not exceed the remaining balance, but it cannot be combined with product discounts or other coupons, used to obtain postpaid products or as a reserved withdrawal balance, or transferred. Unused balances expire, and usage beyond the coupon amount is charged at standard product rates. alibabacloud.com, jp.aliyun.com, and their subordinate pages are excluded.

Read original →

OpenAI offers U.S. college students four months of free ChatGPT Plus

产品应用

OpenAI offers U.S. college students four months of free ChatGPT Plus

OpenAI is offering four months of free ChatGPT Plus to full-time and part-time students enrolled at eligible degree-granting institutions in the United States. The benefit has an officially stated value of $80, and applicants must complete SheerID verification by October 31, 2026.

OpenAI has launched a ChatGPT Plus offer for U.S. college students, requiring eligible students to complete verification through SheerID by October 31, 2026, to claim four months of free access. According to the company, the offer is worth $80 and applies to full-time and part-time students enrolled at degree-granting institutions in the United States. New users must add a payment method when activating ChatGPT Plus, but they will not be charged during the four-month free period.

Read original →

Anthropic upgrades Claude’s streaming renderer

产品应用

Anthropic upgrades Claude’s streaming renderer

Claude’s web and desktop apps have switched to a streaming renderer rebuilt by Anthropic, making long-response output about 4x smoother. Anthropic says stalls on slower laptops are reduced by 9x and the longest freeze time is 4.5x shorter.

Anthropic rebuilt the streaming renderer used by Claude’s web and desktop apps, improving the smoothness of streaming output during long-response generation by about 4x. Anthropic says the new renderer currently processes only content that is still changing, an approach intended to reduce stalls and freezes while long responses are streamed. According to its explanation, long responses have 9x fewer stalls on slower laptops and the longest individual freeze is 4.5x shorter; on a 120Hz MacBook, output can remain at 120fps from the start of generation through completion.

Read original →

Agnes AI open-sources the Agnes 2.5 Pro Alpha multimodal reasoning model

模型发布

Agnes AI open-sources the Agnes 2.5 Pro Alpha multimodal reasoning model

Agnes AI has released the multimodal Agnes 2.5 Pro Alpha model and is offering it through an API. The model is further trained from Qwen3.5-397B-A17B, and its repository uses the Apache License 2.0.

Agnes AI has released Agnes 2.5 Pro Alpha, with the evaluation snapshot dated August 18, 2026, and made the model available through OpenAI-compatible Chat Completions and Responses APIs. It is a post-trained derivative of Qwen/Qwen3.5-397B-A17B, with additional post-training performed by Agnes AI. Both the repository and the upstream model use the Apache License 2.0, and the original copyright notices are retained. According to the company, the model uses extended reasoning for complex tasks and is intended for repository-level coding, technical research, document synthesis, visual analysis, and tool-enabled Agents.

The model was compared with Qwen3.5-397B, Qwen3.7-Max, GLM-5.2-744B, MiniMax-M3-428B, DeepSeek-V4-Pro-1.6T, Claude Opus 4.7, and Claude Opus 4.8, with all results coming from independent Artificial Analysis benchmark measurements. The Quickstart uses 8-GPU tensor parallelism, and the checkpoint is a multi-shard BF16 package that cannot run on a single GPU. The official instructions call for sampling instead of greedy decoding, sufficient max_tokens or max_output_tokens for extended reasoning, and a higher limit if a response ends early. API keys should be stored in environment variables, while image inputs require publicly accessible URLs. Agnes AI says outputs may contain errors, high-impact decisions and tool actions need validation in the application layer, and the Agnes AI service terms should be reviewed before sending sensitive or regulated data.

Read original →

阿里’s Wan3.0 video model officially launches with a limited-time 30% API discount

模型发布

阿里’s Wan3.0 video model officially launches with a limited-time 30% API discount

Alibaba’s Wan3.0 video generation model has exited public testing and officially launched, with individual videos capped at 30 seconds. Its API costs 0.3 to 1.2 yuan per second, is available through eight platforms, and is temporarily offered at 70% of the standard price on some platforms.

Alibaba has moved its Wan3.0 video generation model from public testing to a full release, and users can now access it through eight platforms, including Alibaba Cloud Bailian, the Wanxiang website, and Qwen. A single generation can produce up to 30 seconds of video, and the model supports doc, xls, ppt, pdf, and md document inputs for the first time. According to the company, it has upgraded capabilities covering generation length, all-purpose reference, real-world reproduction, instruction following, and consistency across shots. API pricing is based on resolution and ranges from 0.3 to 1.2 yuan per second, with some platforms temporarily charging 70% of the standard price. Multiple companies, including 井英科技 and Meitu’s RoboNeo, have integrated Wan3.0.

Read original →

字节 integrates TRAE and 扣子 into the 豆包 ecosystem

行业动态

字节 integrates TRAE and 扣子 into the 豆包 ecosystem

ByteDance has integrated the TRAE and Coze teams into the Doubao organization as part of an office AI product team restructuring, according to media reports. The capabilities will be reorganized into workplace and programming product lines, while the company says existing user rights will remain unaffected.

ByteDance has completed the integration of its office AI product teams, moving the entire TRAE and Coze teams into the Doubao organization, according to media reports. The combined teams will report to the head of the Doubao product. At the product level, TRAE Work and Coze will be integrated with Doubao’s workplace capabilities, while TRAE IDE and CLI will form the programming product line, the reports said. Doubao will also launch a unified AI office product called “Doubao Work” and integrate it closely with Feishu. ByteDance officially stated that the restructuring is intended to coordinate resources and will not affect existing user rights.

Read original →

小米 unveils three self-developed 玄戒 chips and showcases an AI Cube prototype

行业动态

小米 unveils three self-developed 玄戒 chips and showcases an AI Cube prototype

Xiaomi released three self-developed chips—Xring O3, O100, and D100—and displayed an AI Cube prototype on August 24, 2026. Lei Jun said all three had completed post-fabrication silicon validation; O3 will debut in the Xiaomi 18 Fold, while the other two are planned for commercialization in 2027.

At its Xring chip technology briefing on August 24, 2026, Xiaomi formally introduced the Xring O3, O100, and D100 and also displayed an AI Cube prototype. Xiaomi founder, chairman, and CEO Lei Jun said the three chips cover on-device AI computing requirements across the company’s “Human × Car × Home” ecosystem and have all completed post-fabrication silicon validation. The Xring O3 will make its debut in Xiaomi’s top-tier new foldable, the Xiaomi 18 Fold, while the Xring O100 and D100 are scheduled for commercial deployment in 2027. According to his explanation, Xring is positioned as the AI computing foundation for Xiaomi’s entire ecosystem, spanning scenarios from pockets, cockpits, and living rooms to factories. It forms part of Xiaomi’s chip strategy and is also described by the company as the physical foundation of its AI strategy.

Read original →

SpaceXAI deploys NVIDIA Vera CPUs and plans to launch a space-based NVL72

行业动态

SpaceXAI deploys NVIDIA Vera CPUs and plans to launch a space-based NVL72

NVIDIA said SpaceXAI will adopt Vera CPUs for its AI infrastructure and plans to bring an optimized Vera Rubin NVL72 into orbit. The first Starmind satellite is scheduled to reach orbit in the fourth quarter of 2027, with deployment scaling in 2028.

NVIDIA announced that SpaceXAI will deploy NVIDIA Vera CPUs to accelerate Agent AI applications and expand the AI infrastructure supporting Grok. SpaceXAI also plans to deploy a Vera Rubin NVL72 optimized for orbital conditions in space. The first Starmind satellite is expected to launch into orbit in the fourth quarter of 2027, followed by a larger deployment in 2028. According to NVIDIA, the space-based system is designed to accommodate orbital conditions while retaining the same architecture and software ecosystem as the ground-based system.

Read original →

NVIDIA releases Vera Rubin NVL72 performance data

行业动态

NVIDIA releases Vera Rubin NVL72 performance data

NVIDIA released early Vera Rubin NVL72 data showing up to 30x the throughput per megawatt of GB300 NVL72 in AgentX tests, while cost per million tokens can fall to as little as 1/35 that of GB300 NVL72.

NVIDIA released current inference-performance data for Vera Rubin NVL72 on Agent workloads, stating that the system delivers up to 30x the throughput per megawatt of GB300 NVL72 in real Agent coding sessions captured by SemiAnalysis AgentX. The workload preserves actual context growth, tool calls and sub-agent spawning. The DeepSeek V4 Pro results are still pending SemiAnalysis review and do not yet include the performance of the Vera CPU when handling tool calling. According to NVIDIA, Vera Rubin NVL72 can reduce cost per million tokens to as little as 1/35 that of GB300 NVL72. On the same model, GB300 NVL72 delivers up to 15x the throughput per megawatt of the NVIDIA Hopper architecture.

OpenRouter data shows that Agentic AI workloads consume 15x as many tokens as a simple chat request. NVIDIA stated that input and output sequences for ordinary chat or document summarization typically range from 1K to 8K tokens, while context in an Agent session can accumulate across multiple steps to reach hundreds of thousands of input tokens. According to the company, the Rubin GPU’s fifth-generation Tensor Cores, third-generation Transformer Engine and NVFP4 quantization accelerate prefill and decode. Sixth-generation NVLink and NVLink Switches provide 10x the packet rate of off-the-shelf Ethernet alternatives while reducing latency to one-third. NVIDIA also said DSX MaxLPS can provision up to 40% more GPUs within the same megawatt budget. Vera Rubin uses a seven-chip architecture that also includes the Vera CPU, Groq 3 LPU, NVLink 6 Switch, BlueField-4 DPU, Spectrum-6 SPX and ConnectX-9 SuperNIC. The company said the platform is in full production.

Read original →

NVIDIA unveils the fully production-ready Groq 3 LPX to scale Vera Rubin inference

行业动态

NVIDIA unveils the fully production-ready Groq 3 LPX to scale Vera Rubin inference

NVIDIA announced that Groq 3 LPX for Vera Rubin NVL72 has entered full production. According to the company, it delivered 3,400 output tokens per second in a 100,000-token test, four times the result of the nearest alternative platform.

NVIDIA Groq 3 LPX, designed for Agentic AI inference, has entered full production and works alongside Vera Rubin NVL72 as a specialized acceleration extension for token generation. According to NVIDIA, an Artificial Analysis test using the open-source Agent model Gemma 4 31B produced 3,400 output tokens per second in a 100,000-token long-context workload, four times the result of the nearest alternative platform. In this architecture, Rubin GPUs handle large-scale context processing, while LPX handles latency-sensitive decode workloads to support per-token responses during Agent reasoning, tool use and system interactions.

According to NVIDIA, Nebius is the first AI cloud to adopt NVIDIA Groq 3 LPX and will add the product to Nebius Token Factory for interactive Agents, coding systems and other real-time AI applications. CoreWeave has deployed Spectrum-X Multiplane in production, connecting Vera Rubin racks through multiple parallel switches to provide a high-bandwidth, flat and lossless AI network. SpaceXAI said its future AI architecture will be built around Vera Rubin, spanning terrestrial data centers and orbital satellites, and that it plans to deploy Vera CPUs for Agentic AI workloads including orchestration, tool use, code execution, data processing and simulation.

Read original →

Nvidia reportedly considers investing billions of dollars in Perplexity AI

行业动态

Media reported on August 24, 2026, that Nvidia was discussing a multibillion-dollar investment in Perplexity AI’s new equity funding round. The round values the company at more than $30 billion, while its annualized revenue has risen above $750 million.

On August 24, 2026, media reported that Nvidia was in talks to participate in Perplexity AI’s latest equity funding round with a multibillion-dollar investment. The two companies were also discussing a technology licensing agreement. According to the report, the round values Perplexity AI at more than $30 billion. Its annualized revenue, which was below $250 million at the start of 2026, has since increased to more than $750 million; both the valuation and revenue figures come from the media report.

Read original →

Mistral AI and HUMAIN move to develop sovereign AI infrastructure in the Middle East

行业动态

Mistral AI and HUMAIN announced a sovereign AI partnership across Saudi Arabia and the Middle East covering infrastructure, model development, and industry deployment. The companies put the collaboration at hundreds of millions of euros, with an initial focus on Arabic, cybersecurity, and voice.

Mistral AI and HUMAIN announced a strategic collaboration across Saudi Arabia and the Middle East covering AI infrastructure, advanced model development, and AI solution deployment, with the companies stating that the collaboration is worth hundreds of millions of euros. They plan to develop and localize advanced AI models, initially focusing on cybersecurity and voice, while also building frontier models with stronger Arabic-language capabilities. Mistral AI will explore using HUMAIN’s data center infrastructure to meet growing local compute demand. The companies also plan to create a joint go-to-market strategy for Saudi Arabia to provide AI solutions to regulated industries. According to their explanation, sovereign AI under this collaboration allows customers to control data, intelligence, compute, and operations, run training and inference on infrastructure and in jurisdictions of their choice, and adapt and own models through open weights. The announcement also states that it contains forward-looking statements and that actual results may differ due to risks, uncertainties, and subsequent commercial agreements.

Read original →

Thinking Machines Lab launches grants for open-source model safety

行业动态

Thinking Machines Lab launches grants for open-source model safety

Thinking Machines Lab has launched Tinker grants offering up to $50,000 in credits per grant for open-weight safety research. Focus areas include hazardous-data identification, tamper-resistant training, and reward hacking.

Thinking Machines Lab has launched Tinker grants for safety research on open-weight models, offering up to $50,000 in credits per grant. The program follows the release of Inkling and Inkling-Small with open weights and targets safety projects that could be accelerated by additional Tinker credits. The listed research directions include measuring capability gains for attackers and defenders separately after defensive fine-tuning to determine whether any gap reflects genuine asymmetry, suppression that disappears under attacker pressure, or sandbagging that persists only during measurement. Another direction is training classifiers that can identify hazardous data at sufficient scale and recall, then using downstream fine-tuning to measure whether filtering reduces hazardous capability uplift.

The research scope also covers safeguards that persist after adversarial fine-tuning and ordinary continued training, tests of generalization against unseen attack strategies, and analysis of how scale and architecture affect robustness. Other questions concern when narrow harmful fine-tuning expands into broad behavioral change, whether models can combine planning, tool use, and domain knowledge into behavior never demonstrated end to end during training, and how reward hacking changes with model capability and optimization pressure. The program also covers whether early signals can predict problems before they become severe, whether strategies for evading or interfering with chain-of-thought monitors can be detected, and whether those strategies transfer across tasks. According to the company, these directions are not an exhaustive list.

Read original →