Daily AI Digest

2026-08-04

Source:橘鸦 AI 早报 · 16 items

2026-08-04
2026-08-04 2026-08-03 2026-08-02 2026-08-01 2026-07-31 2026-07-30 2026-07-29 2026-07-28 2026-07-27 2026-07-26

Alibaba Qwen Releases Qwen3.8-Max Official Version and Plans to Open-Source Model Weights Next Week

要闻

Alibaba Qwen Releases Qwen3.8-Max Official Version and Plans to Open-Source Model Weights Next Week

Alibaba Qwen has released the official version of the Qwen3.8-Max model. Built on the architecture foundation of Qwen 3.5, the model has 2.4T total parameters and 95B active parameters. It is now available on the official API platform and products such as Token Plan, Qwen Studio, and the Qoder series. According to the company, the model has achieved comprehensive improvements in programming, office work, long-horizon tasks, and multimodal agents. In official tests, the model can independently complete long-horizon automated programming lasting approximately 16 days. The model weights will be made public next week, and the Qwen3.8-27B model will also have its weights released.

Alibaba today released the Qwen3.8-Max model, which has been launched on the Qwen AI platform for API services, along with products such as the Token Plan and the Qoder series. The model is built on the Qwen 3.5 architecture, with a total of 2.4 trillion parameters and 95B activated parameters, supporting a context of 1 million tokens. In terms of pricing, domestic input costs ¥12 per million tokens and output costs ¥36; overseas prices are $2 and $6, respectively.

The official statement says the model achieves comprehensive improvements in coding, office work, long-duration tasks, and multimodal agents. According to official tests, it can independently complete long-horizon automated programming tasks lasting up to approximately 16 days. Based on Arena.ai leaderboard data, the model ranks second and fourth in vision and coding capabilities, respectively.

The official announcement states that the model weights will be made public next week, marking the first Max-scale open-weight model in the Qwen series. The Qwen3.8-27B model's weights will also be released. The model is compatible with OpenAI and Anthropic interfaces and supports integration into mainstream agent frameworks.

Read original →

MiniMax-H3 Weights Officially Released with Native ComfyUI Support

要闻

MiniMax-H3 Weights Officially Released with Native ComfyUI Support

MiniMax has officially open-sourced its full-modal generation system MiniMax-H3. The released model weights include two checkpoints of H3-Base: FL2VA and Ref2VA. These weights have received native support in ComfyUI. Its license prohibits usage in countries and regions such as the United States and the European Union.

MiniMax has officially open-sourced its full-modal generation system MiniMax-H3, releasing model weights on HuggingFace, supporting unified understanding of text, images, audio, and video, with the ability to generate videos at up to 2K resolution, 15 seconds in duration, 24FPS, and 32kHz stereo audio. The open-sourced model weights mainly include two task-specific checkpoints: H3-Base-FL2VA and H3-Base-Ref2VA. The official implementations of H3-Context-IR and H3-Regenerate-2K have not been open-sourced yet; they are currently accessible via the official API. For H3-Context-IR, guidance for building an alternative processing pipeline is also provided. ComfyUI has native support from the day the weights were released. The license prohibits use in the United States, the European Union, the United Kingdom, and South Korea. MiniMax employee Ryan Lee stated that the US restriction stems from a copyright lawsuit with a Hollywood production company over generative video, and US users can submit a formal authorization request.

Read original →

Qoder Anniversary: Qwen3.8-Max Launches Across All Platforms, Free Calls for New and Existing Users

开发生态

Qoder Anniversary: Qwen3.8-Max Launches Across All Platforms, Free Calls for New and Existing Users

Qoder announced that Qwen3.8-Max is now officially available across all Qoder platforms. From August 3 to September 3, new and paid users can claim 800 free calls, with an additional 2,000 calls stackable on orders. Additionally, off-peak hours are billed at a 50% discount for both individual and enterprise versions. The China version has also introduced an invitation benefits campaign.

Qoder announced that, to celebrate its first anniversary and the official launch of Qwen3.8-Max across all platforms of Qoder International and China editions, from August 3 to September 3, new registered users and existing paid users can receive 800 free invocations of Qwen3.8-Max. During the campaign period, placing an order can add another 2,000 invocations. The invocations are valid until September 30. Additionally, during off-peak hours from 22:00 to 08:00 the next day (UTC+8), billing is at a 50% discount, applicable to both personal and enterprise editions. The China edition also introduces an activity to earn points by inviting friends.

Read original →

Alibaba Cloud Bailian Token Plan Launches Official Qwen3.8-Max, Resets Weekly Limit

开发生态

Alibaba Cloud Bailian Token Plan Launches Official Qwen3.8-Max, Resets Weekly Limit

Alibaba Cloud Bailian Token Plan announced the launch of the official Qwen3.8-Max version. The previous Qwen3.8-Max-Preview will be taken offline on August 5, and users with active subscriptions will receive a one-time 7-day quota reset.

Alibaba Cloud Bailian platform announced that the official version of Qwen3.8-Max has been launched, and the personal Token Plan now supports this model. Starting from August 5, Beijing time, the official will issue a one-time 7-day quota reset entitlement to all users who previously subscribed. Users can reset their available quota as needed within the validity period. The original Qwen3.8-Max-Preview version will be taken offline on August 5, and the official recommends users to complete the migration as soon as possible to avoid affecting online services. Regarding the Credits consumption standards for Qwen3.8-Max, the official suggests that users monitor usage statistics in the console and plan their usage pace reasonably.

Read original →

Ant Ling Extends Free Access to Ling-3.0-flash on OpenRouter

开发生态

Ant Ling Extends Free Access to Ling-3.0-flash on OpenRouter

Ant Ling announced that the free usage period for the Ling-3.0-flash model on the OpenRouter platform has been extended to 8:00 AM PT on August 6. Users can continue to experience the model for free.

Ant Ling officially announced that the free access period for its Ling-3.0-flash model on the OpenRouter platform has been extended. Users can continue to use Ling-3.0-flash for free on the platform, and this free period will last until 8:00 AM Pacific Time on August 6.

Read original →

TRAE Grants Bonus Points to Existing Users and Advances Multiple Feature Upgrades

开发生态

TRAE Grants Bonus Points to Existing Users and Advances Multiple Feature Upgrades

TRAE, after launching its points system, announced additional general points for existing paid users and doubled the migration bonus points for users who have not upgraded. The company also stated it is working on adjusting image/video generation limits and supporting larger context windows in response to user feedback.

After the launch of its credit system, TRAE announced two additional benefits for existing users. Paid users who have upgraded to the new credit system received an extra batch of universal credits on August 3, based on their membership tier, valid for 31 days. For Speed Pass users who have not yet upgraded, the one-time migration bonus credits have been doubled from the original amount, valid for 31 days from the date of switching to the credit system. Both benefits are universal credits, usable across TRAE Work and TRAE IDE, and require no action upon receipt. The company also stated that features such as adjusting image and video generation limits, supporting larger context windows, supporting model thinking intensity selection, and increasing resources for the Deepseek-V4-Flash official model to cover more users are currently in progress.

Read original →

Kiro Launches Unified Kiro Agent Harness Across All Four Clients

开发生态

Kiro Launches Unified Kiro Agent Harness Across All Four Clients

Kiro has introduced a unified Kiro agent harness, replacing the separate architectures of individual clients. The new suite is now fully available on Kiro IDE, CLI, Web, and iOS, enabling consistent configuration across platforms.

Kiro officially announced the integration of its three previously independent agent suites into a single Kiro agent harness, which is now fully available across four clients: Kiro IDE, CLI, Web, and iOS. The new architecture runs as an independent server process and communicates with clients via the Agent Client Protocol (ACP) and its extensions, ensuring consistent behavior and configuration formats for spec-driven development, custom agents, and Hooks across all clients. Additionally, Kiro introduced a unified capability-level permission model based on Cedar, replacing the previously incompatible permission systems across clients.

Read original →

OpenRouter Releases Ori Eval for Model Comparison and Quantitative Recommendations

开发生态

OpenRouter Releases Ori Eval for Model Comparison and Quantitative Recommendations

OpenRouter has launched Ori Eval, an evaluation tool that helps developers choose the best model for their projects. Ori Eval uses a coding agent to automatically scan codebases and run cross-provider model comparisons.

OpenRouter launched Ori Eval, a CLI tool and evaluation framework now available. Ori Eval helps developers find the best model for their project by testing agents on real prompts, automatically scanning codebases, writing evaluation files, running cross-provider model comparisons, and providing recommendations with scores. Each run uses a fixed harness and model to ensure reproducible results. Developers can start by running a curl command via a coding agent, or manually install the CLI, which requires the Bun runtime and an OpenRouter login. Evaluations call real models and incur costs.

Read original →

Interconnects AI Launches Artifacts Hub and Adoption Dashboard to Analyze Open-Source Model Trends and Adoption

开发生态

Interconnects AI Launches Artifacts Hub and Adoption Dashboard to Analyze Open-Source Model Trends and Adoption

Interconnects AI announced the launch of two free tools to provide insights into the open-source model ecosystem. Artifacts Hub tracks open-source model trends, while Adoption Dashboard displays model adoption data by region and organization.

Interconnects AI announced the launch of two free resources: The Artifacts Hub and the Adoption Dashboard. Artifacts Hub is a curated view that integrates OpenRouter inference token data, Artificial Analysis model intelligence scores, and custom adoption metrics built on Hugging Face data, currently covering 792 models released in the past two years. The Adoption Dashboard is a dynamic dashboard showing download counts and derivative model numbers by geographic region and organization, highlighting the US-China gap and emerging players in the open-source ecosystem.

Read original →

Cloudflare Launches Open-Source Agent Runtime @cloudflare/computer

开发生态

Cloudflare Launches Open-Source Agent Runtime @cloudflare/computer

Cloudflare has released an early preview of its open-source agent runtime @cloudflare/computer. It provides agents with an isolated compute environment, leveraging a SQLite virtual file system and dynamically orchestrating isolates and Linux containers to execute tasks.

Cloudflare has released an early preview of its open-source agent runtime, @cloudflare/computer, designed to provide each agent with an isolated 'computer' environment. At its core, it provides a SQLite-based virtual file system and dynamically orchestrates fast isolates and full Linux containers based on task requirements. The company states that this design aims to optimize performance and cost, intending for agents to handle most work using isolates, with containers needed in fewer than 10% of tasks. Developers can now install the library via npm and try it out.

Read original →

Cloudflare Releases Billable Usage API for Programmatic Cost and Usage Queries on Self-Serve Accounts

开发生态

Cloudflare Releases Billable Usage API for Programmatic Cost and Usage Queries on Self-Serve Accounts

Cloudflare has introduced the new Billable Usage API, available to self-serve accounts, allowing all usage and costs to be queried per product via a single endpoint.

Cloudflare released the Billable Usage API, providing self-serve accounts with a single-endpoint programmatic query capability for cost and usage, covering all usage-based products including Workers, R2, D1, Workers AI, Vectorize, Images, Stream, and more. The API fields align with the FinOps Open Cost and Usage Specification (FOCUS), returning detail rows by product and billing cycle. Usage and cost data are currently updated daily, with enterprise contract support still under development.

Read original →

NousResearch Releases Hermes Agent v0.20.0

产品应用

NousResearch has released Hermes Agent v0.20.0. This version adds features such as streaming conversational voice, agent-to-agent protocol, and desktop app plugin SDK.

NousResearch released Hermes Agent v0.20.0. Compared to the previous version, this release includes approximately 3,650 commits and about 1,400 merged PRs. The core additions in this update include support for streaming synthesis and interruptible real-time conversational voice mode, covering multiple platforms such as CLI, desktop apps, and Feishu. It also introduces the Agent-to-Agent v1.0 protocol, outbound webhooks, fact-checking skills, as well as a plugin SDK for desktop apps and an artifacts sandbox preview. The default iteration limit for tool calls has been raised from 90 to 500, a context micro-compression mechanism has been introduced, and cold start time has been reduced to approximately 1.8 seconds.

Read original →

U.S. Reportedly Completes AI Regulatory Framework, Invites Major Tech Companies to Review

行业动态

U.S. Reportedly Completes AI Regulatory Framework, Invites Major Tech Companies to Review

According to reports, the U.S. government has invited AI companies such as OpenAI, Google, and Anthropic to review a completed voluntary AI regulatory framework on Tuesday. The framework would establish a process for submitting frontier models to the government before release.

According to an exclusive report by The Information, the U.S. government has invited employees from major tech companies including OpenAI, Google, and Anthropic to the White House on Tuesday to review the final version of a completed AI regulatory framework. The framework is voluntary and will establish a process requiring labs to submit frontier models to the government before release.

Read original →

OpenAI Details GPT-Live Architecture: Full-Duplex Model and Asynchronous Delegation

技术与洞察

OpenAI Details GPT-Live Architecture: Full-Duplex Model and Asynchronous Delegation

OpenAI has published details about its third-generation voice system, GPT-Live. It uses a full-duplex model, removing the turn detector and supporting simultaneous listening and speaking. When deep reasoning or tool use is needed, it can asynchronously invoke GPT-5.5 without interrupting the voice stream.

OpenAI published a detailed explanation of the underlying architecture of its third-generation voice system, GPT-Live. The system uses a full-duplex model to support simultaneous listening and speaking, removing traditional turn detectors to reduce latency. When deep reasoning or tool calls are needed, the system can asynchronously invoke frontier models such as GPT-5.5 without interrupting the core voice stream.

To achieve this experience, OpenAI spent six months refactoring the system, separating the core media stream from business logic. The system was rewritten in Go instead of Python, achieving p95 frame smoothness in the new system comparable to p50 in the old system. The team developed the WARP protocol and Instant Connect technology, reducing the number of network startup round trips from six to one. The system supports seamless instance switching and dynamic context compression, and before release, some real traffic was routed through it for silent testing.

Currently, GPT-Live powers ChatGPT Voice, including features such as controlling the computer and coordinating agents, with a dedicated API forthcoming.

Read original →

Google Agent Skills Team Shares Project Building, Testing, and Quality Control Mechanisms

技术与洞察

Google Agent Skills Team Shares Project Building, Testing, and Quality Control Mechanisms

The Google Agent Skills team has shared the project's Skills building and quality mechanisms, requiring Skills to pass automated checks and continuous evaluation to ensure quality at scale.

The Google Agent Skills team detailed the quality control mechanisms used in building, testing, and scaling the project. The project aims to encode Google Cloud domain knowledge into structured open-source instructions to reduce AI agent hallucinations and enforce best practices. To ensure quality at scale, the team established a standardized repository layout for each Skill, with a design that prioritizes referencing remote MCP tools. All Skills must pass an automated CI/CD pipeline, including link checks, before public release, and undergo continuous evaluation for accuracy and efficiency at commit time and weekly.

Read original →

GLM-5.3 References Appear in Zhipu's GitHub Repository and ZCode Documentation

前瞻与传闻

GLM-5.3 References Appear in Zhipu's GitHub Repository and ZCode Documentation

Community users have discovered that Zhipu's ZCode documentation briefly displayed references to GLM-5.3. Meanwhile, related text also appeared in Zhipu's official GitHub repository, but the company has not yet disclosed more information about GLM-5.3.

Community users have discovered a glm-5.3 branch in Zhipu's official GitHub repository zai-org/z-ai-sdk-java, and "GLM-5.3" also briefly appeared on the ZCode documentation page. The model is suspected to be launching soon. The above information comes from community users' observations of the official repository and documentation, and Zhipu has not yet officially announced the release date or availability scope of GLM-5.3.

Note: Content is AI-assisted and may contain hallucinations or errors.

Read original →