Google unveils Gemini 3.8 Flash and a Cyber security edition
要闻
Google released Gemini 3.8 Flash and its Cyber variant. The former scored 54.9% on HLE-Verified, while the latter exceeded a 70% success rate on an internal vulnerability benchmark spanning 20 programming languages.
Google announced Gemini 3.8 Flash and Gemini 3.8 Flash Cyber, its third Flash-series release in six weeks and its first since 3.7 Flash launched three weeks earlier. The two models share the same foundational capabilities and use long-running Agent loops to evaluate and improve the models recursively. According to Google, 3.8 Flash maintains the same speed and low-cost level as 3.7 Flash, surpasses most larger frontier models on DeepSWE v1.1, and outperforms 3.7 Flash and other frontier models on Vals Finance Agent V2 and Harvey's Legal Agent Benchmark. It scored 54.9% on HLE-Verified. On complex tasks, the model adds reasoning steps and calls tools repeatedly, so higher effort levels may consume more tokens. Developers can lower the effort level, while 3.7 Flash will remain fully supported.
Gemini 3.8 Flash Cyber is available only to a selected group of trusted defenders through the Fairwind Program. Google says the model outperformed 3.5 Flash Cyber and larger frontier models on CyberGym's autonomous vulnerability-discovery benchmark. It also exceeded a 70% success rate on an internal vulnerability benchmark covering complex codebases across 20 programming languages. On the CWE-Bench patching benchmark run by Collinear, it recorded a pass@1 of 47.2%, compared with 47.8% for a leading frontier model. Google says it already uses the model to protect internal code and has prioritized vulnerability patching over exploit capabilities. Gemini 3.8 Flash includes misuse safeguards for CBRN and cyber offense, while the Cyber variant applies less restrictive cybersecurity mitigations and is therefore limited to trusted defenders. Gray Swan testing also showed improved prompt injection robustness across the Gemini 3.8 models.
Meta has begun rolling out Muse Spark 1.3 in Muse Code and the Meta Model API. Existing reasoning modes are available, while max reasoning will launch after additional safety testing is completed.
Meta has begun rolling out Muse Spark 1.3 in Muse Code and the Meta Model API. Existing reasoning modes are already available, while max reasoning will be released after additional safety testing is completed. According to the company, the new version improves performance on agentic and coding tasks and strengthens its ability to manage multiple workflows within a single long thread for longer-horizon work. The model can use tools to build context from messy or conflicting sources, correct gaps in its plan, and track information it has gathered. When instructions are ambiguous or execution is blocked, it asks questions or requests help from the user, and it seeks confirmation before taking consequential actions. During long tasks, it can provide ongoing updates or work silently in the background based on user preferences. Meta also says version 1.3 improves constraint retention for complex, long-form instructions, task matching in single-threaded contexts, and recognition of its own capabilities, knowledge boundaries, and execution hurdles to reduce fabricated outcomes.
Qwen releases Qwen3.8-Max-0902, claiming the top spot on CodeArena's overall frontend coding leaderboard
要闻
Qwen has released Qwen3.8-Max-0902, making its API available on the Qianwen AI platform and integrating it into Qianwen Office, Qoder, and the Qianwen app. According to the company, the version ranks first on CodeArena’s overall frontend programming leaderboard.
Qwen released Qwen3.8-Max-0902 one month after Qwen3.8-Max. At launch, the new version became available through an external API on the Qianwen AI platform and was integrated into Qianwen Office, Qoder, and the Qianwen app for enterprises, developers, and users. The update focused its post-training on programming and professional office work. According to Qwen, the model’s overall performance improved, making it more suitable for complex enterprise tasks, scientific research, and long-horizon tasks, while placing first on CodeArena’s overall frontend programming leaderboard.
Multiverse Computing has connected its first large model, Quasar 438B, to the CompactifAI API. The text-only model has 438 billion parameters and a 1M-token context window.
Spanish AI company Multiverse Computing has released the Quasar 438B reasoning model and opened access through the CompactifAI API. It is the company’s first large model, with 438 billion parameters, and targets enterprise-grade Agents, software development, and complex multi-step tasks. It supports English and Spanish, has a 1M-token context window, and accepts and produces text only. The proprietary model has closed weights, while API pricing is $0.60 per 1 million input tokens and $1.80 per 1 million output tokens. According to the company, Quasar 438B scored 43 on Artificial Analysis’s Intelligence Index v4.1.1 and ranked first among the European models included in the comparison.
fal previews MiniMax H3 Max Turbo, doubling speed while halving costs
模型发布
fal has launched the MiniMax H3 Max Turbo preview with text-to-video and image-to-video access. The company says it runs at twice the original model’s speed and half the cost, while promotional 768p output is priced at $0.01 per second for two weeks.
fal has made the MiniMax H3 Max Turbo video model preview available for testing on its platform, with both text-to-video and image-to-video access. Promotional pricing applies for two weeks from launch. According to fal, the version doubles generation speed compared with H3 Max, cuts generation time in half, and reduces cost to half that of the original while retaining its prompt understanding, aesthetics, and overall quality. fal’s evaluation sets the quality target at about 97% of the original model, and the company says it still outperforms H3 and other video models. During the promotion, 768p video output is billed at $0.01 per second.
Claude Cowork and Claude Code desktop apps add support for background computer use
开发生态
Anthropic has enabled background computer use in the Claude Cowork and Claude Code desktop apps. The Beta requires macOS 15 or later and a Pro or Max subscription, with no support for Team or Enterprise plans.
Anthropic has released a Beta of background computer use for the Claude Cowork and Claude Code desktop apps on macOS 15 or later. According to the company, once enabled, Claude operates in a background window by default without taking control of the user’s mouse or keyboard, allowing other work to continue at the same time. The feature is available only to Pro and Max subscribers and does not support Team or Enterprise plans. Computer use is disabled by default and must be turned on manually in settings. Claude requests permission when an app is used for the first time, while some sensitive apps, including investment platforms, are blocked by default. Anthropic says the feature has no sandbox isolation and carries risks including prompt injection, and therefore advises against granting it access to apps containing sensitive data.
Anthropic open-sources a Claude Commerce Agents reference implementation
开发生态
Anthropic has released the Claude commerce agent blueprint and runnable reference implementations. The company says retailers using Claude agents saw carts grow by up to 35% and shoppers become 60% more likely to complete a purchase.
Anthropic has released and made available the Claude commerce agent blueprint, providing complete, working reference implementations of a shopping agent and merchant agent for retail, travel, telecom, and ticketing platforms. The repository includes harnesses, patterns, guardrails, and a Claude Code plugin. Engineering teams can build with the Messages API, Agent SDK, or Claude Managed Agents (beta), then deploy to the Claude API, Amazon Bedrock, Microsoft Foundry, or Google Cloud Vertex AI. Anthropic also provides self-guided demos and engineering implementation details for each vertical, and says teams can get an agent running in days.
The shopping agent can be embedded in an app or website and connected to a catalog, cart, checkout, customer preferences, and order history, while payments are handled through an existing checkout or an agentic payments provider. Its guardrails constrain prices and products to actual catalog data and avoid manipulative upsell practices. The merchant agent provides sales analytics, catalog and inventory management, marketing and promotions, and charts and dashboards. Any change it proactively proposes requires human approval before going live. According to Anthropic, retailers using Claude shopping agents saw carts grow by up to 35%, while shoppers were 60% more likely to complete a purchase.
Perplexity open-sources the Lily local inference engine
开发生态
Perplexity announced that it is open-sourcing Lily, a local inference engine for Apple silicon. Its official benchmarks put average prefill and decode throughput at 1.23× and 1.35× that of MLX-LM, respectively.
Perplexity has announced the open-sourcing of the local inference engine Lily, with a standalone demo available in the perplexityai/pplx-garden repository on GitHub. Built for Perplexity Computer’s hybrid computing system, Lily targets Apple silicon and Qwen3.6-35B-A3B while optimizing prefill and decode as separate workloads. Its single-process implementation combines a Rust runtime, an OpenAI-compatible chat-completions API, and custom Metal kernels, with neither PyTorch nor MLX in the execution path. According to the official benchmark data, Lily outperformed the general-purpose MLX-LM framework at every tested length from 256 to 128K tokens; its average prefill and decode throughput reached 1.23× and 1.35× that of MLX-LM, respectively. The company stated that output quality remained essentially unchanged.
豆包 adds parallel multi-Agent execution and computer control on Mac
产品应用
According to a report dated September 3, 2026, Doubao launched parallel multi-Agent execution and computer control for Mac that week, enabling task decomposition and coordinated processing as well as local GUI operation without MCP, APIs, plugins, or a CLI.
According to the original article published on September 3, 2026, Doubao added two features that week: parallel multi-Agent execution and computer control for Mac. With parallel multi-Agent execution, a primary Agent first breaks down a task, then invokes different sub-Agents to handle separate modules. The feature supports batch processing, modular generation, multidimensional research, and multichannel search. Computer control was previously available on Windows and has now been extended to Mac. According to its description, Doubao can recognize and operate a local computer interface through the GUI, completing relevant actions even without MCP, APIs, plugins, or a CLI.
Anthropic launches Claude Content Checker for detecting file provenance
产品应用
Anthropic has released Claude Content Checker, a browser-based tool that checks whether a file contains a C2PA content credential pointing to Claude. It supports 17 media formats with a maximum file size of 100 MB.
Anthropic has launched Claude Content Checker to inspect C2PA content credentials in file metadata and display a result when a credential points to Claude. The tool supports JPG, PNG, GIF, WEBP, TIFF, HEIC, AVIF, SVG, DNG, JXL, MP4, MOV, AVI, WAV, MP3, M4A, and FLAC, with a file size limit of 100 MB. According to the company, checks run inside the browser and files never leave the device. When Claude creates or processes a supported file type, it adds a small, cryptographically signed note to the metadata. This credential can only indicate whether Claude was involved in producing the file; it cannot determine whether Claude helped create the content and contains no information that identifies the user. C2PA is an open industry standard also used by camera manufacturers and photo-editing software, and any C2PA-aware tool can read it. OpenAI, Google DeepMind, and Gemini also offer similar tools. For watermarks embedded in text, Anthropic separately offers a Detection API that is currently available as a private preview only to eligible organizations, as required under EU law.
Proximal released FrontierSWE v2 with 21 new ultra-long-horizon technical challenges, bringing the total to 34 and setting each evaluation window at 20 hours. Claude Fable 5.1 ranked first.
Proximal has released FrontierSWE v2, adding 21 ultra-long-horizon technical challenges and bringing the total number of tasks to 34. The additions span visual reasoning, graphics, scientific computing, and AI research, including decoding speech from MEG brain recordings, predicting ball trajectories from video, training a 10-day weather forecasting model, building a TORCS racing bot using vision input alone, and reconstructing a flight simulator renderer with OpenGL. The release also removes 4 v1 tasks that had become saturated or could not be scored deterministically: PCQM4Mv2 Molecular Gap Prediction, Pyright Type Checking Optimization, Revideo Rendering Pipeline Optimization, and Dependent Type Checker.
The new version revamps the evaluation harness to measure how far models can progress within a 20-hour window. The system continuously informs an Agent of its remaining time and encourages it to keep working rather than submit a solution prematurely. These mechanisms are integrated into proximus, a minimal coding agent harness built specifically for ultra-long-horizon tasks. The officially published ranking places Claude Fable 5.1 first, followed by GPT-5.6 and GLM-5.3; Opus 5 is used as a fallback for tasks blocked by content filters. According to the official statement, the benchmark is not yet saturated and shows a large gap between models that perform similarly on other benchmarks.
Google launches the Fairwind Program, opening up its network defense capabilities
行业动态
Google’s Fairwind Program has opened access to cyber defense capabilities provided by Gemini 3.8 Flash Cyber. The company says the model can find and fix vulnerabilities while supporting practical vulnerability research, complex code synthesis, and real-time threat intelligence.
Google has launched the Fairwind Program, opening its cyber defense capabilities through Gemini 3.8 Flash Cyber. The company positions the model as an intelligent, cost-effective partner whose role includes not only identifying security flaws but also fixing discovered issues. According to Google, the model can accelerate practical vulnerability research and complex code synthesis while supporting real-time threat intelligence.