Daily AI Digest

2026-07-30

Source:橘鸦 AI 早报 · 19 items

2026-07-30
2026-08-04 2026-08-03 2026-08-02 2026-08-01 2026-07-31 2026-07-30 2026-07-29 2026-07-28 2026-07-27 2026-07-26

OpenAI launches GPT-5.6 Sol efficiency improvements and resets Codex usage limits

要闻

OpenAI launches GPT-5.6 Sol efficiency improvements and resets Codex usage limits

Tibo, head of Codex, announced at noon Beijing time on July 29 that quotas have been reset for all ChatGPT Work and Codex users, along with efficiency improvements. Addressing the issue of Sol consuming quotas too quickly, the update optimizes tool-waiting and web-search scenarios, with typical usage duration expected to extend by about 18%. Meanwhile, the previously suspended five-hour limit will soon be restored. Notably, Tibo later replied to a community user request, hinting that another usage reset would be performed on July 31 local time.

Codex lead Tibo announced at noon Beijing time on July 29 that usage quotas have been reset for all ChatGPT Work and Codex users, along with efficiency improvements, in response to the issue of Sol credits being consumed too quickly. Regarding the reasons, Tibo explained the following: GPT-5.6 Sol is more willing to keep working for longer durations, initiate more tool calls, and coordinate complex workflows across tools and subagents; at the same reasoning intensity, it works harder than previous models, and the High tier consumes more tokens on Sol than on GPT-5.5. Programmatic tool calling gives Sol the flexibility to run tool calls in parallel or continue working while waiting, but this also leads to more turns per reply, more cached input tokens, and higher-than-expected usage, especially when Sol is waiting for tool calls to complete or running a large number of web searches. The impact is not evenly distributed: median users actually find Sol's token efficiency acceptable, while some heavy users handling harder tasks consume quotas much faster; the team focused too much on average and median usage before launch and did not fully account for the long tail. This update optimizes tool waiting and web search scenarios, with typical usage duration expected to extend by approximately 18%. Meanwhile, the previously suspended five-hour limit will resume tomorrow. Notably, Tibo later responded to a community user's request, hinting that another usage reset will be performed on July 31 local time.

Read original →

SpaceXAI launches Grok Voice Think Fast 2.0 voice model

模型发布

SpaceXAI launches Grok Voice Think Fast 2.0 voice model

SpaceXAI has launched its next-generation voice model, Grok Voice Think Fast 2.0, which the company says achieves significantly improved scores on multiple voice benchmarks. The model is now available in the API and Voice Agent Builder.

SpaceXAI released the next-generation voice model Grok Voice Think Fast 2.0, with major improvements in intelligence, transcription accuracy, and conversational capabilities. According to official data, the model achieved significant score improvements on multiple voice benchmarks, and the company claims its transcription accuracy in noisy environments outperforms dedicated speech-to-text models. The model is now available on the API and Voice Agent Builder, priced at $0.08 per minute of audio.

Read original →

Google DeepMind releases music model Lyria 3.5

模型发布

Google DeepMind has released Lyria 3.5, a music generation model, and integrated it into the Google Flow Music platform. The new version improves lyric quality and vocal expressiveness, and supports more complex melodies.

Google DeepMind has released its latest music generation model, Lyria 3.5, which has now been officially integrated into the Google Flow Music platform for users. According to the company, Lyria 3.5 delivers richer arrangements and more complex melodic structures, generates higher-quality lyrics, and significantly improves vocal expressiveness, emotional nuance, and pronunciation naturalness. In terms of creative control, the new model better follows user instructions, supports setting precise BPM and duration controls, and allows exporting stems of full songs.

Read original →

SKT releases 688B-parameter open-source large model A.X K2 and audio version

模型发布

SKT releases 688B-parameter open-source large model A.X K2 and audio version

SKT has released the open-source large language model A.X K2, with a total of 688B parameters and 33B active parameters, trained natively in FP8. It also introduced the A.X K2 ALM audio language model, whose weights are planned for later release.

SKT has released the open-source large language model A.X K2 and its audio version A.X K2 ALM. A.X K2 adopts a MoE architecture with 688B total parameters and 33B activated parameters. It is trained with native FP8 precision and supports a maximum context length of 256K tokens. The company states that the model excels in mathematics and Korean language tasks, and has reached the gold medal threshold in the IMO 2025 test. Currently, the weights of A.X K2 are open-sourced on Hugging Face under the Apache-2.0 license. The A.X K2 ALM audio model is built on the frozen A.X K2 Light backbone, and its weights have not been released yet. The company has announced plans to release them in the future.

Read original →

Cursor officially announces Cursor is coming to iPad

开发生态

Cursor officially announces Cursor is coming to iPad

Cursor officially announced that Cursor is now available on iPad, with all the features of the iPhone version. Both iPhone and iPad versions have added inbox and a full PR review experience.

Cursor officially announced that Cursor is now available on iPad. The iPad version offers all the same features as the iPhone version and provides users with more space to collaborate with agents. Additionally, both the iPhone and iPad versions have added an inbox feature for staying organized, as well as a review experience covering complete PRs, which supports viewing comments, reviewing, and approving.

Read original →

支付宝 upgrades AI payment developer incentive program

开发生态

支付宝 upgrades AI payment developer incentive program

支付宝 has upgraded its AI payment developer incentive program, offering up to 5,660 yuan in Token-specific subsidies. From today until August 21, individual developers can participate by integrating designated payment collection products.

Alipay announced an upgrade to its AI payment developer incentive program, continuing to provide AI developers with a special token subsidy of up to 5,660 yuan, while further expanding the scope of incentives and lowering the participation threshold. From now until August 21, individual developers who integrate any one of Alipay's AI pay-as-you-go, AI web app payment collection, or AI mobile app payment collection products during Vibe Coding can unlock tiered incentives based on the cumulative number of active paying users. In addition, individual developers using AI pay-as-you-go can enjoy a 0% fee rate until December 31, 2026.

Read original →

Qoder open-sources AI coding agent evaluation tool Better Harness

开发生态

Qoder has open-sourced Better Harness, an AI coding agent evaluation tool. It reviews agent workflows across five dimensions, including task understanding, and generates evidence-backed reports and fix suggestions.

Better Harness is an AI coding agent workflow evaluation tool developed by QoderAI, now open-sourced on GitHub under the MIT license. The tool reviews projects across five dimensions: task understanding, controlled execution, change validation, reliable delivery, and learning capture, generating evidence-backed reports and converting findings into prioritized fix recommendations. Better Harness supports Claude Code, Codex Desktop, Codex CLI, Cursor, and Qoder.

Read original →

Perplexity open-sources Numbat: an agent security detection and response suite for client endpoints

开发生态

Perplexity open-sources Numbat: an agent security detection and response suite for client endpoints

Perplexity has open-sourced Numbat, a cross-agent-framework security detection and response layer that enables security teams to monitor agent behavior in real time and block dangerous actions before execution.

Perplexity has officially open-sourced Numbat, an agent detection and response security suite designed for client endpoints. Numbat runs across desktop, command-line, IDE, and gateway agent frameworks, integrating agent hook subsystems, session artifacts, and OTLP telemetry data to provide security teams with real-time monitoring, local detection, pre-execution blocking, and forensic reconstruction capabilities. Numbat includes 52 built-in detection rules covering 11 categories of behavior, including multi-step sequence detection for secrets access, data exfiltration, privilege escalation, and lateral movement. Perplexity internally uses Numbat to protect coding agents such as Claude Code, Codex, OpenCode, and Pi used by its engineers. Numbat is currently available as a single Go binary under the Apache 2.0 license, supporting macOS, Linux, and Windows.

Read original →

OpenAI launches ChatGPT for Academic Researchers program

产品应用

OpenAI has launched the ChatGPT for Academic Researchers program, offering free access to its frontier models for scientists, mathematicians, and engineers. The program initially targets 10,000 researchers and is expected to expand to 100,000 by 2027.

OpenAI announced the launch of a program called "ChatGPT for Academic Researchers," aimed at providing researchers in science, mathematics, and engineering with free access to its frontier models. According to official information, the first phase of the program will cover 10,000 researchers, with plans to scale up to 100,000 by 2027, primarily targeting users at selected academic institutions. The company stated that this initiative aims to help researchers tackle advanced problems, accelerate scientific discovery, and enhance productivity.

Read original →

Google introduces new natural language voice features for the macOS Gemini app

产品应用

Google introduces new natural language voice features for the macOS Gemini app

The macOS version of Gemini has launched new voice capabilities. Users can press and hold the Fn key for intelligent dictation, and with a setting enabled, it can perform complex tasks using on-screen context. It is currently rolling out to English-speaking users worldwide.

Google has officially announced that the Gemini for macOS app now features new natural language voice interaction capabilities, allowing users to create, edit, and summarize content via voice. By pressing and holding the Fn key in any window on the desktop, users can use the default intelligent dictation feature, which automatically filters filler words and captures mid-speech corrections, outputting formatted text directly at the cursor. If users manually enable Gemini's reasoning capabilities in settings, the app can understand the context on the current screen and execute complex tasks such as extracting and summarizing local files, rewriting text, and generating and editing images based on voice commands. This feature is now available to all Gemini for macOS app users worldwide, currently supporting English only, with additional language support to be added later.

Read original →

Replit launches new Replit Design AI design tool

产品应用

Replit launches new Replit Design AI design tool

Replit has launched Replit Design, an AI design tool now available to all users. The suite features Ambient Intelligence and supports one-click application of brand systems.

Replit officially released its new AI design suite, Replit Design, which is now fully available. Replit Design features a built-in capability called Ambient Intelligence that provides one-click adoptable suggestions at every step of the design process, eliminating the need for users to input prompts or master professional design terminology. In terms of model support, users can freely choose from multiple models and compare the outputs of different models. Additionally, the suite directly integrates over 600,000 real app UI screens owned by Mobbin, allowing users to access them without needing a separate Mobbin account.

Read original →

Hermes Agent adds Hey Hermes voice wake word

产品应用

Hermes Agent has introduced a local voice wake feature. Users can say the wake word to start a new session hands-free. Detection is performed entirely on-device and is disabled by default.

Hermes Agent has officially launched a voice wake-up feature, enabling users to start a new voice session hands-free by speaking a specific phrase. The feature supports the command-line interface (CLI), TUI interface, and desktop applications. To ensure privacy, the wake word detection process runs entirely on the local device, and no audio leaves the device until the user speaks an actual command. The feature is disabled by default and is limited to interfaces with a local microphone.

Read original →

OpenAI says two API settings triple GPT-5.6 benchmark scores

技术与洞察

OpenAI says two API settings triple GPT-5.6 benchmark scores

OpenAI stated that enabling two API settings, retained reasoning and compression, increased GPT-5.6 Sol's score on the ARC-AGI-3 benchmark by three-fold to 38.3%, while reducing output tokens by 6x.

The official OpenAI blog explains that by enabling the retained reasoning and compression settings for GPT-5.6 Sol, its performance on the ARC-AGI-3 benchmark improved significantly. According to official data, the model's score under the highest configuration rose from 13.3% to 38.3%, an approximately threefold increase, while output tokens were reduced sixfold. OpenAI noted that the previous test tools' discarded reasoning and rolling truncation mechanisms led to the low scores. The company recommends that API developers use the same settings in their Responses API to maximize performance.

Read original →

OpenAI says applying GPT-5.6 Sol to optimize its own operations cuts service costs by 20%

技术与洞察

OpenAI says applying GPT-5.6 Sol to optimize its own operations cuts service costs by 20%

OpenAI described how it applied GPT-5.6 Sol to optimize its own operational infrastructure. By rewriting production cores and improving speculative decoding, it reduced service costs by 20% and increased token generation efficiency by over 15%. The optimizations are now used in the reasoning stack and Agentic harness.

OpenAI published an article about the efficiency updates for the GPT-5.6 model family. The company stated that they have applied the flagship model GPT-5.6 Sol in Codex to optimize their own infrastructure and performance. According to official data, by rewriting production cores using Triton and Gluon languages, the model reduced end-to-end service costs by 20%; by improving speculative decoding, it increased token generation efficiency by more than 15%. These optimizations are now applied to OpenAI's inference stack and the Agentic harness used by Codex and ChatGPT Work.

Read original →

腾讯混元 open-sources AngelSpec speculative decoding framework

技术与洞察

腾讯混元 open-sources AngelSpec speculative decoding framework

The 腾讯混元 team has open-sourced AngelSpec, a speculative decoding framework that supports training and deployment. The team states that it achieves up to 2.4x end-to-end speedup on Hy3-A21B.

Tencent Hunyuan Team announced the open-source release of the end-to-end speculative decoding framework AngelSpec, publishing version v0.1.0. The framework is based on torch-native, supports MTP and block-parallel speculative decoding training, and unifies the training pipelines for six draft architectures including DFly and DFlash. Official data shows that in tests with the Hy3-A21B model and concurrency levels from 4 to 64, the DFly architecture achieved an end-to-end speedup of 1.98 to 2.40 times compared to autoregressive decoding, with throughput 10.5% to 11.8% higher than DFlash.

Read original →

MiniMax and Fireworks AI co-open-source M3 sparse attention GPU kernels

技术与洞察

MiniMax and Fireworks AI co-open-source M3 sparse attention GPU kernels

MiniMax and Fireworks AI have jointly open-sourced two GPU kernel repositories for MiniMax M3 block-sparse attention.

MiniMax and Fireworks AI have open-sourced the GPU kernel code for the MiniMax M3 block-sparse attention co-developed by both parties, with two repositories now public. fw-ai/minimax-kernels, developed by Fireworks AI, implements KV-outer block-sparse attention based on NVIDIA Blackwell (SM100) CuTe-DSL, licensed under Apache 2.0. The fireworks-msa branch of MiniMax-AI/MSA integrates Fireworks' KV-outer sparse prefill backend into the original MSA kernel, licensed under MIT.

Read original →

月之暗面 completes $3.5 billion funding round at $35 billion valuation

行业动态

According to reports, 月之暗面 has completed a $3.5 billion funding round at a post-money valuation of $35 billion. The company is reportedly approaching investors at a pre-money valuation of $50 billion, with plans to list in Hong Kong as soon as this year.

According to Bloomberg, Moonshot AI has completed a $3.5 billion funding round, far exceeding its original target of $1 billion to $2 billion, with a post-investment valuation of $35 billion. The report noted that this round was driven by the buzz created by Moonshot AI's Kimi K3 model in Silicon Valley. Insiders also said that the company has begun contacting potential investors to seek a new round of funding at a pre-investment valuation of $50 billion, which would be its final round before a Hong Kong listing, with the IPO potentially completed as early as this year.

Read original →

AlphaFold team disbanded, nearly a quarter of researchers leave Google DeepMind

行业动态

According to reports, Google DeepMind has disbanded the AlphaFold team, with nearly a quarter of researchers leaving. Most moved to the Gemini project or Isomorphic Labs.

According to media reports, Google DeepMind has disbanded the AlphaFold team responsible for protein folding. Over the past year, most of the original paper authors have been reassigned to projects around the Gemini large language model, or transferred to Isomorphic Labs, which focuses on drug discovery. Google DeepMind's vice president of research, Pushmeet Kohli, confirmed that the lab's strategy has evolved and is now shifting to developing systems that assist scientists, competing with OpenAI and Anthropic to build frontier AI agents. Additionally, nearly a quarter of the original authors have completely left, including core member John Jumper, who joined Anthropic.

Read original →

SpaceXAI sues to block new law restricting AI "nudification" technology

行业动态

According to reports, SpaceXAI has sued the Minnesota Attorney General in U.S. federal court to block the enforcement of a new law restricting AI "nudification" technology. The company argues that the law is overly broad and does not provide a liability exemption, which would force it to restrict the Grok Imagine feature.

据报道,SpaceXAI已在美国联邦法院起诉明尼苏达州总检察长基思·埃里森,要求阻止该州限制人工智能“裸化”技术(nudification technology)的新法律生效。该法律即众议院第1606号法案(H.F. 1606),定于2026年8月1日生效,禁止网站、应用等服务的所有者允许用户使用裸化功能,违规者每次行为最高面临50万美元民事罚款及私人诉讼责任。SpaceXAI表示不反对禁止色情深伪影像,但认为法律范围过宽,未为采取合理措施的平台提供免责机制,可能覆盖经过同意或具有表达价值的影像。公司称该法将迫使其限制Grok Imagine的图像生成和编辑功能,主张该法违反美国宪法第一修正案。

提示:内容由AI辅助创作,可能存在幻觉和错误。

Read original →