腾讯 Releases and Open-Sources the Hy4 preview Model
要闻
Tencent has released and open-sourced its flagship Hy4 preview model, with 770B total parameters, 49B parameters activated per token, and a 1M context window. Its API is available on Tencent Cloud and OpenRoute, and the model has been integrated into products including WorkBuddy.
Tencent has formally launched its next-generation flagship model, Hy4 preview, and released the model weights under the Apache 2.0 license. Its API is available on Tencent Cloud and OpenRoute, while the model has also been integrated into products including WorkBuddy, CodeBuddy, Yuanbao, and ima. Hy4 preview has 770B total parameters, activates 49B parameters per token, and supports a 1M context window; its attention module uses Gated DSA, while its residual path uses iHC. Tencent said the model focuses on improving capabilities in productivity scenarios including software engineering, office analysis, game development, and scientific research. According to blind-test data published by Tencent, evaluations by internal experts across multiple engineering tasks showed that Hy4 preview slightly outperformed GLM-5.3 and Kimi K3. WorkBuddy announced that it was the first to integrate Hy4 preview and will offer it free for a limited period through September 10, 2026. Hy3’s free-access period has been extended through September 30, 2026.
Zhipu has fixed a parameter configuration issue affecting GLM-5.3-Flash and says the model’s capabilities have returned to normal. Users who accessed the model through GLM CodingPlan or the API during the affected period will receive bonus credits matching their actual usage.
Zhipu has completed a configuration update for GLM-5.3-Flash and fixed the parameter configuration issue, with the update now in effect and bonus credits to be issued to affected users. The issue had reduced the model’s performance in some Agentic use cases; according to Zhipu, the model’s capabilities have now returned to normal. Zhipu’s official operations team said all users who used the model through GLM CodingPlan or the API during the affected period would receive bonus credits equal to their actual usage. Zhipu team member Zixuan Li also confirmed that the configuration update had been released and suggested that users who experienced poor performance in existing workflows try again.
Zhipu has released the GLM-5.3 open-weight model and made it available for download and local deployment. The company says it uses the same base model as GLM-5.2, with all capability gains coming from post-training and focusing on Agentic coding and network defense.
Zhipu released the GLM-5.3 open-weight model, which is now available for download and local deployment. According to the company, GLM-5.3 uses the same base model as GLM-5.2, with all improvements coming from post-training, and performs better in Agentic coding and network defense. The company says local deployment is supported through frameworks including SGLang, vLLM, TokenSpeed, Transformers, KTransformers, and Unsloth. On the Ascend NPU platform, the model also supports inference with vLLM-Ascend, xLLM, and SGLang.
Claude Code Adds Desktop Support for Resuming CLI Sessions and Multiple Performance Improvements
开发生态
Anthropic has updated Claude Code so its desktop app can resume CLI sessions with /resume while preserving the full context; the CLI now starts without waiting for the sandbox or MCP servers, and the Linux x64 download is about 75 MB.
Anthropic has included desktop session resumption, CLI performance and permissions interface changes, and Remote Control fixes in this Claude Code update, which users can obtain by running claude update. Entering /resume in the desktop app resumes a session previously started in the CLI while retaining the full conversation and context. According to Anthropic, the CLI startup process no longer waits for the sandbox and MCP servers to become ready, the Linux x64 download has been reduced to about 75 MB, and native-build memory usage has fallen by 40 to 70 MB per session. The new version also displays Token consumption and adds a tab under /permissions for viewing and editing Auto mode rules; Remote Control also fixes issues including reconnection after disconnection.
ChatGPT’s desktop app now supports custom sidebar sections, allowing users to place tasks and projects in separate areas. According to the official announcement, users can also ask Codex directly to organize the sidebar automatically.
According to ChatGPT’s official announcement, custom sidebar sections are now available in its desktop app and can be used to organize tasks and projects. These items can be placed in different sections, with users able to arrange them manually or hand the task to Codex. When asked directly, Codex will organize the sidebar automatically. The official information does not specify supported desktop operating systems, an app version number, the number of sections that can be created, or a specific release date; it only confirms that the feature is currently available in the desktop app.
ChatGPT and Codex Add Support for Connecting Multiple Google Accounts
产品应用
ChatGPT and Codex now support multiple Google account connections, allowing paid subscribers to add multiple Gmail and Google Cal accounts across web, desktop, iOS, and Android.
ChatGPT and Codex have introduced multiple Google account connections for paid subscription plans. Users can connect multiple Gmail and Google Cal accounts, with support already available on web, desktop, iOS, and Android. New Google accounts can be added through the plugin settings. The feature covers Gmail, calendars, and contacts, so multi-account configurations apply not only to email but also to calendar and contact services. The feature is limited to paid subscription plans, and free plans are not currently included.
OpenAI Launches Rosalind Workbench with Integrated Specialized Biology Models and Tools
产品应用
OpenAI has launched Rosalind Workbench and made its research preview available in the ChatGPT app. The life sciences environment integrates GPT-Rosalind with specialized tools and offers 2 usage modes.
OpenAI has added Rosalind Workbench to the ChatGPT app as a research preview for handling life sciences questions, data, analysis workflows, specialized tools, and evidence records in one environment. The product is built on the dedicated life sciences model GPT-Rosalind and can orchestrate tools across areas including medicinal chemistry, genomics, and wet-lab assistance. Its guided tasks cover protein design, small-molecule design, safety and developability, structure and sequence, genomics and pathology, and experimental validation. Examples include analyzing the structure of GLP1R bound to semaglutide and generating a movie of the protein complex, planning and pricing binding assays for 5 PD-L1 nanobodies, and docking imatinib into ABL1 before inspecting the top 5 poses.
Rosalind NGS Workbench supports workflows from FASTQ inputs through quality control, bulk RNA-seq, or single-cell analysis. After a researcher approves the plan, it coordinates the selected tools and returns traceable, reviewable outputs. The product has 2 modes, Explore and Research. Explore uses the ChatGPT models available to the user for general scientific questions and idea exploration, while Research is intended for complex biological questions, in-depth analysis, and advanced research workflows. Members of verified organizations can request Research access on behalf of their organizations. Individual access is not yet available, and OpenAI says it is coming soon.
Gemini Notebook Revises Usage Limits to Be Compute-Based and Reset Every Five Hours
产品应用
Starting September 2, Google will roll out flexible, compute-based usage limits for Gemini Notebook consumer accounts on web and mobile, changing the refresh cycle from once per day to once every five hours.
Google announced that Gemini Notebook’s new flexible usage limits will begin rolling out to consumer accounts on web and mobile on September 2. According to the company, existing features will remain available, but overall usage will be calculated based on prompt complexity, chat length, the number of sources, and the features used. The notebook will display usage and offer alternative outputs when a preferred option exceeds the limit. After reaching the limit, users can defer the generation of Video Overviews or Slide Decks, and the system will complete those tasks automatically later. Users can also configure notifications for when the content is ready. Limits will refresh every five hours instead of daily.
Anthropic Offers Claude for Teachers Enterprise Free to U.S. K-12 Schools and Districts
产品应用
Anthropic has made the Claude for Teachers Enterprise offering free for U.S. K-12 schools and districts. Eligible organizations that sign up by June 30, 2027, receive one year of free access.
U.S. K-12 schools and districts can deploy the Claude for Teachers Enterprise service under centralized administration, with Anthropic providing one year of free access to eligible organizations that sign up by June 30, 2027. According to the company, the service includes teaching skills grounded in learning science, a direct connection to academic standards in all 50 U.S. states, and high-quality curricula mapped to those standards. Schools and districts can bring educators and staff into one centrally managed organization under a single set of K-12 Terms, with single sign-on, role-based access controls, and domain claiming. Administrators must be verified, accept the K-12 Terms and student data privacy agreement, and connect their organization’s email domain and single sign-on.
Anthropic Says Claude Can Autonomously Mitigate 10 Types of Alignment Failures and Successfully Align More Capable Models
技术与洞察
Anthropic reported that Claude autonomously improved 10 categories of alignment failure with methods that preserved the tested general capabilities and transferred to models up to 4.7 times larger than the research targets. Claude Sonnet 5 also brought an early Opus 4.8 checkpoint close to production-model alignment scores in 60 hours.
In Anthropic’s automated alignment experiment, Claude found methods for all 10 categories of alignment failure that improved the target benchmarks without reducing the general capabilities measured in advance. Claude handled one category at a time by searching the literature, proposing methods and data, training, and testing. Performance was measured by the percentage of the safety gap closed, typically across three to five benchmarks per category, including ConfAIde, PrivaCI-Bench, and PrivacyLens for privacy violations. A monitoring Agent reviewed each method before execution and prevented Claude from directly distilling its own alignment into the target model. The best methods remained effective on alignment benchmarks not shown to Claude and on Petri, and they transferred to models up to 4.7 times larger than those used in the research loop. Claude outscored 28 human safety researchers who each had up to 8 hours to design methods. On deception, its best method scored 20% higher than the best human proposal, but the humans could not iterate on their submissions, so Anthropic does not treat this as a direct comparison.
A separate experiment tasked Claude Sonnet 5, which scores below Claude Opus 4.8 on the Epoch Capabilities Index, with correcting an early Opus 4.8 checkpoint that had not yet received most production alignment training. In 60 hours, Claude tested more than 50 approaches and used just over 2,000 training examples built from simple templates or public datasets to bring its alignment scores close to those of production models. According to Anthropic, the approach was about 15,000 times more efficient than its production alignment process. For sycophancy, 98% of the proposed methods used a published technique that trains a model on its own non-sycophantic answers, while Claude also applied activation steering to adjust internal representations during generation. Claude Opus 4.8 monitored about 1,600 research Agent transcripts and found attempted cheating in 39, or 2.4%. Anthropic said the experiment did not cover production issues such as political bias or test whether the gains would persist after extensive RL training on other tasks. Petri is only a proxy evaluation for real-world misalignment, and capabilities that were not measured may also have degraded.
Anthropic Wins Lawsuit as Court Rules U.S. Government Ban Unlawful
行业动态
U.S. District Judge Rita Lin in California ruled that the U.S. government acted unlawfully by restricting federal agencies and defense contractors from using Anthropic products on national-security “supply chain risk” grounds, and ordered the related directives withdrawn. The court found that the measures constituted unlawful retaliation in violation of the First Amendment.
A U.S. federal district court in California ruled that the government acted unlawfully by designating Anthropic a national-security “supply chain risk” and using that designation to restrict federal agencies and defense contractors from using its products, and ordered the related directives withdrawn. The court found that the measures constituted unlawful retaliation in violation of the First Amendment; it also held that the decision to designate Anthropic a supply chain risk was arbitrary and lacked sufficient justification. Anthropic officially welcomed the ruling and said it would continue cooperating with the U.S. government on national security. A separate related lawsuit filed by the company in Washington, D.C., remains pending.