Mistral Large 4 launches, with model weights to be released at the end of the month
要闻
Mistral has opened the public preview API for Mistral Large 4 and plans to release its weights at the end of the announcement month. The model has 1 trillion total parameters and 49 billion active parameters, with preview access available through Mistral Studio.
Mistral launched the public preview of Mistral Large 4 at the time of its announcement and said it would release the weights by the end of that month; the source does not specify the announcement date. The natively multimodal model has 1 trillion total parameters and 49 billion active parameters. According to the company, it was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters, and the preview API on Mistral Studio runs on the same infrastructure. Its training data covers more than 160 languages, including every official EU language. Ahead of the weight release, Mistral is conducting real-world red-team testing with cybersecurity organizations, vetted partners, and state authorities. Participants use a version of the model with reduced content moderation restrictions and expanded cybersecurity capabilities.
According to evaluation results published by the company, ML4 ranks among the top five models globally on the Artificial Analysis Cyber Index. It scored 82% on a test requiring models to reproduce and patch real vulnerabilities in open-source software, and solved 93% of the challenges in Cybench, which comprises 40 security competition exercises. In software engineering, it scored 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4, with a combined Coding Agent Index score of 49.8%. It scored 59.9% on AutomationBench, which covers 657 business workflows, and reached 1,393 Elo on AA-Briefcase, an evaluation of long-horizon knowledge work. The company said it would publish details of the model architecture, additional benchmarks, and its post-training methodology.
Google releases Nano Banana 2.1 with major upgrades to image generation and editing
要闻
Material from Google’s Gemini Image page lists 2 English poster-generation Prompts covering outdoor sports and travel, with detailed requirements for text, colors, and composition. The supplied content provides no model release date or performance data.
Google presents 2 poster-generation Prompts in material from its Gemini Image page; the supplied content gives no date for their appearance and provides no release timing or feature-upgrade description for Nano Banana 2.1. Both Prompts specify English text and require clear legibility, but use different visual layouts. The outdoor sports poster overlaps a circular photograph of a trail runner with large lettering split across lines, combining an off-white background with orange and black geometric elements, coordinates, an elevation badge, and a barcode. The travel poster uses a 1970s cinematic style, with a cobalt-blue background and cream border framing large lettering. A woman seated on top of a camper van obscures some of the letters, while surfboards, film grain, and warm natural lighting complete the composition.
谷歌 open-sources multimodal embedding model EmbeddingGemma 2
要闻
Google has released EmbeddingGemma 2, a 740 million-parameter model offering unified embeddings for text, code, images, video, and audio under Apache 2.0 for on-device retrieval. Google reports that its MTEB Code score rose from 68.76 to 78.68.
Google has released EmbeddingGemma 2, built on the Gemma 4 architecture with 740 million parameters and an Apache 2.0 license that permits commercial use. It maps text, code, images, video, and audio into a shared embedding space, enabling uses such as finding a video clip from a voice memo or searching audio recordings with a text query. Google says the model retains its predecessor’s multilingual text performance while increasing its MTEB Code score from 68.76 to 78.68, a gain of 9.92 points. According to Google, it can support local codebase indexing, semantic code search, and coding agent retrieval.
According to Google, EmbeddingGemma 2 can generate embeddings locally, support fully offline cross-modal search and retrieval, and work with Gemma 4 in on-device RAG pipelines. The two models share a text tokenizer and audio encoder, which Google says can reduce their combined memory footprint when run together. Google provides Instant Media Search and Video Moments Finder examples in Google AI Edge Gallery for semantic media-library search and locating moments in videos, respectively. Google AI Edge Foresight demonstrates combining local file retrieval with contextual reasoning. Developers can also consult the model card, inference and fine-tuning guides, and an article on building search and RAG systems with LiteRT.
OpenAI shares a collection of mathematical research results from its internal frontier models
要闻
OpenAI has published 722 mathematical manuscripts produced by an internal frontier model, organized into 372 result families, with Lean formalizations for some proofs. Verification progress varies across the results, and the company says those without formalization may contain problems.
OpenAI has released mathematical research results produced by an internal frontier model in a GitHub repository, including 722 manuscripts organized into 372 result families and protocols for paper revisions and citations. The repository also contains Lean formalizations of many proofs, 10 summaries of model reasoning, and statistics such as the number of problems attempted; further formalizations will be added as they become available. The release approach drew on advice and published recommendations from the Institute for Advanced Study’s independent advisory group on mathematics and artificial intelligence. According to OpenAI, the vast majority of the results came from the same unreleased internal model using a fixed process. The model was given approximately 4000 problems during evaluation, with each result consuming an average of approximately 3 hours of ChatGPT Pro thinking compute. The company says the results are at different stages of verification and that results without formalization may contain problems. OpenAI also plans to fund workshops and other activities to help people understand major AI-generated results, and says it is working toward responsibly releasing the model that produced them.
OpenAI makes Auto-review free without counting toward subscription usage limits
开发生态
Tibo, who oversees Codex and ChatGPT at OpenAI, announced that Auto-review is now free for all users signing in with a ChatGPT account and does not count against subscription usage limits. The feature uses a second agent to review the main agent’s actions.
OpenAI’s Tibo, who oversees Codex and ChatGPT, announced that all users signing in with a ChatGPT account can now use Auto-review for free without consuming their subscription usage allowance. It can be enabled under Settings → Permissions → auto-review. According to his explanation, the feature assigns a second agent to check every action taken by the main agent during long-running tasks, with the aim of blocking high-risk actions and actions that stray from the user’s original intent. Tibo said the default sandbox settings require individual approval for every action, which can cause decision fatigue unless users spend time configuring specific rules. He described Auto-review as an improvement to that mechanism and recommended using it instead of full-access permissions, saying that the choice no longer involves a trade-off.
OpenAI opens the Decisions API public beta to all developers
开发生态
OpenAI has opened the Decisions API public beta to all developers, allowing applications to select models, tools, or actions. The company says decisions can be up to 10 times faster than with another calling method, with input priced at $0.1 per 1 million tokens.
OpenAI announced that Decisions API has entered public beta and is available to all developers, enabling applications to select models, tools, or actions in near real time. Powered by GPT-6 Luna, the API accepts text and image inputs. Pricing is $0.1 per 1 million input tokens, with charges limited to input and no charges for caching or output tokens. According to OpenAI, decisions can be up to 10 times faster than when calling GPT-6 Luna through the Responses API.
The API offers three output types: Predicates estimates the probability that a statement is true, Choices selects from predefined options and provides a confidence level, and Scores evaluates inputs within a numerical range. According to the company, uses include routing requests to suitable models, tools, or Agents; generating labels and scores for inputs at scale; analyzing images or determining actions from screenshots; and flagging high-risk tool calls.
OpenAI streamlines paid API tiers from five to three
开发生态
OpenAI announced that its five paid API usage tiers will be consolidated into three: Build, Launch, and Grow. The cumulative payment required to qualify for the highest tier falls from $1000 to $500, and existing paid organizations will migrate automatically.
OpenAI announced a change from five paid API usage tiers to three—Build, Launch, and Grow—but the source does not specify an effective date. The change affects the qualification threshold for accessing higher API rate limits: the highest tier, Grow, requires $500 in cumulative API payments, compared with $1000 for the previous highest tier. According to OpenAI, organizations already on paid tiers will move to the corresponding new tiers without manual action. Details of each tier’s qualification requirements and rate limits have been published on the organization limits settings page of the OpenAI platform.
Space Bunny's free period is ending, with OpenCode sponsoring a few extra free days
开发生态
With Space Bunny’s free access period nearing its end, open-source coding Agent OpenCode has announced that it will fund a few more days of access through its own free tier. OpenCode says the model is still being tuned and will undergo further adjustments before its formal debut.
Open-source coding Agent OpenCode announced on X that it will sponsor a few more days of free Space Bunny access in its free tier, extending access as the model’s current free period approaches its end. The announced arrangement concerns OpenCode’s support for free access to Space Bunny, while the model itself remains in a pre-debut adjustment phase. According to OpenCode, Space Bunny is being tuned, and those adjustments will continue before its formal debut.
Anthropic expands CVP, offering advanced networking capabilities across three tiers
开发生态
Anthropic has expanded CVP into three access tiers, offering qualifying security teams advanced cyber capabilities and model access with fewer blocks. The program covers Claude Opus 5.5, Claude Sonnet 5.5, and Claude Mythos 5.1.
Anthropic has launched an expanded Cyber Verification Program (CVP), combining the previous CVP and Project Glasswing into three access tiers based on the scope of security work, with different verification requirements and security controls for each tier. All three provide access to Claude Opus 5.5, Claude Sonnet 5.5, Claude Mythos 5.1, and future models. Defense Access covers defensive tasks, including incident response, malware reverse engineering, and vulnerability analysis and validation. Eligible applicants include organizations protecting systems they own or maintain, critical infrastructure operators, smaller security firms, open-source maintainers, and individual researchers with a record of reporting vulnerabilities. The company says it aims to respond to applications for this tier within a few days.
Red Team Access adds authorized penetration testing and red-teaming to defensive uses and currently accepts applications only from organizations, not individual researchers. Actions that could cause physical harm or mass disruption, such as deploying ransomware or damaging physical systems, remain subject to real-time blocks. The company expects reviews to take a few weeks, with qualifying organizations able to join Defense Access while waiting. Specialized Access has the fewest cyber blocks and is limited to a small number of verified organizations authorized to test safety systems such as power grids and flight operating systems. Anthropic currently conducts in-depth reviews of each organization in cooperation with the US government. Existing Project Glasswing members will move to this tier without needing reapproval for current models. The program requires data retention to monitor cyber misuse, but organizations that already have zero data retention access to Claude Fable 5.1 or Claude Mythos 5.1 can also use CVP with zero data retention until EFS becomes available.
复旦, 腾讯, and 浙大 open-source Prism for joint audio-video generation at native 2K resolution
模型发布
Teams from Fudan University, Tencent Hunyuan, and Zhejiang University have open-sourced Prism, releasing code and preview weights for native 2K joint audio-video generation under the MIT license. According to official experiments, training is 2.5 times faster than full attention.
Teams from Fudan University, Tencent Hunyuan, and Zhejiang University have released Prism, a dynamic sparse attention framework, alongside its technical report, training and inference code, and preview weights under the MIT license. Based on MOVA-360p, the preview model supports native joint audio-video generation and training at 720p, 1080p, and 2K resolutions, covering both text-to-video and image-to-video generation. The music, sfx, and speech tags in prompts control music, sound effects, and dialogue, respectively. The released weights include the stability-oriented preview-alpha and the motion-oriented preview-beta; Prism-pro has not yet been released. The team states that its experiments show Prism training is 2.5 times faster than full attention and also delivers higher generation quality. Hardware requirements are one GPU with 80GB of VRAM for 720p inference, at least 4 GPUs for 1080p and 2K inference, and at least 64 GPUs for native 2K training.
Amazon’s AGI team has released ALoDLM-8B as open source. Its exit gate represents exit schedules with a variational distribution and uses a truncated-geometric prior with c = 0.4 to avoid exponentially many separate rollouts over all schedules.
Amazon’s AGI team has released ALoDLM-8B as open source, with accompanying documentation describing a method for jointly modelling exit depths. Each token commitment changes the context of the unresolved tokens, so exit depths cannot be treated separately; summing over all exit schedules would require exponentially many separate rollouts. According to the team’s explanation, the model instead uses an exit gate to parameterize a variational distribution qϕ over exit schedules and introduces a truncated-geometric prior π with c = 0.4, making the prior favour shallower exit depths.
Perplexity releases a new version of its decision-making model pplx-decider
模型发布
Perplexity has released pplx-decider-v1.1-27b, an open-weight multimodal decision model with a Decision Index score of 61.56 listed in its model card. The Decision API now costs $0.02 per million input tokens, half the price of v1.
Perplexity announced that the updated pplx-decider-v1.1-27b is now available and that Decision API pricing has dropped to half that of v1, at $0.02 per million input tokens. The model has open weights and supports multimodal decision-making; according to its model card, its backbone remains based on Qwen3.8-27B, with 26B parameters and BF16 precision. The model card lists a Decision Index overall score of 61.56, up from v1’s 56.4 and more than 3.5 points ahead of Jev. Perplexity says the model achieved the highest score on the updated Hugging Face Decision Index 0.3 benchmark, attributing the improvement primarily to removing the causal mask and expanding the training data, including tasksource data.
EmpirioLabs releases decision-making model Aplomb 1
模型发布
EmpirioLabs has made its Aplomb 1 decision model available through its API, Playground and released weights, supporting multiple input types and up to 1,000,000 tokens in one request. The company says fast mode processes a document of that length in about 3 seconds, but is less accurate than a full read on whole-document judgments.
EmpirioLabs’ first decision model, Aplomb 1, was available through its API and Playground as of October 6, 2026, with weights released on Hugging Face under the EmpirioLabs Model License. The model accepts text, JSON, images, audio and video in a single request, with the input state and longest question totaling up to 1,000,000 tokens. For an Agent, it returns a probability for every tool and probability distributions for enum and boolean arguments in one request; the Agent fills in free-text arguments. Up to 128 questions can share a state. According to the company, each is answered independently, so adding a question does not change other answers. Enabling abstain also adds the probability that the input lacks the information needed to answer.
According to the company, the API and Playground default to fast long-context mode for states exceeding 131,072 tokens, processing a 1,000,000-token document in about 3 seconds, versus about 111 seconds for a full read. In its long-context tests, fast mode answered all 525 decisions correctly, while a full read achieved 95.2% accuracy. However, on 24 documents designed to test whole-document judgments such as counts, overall tone or confirmation that something never appears, accuracy was 50% and 61%, respectively. Setting "long_context": "full" in a request switches to a full read, and both modes count the entire state toward usage. Fast mode is available only through the company’s API and Playground; the released weights read every token.
OpenAI's dots team says it rolled out a series of improvements within a week of launch
产品应用
OpenAI dots project member Rohan Varma said the team updated performance, notifications, display, and access features based on user feedback within one week of launch. Its next planned work covers voice reliability, control of the Codex app, connections to multiple computers, and local control on Windows.
OpenAI dots project member Rohan Varma said the team released multiple improvements based on user feedback during the week following the dots launch. According to his account, the team adjusted cloud browsing, sidebar loading, and animation behavior, and reduced pauses while waiting for responses. Fixes covered unnecessary mobile notifications when using ChatGPT, character rendering issues in Safari and Firefox, and Outlook setup getting stuck. The fixes also addressed clipped buttons in the approval interface, oversized Gmail previews obscuring approvals, and enterprise access issues on the Web and when starting voice. The team also adjusted the stability of the Web security banner, clarified plan guidance, and added a way to retry when activity fails to load. Planned work includes continuing to improve voice reliability, expanding dot’s control over the Codex app, enabling dot to connect to multiple computers, and improving support for local computer control on Windows.
ChatGPT launches the Meetings plugin to automatically take meeting notes
产品应用
ChatGPT has launched the Meetings plugin in beta to create meeting notes, personalized summaries, and next steps, saving them to ChatGPT Space. The feature is currently available only to Pro and Business users through the macOS desktop app.
ChatGPT now offers the Meetings plugin in beta for Pro and Business users in the ChatGPT desktop app for macOS. The plugin takes notes during meetings and uses ChatGPT’s knowledge of the user and their previous collaborative work to generate personalized summaries and next steps, saving them to ChatGPT Space. Users can keep notes private or share them with their team, and can subsequently ask ChatGPT to update project plans or draft follow-up content. Installation requires downloading the desktop app first, then searching for Meetings in the plugin directory. The company says support for Enterprise is coming soon.
Claude quietly adds Simplified and Traditional Chinese interfaces to its web and desktop apps
产品应用
Multiple users found that Anthropic’s Claude web and desktop apps now offer Simplified and Traditional Chinese interfaces, both manually selectable in settings. Other users found that subscription upgrade pages in some regions allow Chinese addresses and list support for UnionPay International cards, although a payment attempt was rejected.
According to reports from multiple users, Anthropic’s Claude web and desktop apps have added Simplified and Traditional Chinese interfaces. Both allow users to change the interface language manually in settings, and some users found that their interfaces had switched to Chinese automatically. At roughly the same time, other users upgrading subscriptions on the web found that the page in some regions allowed them to select a Chinese address and displayed support for UnionPay International cards. Users had attempted payment, but orders were rejected; both this payment outcome and the card support information displayed on the page appeared in user reports.
Claude for Google Workspace is now in public beta on all paid Claude plans, allowing users to modify files with Claude inside Docs, Sheets, and Slides. Three beta connectors released alongside it support creating and editing Google files from Claude conversations.
Claude has opened the public beta of Claude for Google Workspace to all paid plans and simultaneously released beta Google Docs, Sheets, and Slides connectors. The two entry points support working in a file’s sidebar or creating and editing Google files from a Claude conversation. According to the company, the sidebar can read the current file and selected text, cells, or slides, and modify content directly. Docs supports targeted rewrites and formatting changes, as well as suggestion cards that users can accept or dismiss. Sheets supports formulas, pivot tables, native charts, and additional tabs, and can write data back after processing it with Python. Slides can generate slides using existing layouts and themes, then check for overlapping elements, content extending beyond slide boundaries, or text that is difficult to read.
The company says the default Ask before edits mode waits for approval before each change, while Accept all edits continues through a task and applies changes. After users sign in with their Claude account, the sidebar can use the account’s models, connectors, and skills. Enterprise controls including the Compliance API, CMEK, and OpenTelemetry audit export also apply to the add-on. Users can install it from Google Workspace Marketplace and open it through Extensions > Claude > Open Claude; administrators can deploy it to an entire domain or selected groups through the Google Admin console. Editing files from a conversation requires enabling the corresponding connectors, with an owner or primary owner required to enable them first on Team and Enterprise plans. File access follows existing Google sharing permissions, and supported setups can open files in a pane beside the conversation.
谷歌's Docs and cloud drive add native .md file support with real-time editing and commenting
产品应用
Google Docs and Google Drive have added native Markdown support, allowing users to open, edit, and comment on .md files directly. The rollout began on October 5, 2026, and some users may need to wait up to 15 days.
Google began rolling out native Markdown support for Google Docs and Google Drive on October 5, 2026, with some users potentially waiting up to 15 days. Opening an .md file displays its rendered contents, including clickable links and structured tables, and users can edit and comment in real time. The source does not state that files need not be converted to another format beforehand. According to Google, the feature enables AI Agents to collaborate on document editing and supports synchronization with Gemini Notebook. Google employee Chandu Thota described Markdown as a common language between humans and AI Agents.
ElevenLabs introduces Architect, a conversational Agent expert
产品应用
ElevenLabs has released Architect, a feature built into ElevenAgents and now available in Alpha, for creating and improving AI Agents through voice or text. The company says it can analyze thousands of conversations and deliver proposed changes that have already been built and validated in simulation tests.
ElevenLabs has released ElevenAgents Architect in Alpha, allowing users to create and improve AI Agents through voice or text conversations within ElevenAgents. The feature can also be called from Claude, Claude Code, ChatGPT, Cursor, and Grok Bot, carrying the Agent’s full context. According to the company, Architect understands all ElevenAgents configuration options and covers best practices for knowledge bases, prompts, workflows, voice settings, procedures, guardrails, tools, and simulation tests. It can analyze conversation records, identify the root causes of problems, and propose changes, as well as search thousands of conversations for improvements a team may have missed. The company says every proposal it returns has already been built and validated in simulation tests.
A practical guide to Claude Code cloud sessions is released
技术与洞察
claude.dev published a guide to Claude Code cloud sessions, covering parallel tasks on separate machines and plan billing rules. Its three example tasks started within 16 seconds of one another and all finished 87 seconds after the first began, with no separate charge for cloud machines.
claude.dev published a practical guide to Claude Code cloud sessions covering task isolation, launch options, and usage restrictions; the source does not specify a publication date. According to the site, each task receives its own virtual machine, with the repository cloned onto a new branch and the environment set up, and its output can become a pull request. Sessions can be launched from claude.ai/code, the Claude mobile app, the Desktop app, a terminal, and Slack. Pro, Max, Team, and Enterprise plans include cloud sessions without a separate machine charge, but sessions share Claude Code usage limits, and some organizations require their owner to enable them first. The guide tested three parallel tasks on a sample Node API repository, addressing a flaky test, outdated API documentation, and string concatenation in the logger. They started within 16 seconds of one another, ran for 61, 65, and 72 seconds respectively, and all finished 87 seconds after the first started. Each session first reconstructed the repository from files supplied in its prompt, a step accounting for roughly one-third to just over half of each run.
Sierra joins Meta, Stripe, 沃尔玛, and others to develop an open standard for personal Agents
技术与洞察
Sierra is developing Personal Agent Protocol with Meta and partners including Stripe and Walmart to govern how personal Agents interact with businesses under user-granted permissions. The standard is open for anyone to implement, and the company plans to publish the v0.1 specification.
Sierra announced that it is developing the open standard Personal Agent Protocol with Meta, Genesys, Instinct, Rocket, Shopify, Stripe, and Walmart; the source does not specify the announcement date. According to the company, the standard is intended to handle authentication, let users determine their personal Agents’ access permissions, and allow businesses to set operational boundaries and gain visibility into Agent activity. Interactions begin on a business’s website, where an Agent discovers services and connection methods before starting a guest session on the user’s behalf to check product availability or return policies. When account access is needed, users can sign in on the business’s page or use credentials already configured with their personal Agent, choosing read-only or write access. Sessions are built on OAuth and preserve interactions across channels before and after sign-in. Businesses can offer websites, APIs, or their own Agents as connection routes. Anyone can implement the standard; the company also plans to publish the v0.1 specification, hold design workshops, and provide a reference implementation.
OpenAI teams up with Ironclad to train AI agents on contract workflows
技术与洞察
OpenAI partnered with Ironclad to turn contract workflows into 11 tasks for AI agent training and evaluation. Official evaluations put GPT-6 Astra’s average score 32% above GPT-5.6 Sol’s, but the results apply only to these research tasks.
OpenAI announced a research partnership with contract management software company Ironclad to use complex contract workflows for AI agent training and evaluation, with Ironclad serving as its first partner. The companies defined 11 tasks spanning legal, sales, and procurement work, including setting up nondisclosure agreements and creating procurement approval workflows. Each task is assessed against 8 to 50 criteria. According to the announcement, experienced users are estimated to need an average of 30 to 40 minutes to complete each task. The partnership aims to improve models’ ability to use professional software to handle complex business problems.
According to the official explanation, GPT-6 Astra is the first frontier model trained on these tasks. In the research evaluation, its average score was 55.0%, compared with 41.6% for GPT-5.6 Sol, a 32% higher score. Estimated time per attempt was 19.2 minutes and 37.0 minutes, respectively, making the former 48% lower than the latter. OpenAI stated that the results cover only the 11 tasks described above and that the times are simulation-based estimates, not measurements of time actually saved by customers. OpenAI is also inviting a small number of software companies to apply for similar research partnerships.
GitHub rebuilds its Git infrastructure for Agent-driven development
技术与洞察
GitHub is rebuilding its Git infrastructure to handle workloads from developers and AI Agents working concurrently. Over the past year, monthly Git events rose from 218.2 billion to 473.3 billion, while the new architecture delivered up to 35 times the previous write throughput in internal tests.
GitHub is overhauling its Git infrastructure while keeping the platform running to support concurrent work by developers and AI Agents, without requiring users to change their existing build workflows. Over the past year, monthly Git events increased from 218.2 billion to 473.3 billion, more than doubling; September recorded 7.38 billion commits, more than five times the figure a year earlier; and pushes grew by 4.9 times year over year to 3.35 billion per month. According to GitHub, the existing Spokes architecture couples persistence with read scaling, so adding read replicas slows writes and limits capacity for highly active repositories. The revised architecture restricts coordination to reference updates that require consistency, moves maintenance such as compression and garbage collection out of the serving path, stores authoritative data in Azure Blob Storage, and supplies read capacity through lightweight cache nodes that scale with traffic. Internal benchmarks showed write throughput of up to 35 times the previous level, with read capacity able to scale independently. GitHub says this foundation is already being implemented and that subsequent articles in the series will describe the architecture in detail.
SemiAnalysis testing finds Anthropic subscriptions deliver roughly five times the value of OpenAI's
技术与洞察
SemiAnalysis has launched an AI subscription monitoring dashboard for Tokenomics Model subscribers, covering providers including OpenAI and Anthropic. It measures allowance consumption by model and token type rather than comparing value solely by monthly fees.
SemiAnalysis has launched Subscriptions Dashboard exclusively for Tokenomics Model subscribers to track AI subscription allowances; the source does not specify a launch date. Coverage includes OpenAI, Anthropic, Meta, SpaceXAI, Cursor, Cognition, Z.ai, MiniMax, and Moonshot. According to SemiAnalysis, the API-equivalent value of the same $200-per-month Claude plan varies with the model and workload, so its monthly fee alone is insufficient to assess its value. The tests isolate Input, Cache write, Cache read, and Output, recording the token counts billed by the provider and changes in its usage meter. Allowance consumption rates are then calculated from complete intervals between meter increments, excluding incomplete intervals at the beginning and end. The article also states that providers may change limits publicly or alter them without an announcement by adjusting credit consumption costs, meaning a single measurement cannot capture ongoing changes.
a16z adds a spending ranking to its Top 100 Consumer AI list for the first time
技术与洞察
a16z added a U.S. consumer spending ranking to its seventh Top 100 Consumer AI Apps report, covering products missed by traffic rankings. In the sample, the top 10% of spenders accounted for roughly half of spending, while ChatGPT had three times as many U.S. paid subscribers as either Claude or Gemini.
a16z introduced a spending ranking with the publication of its seventh Top 100 Consumer AI Apps report, using U.S. consumer card spending observed by YipitData. The new ranking includes dozens of vendors absent from the web and mobile lists, including desktop apps and Agents within existing messaging platforms. The data comes solely from U.S. sample panels, not a census. The existing rankings continue to select 50 web products by monthly visits from Similarweb and 50 mobile apps by monthly active users from Sensor Tower. This edition had 11 first-time entries, the fewest across all seven editions.
According to a16z’s explanation of the sample data, the top 10% of spenders accounted for roughly half of observed spending, with these users more inclined to buy coding, productivity, and creative tools. YipitData’s panel showed that ChatGPT had three times as many U.S. consumer paid subscribers as either Claude or Gemini, and only 8% of U.S. ChatGPT subscribers also subscribed to Claude. Among Claude’s consumer paying users, 7.3% chose Max, which starts at $100 per month. The shares on Google’s and ChatGPT’s corresponding $100-per-month subscriptions were 1.3% and 1.1%, respectively.
Nous Research launches Hermes Index, a leaderboard for agent models
技术与洞察
Nous Research has released Hermes Index, presenting models’ Agent capabilities through mean scores and mean cost per task across four evaluation suites. The page also describes Hermes Bench, which contains 150 tasks across 25 categories and checks the files and state left behind after tasks end.
Nous Research has published the Hermes Index model leaderboard on its portal page, aggregating results from four evaluation suites run inside Hermes Agent into a mean score and mean cost per task. According to the page, higher scores indicate better performance, and users can select up to four models to compare scores and costs across the suites. Each model on the line in the score-versus-cost chart scores higher than every model with a lower cost. The page separately lists the test setup for Hermes Bench: 150 tasks across 25 categories covering Hermes skills, research, diagrams and art, memory, tool use, and safety. Each Agent begins its tasks in a workspace containing real files, some tasks include follow-up turns, and the grader checks the files and state ultimately left behind.
WHO's Africa team uses Google Earth AI to rapidly identify Ebola exposure areas
技术与洞察
During the Ebola outbreak in the Democratic Republic of Congo, the WHO Regional Office for Africa and a local research institution used two Google Earth AI research prototypes. Google says the tools have helped teams identify remote communities at risk of Ebola in just minutes.
The Dakar Emergency Preparedness and Response Hub at the WHO Regional Office for Africa and the Epidemic Modeling and Intelligence Unit (UMIE) at the Democratic Republic of Congo’s National Institute of Biomedical Research (INRB) used two Google Earth AI research prototypes during the country’s Ebola outbreak; the source provides no specific date. The Geospatial Reasoning agent supports spatial mapping through natural-language conversations, while the planetary prediction engine is used for autonomous disease forecasting. Google says these tools combine environmental signals, satellite imagery, mobility data and foundation models to supplement existing health information and bridge reporting gaps. Its Population Dynamics Foundation Model (PDFM) integrates aggregated search trends, population mobility and environmental patterns. The associated embeddings are commercially available in Preview through Population Dynamics Insights on Google Maps Platform. Researchers can request no-cost access for select use cases, and eligible organizations can apply for Google Earth credits through GMP Public Programs. Google.org has also provided funding to INRB to help modernize local testing and disease surveillance capabilities.
Opus 5.5 Agent identifies two room-temperature magnetic semiconductors in three days
技术与洞察
Vals AI said more than 90 Claude Opus 5.5 Agents ran hundreds of cloud-based quantum-mechanical simulations in 3 days, identifying two room-temperature magnetic semiconductor candidates. Their bandgaps and spin-filtering capabilities have not yet been measured.
LLM evaluation organization Vals AI reported in a blog post that more than 90 Claude Opus 5.5 Agents ran hundreds of quantum-mechanical simulations during a 3-day cloud-computing effort, finding two room-temperature magnetic semiconductor candidates for spintronic storage. According to the organization, both are Luttinger compensated magnets with zero net magnetic moment that can nevertheless sort electrons by spin. The existing material KV[Cr(CN)₆] was first synthesized in 1999, when measurements showed that its magnetic order persisted up to 376 K. The latest calculations predict a bandgap of approximately 2.1 eV, with both band edges in the same spin channel. The other candidate, YBaMnFeO₅, is a new design with a predicted bandgap of 2.35 eV and magnetic order persisting up to approximately 420 K. However, the checkerboard atomic arrangement required during synthesis is easily disrupted, potentially making it difficult to produce. Vals AI emphasized that these results are based on ideal crystals and that neither the bandgaps nor spin filtering have been measured. The next step is to resynthesize KV[Cr(CN)₆] and take measurements; all computational data and a one-click verification program have been published on GitHub.
Pentagon confirms it has stopped using all Anthropic products
行业动态
The US Department of Defence confirmed to the BBC that it has stopped using all Anthropic products, after previously setting a six-month phase-out deadline. Multiple people familiar with the matter said Claude had continued to support intelligence analysis and military operations against Iran and was embedded in a Pentagon data platform.
The US Department of Defence confirmed in a statement responding to the BBC that it has stopped using all Anthropic products, without specifying the exact cessation date or explaining the delay. Defence Secretary Pete Hegseth had previously designated Anthropic a supply-chain risk on national security grounds and set a six-month deadline to phase it out. The dispute began when the Pentagon demanded the removal of Claude's safety restrictions so the military could use the company's AI products without limits. Anthropic refused over concerns about mass surveillance and autonomous weapons and sued the Trump administration to overturn the designation. While the litigation has continued, the Pentagon has signed contracts with companies including Google, xAI and OpenAI.
According to multiple people who spoke to the BBC, Claude was still being used before the cessation statement for research, analysis, intelligence gathering and military operations against Iran, and was embedded in the Palantir-operated Maven Smart System. Maven is the Pentagon's primary platform for organising intelligence and other data. Sources said analysts used Maven to access and organise visual data, including satellite imagery and drone footage, before passing it to Claude and other large language models for processing, including identifying potential military targets. Lauren Kahn, a senior research analyst at Georgetown's Center for Security and Emerging Technology, said Claude's deep integration with Maven made rapid removal difficult. An Anthropic spokesperson declined to comment on the cessation statement.
Anthropic is expanding Claude Startups for teams building companies on Claude, offering up to $7,000 in products and credits and up to $45,000 in tool discounts and credits. Eligibility depends on when a startup was founded or received funding.
Anthropic announced an expansion of Claude Startups, opening applications to more founders building companies on Claude. According to the company, members can receive up to $7,000 in Claude products and credits, with benefits including one year of the Claude Team plan, API credits, and office hours with Anthropic’s Applied AI team. The new Claude Startup Stack covers tools for sales, design, and data engineering, among other functions, giving members access to up to $45,000 in discounts and credits from companies building products with Claude. The program also provides support for reaching customers, addressing technical problems, and connecting with other founders. Startups founded within the last five years or funded within the last two years are eligible to apply; applications are available through the Claude Startups page.
The supplied YouTube material’s title says OpenAI plans to merge the Chat and Work toggle, but its body contains only webpage configuration code. It provides no announcement timing, effective date, key figures, or product explanation that could verify the title’s claim.
The supplied YouTube material is titled around OpenAI’s plan to merge the Chat and Work toggle, but does not specify when the plan was announced or when it would take effect. The readable body consists of YouTube page initialization and feature configuration code, with no video transcript, product announcement, or official OpenAI explanation. Consequently, the material supports only reporting the merger plan as stated in the title; it does not establish the specific changes, scope, rollout schedule, or whether the plan has already been implemented.
OpenAI reportedly in talks to raise $30 billion in new funding
前瞻与传闻
OpenAI is discussing a $30 billion funding round with UAE funds including MGX, with their combined contribution under discussion reaching up to $10 billion, according to people familiar with the matter. BlackRock is also in talks to participate, and the financing details could still change.
OpenAI is in talks with several UAE investment funds and BlackRock to raise $30 billion for the ChatGPT developer, according to people familiar with the matter. The UAE funds include Abu Dhabi-based MGX. One person said the funds plan to form an investment syndicate to participate in the round, with a combined contribution of up to $10 billion under discussion. BlackRock is discussing participation alongside that syndicate. The people providing the information requested anonymity and said the fundraising remains ongoing, with specific arrangements subject to change.
DeepSeek reportedly raising at least 80 billion yuan, led by 腾讯 and 宁德时代
前瞻与传闻
DeepSeek is nearing completion of a funding round of at least RMB 80 billion, Bloomberg reported, citing people familiar with the matter. Tencent and CATL rank among the largest investors by committed funding, while signed term sheets indicate the final total could approach RMB 100 billion.
DeepSeek’s new funding round is nearing completion, with at least RMB 80 billion to be raised, according to Bloomberg, citing people familiar with the matter. The company initially sought about RMB 50 billion, but investor demand exceeded expectations after the release of its latest AI model; signed term sheets indicate that the final total could approach RMB 100 billion. Tencent and CATL are among the investors with the largest funding commitments in the round. Bloomberg said the financing would pave the way for DeepSeek’s planned IPO in early 2027. Details have not yet been made public, and Reuters said DeepSeek, Tencent and CATL did not immediately respond to requests for comment.
月之暗面 reportedly completes pre-IPO funding at a valuation of about $50 billion
前瞻与传闻
Moonshot has completed its final pre-IPO private funding round at a valuation of about $50 billion, according to Bloomberg, citing people familiar with the matter. It plans a Hong Kong IPO in the first quarter of 2027 to raise up to $5 billion, with preliminary meetings to gauge investor interest possible as early as October 2026.
Bloomberg reported on October 6, 2026, citing people familiar with the matter, that Moonshot had completed its final pre-IPO private funding round at a valuation of about $50 billion and planned a Hong Kong IPO in the first quarter of 2027 to raise up to $5 billion. According to the report, the company had already submitted a confidential application to the Hong Kong stock exchange and had begun preparing to gauge investor interest, with preliminary meetings possible as early as October 2026. Bank of America is serving as overall coordinator for the listing, while CICC, Deutsche Bank and Goldman Sachs are serving as sponsor banks.
可灵 AI reportedly selects investment banks for a planned Hong Kong IPO to raise at least $1 billion
前瞻与传闻
Kuaishou’s Kling AI has selected underwriting banks for a planned Hong Kong IPO seeking at least US$1 billion and is working with CICC, Goldman Sachs and UBS on the offering, Bloomberg reported, citing people familiar with the matter. Discussions remain ongoing, and the fundraising amount and listing timetable could change.
Kuaishou Technology’s Kling AI had selected underwriting banks by the time the information was disclosed and was preparing a Hong Kong IPO targeting at least US$1 billion, Bloomberg reported, citing people familiar with the matter. The source article puts that amount at approximately RMB6.714 billion. Banks involved in preparations for the potential share offering include CICC, Goldman Sachs Group and UBS Group; the people cited said discussions were continuing and the fundraising amount and listing timetable could still change. Founded in 2024, Kling AI provides AI video services and is pursuing commercialization. It raised US$2.8 billion in a funding round, equivalent to approximately RMB18.801 billion according to the source article, with investors including Alibaba, Tencent and Baidu. The article also states that, following completion of that round, Kling AI’s pre-money valuation was approximately US$15 billion, equivalent to approximately RMB100.717 billion. Its competitors include ByteDance’s Seedance, as well as ShengShu Technology and PixVerse, both of which also plan Hong Kong IPOs.