Daily AI Digest

2026-10-09

Source:橘鸦 AI 早报 · 30 items

2026-10-09
2026-10-09 2026-10-08 2026-10-07 2026-10-06 2026-10-05 2026-10-04 2026-10-03 2026-10-02 2026-10-01 2026-09-30 2026-09-29 2026-09-28

OpenAI introduces Ultrafast mode for GPT-6.1 Sol

开发生态

OpenAI introduces Ultrafast mode for GPT-6.1 Sol

OpenAI is rolling out Ultrafast mode for GPT-6.1 Sol across API, Codex, and ChatGPT Work. The company says it runs at up to 8 times the speed of Sol Standard, with API pricing of $12 per million input tokens and $60 per million output tokens.

OpenAI’s Ultrafast mode for GPT-6.1 Sol is rolling out across API, Codex, and ChatGPT Work and is already available in all supported regions, with support for US and EU data residency. API access is open to all users, priced at $12 per million input tokens and $60 per million output tokens. Access in Codex and ChatGPT Work covers Pro 500, as well as eligible pay-as-you-go Enterprise and quota-billed Edu subscription plans; Enterprise access requires an administrator to enable permissions. According to OpenAI, Ultrafast runs at up to 8 times the speed of Sol Standard and has an intelligence level close to Astra.

Read original →

OpenAI introduces faster steering for Codex in its desktop app

开发生态

OpenAI introduces faster steering for Codex in its desktop app

OpenAI is updating steering for Codex in the ChatGPT desktop app to make it respond faster to additional instructions sent while a task is running. Users can choose whether follow-up messages adjust the current task or wait until the next run.

OpenAI is rolling out faster-response Codex steering in the ChatGPT desktop app, allowing users to send messages during a running task to adjust its execution. Users can configure how follow-up messages are handled under Settings → General → Follow-up behavior: steer the current run or wait until the next run. This lets users correct Codex’s approach, supply missing information, or change the task’s direction. Tibo, a member of the OpenAI Codex team, said the improved steering takes effect immediately, enabling the model to respond faster to adjustments and correct its direction during execution to avoid wasted effort.

Read original →

阶跃星辰 offers Step 5 Preview free for one week across multiple platforms

开发生态

阶跃星辰 offers Step 5 Preview free for one week across multiple platforms

StepFun announced that Step 5 Preview is available on OpenRouter, with OpenCode, Cline, NousResearch, KiloCode and other platforms progressively offering one week of free access starting on the announcement day. Model resources and API documentation were released alongside it.

StepFun announced that Step 5 Preview has launched on OpenRouter and that OpenCode, Cline, NousResearch, KiloCode and other platforms are progressively opening one week of free access from the day of the announcement. The accompanying resources include a model overview, benchmarks, demonstrations, technical specifications and API documentation, and users can switch directly to the model within their existing workflows. According to the company, the model offers flagship-level intelligence for agentic and professional work at substantially lower task costs; the source provides no specific cost figures or percentage reductions.

Read original →

Glyph Cluster launches on Vercel AI Gateway, free during its anonymous phase

开发生态

Glyph Cluster launches on Vercel AI Gateway, free during its anonymous phase

Vercel is offering Glyph Cluster on AI Gateway free during its stealth period to Pro and Enterprise teams that have purchased Gateway credits. The model supports coding reasoning, but ZDR is unavailable, and inputs and responses may be used for training.

Vercel has added Glyph Cluster to AI Gateway in its stealth phase, during which Pro and Enterprise teams with purchased AI Gateway credits can use it for free. Its model name is stealth/glyph-cluster, and it can be accessed through the AI SDK, OpenAI-compatible Chat Completions and Responses APIs, or a coding Agent connected to AI Gateway. For tools such as Claude Code, Codex, and Cursor, the connection process involves installing the latest Vercel CLI, running setup, and then selecting the model in the Agent.

According to the official description, Glyph Cluster is a reasoning model for coding and long-context analysis that can handle planning, synthesis, quantitative reasoning, and comparisons of material across large inputs. It can also review code, explain failures, debug issues, propose changes, and analyze implementation options. The model supports function calling and streaming responses, but accepts only text, not images or files. Tool use is restricted to user-defined function tools, and structured outputs are not supported. ZDR is unavailable, and prompts and responses sent through the model may be used for training and model improvement. Users can also try it in the model playground.

Read original →

Claude introduces two experimental features: Dashboards and Motion

产品应用

Claude introduces two experimental features: Dashboards and Motion

Claude announced two beta features, Dashboards and Motion, available on paid plans and on Team and Enterprise, respectively. Docs, Slides, and Design also left beta and are available on every plan, including Free.

Claude made Dashboards and Motion available in beta at the time of this announcement and moved Docs, Slides, and Design out of beta. Dashboards is available on paid plans, Motion on Team and Enterprise, and the other three features on every plan, including Free. According to the company, users have created more than 45 million documents, presentations, and designs in Claude. Dashboards and Motion are disabled by default for Enterprise; administrators can enable them under Organization settings > Artifacts.

According to the company, Dashboards connects to data platforms such as BigQuery, Databricks, and Snowflake, as well as Salesforce, to generate dashboards from natural-language questions and update them as the data changes. Users can inspect queries and each chart’s last refresh time, or send dashboards to tools such as Amplitude, Grafana, and Hex for further analysis. Motion uses code to animate text, charts, shapes, and images, supports changes in an editor or through requests to Claude, and exports MP4 files. The company says Motion does not use a video generation model and does not generate footage or AI-generated people. Its output can also be further edited in tools such as Adobe, Descript, and Runway.

Read original →

Google Cloud unveils Gemini work agents

产品应用

Google Cloud unveils Gemini work agents

Google Cloud unveiled Gemini, a unified work Agent, at Gemini at Work 2026, bringing question answering, content generation, and code execution into a single API. The company says nearly 80% of Google Cloud customers already use its AI products.

Google Cloud announced Gemini, a unified Agent for work, at Gemini at Work 2026, without disclosing a specific launch date in the article. According to the company, users can submit task objectives through a single interface and a single API, with Gemini planning the work, invoking tools, and connecting to enterprise systems to handle questions, knowledge work, image and media generation, and code writing and execution. It can also work directly within Gmail, Drive, Docs, Slides, Sheets, Chat, and Calendar, carrying over the same memory, skills, and control settings.

According to the company, Gemini executes tasks persistently in the cloud and retains memory and context across devices and channels, allowing work that takes hours or days to continue after users close their computers. It can create temporary sub-agents with separate identities to coordinate parallel or sequential workflows. The underlying model is separate from the Agent, and it can currently orchestrate Gemini-family models and Anthropic’s Claude models. The company also announced governance and cost-control mechanisms including identity and permission management, secure sandboxing, network gateways, Smart Routing, and real-time spending caps. It says nearly 90% of Fortune 100 companies use Gemini Enterprise.

Read original →

腾讯 launches a standalone file browser for WorkBuddy

产品应用

腾讯 launches a standalone file browser for WorkBuddy

Tencent WorkBuddy has launched a standalone file browser that supports viewing and editing multiple local files in tabs within one window, with a right-hand AI chat panel for processing their contents. According to the company, local viewing and editing use no credits, while AI processing is charged according to task complexity.

Tencent WorkBuddy has launched a standalone file browser, accessible by double-clicking a file in the operating system or selecting WorkBuddy through the right-click menu’s open-with option. A single window supports multiple files, each in its own tab. Files opened by double-clicking locally are added to the currently active file window, and closing the last tab closes the entire window. Unsupported file types do not show an option to open them with WorkBuddy. Once opened, files can be viewed and edited directly, with the top toolbar offering different functions depending on the file type. Closing a tab or window with unsaved changes triggers a prompt to save or discard them.

The AI chat panel on the right is collapsed by default and can be expanded using the button in the upper-right corner. Using AI editing also expands it automatically. Users can ask AI to summarize, rewrite, polish or extend the current file, or select and quote text in a question so that revisions are written into the document. According to the company, the browser reads only files that users actively open and does not traverse or scan their disks. Local viewing and editing use no credits, while AI processing consumes credits according to the actual complexity of the task. File contents are used only to process the current task and are not shared with third parties unless users actively share them or upload them to third-party storage.

Read original →

JetBrains open-sources Mellum2.1, focusing on reinforcement learning for coding agents

模型发布

JetBrains open-sources Mellum2.1, focusing on reinforcement learning for coding agents

JetBrains has open-sourced Mellum2.1 for coding Agents, with 12B total parameters and 2.5B active parameters under the Apache 2.0 license. The company says the update primarily uses reinforcement learning to improve the model’s ability to work with code inside repositories.

JetBrains has released Mellum2.1 on Hugging Face; the source does not specify an exact release date. The model retains Mellum2’s mixture-of-experts architecture, with 12B total parameters and 2.5B active parameters, and is open-sourced under the Apache 2.0 license for coding Agents and sub-agents running on users’ own hardware. This update focuses on post-training, primarily reinforcement learning. According to the company, training involved millions of sandboxed runs across thousands of real environments, enabling the model to explore codebases, edit files, and check its changes. JetBrains compared it with Mellum2, Qwen3.5-9B, and Gemma 4 E4B using the same evaluation setup. The company says agentic coding showed the largest improvement over Mellum2; under heavy load, throughput is almost twice that of Qwen3.5-9B, while MTP makes single-request processing about 1.6 times faster. GGUF builds for llama.cpp, Ollama, and LM Studio, along with the MTP head for speculative decoding in vLLM, have yet to be released.

Read original →

Hugging Face open-sources gene annotation model Carbon-A

模型发布

Hugging Face open-sources gene annotation model Carbon-A

Hugging Face has released the gene annotation model Carbon-A and its accompanying database, reporting 566 million new gene candidates across 22,617 species. The model, training data, and technical report are also available, and some predictions have undergone experimental validation.

Hugging Face has released the Carbon Annotation Database and the gene annotation model Carbon-Annotator (Carbon-A); the source does not specify an exact release date. According to the official announcement, the model identified 566 million new gene candidates across 22,617 species, thousands of which had not previously been studied. The database, model, training data, and technical report are available in HuggingFaceBio’s Carbon Annotation Database Collection. The team collaborated with ActiveSite and UCSD to experimentally validate some newly predicted genes.

Carbon-A has 1.2 billion parameters, was trained on RefSeq annotations, and predicts protein-coding regions directly from DNA. It uses a context window of 98,304 base pairs and outputs nucleotide-resolution predictions on both strands, using one model across mammals, other vertebrates, invertebrates, plants, fungi, and protists. According to the official report, it achieved a macro-averaged nucleotide F1 of 0.944 across 42 benchmark genomes and outperformed the evaluated baselines at the nucleotide, exon, and gene levels. On Tetrahymena thermophila, which uses a nonstandard genetic code, its nucleotide F1 was 0.960.

Read original →

LightOnAI open-sources LightOnOCR-3

模型发布

LightOnAI open-sources LightOnOCR-3

LightOnAI has open-sourced LightOnOCR-3 to turn complex documents into structured content that applications can use directly. According to the company, the model family supports layout element localization, image descriptions and chart data extraction alongside text transcription; the earlier LightOnOCR has exceeded 4.5 million downloads on Hugging Face.

LightOnAI has open-sourced LightOnOCR-3, a family of end-to-end OCR models for converting complex documents into structured content that applications can use directly; the source does not specify a release date. According to the company, the models go beyond text transcription to identify and locate layout elements, generate image descriptions, and extract data from charts and scientific figures. LightOnOCR-3 builds on the earlier LightOnOCR, which has exceeded 4.5 million downloads on Hugging Face; that download figure refers to LightOnOCR, not the newly released LightOnOCR-3.

Read original →

Odyssey unveils world model Odyssey-3

模型发布

Odyssey unveils world model Odyssey-3

Odyssey has released its Odyssey-3 world model and opened a research preview. According to the company, Pro achieved the highest reported score of 66.1 on Physics-IQ Verified’s video-to-video test, while the model can also support environment simulation and control-policy training.

Odyssey has released the Odyssey-3 world model, with its research preview now available. Built as an autoregressive diffusion transformer, the model uses previous observations and the latest inputs to predict object motion, interactions, and how situations change over time. Developers can generate environments from a prompt, move through them or introduce events during generation, and observe the model’s responses. The preview supports first-person and third-person navigation as well as independent camera movement. According to the company, Odyssey-3 Pro achieved the highest reported score of 66.1 on Physics-IQ Verified’s video-to-video test and scored 54.7 on its image-to-video test. The benchmark evaluates model predictions using videos of real physical experiments across fluid dynamics, optics, solid mechanics, magnetism, and thermodynamics.

In the company’s evaluation using WorldMark’s own captions and the mean of its 13 metric scores, Odyssey-3 ranked 1st in 3 of the 4 categories: first-person stylized, third-person real, and third-person stylized environments. The company also states that applying the model to a specific physical system still requires evaluating the behaviors needed for that machine and its tasks. Adaptation methods include training an action decoder or policy on paired observations and actions. According to the company, the model completed manipulation tasks with only tens of hours of robot demonstrations and exhibited recovery behaviors not present in those demonstrations, including reorienting a gripper after a missed grasp and retrieving an object dropped in an unusual position. For driving on real roads in India, the team froze the Odyssey-3 backbone and trained a policy on just 20 hours of driving data, enabling closed-loop driving by predicting waypoints ahead of the vehicle.

Read original →

Grok Imagine Video 1.5 Lite launches at $0.14 per second for 1080p

模型发布

Grok Imagine Video 1.5 Lite launches at $0.14 per second for 1080p

Grok Imagine has released Video 1.5 Lite through its API, supporting text-to-video and image-to-video generation. Pricing is per second and varies by resolution: $0.02 for 480p, $0.03 for 720p, and $0.14 for 1080p.

Grok Imagine announced that its Video 1.5 Lite video generation model is now available on the Grok Imagine API, offering two functions: text-to-video and image-to-video generation. Charges depend on video resolution and are calculated per second: $0.02 per second for 480p, $0.03 per second for 720p, and $0.14 per second for 1080p. The source does not specify an exact launch date or list prices for resolutions other than these three.

Read original →

Ghostty creator releases terminal protocol OSC 7501

技术与洞察

Ghostty creator releases terminal protocol OSC 7501

Read original →

Study finds periodic weaknesses in DeepSeek-V4's long-context processing

技术与洞察

Study finds periodic weaknesses in DeepSeek-V4's long-context processing

Read original →

Epoch AI launches a real-world work evaluation

技术与洞察

Epoch AI launches a real-world work evaluation

Read original →

Human Mathematical Association calls for a boycott of OpenAI

技术与洞察

Human Mathematical Association calls for a boycott of OpenAI

Read original →

谷歌's AMIE research published in The Lancet

技术与洞察

Read original →

Updated Claude usage policy: persistent abuse of Claude may lead to a ban

行业动态

Updated Claude usage policy: persistent abuse of Claude may lead to a ban

Read original →

Anthropic launches Cyber Mission to protect critical infrastructure and open-source software

行业动态

Anthropic launches Cyber Mission to protect critical infrastructure and open-source software

Read original →

U.S. federal research initiative Genesis Mission secures $2.4 billion in commitments from 11 industry partners

行业动态

U.S. federal research initiative Genesis Mission secures $2.4 billion in commitments from 11 industry partners

Read original →

OpenAI's annualized revenue reportedly nears $50 billion, about $20 billion below earlier reports

行业动态

OpenAI's annualized revenue reportedly nears $50 billion, about $20 billion below earlier reports

Read original →

Safety researcher fired by OpenAI denies misconduct in an open letter

行业动态

Safety researcher fired by OpenAI denies misconduct in an open letter

Read original →

OpenAI bans Russian and Iranian accounts spreading propaganda through fake journalists and think tanks

行业动态

Read original →

Manus parent company 蝴蝶效应 raises over $500 million

行业动态

Manus parent company 蝴蝶效应 raises over $500 million

Read original →

Arena raises $200 million in Series B funding and launches an alignment index

行业动态

Arena raises $200 million in Series B funding and launches an alignment index

Read original →

Suspected Chinese attackers use AI tools to attack South Korean banks and steal data

行业动态

Suspected Chinese attackers use AI tools to attack South Korean banks and steal data

Read original →

Ecosia switches to Chinese open-source models over dissatisfaction with Mistral's lagging quality

行业动态

Read original →

Harness acquires some of Augment Code's assets

行业动态

Harness acquires some of Augment Code's assets

Read original →

Isomorphic Labs reportedly in talks for new funding at a valuation of at least $40 billion

前瞻与传闻

Isomorphic Labs reportedly in talks for new funding at a valuation of at least $40 billion

Read original →

阶跃终端 sets October 13 launch for its first agent-powered smartphone

前瞻与传闻

阶跃终端 sets October 13 launch for its first agent-powered smartphone

Read original →