OpenAI is rolling out Ultrafast mode for GPT-6.1 Sol across API, Codex, and ChatGPT Work. The company says it runs at up to 8 times the speed of Sol Standard, with API pricing of $12 per million input tokens and $60 per million output tokens.
OpenAI’s Ultrafast mode for GPT-6.1 Sol is rolling out across API, Codex, and ChatGPT Work and is already available in all supported regions, with support for US and EU data residency. API access is open to all users, priced at $12 per million input tokens and $60 per million output tokens. Access in Codex and ChatGPT Work covers Pro 500, as well as eligible pay-as-you-go Enterprise and quota-billed Edu subscription plans; Enterprise access requires an administrator to enable permissions. According to OpenAI, Ultrafast runs at up to 8 times the speed of Sol Standard and has an intelligence level close to Astra.
OpenAI introduces faster steering for Codex in its desktop app
开发生态
OpenAI is updating steering for Codex in the ChatGPT desktop app to make it respond faster to additional instructions sent while a task is running. Users can choose whether follow-up messages adjust the current task or wait until the next run.
OpenAI is rolling out faster-response Codex steering in the ChatGPT desktop app, allowing users to send messages during a running task to adjust its execution. Users can configure how follow-up messages are handled under Settings → General → Follow-up behavior: steer the current run or wait until the next run. This lets users correct Codex’s approach, supply missing information, or change the task’s direction. Tibo, a member of the OpenAI Codex team, said the improved steering takes effect immediately, enabling the model to respond faster to adjustments and correct its direction during execution to avoid wasted effort.
阶跃星辰 offers Step 5 Preview free for one week across multiple platforms
开发生态
StepFun announced that Step 5 Preview is available on OpenRouter, with OpenCode, Cline, NousResearch, KiloCode and other platforms progressively offering one week of free access starting on the announcement day. Model resources and API documentation were released alongside it.
StepFun announced that Step 5 Preview has launched on OpenRouter and that OpenCode, Cline, NousResearch, KiloCode and other platforms are progressively opening one week of free access from the day of the announcement. The accompanying resources include a model overview, benchmarks, demonstrations, technical specifications and API documentation, and users can switch directly to the model within their existing workflows. According to the company, the model offers flagship-level intelligence for agentic and professional work at substantially lower task costs; the source provides no specific cost figures or percentage reductions.
Glyph Cluster launches on Vercel AI Gateway, free during its anonymous phase
开发生态
Vercel is offering Glyph Cluster on AI Gateway free during its stealth period to Pro and Enterprise teams that have purchased Gateway credits. The model supports coding reasoning, but ZDR is unavailable, and inputs and responses may be used for training.
Vercel has added Glyph Cluster to AI Gateway in its stealth phase, during which Pro and Enterprise teams with purchased AI Gateway credits can use it for free. Its model name is stealth/glyph-cluster, and it can be accessed through the AI SDK, OpenAI-compatible Chat Completions and Responses APIs, or a coding Agent connected to AI Gateway. For tools such as Claude Code, Codex, and Cursor, the connection process involves installing the latest Vercel CLI, running setup, and then selecting the model in the Agent.
According to the official description, Glyph Cluster is a reasoning model for coding and long-context analysis that can handle planning, synthesis, quantitative reasoning, and comparisons of material across large inputs. It can also review code, explain failures, debug issues, propose changes, and analyze implementation options. The model supports function calling and streaming responses, but accepts only text, not images or files. Tool use is restricted to user-defined function tools, and structured outputs are not supported. ZDR is unavailable, and prompts and responses sent through the model may be used for training and model improvement. Users can also try it in the model playground.
Claude introduces two experimental features: Dashboards and Motion
产品应用
Claude announced two beta features, Dashboards and Motion, available on paid plans and on Team and Enterprise, respectively. Docs, Slides, and Design also left beta and are available on every plan, including Free.
Claude made Dashboards and Motion available in beta at the time of this announcement and moved Docs, Slides, and Design out of beta. Dashboards is available on paid plans, Motion on Team and Enterprise, and the other three features on every plan, including Free. According to the company, users have created more than 45 million documents, presentations, and designs in Claude. Dashboards and Motion are disabled by default for Enterprise; administrators can enable them under Organization settings > Artifacts.
According to the company, Dashboards connects to data platforms such as BigQuery, Databricks, and Snowflake, as well as Salesforce, to generate dashboards from natural-language questions and update them as the data changes. Users can inspect queries and each chart’s last refresh time, or send dashboards to tools such as Amplitude, Grafana, and Hex for further analysis. Motion uses code to animate text, charts, shapes, and images, supports changes in an editor or through requests to Claude, and exports MP4 files. The company says Motion does not use a video generation model and does not generate footage or AI-generated people. Its output can also be further edited in tools such as Adobe, Descript, and Runway.
Google Cloud unveiled Gemini, a unified work Agent, at Gemini at Work 2026, bringing question answering, content generation, and code execution into a single API. The company says nearly 80% of Google Cloud customers already use its AI products.
Google Cloud announced Gemini, a unified Agent for work, at Gemini at Work 2026, without disclosing a specific launch date in the article. According to the company, users can submit task objectives through a single interface and a single API, with Gemini planning the work, invoking tools, and connecting to enterprise systems to handle questions, knowledge work, image and media generation, and code writing and execution. It can also work directly within Gmail, Drive, Docs, Slides, Sheets, Chat, and Calendar, carrying over the same memory, skills, and control settings.
According to the company, Gemini executes tasks persistently in the cloud and retains memory and context across devices and channels, allowing work that takes hours or days to continue after users close their computers. It can create temporary sub-agents with separate identities to coordinate parallel or sequential workflows. The underlying model is separate from the Agent, and it can currently orchestrate Gemini-family models and Anthropic’s Claude models. The company also announced governance and cost-control mechanisms including identity and permission management, secure sandboxing, network gateways, Smart Routing, and real-time spending caps. It says nearly 90% of Fortune 100 companies use Gemini Enterprise.
腾讯 launches a standalone file browser for WorkBuddy
产品应用
Tencent WorkBuddy has launched a standalone file browser that supports viewing and editing multiple local files in tabs within one window, with a right-hand AI chat panel for processing their contents. According to the company, local viewing and editing use no credits, while AI processing is charged according to task complexity.
Tencent WorkBuddy has launched a standalone file browser, accessible by double-clicking a file in the operating system or selecting WorkBuddy through the right-click menu’s open-with option. A single window supports multiple files, each in its own tab. Files opened by double-clicking locally are added to the currently active file window, and closing the last tab closes the entire window. Unsupported file types do not show an option to open them with WorkBuddy. Once opened, files can be viewed and edited directly, with the top toolbar offering different functions depending on the file type. Closing a tab or window with unsaved changes triggers a prompt to save or discard them.
The AI chat panel on the right is collapsed by default and can be expanded using the button in the upper-right corner. Using AI editing also expands it automatically. Users can ask AI to summarize, rewrite, polish or extend the current file, or select and quote text in a question so that revisions are written into the document. According to the company, the browser reads only files that users actively open and does not traverse or scan their disks. Local viewing and editing use no credits, while AI processing consumes credits according to the actual complexity of the task. File contents are used only to process the current task and are not shared with third parties unless users actively share them or upload them to third-party storage.
JetBrains open-sources Mellum2.1, focusing on reinforcement learning for coding agents
模型发布
JetBrains has open-sourced Mellum2.1 for coding Agents, with 12B total parameters and 2.5B active parameters under the Apache 2.0 license. The company says the update primarily uses reinforcement learning to improve the model’s ability to work with code inside repositories.
JetBrains has released Mellum2.1 on Hugging Face; the source does not specify an exact release date. The model retains Mellum2’s mixture-of-experts architecture, with 12B total parameters and 2.5B active parameters, and is open-sourced under the Apache 2.0 license for coding Agents and sub-agents running on users’ own hardware. This update focuses on post-training, primarily reinforcement learning. According to the company, training involved millions of sandboxed runs across thousands of real environments, enabling the model to explore codebases, edit files, and check its changes. JetBrains compared it with Mellum2, Qwen3.5-9B, and Gemma 4 E4B using the same evaluation setup. The company says agentic coding showed the largest improvement over Mellum2; under heavy load, throughput is almost twice that of Qwen3.5-9B, while MTP makes single-request processing about 1.6 times faster. GGUF builds for llama.cpp, Ollama, and LM Studio, along with the MTP head for speculative decoding in vLLM, have yet to be released.
Hugging Face open-sources gene annotation model Carbon-A
模型发布
Hugging Face has released the gene annotation model Carbon-A and its accompanying database, reporting 566 million new gene candidates across 22,617 species. The model, training data, and technical report are also available, and some predictions have undergone experimental validation.
Hugging Face has released the Carbon Annotation Database and the gene annotation model Carbon-Annotator (Carbon-A); the source does not specify an exact release date. According to the official announcement, the model identified 566 million new gene candidates across 22,617 species, thousands of which had not previously been studied. The database, model, training data, and technical report are available in HuggingFaceBio’s Carbon Annotation Database Collection. The team collaborated with ActiveSite and UCSD to experimentally validate some newly predicted genes.
Carbon-A has 1.2 billion parameters, was trained on RefSeq annotations, and predicts protein-coding regions directly from DNA. It uses a context window of 98,304 base pairs and outputs nucleotide-resolution predictions on both strands, using one model across mammals, other vertebrates, invertebrates, plants, fungi, and protists. According to the official report, it achieved a macro-averaged nucleotide F1 of 0.944 across 42 benchmark genomes and outperformed the evaluated baselines at the nucleotide, exon, and gene levels. On Tetrahymena thermophila, which uses a nonstandard genetic code, its nucleotide F1 was 0.960.
LightOnAI has open-sourced LightOnOCR-3 to turn complex documents into structured content that applications can use directly. According to the company, the model family supports layout element localization, image descriptions and chart data extraction alongside text transcription; the earlier LightOnOCR has exceeded 4.5 million downloads on Hugging Face.
LightOnAI has open-sourced LightOnOCR-3, a family of end-to-end OCR models for converting complex documents into structured content that applications can use directly; the source does not specify a release date. According to the company, the models go beyond text transcription to identify and locate layout elements, generate image descriptions, and extract data from charts and scientific figures. LightOnOCR-3 builds on the earlier LightOnOCR, which has exceeded 4.5 million downloads on Hugging Face; that download figure refers to LightOnOCR, not the newly released LightOnOCR-3.
Odyssey has released its Odyssey-3 world model and opened a research preview. According to the company, Pro achieved the highest reported score of 66.1 on Physics-IQ Verified’s video-to-video test, while the model can also support environment simulation and control-policy training.
Odyssey has released the Odyssey-3 world model, with its research preview now available. Built as an autoregressive diffusion transformer, the model uses previous observations and the latest inputs to predict object motion, interactions, and how situations change over time. Developers can generate environments from a prompt, move through them or introduce events during generation, and observe the model’s responses. The preview supports first-person and third-person navigation as well as independent camera movement. According to the company, Odyssey-3 Pro achieved the highest reported score of 66.1 on Physics-IQ Verified’s video-to-video test and scored 54.7 on its image-to-video test. The benchmark evaluates model predictions using videos of real physical experiments across fluid dynamics, optics, solid mechanics, magnetism, and thermodynamics.
In the company’s evaluation using WorldMark’s own captions and the mean of its 13 metric scores, Odyssey-3 ranked 1st in 3 of the 4 categories: first-person stylized, third-person real, and third-person stylized environments. The company also states that applying the model to a specific physical system still requires evaluating the behaviors needed for that machine and its tasks. Adaptation methods include training an action decoder or policy on paired observations and actions. According to the company, the model completed manipulation tasks with only tens of hours of robot demonstrations and exhibited recovery behaviors not present in those demonstrations, including reorienting a gripper after a missed grasp and retrieving an object dropped in an unusual position. For driving on real roads in India, the team froze the Odyssey-3 backbone and trained a policy on just 20 hours of driving data, enabling closed-loop driving by predicting waypoints ahead of the vehicle.
Grok Imagine Video 1.5 Lite launches at $0.14 per second for 1080p
模型发布
Grok Imagine has released Video 1.5 Lite through its API, supporting text-to-video and image-to-video generation. Pricing is per second and varies by resolution: $0.02 for 480p, $0.03 for 720p, and $0.14 for 1080p.
Grok Imagine announced that its Video 1.5 Lite video generation model is now available on the Grok Imagine API, offering two functions: text-to-video and image-to-video generation. Charges depend on video resolution and are calculated per second: $0.02 per second for 480p, $0.03 per second for 720p, and $0.14 per second for 1080p. The source does not specify an exact launch date or list prices for resolutions other than these three.