Gemini Omni 1.1 Flash launches with start/end frame controls and 4K upscaling
要闻
Google has released Gemini Omni 1.1 Flash for developers. According to the company, it supports scene extension, start and end frame control, and 1080p or 4K output, while 360p previews are up to 60% faster and cost one-third as much as standard 720p generation.
Google has released Gemini Omni 1.1 Flash and says the version is ready for production use, with access now available through the Gemini API in Google AI Studio or the Gemini Enterprise Agent Platform. According to the official description, the new version can analyze up to 10 seconds of preceding context when extending a scene, whereas previous models referenced only the final second. Videos can be extended in 10-second increments to a cumulative maximum of 40 seconds. Developers can also specify the starting and ending frames of a shot, allowing the model to generate continuous video between two keyframes for camera orbits, zoom transitions, or seamless loops. Compared with standard 720p, lightweight 360p previews generate up to 60% faster based on system throughput and cost one-third as much; finished videos can be output in 1080p or 4K.
Tibo says usage limits have been reset for all ChatGPT Work and Codex users
要闻
Codex head Tibo said usage allowances had been reset for every ChatGPT Work and Codex user across both services. A user also called for the existing 5-hour usage limit to be removed.
Codex head Tibo announced that the usage allowances for all ChatGPT Work and Codex users had been reset, with the action covering every user of both services. User feedback focused on the existing 5-hour usage limit. One user said the restriction had prevented them from using their remaining allowance before the reset and asked for the limit to be removed. Tibo’s announcement confirmed only that the allowance reset had been completed; the source did not say whether the 5-hour limit would be removed or provide a timetable for such a change.
OpenAI has opened Luna Reserve to some individual ChatGPT Plus and Pro accounts. Once regular usage is exhausted, it provides a separately metered, capped reserve available only for GPT-5.6 Luna.
OpenAI has introduced the Luna Reserve backup mechanism in Codex and ChatGPT Work for some individual ChatGPT Plus and Pro accounts, providing additional usage after an account exhausts its regular allowance. Luna Reserve is metered separately from the regular allowance, has its own cap, and can be used only with GPT-5.6 Luna; it neither restores exhausted regular usage nor unlocks more capable models. The mechanism is currently unavailable in ChatGPT Business and Enterprise workspaces and does not apply to regular Chat or the API.
ChatGPT updates Temporary Chat with memory and plugin access and conversion to regular chats
要闻
OpenAI announced new controls for ChatGPT temporary chats. When creating a chat, users can enable personalization to access existing memories, custom instructions, and plugins, or save the conversation and convert it into a regular chat governed by their account settings.
OpenAI says ChatGPT is rolling out controls for temporary chats, allowing users to enable personalization or save the conversation when creating a temporary chat. Once saved, the conversation is converted into a regular chat governed by account-level settings. With personalization enabled, a temporary chat can access existing memories, custom instructions, and plugins from the account, but it does not create new memories during the interaction. Personalization can only be configured when the temporary chat is created and cannot be changed after the conversation begins. A personalized temporary chat does not appear in the sidebar or chat history until it is saved, after which it appears as a regular chat.
OpenAI joins Anthropic and others in calling for stronger cyber defenses
要闻
OpenAI joined Anthropic, AWS, Google, Microsoft, Oracle, and others in a collective call for stronger cyber defense, setting out 3 principles and asking frontier AI companies to make Agentic identities traceable and accountable.
OpenAI, Anthropic, AWS, Google, Microsoft, Oracle, and other companies jointly released an action initiative to strengthen cyber defense, and OpenAI stated that the window for responding to AI-driven cyberattacks is limited. The initiative sets out 3 principles: recognizing that current security measures are insufficient, empowering defenders with AI that has cyber capabilities, and mobilizing a collective global response. According to OpenAI, the proposed actions address organizations, cybersecurity companies, governments, and frontier AI companies. Frontier AI companies are asked to ensure that Agentic identities are traceable and accountable, and to provide model access and funding to under-resourced defenders of critical infrastructure.
Anthropic releases a research preview of Model Hardware Standard
要闻
Anthropic has opened the MHS research preview to an initial group of research labs and advanced manufacturers. The company says the specification lets AI Agents operate microscopes, liquid handlers, and robotic arms in parallel while reducing device integration from weeks or months to hours or minutes.
Anthropic has provided an early version of the Model Hardware Standard (MHS) to an initial group of partners in science, robotics, electronics, and manufacturing to test AI Agent operation of physical equipment and jointly develop safety evaluations and best practices before making the standard open source. Development of MHS began as a collaboration between Anthropic and HHMI Janelia Research Campus. It works with any device that has a programmable interface, is model-agnostic, and allows an Agent harness to connect through standard protocols such as the Model Context Protocol. Anthropic says Agents can use it to control microscopes, liquid handlers, robotic arms, and other equipment in parallel for tasks including drug discovery experiments and laser calibration on a quantum computer, while reducing hardware integration from weeks or months to hours or minutes.
MHS uses a standardized driver to unify interfaces across different devices, applying primitives such as read and write to retrieve or set parameters. It also publishes device information in a common format so devices and Agents can discover one another and communicate across networks without a custom translator. Users can add information such as weight, measurable properties, adjustable parameters, and safety limits through natural-language tags, which the driver uses to generate a reference file automatically. Equipment can be controlled through three mechanisms: MCP, the command line interface, and code files (APIs). An Agent can orchestrate steps, monitor results, adjust parameters in real time, and place long-running or high-speed tasks in code files. Anthropic says Claude repeatedly adjusted a laser and observed the results through a camera during testing, then generated a deterministic script that completed laser alignment with a single command.
蚂蚁百灵 releases finance-enhanced model Ling-3.0-flash-Fin
模型发布
Ant Ling has released the finance-focused Ling-3.0-flash-Fin model with 124B total parameters and 5.1B activated parameters. It is temporarily available for free on two platforms, while its weights are scheduled to be open-sourced in the week following release.
Ant Ling has released the finance-focused Ling-3.0-flash-Fin model, although the source text does not specify an exact release date. Built on Ling-3.0-flash, the model has 124B total parameters and activates 5.1B parameters per run. According to the company, it performed strongly across multiple financial benchmarks and also improved its general-capability score. The model is now available through Vercel AI Gateway and OpenRouter. Free access on Vercel AI Gateway runs until September 25 as stated in the source, which does not specify the year, while OpenRouter offers a one-month free trial. The company also said the model weights are scheduled to be open-sourced in the week following release.
fal launches H3 Max, a post-trained model based on MiniMax H3
模型发布
fal has released H3 Max, a post-trained video model based on MiniMax H3. According to the company, it generates a 5-second video in under 3 seconds, delivers roughly 35 times the throughput of the official H3 endpoint, and ranks first in all three human-preference evaluation categories.
fal now offers H3 Max through its platform. The post-trained video model was developed by fal Research from the open-weights MiniMax H3 and optimized by fal’s inference team. According to fal, the team added new data during post-training, focusing on prompt adherence and visual quality while iterating on the model and inference stack together. The model was trained and served on NVIDIA GB200 NVL72 systems, which the company says deliver up to 2 times the per-chip performance of the previous-generation accelerators it had used. H3 Max was also trained on fal Serverless.
fal’s quality evaluation primarily used head-to-head human-preference comparisons across overall quality, prompt understanding, and aesthetics, with results aggregated through Bayesian Elo ratings with 95% confidence intervals. The 12 video models evaluated included the official MiniMax H3 endpoint, Gemini Omni Flash, Wan 3.0, Seedance 2.5, Kling 3, and Veo 3.1. According to the company, H3 Max ranked first in all three categories, won a majority of head-to-head comparisons against every tested model, and also placed first in independent rankings from Artificial Analysis and Design Arena. fal says the model generates a 5-second video in under 3 seconds, provides roughly 35 times the throughput of the official H3 endpoint, and is 15 times faster on average than models with comparable quality. It is available through Playground, fal Agent, and the API, with a 50% discount during the first week.
Midjourney V8.2 image editing model enters testing
模型发布
Midjourney has opened testing of its first V8.2 image edit model to all users and updated the UIs on its main and alpha sites. The company is collecting edge cases and interface feedback and says rapid follow-up updates will continue.
Midjourney’s first V8.2 image edit model has entered testing for all users, and the UIs on midjourney.com and alpha.midjourney.com have also been updated. According to the company, the alpha site will continue running experiments around redesigning the interface for the model, and the test is expected to surface edge cases where the model does not work as intended. If an image produces poor results, users can click the sad face emoji on the right side of the lightbox or submit the corresponding jobid or URL while the image remains open in the lightbox. Interface suggestions and model edge cases can be posted in #ideas-and-features, while results made with the model can be discussed or shared in #edit-showcase. The company says rapid follow-up updates will continue, with this round of testing involving a broader community.
BreezeBlue open-sources the Breeze TTS 2 speech model
模型发布
BreezeBlue has released Breeze TTS 2, an open-weight speech model for real-time interaction. The company says it ranks No. 1 among open-weight models on the Artificial Analysis TTS leaderboard and supports streaming 24 kHz PCM output.
BreezeBlue has launched Breeze TTS 2, an open-weight text-to-speech model for real-time interaction. According to the company, the model ranks No. 1 among open-weight models on the Artificial Analysis TTS leaderboard and outperforms frontier proprietary systems. Natural-language instructions can design a voice without reference audio or direct a reference speaker’s tone, emotion, pace, and delivery. Using --cfg-scale 4 strengthens instruction-following. Speaker cloning requires clear reference speech with as little background noise as possible and an exact transcript.
The Breeze TTS 2 CLI and single-concurrency streaming API use PyTorch eager streaming by default and skip graph warmup. The API returns mono 24 kHz signed 16-bit little-endian PCM, while --fast-all enables the best configuration for every inference stage at the cost of additional cold-start time. The default Docker image targets H100/Hopper(sm90), with a separate build method for A100, and all required model components are included in the checkpoint. The source code uses Apache License 2.0, while model weights, checkpoints, adapters, derivative models, and self-hosted outputs are separately governed by the BreezeBlue Research and Non-Commercial License. Commercial use requires written authorization from RESONIA, INC.; the contact address is contact@breeze.blue.
Cohere has officially launched Cohere Parse 5, a document parsing model that handles text, tables, forms, and images. According to the company, it scores 79.2 on ParseBench and costs $1.50 per 1,000 pages.
Cohere has officially released and launched Cohere Parse 5 for building RAG systems and automating document processing. The model recognizes text, tables, forms, and images in documents, then produces machine-readable documents with bounding boxes. According to the company, it scored 79.2 on ParseBench, outperforming competitors including Mistral OCR 4. Cohere has priced the service at $1.50 per 1,000 pages; according to its explanation, this is up to 95% cheaper than frontier large models.
New Qoder launches with dual coding and general-purpose modes plus a desktop app
开发生态
Qoder has launched a new Agent workspace and desktop app with separate Coding and general-purpose modes. Its international and China editions are now available to download, while its global user base has exceeded 6 million and the system supports more than 40 connectors and over 70 plugins.
Qoder has released its next-generation Agent workspace and desktop client, with downloads now available for both the international and China editions. The product is built around Coding capabilities and includes programming and general-purpose modes. According to the company, the system includes leading global models and an Auto intelligent scheduling engine, while Harness technology supports complex, long-chain autonomous tasks. Qoder also says the platform can connect to work systems through more than 40 connectors and over 70 plugins, with voice activation and proactive reminders added as new features. Qoder’s cumulative global user count has exceeded 6 million, and the company is also offering limited-time benefits including dedicated Credits and model discounts.