Qwen open-sources Qwen-Image-2.1, an image generation and editing model
要闻
Qwen released Qwen-Image-2.1 and opened its model weights, combining text-to-image generation and image editing in one model. Its visual generation component uses a 7B-parameter, 32-layer Single-Stream DiT and supports up to 10 reference images.
Qwen opened the model weights when it released Qwen-Image-2.1, placing text-to-image generation and image editing in a single model. The visual generation component uses a 7B-parameter, 32-layer Single-Stream DiT. According to Qwen, the new version uses mixed-granularity attention and KV Cache reuse to reduce repeated computation for multi-image inputs. It can generate and edit transparent RGBA images, as well as extract subjects from ordinary photos to create transparent layers. Editing accepts up to 10 reference images and provides local controls including region selection, painting, and separate masks. Qwen said the model improves consistency for portrait identity and for product text, textures, and shapes, while also improving typography, lighting and shadows on people, and visual details. It also covers panorama, infographic, and storyboard generation. Diffusers, ComfyUI, vLLM-Omni, SGLang, and LightX2V all added support on the day of release.
阶跃 releases Step 5 Preview, now available to all users
要闻
StepFun has fully opened Step 5 Preview for real-world Agentic tasks. The model uses a sparse MoE architecture with 600B total parameters and 27B active parameters, and supports a 1 million-Token context window.
StepFun has officially released and fully opened Step 5 Preview, its flagship foundation model for real-world Agentic tasks. Its target areas include AI coding, software engineering, professional knowledge work, and finance. According to the company, the model uses a sparse MoE architecture with 600B total parameters and 27B active parameters, accepts text and visual input, and supports a 1 million-Token context window. It is designed for long-context workloads, multi-turn tool calls, and sustained task execution. The company said Step 5 Preview scored 44 on the Artificial Analysis Intelligence Index, ranked among the world’s top three open models, and costs one-eighth as much per task as Claude Opus 5. Users can access it through APIs on domestic and international open platforms, try it online in Studio, or subscribe to Step Plan; the model weights are scheduled to be released on October 15.
TypeSafe AI opens Jev to all users and gives new users $5 in credits
开发生态
TypeSafe AI has opened Jev to all users through its online console and removed the waitlist requirement, allowing immediate access. Every user can receive $5 in initial credit, which the company says is enough for about 120 million tokens.
TypeSafe AI has now made Jev available to all users through its online console, and the revised access policy is already in effect. Users no longer need to join a waitlist in advance and can start using Jev directly from the console. The previous waiting restriction has been removed, so waitlist status is no longer a condition for access. All users can receive $5 in initial credit; according to TypeSafe AI, that amount can support approximately 120 million tokens.
Google open-sources Agentic orchestrator AX, supporting pause and resume for stateful tasks
开发生态
Google has open-sourced AX, an Agentic orchestrator for Kubernetes clusters that declaratively runs billions of autonomous Agent tasks per cluster and supports pausing and resuming stateful tasks. Its specification is currently at v1alpha1, and major breaking changes may arrive before a stable release.
Google has open-sourced AX under the Apache License 2.0 as a declarative Agentic orchestrator for Kubernetes clusters, designed to execute stateful Agent tasks in isolation and manage their suspension and resumption. Built on Agent Substrate, AX configures tasks through workspace and gateway specifications, connects their workspaces, restricts network access, and describes resources uniformly with ax.io/v1alpha1 manifests. According to the project, each cluster can run billions of autonomous Agent workloads. Its core concepts, protocols, and specifications remain under refinement, and major breaking changes may be introduced before a stable release.
The ax CLI connects to the control plane over gRPC and follows a kubectl-like command model with apply, get, describe, watch, delete, and several Agent-specific operations. It follows the active Kubernetes context, resolving and establishing a tunnel to the relevant cluster’s control plane in the background. Deployment requires a Kubernetes cluster, ko, a container registry that the cluster can pull from, and a reachable Agent Substrate Control API; the default in-cluster address is api.ate-system.svc.cluster.local:443. The deployment process installs Redis first, then uses ko to build and deploy the control plane images, with all resources placed in the ax-system namespace. The ./demo.sh script applies a custom workspace, waits for readiness, runs commands through ax ssh, and suspends the task.
腾讯 open-sources WeVisDoc, an end-to-end document parsing model
模型发布
Tencent has released WeVisDoc on GitHub and Hugging Face, providing 2B and 4B weights, code, and tutorials. The model converts page images with different layouts and capture conditions directly into structured Markdown.
Tencent has made WeVisDoc’s 2B and 4B weights, supporting code, and tutorials available on GitHub and Hugging Face. The end-to-end document parsing model converts page images with different layouts and capture conditions directly into structured Markdown. The two versions are fine-tuned from Qwen3-VL-2B-Instruct and Qwen3-VL-4B-Instruct, respectively, and can preserve text, LaTeX formulas, HTML tables, and reading order in a single output sequence. According to the company, WeVisDoc-4B outperformed the compared end-to-end parsing models across four settings: OmniDocBench v1.6 and the Clean, Digital Degraded, and Real Degraded settings of PureDocBench.
WebCraftBench evaluates AI web apps using real-world interactions and code coverage
技术与洞察
The arXiv page title says WebCraftBench evaluates AI web applications through real interactions and code coverage. However, the supplied text only introduces arXivLabs and its four partnership values, providing no benchmark version, data, or evaluation results.
The supplied arXiv page title describes WebCraftBench as a method for evaluating AI web applications through real interactions and code coverage, but the body provides no authors, publication date, version number, dataset size, metric definitions, or experimental results. The body instead introduces arXivLabs, a framework that allows collaborators to develop and share new features directly on the arXiv website. According to the page, individuals and organizations working with arXivLabs accept four values: openness, community, excellence, and user data privacy. arXiv works only with partners that adhere to these values and directs people with project ideas that could add value to the arXiv community to learn more about arXivLabs.
硅基流动 completes the second tranche of its Series B+ and a Series C funding round
行业动态
SiliconFlow announced the completion of the second tranche of its Series B+ and its Series C financing, bringing its cumulative equity financing in fiscal 2026 to nearly RMB 2.9 billion. The funds will support R&D in inference engines, heterogeneous compute scheduling, and model-chip adaptation, as well as global market expansion.
SiliconFlow has announced the completion of the second tranche of its Series B+ financing and its Series C financing, bringing its cumulative equity financing in fiscal 2026 to nearly RMB 2.9 billion. Participating institutions included the China Internet Investment Fund, Guoxin Fund, and China Mobile Chain Leader Fund, while some existing shareholders made additional investments. According to the company, the proceeds will increase R&D investment in inference engines, heterogeneous compute scheduling, and model-chip adaptation, while strengthening its Token supply platform and further expanding into global markets.
智谱MaaS platform to launch a no-retention feature for data and content
行业动态
Zhipu’s MaaS platform will add a data non-retention feature that any user may request. Once enabled, inputs and outputs will be used only for the current model call, while Batch API and File API are excluded, and relevant data may still be retained for 30 days or longer to meet legal requirements or investigate violations and abuse.
Zhipu’s MaaS platform will introduce a data non-retention feature, with the specific activation date and scope subject to the platform’s confirmation. Any user may apply, and enterprises and developers can submit requests through the MaaS console. According to its description, once the feature takes effect, the platform will not store user inputs or outputs statically, and the related data will be used only to complete the current model call. Features requiring persistent storage of tasks or files, including Batch API and File API, are not covered. The platform may still retain related data for 30 days or longer when required by laws and regulations or when investigating suspected violations or abuse.