Anthropic releases Claude Haiku 5.5, cutting average costs by about 75%
要闻
Anthropic has released Claude Haiku 5.5, saying its average running cost is around 75% lower than Haiku 4.5’s. Alongside the release, it halved Sonnet 5.5’s cache reads pricing and introduced a monthly API credit for Claude Max and Team subscribers.
With this release, Anthropic introduced Claude Haiku 5.5 and changed Sonnet 5.5 pricing and subscribers’ API benefits. According to the company, Haiku 5.5 costs around 75% less to run on average than Haiku 4.5; halving Sonnet 5.5’s cache reads pricing reduces its running cost by around 20% for most agentic work. The new monthly API credit is for Claude Max and Team subscribers and is intended to support building agents and applications on the Claude Platform. The source does not specify the credit amount.
According to Anthropic, Haiku 5.5 is designed for high-volume, cost-sensitive tasks and can handle summaries, context compaction, database queries, and classification requests. It can also work as a subagent alongside Opus 5.5 and Sonnet 5.5 on coding tasks. The company describes it as its fastest model to date and says it suits speed-sensitive uses such as live customer support and browser use. It is also the first Haiku-series model with an adjustable effort setting, allowing users to choose whether to prioritize optimizing for cost or intelligence.
OpenAI rolls out GPT-6 to all users and launches an intelligent interaction interface
要闻
OpenAI is rolling out GPT-6 and Intelligent UI in ChatGPT, with paid users receiving the update on launch day and Free and Go users following the next day. According to its evaluations, Instant began answering 44% earlier in web-search scenarios.
OpenAI began rolling out GPT-6 and Intelligent UI to ChatGPT Plus, Pro, Business, and Enterprise users on launch day, with Free and Go users receiving the update the following day; the source does not specify a calendar date. The release changes only the Chat experience, leaving the models used by Work and Codex unchanged. Paid plans use GPT-6 Sol, while Free and Go use GPT-6 Luna. According to OpenAI, GPT-6 can progressively output answers while continuing to think. In its internal evaluations, GPT-6 Extra High took about as long to begin answering as GPT-5.6 Medium, while scoring higher overall than GPT-5.6 Extra High. For queries involving web search, GPT-6 Instant began answering an average of 44% earlier than GPT-5.6 Instant.
The interaction updates include combining text, images, charts, buttons, forms, and interactive components according to the question, as well as directly generating task tools such as calculators and mini-games. OpenAI says Intelligent UI uses native components and a compiler that support streaming output, allowing the interface to build progressively alongside the answer rather than waiting for the entire response to finish. Model safety training has also been updated; according to the company, training against high-risk misuse, including cyberattacks, biological threats, and violence, has been strengthened.
Codex reset usage limits yesterday based on a vote and grants extra manual resets today
要闻
OpenAI Codex lead Tibo announced that all paid accounts will receive one additional usage-limit reset that can be saved for later, expected by the end of the campaign’s third day in Pacific Time. He had already confirmed an earlier reset after 76% of voters supported it.
OpenAI Codex lead Tibo announced on day three of the 28-day update campaign that all paid accounts would receive one additional usage-limit reset that can be saved for later, with delivery expected by the end of that day in Pacific Time. According to his explanation, the offer celebrates GPT-6 arriving in Chat and the combined active user count for Codex and ChatGPT Work reaching a new high of 40 million. Earlier, after releasing four updates on day two of the campaign, Tibo held a vote on whether to reset usage limits, with 76% of voters choosing a reset. He subsequently confirmed that the reset had been carried out and that all four improvements remained in place.
Anthropic has cut Claude Sonnet 5.5’s cache-read price by 50% to $0.10 per million tokens. The company estimates that this change reduces running costs by approximately 20% for most long-running tasks.
Anthropic announced that it is halving Claude Sonnet 5.5’s cache-read price, bringing it to $0.10 per million tokens. The announced price change applies to cache reads, and the company also provided an estimate of its impact on task running costs: Sonnet 5.5 is expected to cost approximately 20% less to run for most long-running tasks. This cost reduction is the company’s estimate and applies to most long-running tasks, rather than all tasks.
Claude grants monthly API credits to Max and Team subscribers
要闻
Claude’s Max and Team subscriptions include monthly API credits that users can claim by linking a Claude Console organization, without adding a payment method. Team credits are pooled across all seats, with a monthly cap of $500 USD.
Claude now provides monthly Claude API credits with Max and Team subscriptions, which users must link to a Claude Console organization to claim. According to the official documentation, credits are issued each billing cycle, or monthly for annual subscriptions, and can be used to build and run users’ own applications and Agents without adding a payment method to Claude Platform. Team allocations depend on the seat configuration at the start of each billing month, including the first allocation: 3 Standard seats and 2 Premium seats receive $260 USD per month, and adding 1 Premium seat brings the following month’s credits to $360 USD. The credits can only be used on Claude Platform and do not change the rules for rate limits or spend limits. After claiming them, users can check Promotional credits under Settings > Billing in the Console, then create an API key to send requests.
谷歌 opens its SynthID detection website to users worldwide
要闻
Google has opened SynthID Detector to users worldwide after previously limiting access to media professionals. It is currently available only in English and lets users upload images, videos or audio to check for invisible watermarks, though detection results are not 100% accurate.
Google announced that its AI-generated content detection platform, SynthID Detector, is now available to users worldwide, with an English version currently offered; the early version was restricted to media professionals. The platform accepts image, video and audio uploads and checks whether files contain invisible SynthID watermarks from Google or partners including OpenAI, NVIDIA and Kakao, with Apple also set to add support soon. According to Google, since SynthID launched in 2023, it has watermarked more than 180 billion images and videos and 240,000 years of audio content; verification features built into Search, the Gemini app and Chrome currently process more than 1 million requests per day. Watermarks remain identifiable after edits such as adding filters, but the tool’s detection results are not 100% accurate.
Claude Code team member launches html-plan skill for users to try
开发生态
Anthropic Claude Code team member Thariq has made the html-plan skill available to try through the community plugin marketplace. It generates HTML plan pages containing code snippets, questions needing clarification, and mockups, with user feedback being collected ahead of a wider rollout.
Anthropic Claude Code team member Thariq announced on X that his html-plan skill is now available to install through the community plugin marketplace. Users first run `claude plugin marketplace add anthropics/claude-plugins-community` to add the marketplace, then run `claude plugin install html-plan@claude-community` to install it. According to his explanation, the skill presents plans in concise language alongside code snippets, questions needing clarification, and mockups; an accompanying lint mechanism is intended to reduce common Claude failures, with some rules targeting chart and flowchart generation. Thariq said he particularly likes the call-stack display and code-snippet annotation features, noting that the call-stack display draws on developer @dillon_mulroy’s approach. He wants to collect user feedback before a wider rollout and turn it into a plugin that replaces Claude Code’s default planning mode with HTML plans.
Claude SDK adds built-in computer use and browser use toolsets
开发生态
Anthropic has provided minimal Python and TypeScript examples of Claude’s computer toolset on GitHub, using VNC to execute desktop actions through a driver that implements 10 tools. The examples demonstrate the interface and are not production code.
Anthropic has provided a minimal VNC example of computer_toolset_20260801 in its claude-quickstarts repository on GitHub, with Python and TypeScript implementations. Developers declare one toolset entry in tools, and the model issues desktop actions such as screenshot, left_click, and type as separate tool calls that the driver executes. The driver consists of a few hundred lines, controls one desktop through VNC’s RFB protocol, and implements 10 tools covering screenshots, clicks, mouse movement, scrolling, typing, key input, and waiting. The SDK provides tool schemas, rejection handling for unimplemented tools, and the tool runner. According to the official explanation, the code demonstrates the interface and is not production code.
Execution safety and call verification are also part of the example. According to the official explanation, screen content influences the model’s next actions, and a focused terminal executes what the model types, so the execution environment should be a disposable desktop inside a sandbox. The VNC server has no password, connections must be restricted to localhost, and container ports must be published only to the host’s loopback interface. Before executing any action other than screenshot, run uses confirm to require the user to enter y; exercise instead approves calls automatically and requires no API key. exercise checks 6 valid calls and verifies that both a left_click one pixel outside the screen and the unimplemented zoom return is_error. It exits with a non-zero status if any result differs from the expected outcome.
OpenRouter adds the full ElevenLabs voice model lineup at 50% off for a limited time
开发生态
OpenRouter launched nine ElevenLabs Text to Speech models and two Speech to Text models on October 7, 2026. All users can access these models at 50% of the platform’s list price during a limited-time promotion.
OpenRouter announced the addition of nine ElevenLabs Text to Speech models and two Speech to Text models on October 7, 2026, with a 50% discount off list prices for all OpenRouter users running through October 19, 2026, at 8 a.m. Pacific Time. According to the company, developers do not need a separate ElevenLabs plan and can call POST /api/v1/audio/speech and POST /api/v1/audio/transcriptions using an OpenRouter API key. The launch does not include Realtime Scribe v2 over WebSocket. Users with an existing ElevenLabs account can add their own key through BYOK to access their own cloned voices, pronunciation dictionaries, and higher-quality formats such as mp3_44100_192.
Capabilities and request limits vary by model. According to the company, Eleven v4 and v3 can interpret audio tags as instructions for tone and delivery, with per-request limits of 10,000 and 5,000 characters respectively; Flash v2.5 supports up to 40,000 characters per request. Scribe v2 can produce transcripts with timestamps and speaker labels, while Scribe v2 Medical is fine-tuned for medical terminology. Transcription uploads are capped at 25 MB, equivalent to about 27 minutes of 128 kbps MP3 audio. Providing a public URL bypasses that upload cap, but long files may still reach the 180-second upstream request timeout. Text to Speech is billed by input Unicode code points, with audio tags included in the character count. Scribe is billed per second of audio, and usage.cost in the transcription response reports the cost of the request.
GitHub Copilot opens access to a local sandbox, with hybrid inference preview coming at month's end
开发生态
GitHub Copilot will preview automatic routing between local and cloud inference by the end of the month, while also allowing developers to select local models directly. Microsoft says the quantized on-device version of MAI Code 1.1 Flash is 53GB, an 80% reduction in size compared with the Bfloat16 cloud version.
GitHub Copilot will offer automatic routing between local and cloud inference by the end of the month; the source does not specify the month or year. GitHub Copilot CLI, the Copilot app, and VS Code will offer two ways to use local models: letting Auto decide whether a task uses local or cloud inference, or having developers explicitly select a local model. According to the official explanation, Auto considers task context and cache state during multi-turn sessions, retaining reusable cached work as it switches models. Direct selection supports MAI Code 1.1 Flash through the Windows ML provider, as well as connections to OpenAI-compatible local endpoints and selection of the models they expose. Execution permissions are separately constrained by sandboxing. Copilot uses Microsoft Execution Containers (MXC), an open-source library from the Windows team, to control access to files, networks, credentials, system capabilities, and other resources by processes and local services launched by an Agent.
The local model was developed by Microsoft AI using a mixture-of-experts architecture, with 137 billion total parameters and 6.8 billion active parameters, and uses quantization and speculative decoding. Microsoft says the quantized on-device version is 53GB, an 80% reduction in size compared with the Bfloat16 cloud version. Its published figures for the first shipping version on Surface Laptop Ultra show peak memory usage of 75.5GB at 256k context, with prompt-processing throughput of 923.5 and 769.8 tokens per second at 64k and 128k context, respectively. The device is built around NVIDIA RTX Spark, with up to 128 GB of unified memory and up to 1 petaflop of AI compute. The source also explains that the operating system, applications, inference runtime, and key-value cache require memory, and that keeping a model loaded does not guarantee constant response times.
微软 MXC becomes generally available, defining execution permissions for AI agents
开发生态
Microsoft announced general availability of Microsoft Execution Containers (MXC) on October 7, 2026, enabling developers and IT administrators to set policy-based limits on Agent access to files and networks. Windows 365 support is also generally available.
Microsoft announced general availability of Microsoft Execution Containers (MXC) and its Windows 365 support on October 7, 2026. MXC targets untrusted code and dynamically generated workloads and can contain model-generated output, plugins, tools, an Agent harness, or an entire Agent. According to Microsoft, developers and IT administrators can define resources that a workload may access, including files and network destinations, and MXC enforces those policies at runtime using the appropriate container. The policies remain outside the Agent workload’s control, so an Agent or generated code cannot grant itself additional access.
Developers integrate MXC through a unified JSON configuration schema and a multi-language SDK, with MXC mapping the required controls to backends on Windows, macOS, or Linux. Lightweight containment uses AppContainer, Seatbelt, and Bubblewrap, respectively. The session container and WSL container support Windows 11 only, while MicroVM supports Windows 11 and Linux and remains experimental. Microsoft also said Windows will support distinguishing Agent activity from user activity through Microsoft Entra and extend Microsoft Agent 365 controls to local Agents on devices. These capabilities are intended to let IT teams manage MXC containers, apply policies, and monitor Agent activity, but the source provides no specific launch dates for these forthcoming features.
Liquid AI open-sources decision models d1-3B and d1-omni-600M
模型发布
Liquid AI has released the open-weight decision models d1-3B and d1-omni-600M on Hugging Face. The company reports scores of 48.57 and 15.95, respectively, on the public split of Decision Index v0.2.1.
Liquid AI has released two open-weight decision models, d1-3B and d1-omni-600M, on Hugging Face, with llama.cpp support available at launch. According to the company, d1 returns an answer in a single forward pass rather than generating output tokens. The base model for d1-3B was created by averaging the weights of LFM2.5-2.6B and the text backbone of LFM2.5-VL-3B, followed by fine-tuning and merging. d1-omni-600M is an experimental checkpoint based on LFM2.5-Encoder-350M, with audio and vision capabilities added in stages. It supports text, image, and audio inputs and remains under development.
The company reports that d1-3B scored 48.57 on the public split of Decision Index v0.2.1, matching Decider 35B-A3B, while d1-omni-600M scored 15.95. Across seven public text benchmarks, their mean scores were 82.9 and 78.4, respectively. The company measured d1-3B's input-to-output latency with one request at a time: a single question took 8 ms on an NVIDIA GeForce RTX 4090, 16 ms on a Jetson AGX Thor, 26 ms on a Jetson AGX Orin, and 50 ms on a Jetson Orin Nano. This release did not report results for the private vision split of Decision Index v0.3 or inference latency data for d1-omni-600M.
Perplexity releases two multimodal embedding models
模型发布
Perplexity has made the 0.6B and 9B multimodal embedding models in its pplx-embed-v2-late series available on Hugging Face. They support retrieval of text, images, and visual documents and share an embedding space, allowing queries across models’ indexes.
Perplexity has released the pplx-embed-v2-late series, comprising 0.6B and 9B multimodal late-interaction embedding models, now available on Hugging Face. Both are built on Qwen3.5, output a 128-dimensional vector for each token, and use MaxSim for scoring. They support retrieval of text, images, and visual documents, with no OCR required to process PDFs, slides, or scanned documents. Both were distilled from the same internal 18B ColBERT teacher model and share an embedding space, so the 0.6B model can directly query an index built by the 9B model. According to the official evaluation, the 9B model achieved a Recall@1000 of 74.8% on Q2D-Web, compared with 69.3% for nemotron-embed-8b.
生数科技 releases flagship video model Vidu Q4 Preview
模型发布
Shengshu Technology has released its flagship audio and video model Vidu Q4 Preview through its web product and API, with launch promotional pricing starting at $0.014 per second. The model supports up to 15 image references, 3 audio references, and 2K and 4K output.
Shengshu Technology launched Vidu Q4 Preview, the first public preview of its next-generation flagship audio and video model, on October 7, 2026, making it available through Vidu’s web product and API platform that day. Launch promotional pricing starts at $0.014 per second, with output resolutions including 540p, 720p, 1080p, 2K and 4K, alongside support for up to 15 image references, 3 audio references and 10-bit color depth. According to the company, comparable output specifications and billing conditions allow up to 5 times as much output for the same budget. Final pricing, supported resolutions, feature availability and usage terms may vary by plan and region.
According to the company, the model focuses on improving coordination among character expressions, emotions, body movements and voices, as well as between camera movements, editing and on-screen action, while making effects such as explosions, fireworks and particles fit more naturally into scenes. Its target users include independent creators, small studios and production teams working on narrative and commercial projects. The preview precedes the full Q4 model and lets creators use it in real projects. Feedback on performance, creative control and everyday production needs will inform subsequent improvements.
Musk: Grok Bot will no longer use only in-house models
产品应用
Musk announced that SpaceXAI’s Grok Bot will move beyond in-house models and select backends by task, with access to third-party APIs including Claude Opus 5.5, Midjourney and Suno. Models will also be assigned according to question complexity.
Musk announced on X that SpaceXAI’s Grok Bot will select backend models by task going forward, rather than remain limited to its in-house Grok series; the source did not specify an effective date. According to his explanation, available third-party APIs include Claude Opus 5.5, Midjourney and Suno, with selection based on which model is most likely to deliver the best result for the user. He subsequently added that simple questions will go to small, fast models, while only questions requiring complex answers will be assigned to large models.
Grok Bot 0.68.1 released with slide creation and formatted email support
产品应用
SpaceXAI has released Grok Bot 0.68.1 with slide creation and formatted email sending. It can deliver presentations as PowerPoint files or Google Slides, and the screen resolution used for computer operation has increased from 1280×800 to 1920×1200.
SpaceXAI has released Grok Bot 0.68.1 with updates to slides, email, and computer operation. Users can work with the Bot to create slides, choose from examples it presents, and receive the results as PowerPoint files or Google Slides. Email features support sending formatted messages containing headings, bold text, and lists from draft cards, as well as searching Gmail’s spam folder and moving messages back to the inbox. According to the company, the Bot operates computers faster, and the screen resolution it uses has changed from 1280×800 to 1920×1200. In one-on-one chats, user messages take on the Bot’s color. This feature requires Grok Bot 0.64.0 or later and can be disabled by setting the accent color to black in settings.
Grok Bot can now search and read X, available to all users with no setup required
产品应用
Grok Bot announced that all users can now search, read, and monitor X content without setting up an X connector. A team member also demonstrated a workflow that routes feature requests and bug reports from product feedback to development tools.
Grok Bot announced that its X content search, reading, and monitoring capabilities are now available to all users, with no X connector setup required. For handling product feedback, team member lauren demonstrated a workflow in which Grok Bot monitors user feedback about her product on X, sends feature requests to an issue tracker, and passes bug reports to Cursor cloud agent or a project for classification and fixes. Other uses listed by the official team include following breaking news and compiling weekly summaries of industry developments.
SpaceXAI Grok Bot team member Larsen Cundric said X has launched @bot mentions, allowing instructions in posts to be sent to a user’s Grok Bot. Users must first link their X account to a grok.com account with bots enabled.
SpaceXAI Grok Bot team member Larsen Cundric said the @bot mention feature on X is now live. According to his explanation, users must first link their X account to a grok.com account with bots enabled. They can then mention @bot and include instructions in replies, posts, or quote posts, sending tasks directly to their Grok Bot. Example instructions include adding content to a Notion reading list, setting a reminder to read it later, summarizing a post thread, and drafting a reply.
OpenAI to launch college application planning tools for ChatGPT for Teens
产品应用
OpenAI announced that College Planner is coming to ChatGPT for Teens, which applies automatically to accounts belonging to users under 18. The tool will consolidate college application and financial aid tasks, initially serving US students in grades 10–12 who plan to attend four-year colleges.
OpenAI outlined updates to ChatGPT for Teens, with the college application planning tool College Planner set to launch; no specific launch date was announced. The tool will bring school application requirements, deadlines, tasks, and financial aid application steps into a plan that can be continually updated. It will initially serve US high school students in grades 10–12 preparing to apply to four-year colleges. According to the company, coverage will later expand to more countries and to institutions including two-year colleges and technical colleges. ChatGPT for Teens applies automatically to accounts belonging to users under 18.
Among the learning features, iOS now supports photographing multiple pages in succession and combining them into a PDF, while the Android version remains in development. Flashcards and easier-to-create quizzes have also been introduced. OpenAI said teenagers spend an average of less than 15 minutes per day using the service, with nearly 1.2 million teenagers using learning visualization features and more than 180,000 using study mode within one week. OpenAI will also support College Advising Corps in expanding to additional states and establish an AI scholarship program. It will support the student advisory council at Boston Children’s Hospital’s Digital Wellness Lab over the next three years.
OpenAI launched Codex Cloud with support for Tailscale, currently its only VPN provider, without a joint integration project between the companies. The feature lets cloud-based Codex access resources inside a tailnet, but supports only specific connection methods.
OpenAI included Tailscale support when Codex Cloud launched and is gradually making the product available to paid ChatGPT accounts on Plus and higher tiers. Codex Cloud works on GitHub repositories in an environment that is not tied to a particular device; connecting Tailscale also enables access to internal APIs and testing services in a tailnet. According to Tailscale, there was no partnership agreement, custom API, or integration project between the companies, and OpenAI used the same Tailscale available to all users. It is currently the only VPN provider supported by Codex Cloud.
Configuration is available under Advanced > VPN in the cloud environment, and users must generate an auth key in the Tailscale admin console. OpenAI’s instructions require enabling Reusable and Ephemeral and tagging the key so that Codex VMs join the network with tag:codex and receive the corresponding rules rather than inheriting the permissions of the user who generated the key. Supported destinations currently include HTTP, HTTPS, and arbitrary TCP through HTTP CONNECT on proxy:8088. Private IPv4 subnet routes are supported, but require IPv4 addresses or hostnames that resolve to IPv4. MagicDNS, split DNS, UDP, ICMP, inbound connections, SSH, and Tailscale SSH are not supported. Tailscale says that, by default, Codex can access nearly all resources in a tailnet.
Google launches experimental game generation platform Playground
产品应用
Google has launched Playground, an experimental game-generation platform, for U.S. users aged 18 and older, enabling game creation, testing and sharing through text. Creation access will roll out in tiers based on Google AI subscriptions; Unity Spark integration is not yet available.
Google has launched Playground at playground.google for U.S. users aged 18 and older, with creation access rolling out in tiers based on Google AI subscriptions. According to Google, the experimental platform requires no coding experience: users can enter text through a conversational interface, start with a blank project, starter prompts or guided support, and immediately test their games. They can also adjust physics, rules, characters and environments. The platform runs in a browser, allowing users to play community games on a phone or laptop. Games can remain private, be shared through a link or be published to the Playground Explore gallery.
Community features include in-game leaderboards and multiplayer gameplay in select genres. A Play Games profile lets users set a custom handle, like games, compete in rankings and follow creators. Google says the gallery recommends games based on player ratings and play activity, and every published game undergoes safety screening aligned with its Community Guidelines, alongside user reporting. Playground will later integrate with Unity Spark, which Google says provides professional-level game mechanics, high-fidelity 3D capabilities and the flexibility of the Unity runtime. Unity Spark is currently in testing, with a closed beta coming soon.
谷歌 launches AI Edge Foresight, an on-device AI meeting assistant for Mac
产品应用
Google has launched AI Edge Foresight, an on-device meeting assistant for Mac that uses EmbeddingGemma 2 and Gemma 4 for meeting and file retrieval. EmbeddingGemma 2 has 740M parameters and, according to Google, maps text, images, video frames, and audio into a unified vector space.
Google has launched Google AI Edge Foresight, an experimental meeting assistant for Mac; the source does not specify an exact launch date. According to Google, the app runs on EmbeddingGemma 2 and Gemma 4, handling note-taking assistance and the indexing and retrieval of conversation transcripts and private files locally on the device. It connects directly to system audio and the microphone, works with any meeting platform, and can operate entirely offline. Google says Foresight maps text, audio, and visual data into a unified vector space to search a personal knowledge library, with sensitive data remaining on the device. The open-weight EmbeddingGemma 2 model, released alongside it by Google DeepMind, has 740M parameters. According to the company, active RAM requirements on a Google Pixel 11 Pro can be as low as approximately 191MB for text-only weights and approximately 567MB for the full multimodal model.
Google AI Edge Gallery has added two showcases using the same model: Instant Media Search and Video Moments Finder. The former searches local images and videos using natural language or example images, stores media vectors in a local SQLite database, and returns results based on cosine similarity. The latter locates visual moments in local videos without transcribing audio or generating intermediate text descriptions. Google says it will provide the model as a service on Android through ML Kit in the coming weeks, supporting NPU acceleration on devices where available and automatic model updates. MediaPipe Tasks is adding EmbeddingGemma 2 support to Embedder and Semantic Retriever to abstract media preprocessing and data transformations. MediaPipe Decision Task can use the model to classify images, text, or audio on-device without fine-tuning.
DGX Station to support Windows and run trillion-parameter-scale models locally
产品应用
NVIDIA and Microsoft previewed DGX Station with Windows support. According to the companies, its 748GB of memory and up to 20 petaFLOPS of FP4 AI compute enable local execution of trillion-parameter-scale models, allowing developers to handle large AI workloads within Windows.
NVIDIA and Microsoft previewed DGX Station for Windows at the Windows AI and Surface event in San Francisco, bringing the previously Linux-based DGX Station into the Windows environment. The system uses the GB300 Grace Blackwell Ultra Desktop Superchip and provides 748GB of coherent memory and up to 20 petaFLOPS of FP4 AI compute; according to the companies, this configuration can run models of up to a trillion parameters locally. They say developers and researchers can fine-tune and run inference on large models on their primary machine, while running always-on AI Agents that connect directly to existing Windows applications and infrastructure. Linux AI toolchains remain available through WSL when needed. The previous deployment arrangement required enterprise developers to maintain separate environments: Linux for heavy AI workloads and Windows for productivity tools, applications and workflows.
Surface Laptop Ultra opens for preorders, starting at $2,599
产品应用
Microsoft has opened pre-orders for Surface Laptop Ultra and Surface RTX Spark Dev Box, with the laptop starting at $2599. According to the company, both devices use NVIDIA RTX Spark and support running AI workloads locally.
Microsoft opened pre-orders for Surface Laptop Ultra and Surface RTX Spark Dev Box on October 7, 2026, with the former starting at $2599. Both devices are built around NVIDIA RTX Spark Superchip and offer up to 128 GB of unified memory. Laptop Ultra includes an NVIDIA Blackwell RTX GPU with up to 6,144 cores and an NVIDIA Grace CPU with up to 20 cores. According to Microsoft, the laptop can run AI models exceeding 120B parameters locally, delivers up to 1 petaflop of AI performance, and supports Agent coding tasks using local models with GitHub Copilot. It has a 15-inch PixelSense Ultra touchscreen and user-removable storage. Designed for desktop development, Dev Box belongs to the Project Zenith device family and comes with preinstalled tools and environments including Visual Studio Code, Git, GitHub CLI, GitHub Copilot, WSL, Python, and Node.