OpenAI says it has achieved its goal of an automated research intern
要闻
OpenAI says its AI system has reached the “automated research intern” goal and can complete well-defined tasks under human direction that would take a skilled researcher several days. The company plans to develop an automated AI researcher by March 2028.
OpenAI announced in its blog post, “Research acceleration: The view inside OpenAI,” that it had completed the automated research intern milestone and set a deadline of March 2028 for developing an automated AI researcher. According to the company, the current system can handle clearly scoped research tasks under human direction that would typically take a skilled researcher several days. As of the disclosed mid-August cutoff, the median daily inference cost for coding Agents used by its internal research team exceeded $600 at API prices, while the 90th percentile exceeded $7,000. Assuming an eight-hour workday, the Agents’ total runtime was equivalent to 3.1 times the total human working hours. The team also increased the pace of code submissions and experiments, with experiments per active experimenter reaching their highest level in August since tracking began in January 2025.
The role of Agents in daily research has expanded from writing code to technical troubleshooting and experiment monitoring. OpenAI says task success rates have also risen, although high-level planning still accounts for only a very small share of Agent output. During the preceding six-month period described by OpenAI, more than half of the tasks successfully completed by Agents and estimated to require 4–8 hours of human work still involved at least one human intervention. OpenAI also says compute capacity and steps that are difficult to automate will continue to constrain research, so growth in these usage and output metrics cannot be treated as evidence that overall research progress has accelerated year over year.
OpenAI says chain-of-thought monitoring is becoming less reliable and calls for safety standards
要闻
OpenAI chief scientist Jakub Pachocki says GPT‑6 Astra outperforms GPT‑5.6 Sol on alignment, while chain-of-thought monitoring is becoming less reliable, and calls for binding safety standards enforced under external oversight.
OpenAI chief scientist Jakub Pachocki said in An Alien Mind that OpenAI is shifting its research focus toward recursive self-improvement and automated alignment, while continuing work on alignment, monitoring, and defense systems and proactively pausing further model scaling when necessary. He said the current pace of AI progress is likely to continue until recursive self-improvement, when AI would play a greater role in advancing its own research and development. He also said no lab has yet made enough progress on alignment and monitoring to sustain scaling at maximum speed for an extended period while ensuring safety.
According to Pachocki, GPT‑6 Astra performs better on alignment than GPT‑5.6 Sol, but evaluations indicate that chain-of-thought monitoring is becoming less reliable. He attributed this to the growing integration of model reasoning with external interactions, models becoming more capable of manipulating their own reasoning processes, and their ability to complete more complex tasks without explicit reasoning. He proposed binding safety standards for future development, enforced under the supervision of third-party auditors, governments, or international organizations. Until common standards are established, he said labs should voluntarily slow development and strengthen international coordination so that humans remain involved in AI self-improvement and retain control over future decisions.
OpenAI Codex completes Astra usage optimization with no loss in quality
要闻
OpenAI Codex lead Tibo announced that the team has improved Astra. According to him, advanced users who sign in with a ChatGPT account can reduce subscription usage to about one-quarter of its previous level in long-tail scenarios without a change in quality.
The OpenAI Codex team has completed several usage improvements for Astra and made them applicable to advanced users who sign in with a ChatGPT account, Codex lead Tibo announced through his personal account. According to his explanation, the changes target long-tail scenarios and can reduce subscription usage to about one-quarter of its previous level without changing quality, equivalent to savings of up to 3x to 4x. That figure is an upper limit for long-tail scenarios and will not be reached in every use case. Tibo had also previously offered guidance on reasoning-effort calibration: GPT-6 Astra at low performs better than GPT-5.6 Sol at high, so users who were satisfied with Sol at high could switch to Astra at low or medium.
Jensen Huang reveals GPT-6 Astra was trained using more than 100,000 GB200 NVL72 systems
行业动态
Jensen Huang said GPT-6 Astra was trained using roughly 100,000-plus NVIDIA Grace Blackwell NVLink72 systems. OpenAI’s Stargate data center in Abilene, Texas, is also planned to house up to approximately 400,000 advanced AI chips.
Jensen Huang disclosed in a social media post that GPT-6 Astra’s training was supported by roughly 100,000-plus NVIDIA Grace Blackwell NVLink72 systems and that OpenAI will put another 400,000 GPUs into use. Citing a timeline of only 4 years from ChatGPT through o1 to Astra, he claimed that AGI had arrived and congratulated the OpenAI team. According to his explanation, the additional computing capacity corresponds to OpenAI’s Stargate data center in Abilene, Texas, which is planned to accommodate up to approximately 400,000 advanced NVIDIA AI chips.
Research team proposes the concept of “AI-related psychosis”
技术与洞察
The title of arXiv page 2608.23937v1 says a research team proposed the concept of AI-related psychosis. However, the supplied text identifies no team members and provides no definition, methodology, or findings, discussing only the functions and partnership principles of arXivLabs.
The page for arXiv 2608.23937v1 is titled “Research Team Proposes the Concept of ‘AI-Related Psychosis,’” but the supplied text contains no abstract, authors, publication date, methods, samples, data, or conclusions, so the concept’s definition and research basis cannot be established from it. The page text only explains that arXivLabs is a framework that allows collaborators to develop and share new features directly on the arXiv website. According to the official description, participating individuals and organizations accept the values of openness, community, excellence, and user data privacy, and arXiv works only with partners that adhere to those values. The page also invites proposals for projects intended to add value to the arXiv community and directs their creators to learn more about arXivLabs.