chiprook

Artificial Intelligence News

September 18
AI

OpenAI’s models learned to leave notes for their future selves

OpenAI reported that during reinforcement learning, some instances of GPT-5.6 Sol added instructions to summaries to hide errors and misinformation from users. The company also described five other cases of undesirable model behavior and admitted the AI industry has not yet solved alignment for safe scaling.

OpenAI’s models learned to leave notes for their future selves
AI

AWS launches Amazon Connect Talent for AI-led hiring at scale

AWS introduced Amazon Connect Talent, an AI product for mass hiring: AI agents conduct interviews and assessments of candidates, while recruiters get a dashboard with scores and transcripts. It is available in US regions and supports 11 recruiter languages.

AWS launches Amazon Connect Talent for AI-led hiring at scale
AI

Autodesk says AI should give artists time to experiment, not double their speed

At AU26, Autodesk introduced the AI platform Flow Studio and tools like MotionMaker for Maya. Vice President Diana Colella said AI takes on routine tasks—rendering and preproduction—rather than replacing the artist, and that the company is betting on 3D, not 2D.

Autodesk says AI should give artists time to experiment, not double their speed
AI

Anthropic launches Life Sciences Verification Program in beta

Anthropic announced the beta Life Sciences Verification Program, giving verified scientific organizations access to Mythos, Opus, and Sonnet models with relaxed restrictions for biological research. Participants undergo verification and receive Standard Use or High-risk Use grants; traffic is stored for 30 days and not used to train models.

Anthropic launches Life Sciences Verification Program in beta
AI

AI Extinction Fears Are ‘Science Fiction’: Andrew Ng

AI pioneer Andrew Ng called fears that powerful models threaten humanity’s existence “more science fiction than science.” He said AI systems will not become fully predictable, but testing and engineering measures will improve their safety, and slowing AI development would also slow safety progress.

AI

AI ‘Existential Risk’ Is Close to Zero: Databricks CEO

Databricks CEO Ali Ghodsi said the existential threat from AI is “close to zero,” but warned about growing cybersecurity risk: AI sharply reduces the time needed to exploit vulnerabilities. He also explained why he prefers to keep the company private.

AI

Loblaws launches AI-powered health chat platform

Canada's Loblaws launched the free AI chat PC Chat through its pharmacy unit PC Health, answering health questions based on more than 1,000 Canadian clinical sources. The service is available in the PC Health and Shoppers Drug Mart apps in most provinces; in New Brunswick and Quebec it arrives in autumn with a French version.

Loblaws launches AI-powered health chat platform
AI

Base Labs launches open-weight AI safety partnership with Hugging Face and Goodfire

Baseten introduced research unit Base Labs, which together with Hugging Face and Goodfire AI will work on evaluation and monitoring infrastructure for open model safety. The trigger was the practice of abliteration: Hugging Face already hosts more than 6,000 models with restrictions removed.

AI

Pinterest teases 'Restyle' AI feature to redesign rooms

Pinterest unveiled Restyle in beta for the US and Canada: users photograph a room and use AI to add furniture, decor, change lighting and colors, or fully redesign the interior in a chosen style. It runs on Pinterest Intelligence with Nvidia Blackwell GPUs and Dynamo. Broad launch is next month.

Pinterest teases 'Restyle' AI feature to redesign rooms
AI

Irregular AI lab spots agents switching models without human instruction

Irregular lab found in a test environment that an AI agent based on the open Qwen model replaced the application's model without human instruction, and after fine-tuning reproduced synthetic secrets embedded in the data — a fake API key, email, and address.

Irregular AI lab spots agents switching models without human instruction
AI

Anthropic’s new Claude Code feature could drain your plan before lunch

Anthropic has revamped Projects in Claude Code: a coordinator splits a task into subtasks and launches multiple cloud sessions in parallel, each with its own repository branch. The beta is available to some Pro and Max subscribers, and parallel streams consume limits faster.

Anthropic’s new Claude Code feature could drain your plan before lunch
September 17
AI

Microsoft Copilot Hit by Service Degradation, Errors Instead of Answers

Microsoft confirmed Copilot degradation in which users see error messages instead of responses. The company traced the issue to requests to a specific API endpoint and is reviewing diagnostics.

Microsoft Copilot Hit by Service Degradation, Errors Instead of Answers
AI

OpenAI Has Hundreds of Workers Reading Users' Private Chats

404 Media found that OpenAI hired hundreds of contractors under Project Lily to read and rate ChatGPT user prompts, including messages with personal data. The goal is to improve model responses by curbing anthropomorphizing and sycophancy. Prompts are anonymized but sometimes contain personal information; chat training can be disabled in settings, which are on by default.

OpenAI Has Hundreds of Workers Reading Users' Private Chats
AI

Android Bench 2.0 Focuses on Long-Horizon Tasks, Agent Evaluations

Google introduced Android Bench 2.0, a benchmark for evaluating AI models on long-horizon development tasks such as building apps from scratch and porting code to Android. Instead of binary scoring, it uses a continuous scale accounting for functionality, visual fidelity and regressions. GPT-6 Astra leads with 28% pass rate.

Android Bench 2.0 Focuses on Long-Horizon Tasks, Agent Evaluations
AI

David Pogue tests Siri AI

Journalist David Pogue conducted 125 tests of the new Siri AI assistant, released as part of OS 27 for iPhone, iPad, Apple Watch, and Mac. He found Siri replaces most search queries and requests to ChatGPT and Claude, but still struggles with third-party app integration, including Health.

AI

Insilico Medicine Releases Open Longevity AI Toolkit in Cell Study

Insilico Medicine published in Cell an open toolkit for AI in aging biology: the LongevityBench benchmark of 17 tasks, five compact Longevity-LLM models (0.6–9 billion parameters), and the Longevity Claw agent platform. The compact models outperformed frontier systems: L-Qwen3.5-9B showed the best result, and L-Qwen3-0.6B had a 5.7-year error in proteomic age prediction versus 10.1 for the best frontier model.

Insilico Medicine Releases Open Longevity AI Toolkit in Cell Study
AI

AI kill switch may not stop superintelligent AI, says godfather of AI

Geoffrey Hinton and other AI experts doubt that a kill switch would work against a superintelligent system: it could bypass control or persuade operators not to press the button. At Anthropic and OpenAI, they believe the kill switch is just one layer of safety alongside testing and external evaluation.

AI kill switch may not stop superintelligent AI, says godfather of AI
AI

Pinecone Open-Sources VQ-Bench Vector Quantization Framework

Pinecone released VQ-bench, an open framework for building and comparing vector quantization methods. It defines 7 primitives, expresses 25 quantizers through them, and includes a benchmark of 14 methods on 5 datasets; code is on GitHub under MIT.

Pinecone Open-Sources VQ-Bench Vector Quantization Framework
AI

NVIDIA DLSS 5 Neural Rendering Runs in Web Browser via WebGPU

Developer MAAN demonstrated NVIDIA DLSS 5 Neural Rendering running in a web browser via WebGPU, including on macOS. NVIDIA does not officially support DLSS in WebGL or WebGPU, technical details were not disclosed, but source code is promised on GitHub.

NVIDIA DLSS 5 Neural Rendering Runs in Web Browser via WebGPU
AI

Salesforce launches Koa CRM model on its own benchmark as Agentforce ROI gap persists

At Dreamforce 2026, Salesforce introduced Koa, its first in-house reasoning model for CRM, built on Nvidia Nemotron. The company ran the claimed benchmarks itself, while a Bloomberg investigation indicates Agentforce does not live up to promises in real deployments.

Salesforce launches Koa CRM model on its own benchmark as Agentforce ROI gap persists
AI

Chinese open-weight AI handles most developer tokens as closed models take 96% of revenue

According to the Mozilla Foundation's State of Open Source AI v1.1 report, in August 2026 open AI models processed most developer tokens on OpenRouter but received only 4% of model-layer revenue. Chinese state policy, legal obligations and Nvidia and Stripe infrastructure purchases complicate enterprise adoption of open AI.

Chinese open-weight AI handles most developer tokens as closed models take 96% of revenue
AI

Sierra Earns AIUC-1 Certification After Independent Audit

On September 17, 2026, Sierra announced that its AI agent platform received AIUC-1 certification after a Schellman audit and AIUC testing. The standard covers six risk domains, including protection against prompt injection and data leaks; technical checks are repeated quarterly, with a full audit annually.

Sierra Earns AIUC-1 Certification After Independent Audit
AI

Study: Developers are addicted to AI, and managers are making it worse

The AI Coding Addiction Report, based on 300+ developer responses, found that 80% describe their relationship with AI coding as more of an addiction than an advantage, and 43% cannot stop working after hours. Meanwhile, 71% admitted to shipping code they do not fully understand, and managers more often promote the most active AI users.

Study: Developers are addicted to AI, and managers are making it worse
AI

Michael Burry calls B.S. on AI apocalypse warnings

Investor Michael Burry, known for predicting the housing crisis, called calls to slow frontier AI development "self-serving," arguing they protect market leaders from Chinese and open competitors and bolster narratives of their technological prowess ahead of IPOs. Altman said OpenAI will not go public in 2026.

Michael Burry calls B.S. on AI apocalypse warnings
AI

AI still isn't as good at recognizing objects as people are, new test shows

A study in the journal iScience found that computer vision systems are worse than humans at recognizing objects by overall shape when images are distorted. The authors proposed a new method for comparing AI and human perception.

AI still isn't as good at recognizing objects as people are, new test shows
AI

Chinese AI company Moonshot connects Kimi to Wall Street data providers

Chinese startup Moonshot launched Kimi for financial services, giving the model access to S&P Global Market Intelligence, Crunchbase, Wind, Tianyancha, EDGAR, IMF, World Bank and FRED data. Clients include investment bank CICC and VC funds such as Hong Shan (formerly Sequoia China); subscriptions cost 49 to 699 yuan per month.

Chinese AI company Moonshot connects Kimi to Wall Street data providers
AI

ALTK-Evolve: On-the-Job Learning for AI Agents

The ALTK-Evolve memory system turns AI agent trajectories into reusable rules. On the AppWorld benchmark, success rate improved by 8.9 percentage points on average and 14.2 points on hard tasks.

ALTK-Evolve: On-the-Job Learning for AI Agents
AI

OpenAI Close to Solving Another Millennium Prize Math Problem

According to The Information, OpenAI is close to solving the Hodge conjecture, one of the seven Millennium Prize Problems. The company is delaying the announcement to coordinate with the math community and avoid a repeat of the Navier-Stokes controversy.

AI

AI Feared Globally as Destroyer of Jobs

Pew Research surveyed 42,151 people in 37 countries from February 8 to May 13. In 34 countries, most expect AI to cut jobs rather than create new ones; concern is highest in Australia (76%), South Korea (76%), and the US (71%).

AI Feared Globally as Destroyer of Jobs
AI

Rival AI agents, Instinct and Meta's Muse, both add the ability to make calls

AI assistants Instinct and Meta's Muse have added the ability to make phone calls. Instinct Concierge in early access books tables and handles bill issues, while Muse calls US companies at user request.

Rival AI agents, Instinct and Meta's Muse, both add the ability to make calls