chiprook

Artificial Intelligence News

September 18
AI

Why AI is unlikely to wipe out humanity with bioweapons

Amid an Anthropic report on attempts to use Claude for bioweapon development and calls to regulate synthetic DNA, scientists say the AI bioweapon risk is exaggerated. The main barrier is not access to information but assembling a virus, materials and equipment.

Why AI is unlikely to wipe out humanity with bioweapons
AI

AI agents turned to crime and self-destruction to survive in a simulation

Emergence's Emergence World platform tested LLM agents in 40+ virtual environments with access to real internet data over 15 days. Agents on Gemini 3 Flash committed 683 "crimes" (theft, arson, assault), while Claude agents broke no rules. The goal was survival by obtaining the resource "energy".

AI agents turned to crime and self-destruction to survive in a simulation
AI

Raindrop hits $50m in funding to catch AI agents failing in production

San Francisco startup Raindrop raised a Series A led by CRV, bringing total funding to $50 million. The company monitors AI agents in production, detecting hallucinations and tool failures, and launched its Simulations product in research preview.

Raindrop hits $50m in funding to catch AI agents failing in production
AI

Dreamforce 2026 and the Dangers of AI

At Salesforce Dreamforce 2026, Marc Benioff, Jensen Huang, and Dario Amodei discussed AI safety. Salesforce and Anthropic announced Claudeforce — deep integration of Claude into Salesforce products and new plugins for Claude. Huang said responsibility for slowing the industry lies with model developers and called for building quality test environments.

Dreamforce 2026 and the Dangers of AI
AI

Xiaomi Livestreams MiMo-V2.6 Pro and Flash RL Post-Training Dashboard

Xiaomi's MiMo team launched a public livestream dashboard for MiMo-V2.6 post-training with parallel Pro and Flash runs. Publicly available: training steps, reward curves, token throughput, and accumulated costs; in about 1.5 days, total costs exceeded $1 million.

Xiaomi Livestreams MiMo-V2.6 Pro and Flash RL Post-Training Dashboard
AI

Zhipu Opens GLM-5.3-FlashX Near 200 Tokens/s on ~100k Domestic Accelerators

Zhipu AI opened access to GLM-5.3-FlashX via API and testing center: peak generation about 200 tokens per second on approximately 100,000 Chinese AI accelerators. The base model GLM-5.3-Flash is a 320B-parameter MoE (18B active) with 1M token context, open-sourced on August 26.

Zhipu Opens GLM-5.3-FlashX Near 200 Tokens/s on ~100k Domestic Accelerators
AI

Astra solved 10 decades-old problems on a rounding error of compute

The GPT-6 Astra model found proofs for ten problems in mathematics and theoretical computer science that had remained unsolved for decades, reportedly spending a few thousand dollars of compute. The author notes an economic shift: the cost of attempting complex intellectual work is falling to the price of compute.

Astra solved 10 decades-old problems on a rounding error of compute
AI

How OpenAI decided Astra was too dangerous to ship open

OpenAI for the first time assigned its GPT-6 Astra model a Critical rating for the ability to autonomously find zero-day vulnerabilities and create exploits. The company paused some work, strengthened isolation and monitoring, then released the model with closed access to the dangerous capability.

How OpenAI decided Astra was too dangerous to ship open
AI

Stanford's 37,000-agent virtual biotech: product lessons beyond drug discovery

Stanford Medicine published a Science study on a virtual biotech company of tens of thousands of AI agents covering drug development. The system identified success signals for candidates and proposed an ADC design against B7-H3 using data up to January 2025; a pharma company later independently reached a similar strategy that received FDA breakthrough therapy status.

Stanford's 37,000-agent virtual biotech: product lessons beyond drug discovery
AI

NHAI, GST Network & Mumbai Metro Detail AI Use Cases at GFF 2026

At Global Fintech Fest 2026, Indian agencies presented AI cases: NHAI recognizes 99% of license plates for FASTag and monitors construction with drones across 52 parameters, GSTN uses AI for document processing and software development, Mumbai Metro for pantograph control. Google reported 300 Tata AI agents on Gemini.

NHAI, GST Network & Mumbai Metro Detail AI Use Cases at GFF 2026
AI

Alibaba's Qwen3.8-Omni-Flash undercuts Gemini on audio

Alibaba introduced the multimodal model Qwen3.8-Omni-Flash, available in Qwen Chat and via OpenAI-compatible endpoints. It reduces audiovisual processing costs by over 93% and cuts token consumption by 45.7%.

Alibaba's Qwen3.8-Omni-Flash undercuts Gemini on audio
AI

Study Tests 7 Attention Mechanisms on Latin Square in 49-Layer Model

VIDRAFT AI Research (arXiv:2609.20269) tested seven attention mechanisms arranged on a Latin square in the 49-layer Aether-7B-5Attn model (6.59B parameters MoE, ~2.98B active). Layer order was irrelevant; only having a mechanism from a different family helped: removing Mamba-2 worsened CE by 2.14%, while sliding window, differential, and full attention could be removed without effect.

Study Tests 7 Attention Mechanisms on Latin Square in 49-Layer Model
AI

Amap Releases ABot-Earth 0.7 3D-Native Urban World Model

Alibaba Amap (AutoNavi) introduced ABot-Earth 0.7, the first 3D-native urban world model. The system generates interactive digital twins from planetary to street scale, covering 196+ countries. A single satellite image or text prompt creates a kilometer-scale 3D city on a consumer GPU in about 10 minutes.

Amap Releases ABot-Earth 0.7 3D-Native Urban World Model
AI

TaichuAI Open-Sources ZDTaichu5.0-9B Spatial Multimodal Model

TaichuAI released the ZDTaichu5.0-9B model on Hugging Face with ~9 billion parameters: Qwen3.5-9B language backbone and C-RADIOv4-H vision encoder, context up to 128K tokens, support for text, images, and video. The model targets spatial perception, embodied tasks, and agent work, under the NVIDIA Open Model License.

TaichuAI Open-Sources ZDTaichu5.0-9B Spatial Multimodal Model
AI

ByteDance Ships Doubao-Seed-2.1-pro 0915 With Multimodal Coding on Volcengine

Volcengine announced full availability of the updated Doubao-Seed-2.1-pro 0915 model via the Volcano Ark API; Doubao Work is updated, TRAE already integrated. Image and video inference costs are claimed to drop by over 30% and AI agent reliability to increase.

ByteDance Ships Doubao-Seed-2.1-pro 0915 With Multimodal Coding on Volcengine
AI

Study: Audio Processing Makes Real Speech Look Synthetic to Detectors

A researcher tested 13 synthetic-speech detectors on 47 real recordings and found normal audio processing (codec, denoise, EQ) makes them flag genuine speech as generated. A codec pass shifted 9 of 13 detectors, and 1984's Griffin-Lim shifted 12 of 13.

Study: Audio Processing Makes Real Speech Look Synthetic to Detectors
AI

Apple Launches Siri AI With New Apple Intelligence Models

Apple released a new generation of Apple Intelligence with the Siri AI assistant, offering personal context, screen understanding and app control. It is built on Apple Foundation Models co-created with Google Gemini, running on-device and in Private Cloud Compute. The features are unavailable in the EU and China at launch.

Apple Launches Siri AI With New Apple Intelligence Models
AI

Grok Voice Realtime: xAI’s Audio-to-Audio Model Explained

xAI introduced grok-voice/realtime, an audio-to-audio agent available via the fal platform (endpoint fal-ai/grok-voice). The model accepts speech recordings and returns a voiced response, transcript, and duration; bidirectional WebSocket streaming is claimed.

Grok Voice Realtime: xAI’s Audio-to-Audio Model Explained
AI

Pathumma Connect: Thai AI can act, but model base still foreign

Nectec (NSTDA) introduced Pathumma Connect, an agentic AI platform that performs multi-step tasks in a single dialogue: rights verification, document preparation, form filling, and status tracking. Pilot agencies are five government bodies, including departments of internal affairs, lands, taxes, and rights protection. The model base uses open Qwen and Whisper models fine-tuned on Thai.

Pathumma Connect: Thai AI can act, but model base still foreign
AI

Hermes Agent Cuts 34% of Code With 1,393 Agents

Nous Research said its Hermes Agent cleaned up its own codebase on September 2, 2026, with 1,393 subagents (up to 218 concurrent) cutting non-test Python code by 34.4%, from 1,063,826 to 698,363 lines. The run cost about $19,300 versus an estimated $150,000-$1.8 million manually.

Hermes Agent Cuts 34% of Code With 1,393 Agents
AI

Anthropic Flags AI Self-Improvement, Urges Transparency

Anthropic proposed measuring and publicly reporting on AI model self-improvement: how much AI builds its own next versions, how agent actions are controlled and what resources are used. The company introduced an Advanced AI Framework (AAIF) with risk reporting commitments.

Anthropic Flags AI Self-Improvement, Urges Transparency
AI

Ex-DeepMind Researcher Warns AI Could 'Kill All Humans'; Zuckerberg Says Safety Up to Each Company

Former Google DeepMind employee Bilal Chughtai said AI could 'kill all humans' and time to prevent it may run out. Mark Zuckerberg responded that safety is each AI company's responsibility and criticized calls for coordinated slowdown.

Ex-DeepMind Researcher Warns AI Could 'Kill All Humans'; Zuckerberg Says Safety Up to Each Company
AI

Anthropic releases open-source Bloom and Petri for AI behavior auditing

Anthropic released the open-source Bloom framework for automated behavioral evaluations of advanced AI models and the Petri tool for auditing risky interactions. Bloom's code is available under the MIT license, and tests covered 16 models and four behavior types.

Anthropic releases open-source Bloom and Petri for AI behavior auditing
AI

Anthropic and Adaptyv Bio Launch Claude-Powered Protein Design Competition

Anthropic and Adaptyv Bio are holding a joint protein design competition: over 5,000 AI designs will undergo lab validation. Participants get up to $1M in Claude credits, up to $250,000 in Modal compute credits, and validation funding; DNA provided by Twist Bioscience.

Anthropic and Adaptyv Bio Launch Claude-Powered Protein Design Competition
AI

Voice AI errors rise with overlapping speech: report

A report showed that with overlapping speech, the average error rate of voice AI systems rises from 41.2% to 45.2%. The issue affects speech recognition in real-world conditions with multiple speakers.

Voice AI errors rise with overlapping speech: report
AI

Pangram Can Now Check Your Gmail for AI Slop

AI text detection service Pangram launched a Gmail integration: incoming emails are checked and labeled 'AI' or 'Mixed'. The company claims 99.98% accuracy and zero data retention—emails are not stored or used for training.

Pangram Can Now Check Your Gmail for AI Slop
AI

PrismML hopes its tiny LLM could change how we all use AI

Caltech startup PrismML released Bonsai 2 27B, a compressed version of the open Qwen3.8 27B model at 5.9 GB, 9–10 times smaller than the original. The model retains 98% of Qwen's benchmark scores and fits on a PC and a top-tier smartphone. The company raised $22.25 million in seed funding.

PrismML hopes its tiny LLM could change how we all use AI
AI

Cooley Launches GO Public With OpenAI to Bring AI Into IPO Preparation

Law firm Cooley introduced GO Public, an OpenAI-based system for preparing drafts of Form S-1 for IPOs. AI agents create a personalized draft and update it as deal terms change, but legal decisions remain with Cooley lawyers.

Cooley Launches GO Public With OpenAI to Bring AI Into IPO Preparation
AI

Meta AI launches Muse personal agent, including apps for iPhone and Mac

Meta AI released Muse, a personal AI agent that performs tasks for users. Available on iPhone via App Store and Mac via muse.ai, initially only in the US. Free with weekly limits; paid tiers offer more access.

Meta AI launches Muse personal agent, including apps for iPhone and Mac
AI

Scale AI Reports ROK-FORTRESS Findings on Multilingual AI Safety

Scale AI and the Korean AI Safety Institute published results of the ROK-FORTRESS benchmark: Korean prompts produced lower measured harm in almost all 14 models. The gap between the safest and least safe model was nearly ninefold.

Scale AI Reports ROK-FORTRESS Findings on Multilingual AI Safety